Your update succeeds. The background indexer has not caught up. A query reaches a different node whose cache still contains the old document. Which version should the query return?
That question is about authority and time before it is about nearest-neighbor search. If a query can consult the committed update, it can use the new version without waiting for index construction. If a later deletion is ignored, the old index can resurrect a document that should have disappeared.
The primer’s time and disclosure vocabulary separates acknowledgment, durability, indexing, query visibility and authorization. Here we follow those boundaries through an update, deletion and recovery.
We continue the synthetic multi-tenant engineering-document service from the eligible-corpus chapter. The added requirement is explicit: writes may be acknowledged before asynchronous indexing finishes, and queries must declare their freshness behavior. A crash now interrupts that sequence. Reading time covers the prose; leave extra time for the state tables and recovery exercise.
Evidence boundary: The ordered records, cursors, request IDs and recovery rules below are a stipulated teaching model. They do not disclose turbopuffer’s implementation. Public API and storage contracts are attributed separately to documentation inspected on October 1, 2026. No live service, crash injection or benchmark was run.
Four states that a successful write does not collapse
A write-ahead log, or WAL, records changes durably so they can be recovered after a crash. An index is derived state: a structure built from document changes to make reads efficient. A cache is a local copy used to avoid fetching that structure again.
turbopuffer’s architecture documentation says successful writes are durably written to object storage. Its diagram distinguishes an index cursor from a commit point and from written-but-uncommitted data. Committed data is indexed asynchronously; recent unindexed data remains searchable through a slower exhaustive path. This establishes a useful architectural distinction, without specifying its complete recovery protocol.
| State | What it establishes | What remains to be established |
|---|---|---|
| Acknowledged | The service has returned the promised success response. In the cited API, that promises durable object-storage commit. | Which query mode includes the write. |
| Indexed | A published derived index incorporates the change. | Whether a particular reader selects that index and snapshot. |
| Cached | A node holds some copy of index/document data. | Whether the copy covers the selected read boundary. |
| Query-visible | The selected read path includes the change under its declared contract. | Whether the live document ranks in top-k, is relevant, or may be disclosed. |
A document can be query-visible through the recent-write path while unindexed and absent from a node’s old cache. Conversely, having bytes in a cache does not prove those bytes represent the selected version. Visibility also does not guarantee a top-k slot: a visible replacement can move farther from the query.
Before reading further, predict two failures. What happens if we search only an old base index after a replacement? What happens if we merge its hits with a deletion record but treat the deletion as an ordinary candidate with no vector score?
Give the paper model a clock
We stipulate one writer assigning an ordered stream of durable records. Each committed change has a unique increasing commit position. A document’s version increases on replacement or deletion. Each replacement supplies its complete searchable state; a deletion supplies a tombstone, a record saying that the document is deleted. These rules avoid having to invent a concurrent-writer ordering protocol.
Use three different positions:
- D, durable commit position: the end of the authoritative committed prefix. Complete-looking files beyond D are not committed by this model.
- I, published index cursor: an immutable base represents the latest state of each document through I. A new base is built separately, then published with its matching cursor.
- S, selected query snapshot: the query observes committed changes through S. For this model,
I ≤ S ≤ D; older snapshots use a retained compatible base.
The tail for this query is the committed interval after I through S. A snapshot is a stable logical read boundary, not the time at which each remote request happens. This model requires the reader to keep the chosen base and boundary together throughout candidate generation and projection, even if indexing advances during the query.
Write and derived index
- Commit the replacement
Durable authority advances to D=11; doc7 v2 is committed.
- Acknowledge, then crash
Success has returned. The published base remains I=10; losing worker memory does not erase position 11.
- Recover and publish
Rebuild from committed authority. Publish a complete replacement base with I=11 when ready.
- Commit a later delete
D=12 includes doc7's v3 tombstone. Index and cache may still lag.
Query while the base lags
- Pin base I=10
Keep this base and its document versions available for the read.
- Select S=12
Read the complete committed tail (10,12]; exclude the written-only position 13.
- Resolve document versions
doc7 v3 wins over v2 and the base's v1. The tombstone suppresses all older hits.
- Rank and project
Rank live eligible latest versions. Fetch the selected version's content; authorize before disclosure.
Stipulated single-writer model. A crash after acknowledgment leaves the committed record authoritative. Publication changes I; query selection chooses S. These are separate boundaries.
The trace in words: position 11 becomes durable before success is returned. A crash leaves I at 10, so recovery still needs position 11. Once a complete base is published at I=11, queries can omit that position from the tail. A later delete at 12 must still override older copies. A query using I=10 and S=12 needs positions 11 and 12 regardless of which node has cached the base. The table below provides every record needed to reproduce that query.
Reconcile the document before ranking its copies
Each dense query selects the body field with one vector per live document. Saved distances are squared Euclidean distances; lower is better, with ascending doc ID for ties. The trusted tenant is A unless a case says otherwise. Search eligibility is tenant membership and live state at S. Exact-neighbor fidelity, judged relevance and authorization under current policy remain separate tests.
Here is the complete synthetic record stream. The base at I=10 contains positions 8–10. A dash means a deletion has no searchable vector.
| position | doc_id | version | tenant | operation | distance | authority | request_id |
|---|---|---|---|---|---|---|---|
| 8 | 7 | 1 | A | replace | 0.10 | committed | seed7 |
| 9 | 8 | 1 | A | replace | 0.20 | committed | seed8 |
| 10 | 9 | 1 | A | replace | 0.20 | committed | seed9 |
| 11 | 7 | 2 | A | replace | 0.90 | committed | update7 |
| 12 | 7 | 3 | A | delete | - | committed | delete7 |
| 13 | 7 | 4 | A | replace | 0.01 | written-only | pending7 |
At S=11, reconcile doc7’s base v1 with its tail v2 and keep v2. Its current distance is 0.90, so doc8 and doc9 now precede it. At S=12, v3 wins and suppresses doc7 completely. The very attractive v4 at position 13 has no authority to participate. Taking the maximum version across arbitrary files would therefore be wrong: restrict to committed records within the snapshot first.
For the exact paper oracle, the steps are:
- Select a compatible published base at I and all committed tail records in
(I,S]. - For each doc ID, resolve its latest version at S, including tombstones. An older hit cannot survive because the replacement ranks worse or has no vector.
- Apply tenant/live eligibility to that resolved state, compute its selected-field distance, then sort distinct documents by
(distance, doc_id). - Take up to k and project content from those same selected versions.
An efficient engine can implement different physical plans. It still needs an equivalent version-resolution rule and enough candidates to honor its stated retrieval contract. Our exhaustive toy calculates exact neighbors; including all committed writes does not turn a production ANN algorithm into exact search.
The following selected snapshots are deliberate historical reads with D=12. They do not describe fresh strong queries starting after position 12. Predict doc7’s winning version and the exact result before opening the answer.
| case | I | D | S | tenant | k |
|---|---|---|---|---|---|
| original | 10 | 12 | 10 | A | 2 |
| replacement | 10 | 12 | 11 | A | 2 |
| deletion | 10 | 12 | 12 | A | 2 |
| newer-base | 11 | 12 | 12 | A | 2 |
| caught-up | 12 | 12 | 12 | A | 2 |
| all-live | 10 | 12 | 11 | A | 5 |
| empty-tenant | 10 | 12 | 12 | C | 2 |
| incompatible-base | 11 | 12 | 10 | A | 2 |
Check version resolution, ties, and compatible bases
| case | doc7_version | doc7_state | exact_ids |
|---|---|---|---|
| original | 1 | live | [7,8] |
| replacement | 2 | live | [8,9] |
| deletion | 3 | deleted | [8,9] |
| newer-base | 3 | deleted | [8,9] |
| caught-up | 3 | deleted | [8,9] |
| all-live | 2 | live | [8,9,7] |
| empty-tenant | 3 | deleted | [] |
| incompatible-base | - | invalid | ERROR |
The original snapshot uses v1 even though D has advanced. Replacement makes v2 visible but moves doc7 out of top-2. The delete is effective before it reaches the base; advancing I changes the amount of tail work, not the resolved answer at S=12. Doc8 wins the distance tie with doc9 by ID.
Requesting five at S=11 returns the three live eligible documents. Tenant C has none, so its answer is empty. I=11 cannot answer S=10 from a latest-state-only base: v1 has already been replaced there. Select a retained base at or before S, or reject the historical read. Do not silently use the future base. A pass includes every winning version, the selected boundary, tie order and the incompatibility explanation.
Guided failure: delete after truncation
At I=10, the stale base’s exact top-2 is [7,8]. At S=12, suppose a flawed plan takes just those two hits, checks their current versions and drops doc7. It returns [8]. That is eligible but incomplete: doc9 exists with distance 0.20 and belongs in the exact top-2.
Check why deletion needs candidate replenishment
The correct result is [8,9]. Deleting stale doc7 from an already truncated list cannot recover doc9. Likewise, replacing doc7’s distance with 0.90 and sorting only [7,8] at S=11 misses doc9. Version checks are necessary; checking just k old hits is insufficient.
For this small exact model, scan and reconcile all records before top-k. An indexed implementation needs to obtain additional candidates, continue a search, or use another plan with justified coverage. A fixed overfetch factor has no universal sufficiency proof when many old hits can be invalidated. Return fewer than k because the eligible corpus has fewer items, or disclose a retrieval limitation; do not confuse the two. A pass names the missing doc9 and explains how candidate coverage will be restored.
This extends the candidate-truncation lesson with temporal invalidation. The mask now changes as versions arrive, even when the query’s tenant and k stay fixed.
Read the named consistency mode precisely
turbopuffer’s Query API defines these modes. The names below are its contract, not names for our historical S=10/S=11 exercises.
| Query mode | Documented recent-write behavior | Consequence for the application |
|---|---|---|
strong (default) | Searches all unindexed writes and updates the cache, including all data written before the query started. | A prior successful update/delete need not wait for background index publication to become visible. |
eventual | Searches at most 128MiB of unindexed writes; accepts stale results for higher throughput. | Do not promise that every acknowledged recent change is reflected. The API does not give this chapter a deterministic rule for which omitted records a toy should choose. |
The byte bound is not a time bound and not a document-count bound. A tiny deletion after a large backlog cannot be assumed visible merely because the deletion record itself is small. The documentation describes potentially long staleness after significant writes; those conditions belong to the eventual behavior, rather than weakening the separately stated strong-mode rule. The Guarantees summary phrases staleness more generally; the mode-specific Query contract is the authority for this distinction. Vendor percentages and example times are not a measurement of our service.
A multi-query runs multiple subqueries in one namespace, the service’s named collection of documents. The API says all its reads use the same consistent snapshot and share the root request’s consistency level. That prevents independently issued lexical and dense reads from accidentally choosing different boundaries within this operation. It does not make their scores comparable, settle ANN fidelity, or establish a general read-write transaction.
The Guarantees page also states that strong queries on a sharded namespace read one snapshot across shards; eventual queries may reflect different times per shard. This is a named product guarantee. Shared snapshots alone are not a proof of general serializable cross-shard transactions. The page promises atomic application of an upsert batch and atomic evaluation of a conditional write, while explicitly excluding general-purpose read-write transactions. Its filter-based patch/delete operations first find IDs at a snapshot, then recheck those IDs atomically: newly qualifying documents between phases can be missed. Those scopes matter when proposing an application invariant.
For our service, choose strong behavior when a user must immediately observe a successful edit or deletion. If the authoritative read path is unavailable, return an explicit error under that requirement; do not label an old cache result as fresh success. A separately chosen stale-read experience needs visible semantics, and current-policy disclosure checks still happen before content leaves the application. A historical eligible hit is not permission to disclose it.
Publication and reclamation are different proofs
A manifest describes which files form a published database state and where recovery begins. SlateDB’s overview explains WAL replay, derived sorted tables and manifest state. It is an accessible design example; it does not establish turbopuffer’s manifest format or publication algorithm.
In our stipulated model, publication must make a complete base and its correct cursor discoverable as one valid state. Moving I to 12 while the selected files still omit the deletion would tell readers to skip needed tail data. Building files is therefore insufficient evidence that a new base has been published correctly. After a crash, unused candidate files may exist beside the last valid state; their existence alone cannot choose authority.
AWS’s current S3 consistency model promises strong read-after-write for object PUT/DELETE and atomic updates to a single key: a concurrent read gets old or new contents, never a partial object. It explicitly does not provide atomic updates across keys. Multiple individually complete objects can still be an inconsistent set for an index. The storage guarantee is a building block; the database must define its own valid publication and recovery boundary.
A newly published base also does not prove all older data can be freed. SlateDB’s snapshot documentation says a snapshot pins state so it cannot be garbage-collected while alive. In our exercise, an existing S=10 reader still needs doc7 v1 even after a new base includes v3. Meanwhile, a newer read using an old base still needs the v3 tombstone to prevent v1 from returning.
Before reclaiming old data or tombstones, require evidence about all relevant readers, references and recovery paths:
| Proposed reclamation | Evidence required in this model |
|---|---|
| Remove an old base/version | No allowed active snapshot or retained read state still needs it; no valid published state references it. |
| Remove a deletion tombstone | No readable/replayable older live copy can reappear without an equivalent deletion/version barrier, and no required snapshot needs the tombstone’s state. |
| Truncate committed recovery records | A durable, valid published checkpoint represents their effects, and every required reader/recovery/retry path remains covered. |
Index catch-up, a warm cache, or elapsed time alone supplies none of these proofs. A production algorithm must track the relevant reader lifetime, references and recovery/retry horizon. We have stated what must be established, without asserting an implementation for doing it.
Independent recovery problem: the late retry
Now use this changed input rather than memorizing the previous IDs. The writer crashes after acknowledging the replacement at position 23 but before publishing a new base. It recovers with I=22, then a delete commits at 24. A delayed client retries its original replacement request. A complete-looking record at 25 has never committed.
For this exercise only, authority durably includes a request receipt mapping request ID to document, version, operation/payload fingerprint and original commit position. Retry commands below replace doc7 for tenant A. Fingerprint tags stand for equality of the complete synthetic payload, not equality of its distance; equal distances alone cannot identify identical vectors or contents. The retry rule is stipulated: the same ID with identical contents returns its original receipt without a new change; different contents under the same ID are rejected. A new ID with a version at or below the latest document version is rejected. An intentionally new higher version needs a new operation and commit. This is an idempotence/version contract for the toy, not a claim about a vendor’s request-ID API.
| position | doc_id | version | tenant | operation | distance | authority | request_id | payload_fingerprint |
|---|---|---|---|---|---|---|---|---|
| 20 | 7 | 1 | A | replace | 0.10 | committed | initial7 | Hinitial7 |
| 21 | 8 | 1 | A | replace | 0.20 | committed | initial8 | Hinitial8 |
| 22 | 9 | 1 | A | replace | 0.20 | committed | initial9 | Hinitial9 |
| 23 | 7 | 2 | A | replace | 0.05 | committed | blue7 | Hblue |
| 24 | 7 | 3 | A | delete | - | committed | remove7 | Hremove |
| 25 | 7 | 4 | A | replace | 0.01 | written-only | green7 | Hgreen |
The same physical table can be inspected at different durable boundaries. Use all committed records at or before the stated D, not future rows. All queries ask for exact body top-2 as trusted tenant A.
| case | I | D | S | tenant | k | retry_doc_id | retry_id | retry_version | retry_distance | retry_fingerprint |
|---|---|---|---|---|---|---|---|---|---|---|
| recover-ack | 22 | 23 | 23 | A | 2 | 7 | blue7 | 2 | 0.05 | Hblue |
| late-retry | 22 | 24 | 24 | A | 2 | 7 | blue7 | 2 | 0.05 | Hblue |
| changed-content | 22 | 24 | 24 | A | 2 | 7 | blue7 | 2 | 0.07 | Hblue-changed |
| stale-new-id | 22 | 24 | 24 | A | 2 | 7 | another7 | 2 | 0.05 | Hblue |
| retained-snapshot | 22 | 24 | 22 | A | 2 | 7 | blue7 | 2 | 0.05 | Hblue |
Produce a version/visibility table for these cases. Decide the retry outcome and the original receipt position, then name the first recovery action, what can be rebuilt and what remains unknown. Explain why the position-25 record and stale base hit cannot justify resurrecting doc7. Finally, state the response semantics if the service cannot establish the required committed boundary.
Check the crash, retry and changed-snapshot decision
| case | doc7_version | doc7_state | exact_ids | retry_outcome | receipt_position |
|---|---|---|---|---|---|
| recover-ack | 2 | live | [7,8] | original-receipt | 23 |
| late-retry | 3 | deleted | [8,9] | original-receipt | 23 |
| changed-content | 3 | deleted | [8,9] | reject-mismatch | - |
| stale-new-id | 3 | deleted | [8,9] | reject-version | - |
| retained-snapshot | 1 | live | [7,8] | original-receipt | 23 |
First establish the authoritative committed boundary and last valid published base/checkpoint, preserving records needed for replay. Rebuild the lost in-memory/derived state from that authority; do not begin by trusting cached hits or publishing the newest-looking files. At D=23, v2 is recoverable although unindexed. The same request returns position 23 without another mutation. At D=24, that receipt acknowledges the old operation; it does not assert that v2 is still the latest document state. The delete remains effective.
The receipt is resolved against durable authority through D even when a query intentionally uses an older S. Thus the retained S=22 query sees v1 while the retry still returns receipt 23. A mismatch must fail, and a fresh request ID cannot turn stale v2 into a valid new version. Written-only v4 is excluded until a valid new operation commits; the exercise supplies no authority to advance D to 25.
At S=24, base-only [7,8] is the stale deletion counterexample. Dropping 7 without replenishing returns the incomplete [8]; the exact answer is [8,9]. Changing S to 22 legitimately changes the answer to [7,8] with v1, provided that snapshot is retained. This is a historical read choice, not compliance with a fresh read-after-delete requirement.
For a fresh read that must include a prior acknowledgment, serve it only after establishing the required committed boundary and reconciling through it; otherwise return an explicit failure. A stale fallback must be separately allowed and labeled. Current authorization still precedes content disclosure. The toy leaves writer takeover, fencing, concurrent retries, object failure handling and receipt-retention enforcement unspecified; it proves no proprietary failure behavior.
A pass includes all five rows, durable receipt/version rules, recoverable authority, the first recovery action, response semantics and a reclamation condition covering readers and replay. These calculations check the model’s answers, not learner mastery or production correctness.
Defend the freshness decision
Keep a short artifact for the service: D, I and selected S; the winning version or tombstone for each result ID; whether the query met its recent-write promise; the action after an interrupted index publication; and the evidence required before reclaiming old state. Specify current-policy authorization separately from snapshot eligibility and exact-neighbor fidelity.
The decision has a cost. Searching a growing committed tail adds exhaustive work, even when it protects read-after-write behavior. Waiting for indexing shifts that work into latency or delayed visibility. Accepting stale reads changes what a successful query means. In the storage and shards chapter, we will attach bytes, remote dependencies and resource budgets to those choices.
For selected reading, use Architecture for committed versus indexed data; Query for the precise strong/eventual and multi-query contracts; and Guarantees for batch, conditional and sharded scope. Read SlateDB’s overview for WAL/manifest reconstruction and its snapshot page for why readers keep old state alive. Read the S3 consistency section to separate single-object atomicity from multi-object publication. These are rolling documents with no pinned implementation version in this lesson.
Designing Data-Intensive Applications, second edition, by Martin Kleppmann and Chris Riccomini (March 2026), is a study bridge for derived data, consistency and recovery. Its full text was not accessed here, so no unread chapter is being used as authority. Bring this completed trace to the edition you obtain and check its treatment of those topics against your service’s required response semantics.