A retrieval change raises the planned-workload relevance score from 66.25% to 95%. Its successful requests also have a smaller elapsed-duration maximum. The team proposes a broad launch. What evidence would make you reject it, and what uncertainty would make you continue testing?
The filtering lesson gave us an exact eligible dense answer. The visibility lesson separated committed authority from the selected snapshot. The resource capstone counted bytes, requests and dependency rounds. Here those obligations meet a new question: whether the evidence supports a particular exposure decision.
Your exit artifact is a one-page decision that names the required slices, evidence limits, guardrails and next distinguishing experiment. Reading time estimates the prose; allow extra time for the complete input appendix, predictions and your own record.
Evidence boundary: This is an openly synthetic evidence-reading fixture. All document titles/vectors, relevance judgments, lexical scores, planned weights, arrivals, outcomes, durations and costs are author-stipulated. No search engine, assessor study, load test or deployment ran. Accessible primary passages were inspected on October 1, 2026; they support the evaluation concepts within their stated scope. Model checks cannot establish production performance, current-policy implementation safety or learner competence.
Make the proposal precise enough to reject
The service retrieves engineering documents for declared information needs. It compares a lexical baseline, dense candidates and one frozen equal-weight fusion. The proposed launch serves both required slices: common tenant-A queries and rarer tenant-B incident queries. Every required slice must avoid judged-relevance regression from lexical. A favorable average cannot waive that requirement.
All ranking comparisons use corpus C40, snapshot S40, live version-1 documents, the named body vector field and k=2. Tenant comes from trusted identity context. Common queries require tenant A with kind unrestricted; incident queries require tenant B AND kind=incident. Current payload permission is a separate trusted decision before content leaves. The primer’s answer and permission contracts explain why snapshot predicate membership is insufficient permission to disclose.
Our eight-document corpus contains tenant-A guides 101–104, tenant-B incidents 201–203 and tenant-B guide 204. The incident predicate excludes 204. There is one eight-dimensional unit-basis vector per document. Dense ranking uses squared Euclidean distance ascending, then document ID ascending. Lexical ranking uses explicitly stipulated scores descending, then ID ascending. These scores are not a BM25 execution, and the vector geometry is not an ANN traversal simulation.
Each query has an overt information need and binary labels for all eight documents. Relevant means useful for that need, not merely containing the query words. The textbook’s system-evaluation passage calls for a document collection, information needs expressible as queries and relevance judgments. Our labels are authored examples of that structure; they are not independent assessor observations.
Opening prediction: does 95% support the proposed broad launch?
It supports investigating the change. It does not satisfy the proposal by itself. We still need the required-slice comparison, full request outcomes, compatible-reference/freshness/disclosure checks and evidence appropriate to a real deployment. A confirmed required-slice regression rejects this proposal; missing consequential evidence calls for further testing.
Name the denominator before comparing the score
There are six distinct information needs. q1–q4 form the common slice; q5–q6 form the required incident slice. Every query has two judged-relevant eligible documents. Binary judged recall@2 counts relevant eligible IDs in the returned top-2, divided by all judged-relevant eligible IDs. Precision@2 divides those hits by requested k=2. They happen to coincide in these complete quality rows.
Dense exact-neighbor fidelity asks a different question: how many returned IDs belong to the exact eligible dense top-2? Its reference fixes corpus, snapshot, field, metric, predicate and tie order. In this fixture the saved dense windows contain the first three exact eligible neighbors, so dense top-2 fidelity is 1 for every query, while judged recall is 1/2 for every query. A faithful dense answer can serve the information need poorly.
The pgvector v0.8.6 monitoring passage recommends comparing approximate output with exact search to monitor recall. That is agreement with the defined dense answer. It supplies no relevance judgments for our users.
The planned workload is stipulated as 90% common and 10% incident. Each common need has weight 9/40; each incident need has weight 1/20. The macro mean instead weights all six needs equally. Those are legitimate different summaries of the same frozen query evidence, so label them separately.
| Method | Common recall@2 | Required incident recall@2 | Six-need macro mean | Planned-workload mean, 90/10 |
|---|---|---|---|---|
| Lexical baseline | 5/8 | 1 | 3/4 | 53/80 = 66.25% |
| Dense candidates | 1/2 | 1/2 | 1/2 | 1/2 = 50% |
| Equal-weight fusion | 1 | 1/2 | 5/6 | 19/20 = 95% |
For lexical, 0.9 × 5/8 + 0.1 × 1 = 53/80. For equal fusion, 0.9 × 1 + 0.1 × 1/2 = 19/20. The gain is real within this authored reference. The incident loss is also real within that reference, and it violates a required condition of the all-slice proposal.
The request schedule later contains repeated q1/q2 and is 75/25, not 90/10. Do not overwrite the planned weights with those sample frequencies or compute query quality only on requests that succeeded.
Trace the rank fusion that causes the incident loss
Our fusion uses the union of each input’s first three IDs, deduplicated by stable identity. Each input rank is a unique one-based position after its total order. An absent item contributes zero from that list. With lexical weight wL, dense weight wD and constant 60:
fusion score = wL/(60 + lexical position) + wD/(60 + dense position)
Sort exact rational scores descending, then document ID ascending. The frozen comparison uses wL=wD=1. This finite-window fusion reference is its own answer function; it is not the exact dense oracle or a guarantee about documents outside the windows.
The inspected official pgvector-python example combines reciprocal-rank terms with zero for list absence. Its cosine metric, SQL RANK() ties and input limits differ from our squared-distance, unique-position, three-item contract. The rolling example is reading evidence, not a pinned runtime or proof of relevance superiority.
q5 asks: diagnose a live cutover that missed a replacement or deletion, using incident evidence. Its relevant eligible documents are 201 and 202. Lexical order is 201, 202, 203; dense order is 203, 202, 201. The three eligible squared distances are 9 for 203, 11 for 202 and 13 for 201. The frozen fusion contributions are:
| ID | Relevant? | Lexical position / contribution | Dense position / contribution | Equal-weight total |
|---|---|---|---|---|
| 201 | yes | 1 / 1/61 | 3 / 1/63 | 124/3843 |
| 202 | yes | 2 / 1/62 | 2 / 1/62 | 1/31 |
| 203 | no | 3 / 1/63 | 1 / 1/61 | 124/3843 |
Before opening the answer, predict the top-2. Check the exact scores rather than rounding them to the same two decimals.
Check q5’s score and tie boundary
124/3843 > 1/31: cross-multiplying gives 124 × 31 = 3844 > 3843. The difference is 1/119133. IDs 201 and 203 tie for the larger score; ID order puts 201 first. The result is 201, 203, so one of two relevant eligible documents is retrieved: recall@2=1/2. Lexical returned 201, 202, giving recall 1.
q6 exposes the same required-slice problem from another need. Relevant IDs are 202, 203; lexical orders 202, 203, 201 and dense orders 201, 203, 202. Equal fusion returns 201, 202 and misses relevant 203. Both incident queries lose, so the slice falls from 1 to 1/2.
Inspect every outcome before reading duration
The fixture schedules eight independent arrivals at 0, 5, 10, 15, 20, 25, 30 and 35 ms: q1, q2, q3, q4, q1, q2, q5, q6. The next scheduled arrival does not wait for a prior completion. The schedule contains six common attempts and two incident attempts. These durations are stipulated scheduled-to-complete elapsed durations; we have no queue/service/network-stage measurement.
The lexical schedule has eight OK responses. Equal fusion also has eight. Dense has six complete OK responses, one RETRIEVAL_LIMIT response containing one inspected eligible candidate, and one UNAVAILABLE at the 150 ms deadline containing no IDs. A capped or failed response is not a complete empty result; these requests have eligible documents.
Complete-response coverage counts complete responses over all scheduled attempts. Material-success coverage additionally requires a compatible reference, snapshot-eligible output, freshness and current disclosure. Both have a stipulated worksheet goal of 99%. The eight authored requests cannot establish a real 99% success rate or a confidence bound.
| Plan | Material successes / attempts | Capped | Failed | Nominal-OK synthetic durations, ms | Nominal-OK / all-attempt synthetic maximum, ms |
|---|---|---|---|---|---|
| Lexical | 8/8 | 0 | 0 | 70,75,80,85,75,80,100,110 | 110 /110 |
| Dense | 6/8 | 1 | 1 | 20,25,22,26,30,32 | 32 /150 |
| Equal fusion | 8/8 | 0 | 0 | 40,42,44,46,42,45,60,62 | 62 /62 |
Dense looks attractive if you retain only its six nominal-OK durations. Its coverage is 75%, and its common slice is 4/6 rather than its incident slice’s 2/2. Completion coverage is an independent gate; a small conditional duration does not replace it. No number in this table is measured latency or p99.
Use the view below to predict, reveal, select a slice and inspect every request row. The original reference inputs, complete judgments, distances and exact fusion contributions are available in its closed appendix.
Choose a slice, then defend the decision
Author-stipulated teaching evidence. C40/S40, body field, one vector per document, k=2. Scores, labels, arrivals, durations and costs are synthetic. The view selects saved evidence; it does not run a search engine.
Predict before revealing. Changing either input hides the current answer. The full native-details transcript below works without JavaScript and keeps every checked mode and slice available.
All 18 evidence views and their request rows — complete static transcript
Each view compares the same frozen information needs. Overall quality uses the stipulated 90% common / 10% critical workload. The eight-request schedule is 6 common / 2 critical, or 75% / 25%. A selected slice keeps its own needs, attempts, outcome counts and durations. The proposed launch still covers both required slices.
Planned workload — All scheduled attempts
Does an aggregate relevance gain justify launching equal-weight fusion? Selected population: All scheduled attempts .
Overall proposal evidence: Weighted recall rises from 53/80 (66.25%) to 19/20 (95%). The weights describe a stipulated 90/10 workload, not the request schedule.
An aggregate conceals required-slice obligations. Continue to the slice evidence before deciding.
All scheduled attempts: 6 distinct information needs; planned-workload share 1 (100%). 8 scheduled attempts; request-schedule share 1 (100%). These are separate denominators.
Frozen C40/S40 reference audit
Binary judged recall@2 measures relevant eligible hits over all judged-relevant eligible IDs. Precision@2 uses requested k=2; it equals recall in these six full quality rows. Dense fidelity measures agreement with the exact eligible dense top-2, separately from relevance. A selected-slice quality pass cannot approve an all-slice proposal.
Lexical baseline
- Frozen judged recall@2
- 53/80 (66.25%)
- Dense exact-neighbor fidelity
- not applicable; distinct answer function
- All-required-slice quality gate
- pass
- Material-success coverage
- 8 /8; 1 (100%) — pass against 99%
Continue testing. The fixture cannot establish a real deployment approval.
Dense candidates
- Frozen judged recall@2
- 1/2 (50%)
- Dense exact-neighbor fidelity
- 1 (100%)
- All-required-slice quality gate
- FAIL
- Material-success coverage
- 6 /8; 3/4 (75%) — FAIL against 99%
Reject the proposed all-slice launch. The fixture cannot establish a real deployment approval.
Equal-weight fusion
- Frozen judged recall@2
- 19/20 (95%)
- Dense exact-neighbor fidelity
- not applicable; distinct answer function
- All-required-slice quality gate
- FAIL
- Material-success coverage
- 8 /8; 1 (100%) — pass against 99%
Reject the proposed all-slice launch. The fixture cannot establish a real deployment approval.
Lexical-weight-2 fusion (post-hoc)
- Frozen judged recall@2
- 1 (100%)
- Dense exact-neighbor fidelity
- not applicable; distinct answer function
- All-required-slice quality gate
- pass
- Operating evidence
- unknown — post-hoc weights have no operating schedule
Continue testing. Quality was repaired after inspecting this fixture; freeze a new hypothesis and test fresh held-out needs.
Material success means the full response passes compatible-reference, snapshot eligibility, freshness and current disclosure requirements. Its denominator includes every scheduled attempt. Nominal OK and complete count can remain unchanged when a gate fails. Durations above are synthetic elapsed durations, not measured latency or p99.
Every selected request, including cap/failure and gate violations
| Plan / request / query | Slice | Scheduled / duration / completion ms | Outcome / IDs sent | Complete / material | S / reference | Snapshot eligibility | Freshness | Current disclosure | Bytes / requests / rounds |
|---|---|---|---|---|---|---|---|---|---|
| Lexical baseline r1 / q1 | common | 0 / 70 / 70 | OK 101, 103 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r2 / q2 | common | 5 / 75 / 80 | OK 101, 102 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r3 / q3 | common | 10 / 80 / 90 | OK 102, 101 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r4 / q4 | common | 15 / 85 / 100 | OK 103, 104 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r5 / q1 | common | 20 / 75 / 95 | OK 101, 103 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r6 / q2 | common | 25 / 80 / 105 | OK 101, 102 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r7 / q5 | critical | 30 / 100 / 130 | OK 201, 202 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 2 / 2 |
| Lexical baseline r8 / q6 | critical | 35 / 110 / 145 | OK 202, 203 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 2 / 2 |
| Dense candidates r1 / q1 | common | 0 / 20 / 20 | OK 102, 104 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r2 / q2 | common | 5 / 25 / 30 | OK 103, 104 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r3 / q3 | common | 10 / 100 / 110 | RETRIEVAL_LIMIT 104 | incomplete / NOT material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r4 / q4 | common | 15 / 150 / 165 | UNAVAILABLE none | incomplete / NOT material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r5 / q1 | common | 20 / 22 / 42 | OK 102, 104 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r6 / q2 | common | 25 / 26 / 51 | OK 103, 104 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r7 / q5 | critical | 30 / 30 / 60 | OK 203, 202 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r8 / q6 | critical | 35 / 32 / 67 | OK 201, 203 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Equal-weight fusion r1 / q1 | common | 0 / 40 / 40 | OK 101, 102 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r2 / q2 | common | 5 / 42 / 47 | OK 101, 103 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r3 / q3 | common | 10 / 44 / 54 | OK 102, 104 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r4 / q4 | common | 15 / 46 / 61 | OK 104, 103 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r5 / q1 | common | 20 / 42 / 62 | OK 101, 102 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r6 / q2 | common | 25 / 45 / 70 | OK 101, 103 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r7 / q5 | critical | 30 / 60 / 90 | OK 201, 203 | complete / material success | S40 / compatible | pass | pass | pass | 20,480 / 3 / 2 |
| Equal-weight fusion r8 / q6 | critical | 35 / 62 / 97 | OK 201, 202 | complete / material success | S40 / compatible | pass | pass | pass | 20,480 / 3 / 2 |
CAPPED means explicit RETRIEVAL_LIMIT with one inspected eligible ID. FAILURE means UNAVAILABLE at the 150 ms deadline with no delivered IDs. Neither is a complete empty answer; these queries have eligible documents. Weight-2 fusion has no manufactured request rows.
Selected frozen query answers
q1: Understand safe tenant-scoped reads immediately after an acknowledged update.
Eligible IDs 101, 102, 103, 104; judged relevant eligible IDs 101, 102. These are the C40/S40 reference.
- Lexical baseline: selected 101, 103; judged recall 1/2 (50%); precision 1/2 (50%) .
- Dense candidates: selected 102, 104; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 101, 102; judged recall 1 (100%); precision 1 (100%) .
- Lexical-weight-2 fusion (post-hoc): selected 101, 102; judged recall 1 (100%); precision 1 (100%) .
q2: Route a tenant request safely when its warm cache node is replaced.
Eligible IDs 101, 102, 103, 104; judged relevant eligible IDs 101, 103. These are the C40/S40 reference.
- Lexical baseline: selected 101, 102; judged recall 1/2 (50%); precision 1/2 (50%) .
- Dense candidates: selected 103, 104; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 101, 103; judged recall 1 (100%); precision 1 (100%) .
- Lexical-weight-2 fusion (post-hoc): selected 101, 103; judged recall 1 (100%); precision 1 (100%) .
q3: Preserve fresh results when ingestion backlog competes with finite query budgets.
Eligible IDs 101, 102, 103, 104; judged relevant eligible IDs 102, 104. These are the C40/S40 reference.
- Lexical baseline: selected 102, 101; judged recall 1/2 (50%); precision 1/2 (50%) .
- Dense candidates: selected 104, 103; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 102, 104; judged recall 1 (100%); precision 1 (100%) .
- Lexical-weight-2 fusion (post-hoc): selected 102, 104; judged recall 1 (100%); precision 1 (100%) .
q4: Plan admission during a cold-cache replacement and traffic burst.
Eligible IDs 101, 102, 103, 104; judged relevant eligible IDs 103, 104. These are the C40/S40 reference.
- Lexical baseline: selected 103, 104; judged recall 1 (100%); precision 1 (100%) .
- Dense candidates: selected 104, 102; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 104, 103; judged recall 1 (100%); precision 1 (100%) .
- Lexical-weight-2 fusion (post-hoc): selected 103, 104; judged recall 1 (100%); precision 1 (100%) .
q5: Diagnose a live cutover that missed a replacement or a deletion; require incident evidence.
Eligible IDs 201, 202, 203; judged relevant eligible IDs 201, 202. These are the C40/S40 reference.
- Lexical baseline: selected 201, 202; judged recall 1 (100%); precision 1 (100%) .
- Dense candidates: selected 203, 202; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 201, 203; judged recall 1/2 (50%); precision 1/2 (50%) .
- Lexical-weight-2 fusion (post-hoc): selected 201, 202; judged recall 1 (100%); precision 1 (100%) .
q6: Diagnose a deleted result reappearing after cache-generation rerouting; require incident evidence.
Eligible IDs 201, 202, 203; judged relevant eligible IDs 202, 203. These are the C40/S40 reference.
- Lexical baseline: selected 202, 203; judged recall 1 (100%); precision 1 (100%) .
- Dense candidates: selected 201, 203; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 201, 202; judged recall 1/2 (50%); precision 1/2 (50%) .
- Lexical-weight-2 fusion (post-hoc): selected 202, 203; judged recall 1 (100%); precision 1 (100%) .
Planned workload — Common tenant-A queries
Does an aggregate relevance gain justify launching equal-weight fusion? Selected population: Common tenant-A queries .
Overall proposal evidence: Weighted recall rises from 53/80 (66.25%) to 19/20 (95%). The weights describe a stipulated 90/10 workload, not the request schedule.
An aggregate conceals required-slice obligations. Continue to the slice evidence before deciding.
Common tenant-A queries: 4 distinct information needs; planned-workload share 9/10 (90%). 6 scheduled attempts; request-schedule share 3/4 (75%). These are separate denominators.
Frozen C40/S40 reference audit
Binary judged recall@2 measures relevant eligible hits over all judged-relevant eligible IDs. Precision@2 uses requested k=2; it equals recall in these six full quality rows. Dense fidelity measures agreement with the exact eligible dense top-2, separately from relevance. A selected-slice quality pass cannot approve an all-slice proposal.
Lexical baseline
- Frozen judged recall@2
- 5/8 (62.5%)
- Dense exact-neighbor fidelity
- not applicable; distinct answer function
- All-required-slice quality gate
- pass
- Material-success coverage
- 6 /6; 1 (100%) — pass against 99%
Continue testing. The fixture cannot establish a real deployment approval.
Dense candidates
- Frozen judged recall@2
- 1/2 (50%)
- Dense exact-neighbor fidelity
- 1 (100%)
- All-required-slice quality gate
- FAIL
- Material-success coverage
- 4 /6; 2/3 (66.67%) — FAIL against 99%
Reject the proposed all-slice launch. The fixture cannot establish a real deployment approval.
Equal-weight fusion
- Frozen judged recall@2
- 1 (100%)
- Dense exact-neighbor fidelity
- not applicable; distinct answer function
- All-required-slice quality gate
- FAIL
- Material-success coverage
- 6 /6; 1 (100%) — pass against 99%
Reject the proposed all-slice launch. The fixture cannot establish a real deployment approval.
Lexical-weight-2 fusion (post-hoc)
- Frozen judged recall@2
- 1 (100%)
- Dense exact-neighbor fidelity
- not applicable; distinct answer function
- All-required-slice quality gate
- pass
- Operating evidence
- unknown — post-hoc weights have no operating schedule
Continue testing. Quality was repaired after inspecting this fixture; freeze a new hypothesis and test fresh held-out needs.
Material success means the full response passes compatible-reference, snapshot eligibility, freshness and current disclosure requirements. Its denominator includes every scheduled attempt. Nominal OK and complete count can remain unchanged when a gate fails. Durations above are synthetic elapsed durations, not measured latency or p99.
Every selected request, including cap/failure and gate violations
| Plan / request / query | Slice | Scheduled / duration / completion ms | Outcome / IDs sent | Complete / material | S / reference | Snapshot eligibility | Freshness | Current disclosure | Bytes / requests / rounds |
|---|---|---|---|---|---|---|---|---|---|
| Lexical baseline r1 / q1 | common | 0 / 70 / 70 | OK 101, 103 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r2 / q2 | common | 5 / 75 / 80 | OK 101, 102 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r3 / q3 | common | 10 / 80 / 90 | OK 102, 101 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r4 / q4 | common | 15 / 85 / 100 | OK 103, 104 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r5 / q1 | common | 20 / 75 / 95 | OK 101, 103 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r6 / q2 | common | 25 / 80 / 105 | OK 101, 102 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Dense candidates r1 / q1 | common | 0 / 20 / 20 | OK 102, 104 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r2 / q2 | common | 5 / 25 / 30 | OK 103, 104 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r3 / q3 | common | 10 / 100 / 110 | RETRIEVAL_LIMIT 104 | incomplete / NOT material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r4 / q4 | common | 15 / 150 / 165 | UNAVAILABLE none | incomplete / NOT material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r5 / q1 | common | 20 / 22 / 42 | OK 102, 104 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r6 / q2 | common | 25 / 26 / 51 | OK 103, 104 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Equal-weight fusion r1 / q1 | common | 0 / 40 / 40 | OK 101, 102 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r2 / q2 | common | 5 / 42 / 47 | OK 101, 103 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r3 / q3 | common | 10 / 44 / 54 | OK 102, 104 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r4 / q4 | common | 15 / 46 / 61 | OK 104, 103 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r5 / q1 | common | 20 / 42 / 62 | OK 101, 102 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r6 / q2 | common | 25 / 45 / 70 | OK 101, 103 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
CAPPED means explicit RETRIEVAL_LIMIT with one inspected eligible ID. FAILURE means UNAVAILABLE at the 150 ms deadline with no delivered IDs. Neither is a complete empty answer; these queries have eligible documents. Weight-2 fusion has no manufactured request rows.
Selected frozen query answers
q1: Understand safe tenant-scoped reads immediately after an acknowledged update.
Eligible IDs 101, 102, 103, 104; judged relevant eligible IDs 101, 102. These are the C40/S40 reference.
- Lexical baseline: selected 101, 103; judged recall 1/2 (50%); precision 1/2 (50%) .
- Dense candidates: selected 102, 104; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 101, 102; judged recall 1 (100%); precision 1 (100%) .
- Lexical-weight-2 fusion (post-hoc): selected 101, 102; judged recall 1 (100%); precision 1 (100%) .
q2: Route a tenant request safely when its warm cache node is replaced.
Eligible IDs 101, 102, 103, 104; judged relevant eligible IDs 101, 103. These are the C40/S40 reference.
- Lexical baseline: selected 101, 102; judged recall 1/2 (50%); precision 1/2 (50%) .
- Dense candidates: selected 103, 104; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 101, 103; judged recall 1 (100%); precision 1 (100%) .
- Lexical-weight-2 fusion (post-hoc): selected 101, 103; judged recall 1 (100%); precision 1 (100%) .
q3: Preserve fresh results when ingestion backlog competes with finite query budgets.
Eligible IDs 101, 102, 103, 104; judged relevant eligible IDs 102, 104. These are the C40/S40 reference.
- Lexical baseline: selected 102, 101; judged recall 1/2 (50%); precision 1/2 (50%) .
- Dense candidates: selected 104, 103; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 102, 104; judged recall 1 (100%); precision 1 (100%) .
- Lexical-weight-2 fusion (post-hoc): selected 102, 104; judged recall 1 (100%); precision 1 (100%) .
q4: Plan admission during a cold-cache replacement and traffic burst.
Eligible IDs 101, 102, 103, 104; judged relevant eligible IDs 103, 104. These are the C40/S40 reference.
- Lexical baseline: selected 103, 104; judged recall 1 (100%); precision 1 (100%) .
- Dense candidates: selected 104, 102; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 104, 103; judged recall 1 (100%); precision 1 (100%) .
- Lexical-weight-2 fusion (post-hoc): selected 103, 104; judged recall 1 (100%); precision 1 (100%) .
Planned workload — Required tenant-B incident queries
Does an aggregate relevance gain justify launching equal-weight fusion? Selected population: Required tenant-B incident queries .
Overall proposal evidence: Weighted recall rises from 53/80 (66.25%) to 19/20 (95%). The weights describe a stipulated 90/10 workload, not the request schedule.
An aggregate conceals required-slice obligations. Continue to the slice evidence before deciding.
Required tenant-B incident queries: 2 distinct information needs; planned-workload share 1/10 (10%). 2 scheduled attempts; request-schedule share 1/4 (25%). These are separate denominators.
Frozen C40/S40 reference audit
Binary judged recall@2 measures relevant eligible hits over all judged-relevant eligible IDs. Precision@2 uses requested k=2; it equals recall in these six full quality rows. Dense fidelity measures agreement with the exact eligible dense top-2, separately from relevance. A selected-slice quality pass cannot approve an all-slice proposal.
Lexical baseline
- Frozen judged recall@2
- 1 (100%)
- Dense exact-neighbor fidelity
- not applicable; distinct answer function
- All-required-slice quality gate
- pass
- Material-success coverage
- 2 /2; 1 (100%) — pass against 99%
Continue testing. The fixture cannot establish a real deployment approval.
Dense candidates
- Frozen judged recall@2
- 1/2 (50%)
- Dense exact-neighbor fidelity
- 1 (100%)
- All-required-slice quality gate
- FAIL
- Material-success coverage
- 2 /2; 1 (100%) — pass against 99%
Reject the proposed all-slice launch. The fixture cannot establish a real deployment approval.
Equal-weight fusion
- Frozen judged recall@2
- 1/2 (50%)
- Dense exact-neighbor fidelity
- not applicable; distinct answer function
- All-required-slice quality gate
- FAIL
- Material-success coverage
- 2 /2; 1 (100%) — pass against 99%
Reject the proposed all-slice launch. The fixture cannot establish a real deployment approval.
Lexical-weight-2 fusion (post-hoc)
- Frozen judged recall@2
- 1 (100%)
- Dense exact-neighbor fidelity
- not applicable; distinct answer function
- All-required-slice quality gate
- pass
- Operating evidence
- unknown — post-hoc weights have no operating schedule
Continue testing. Quality was repaired after inspecting this fixture; freeze a new hypothesis and test fresh held-out needs.
Material success means the full response passes compatible-reference, snapshot eligibility, freshness and current disclosure requirements. Its denominator includes every scheduled attempt. Nominal OK and complete count can remain unchanged when a gate fails. Durations above are synthetic elapsed durations, not measured latency or p99.
Every selected request, including cap/failure and gate violations
| Plan / request / query | Slice | Scheduled / duration / completion ms | Outcome / IDs sent | Complete / material | S / reference | Snapshot eligibility | Freshness | Current disclosure | Bytes / requests / rounds |
|---|---|---|---|---|---|---|---|---|---|
| Lexical baseline r7 / q5 | critical | 30 / 100 / 130 | OK 201, 202 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 2 / 2 |
| Lexical baseline r8 / q6 | critical | 35 / 110 / 145 | OK 202, 203 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 2 / 2 |
| Dense candidates r7 / q5 | critical | 30 / 30 / 60 | OK 203, 202 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r8 / q6 | critical | 35 / 32 / 67 | OK 201, 203 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Equal-weight fusion r7 / q5 | critical | 30 / 60 / 90 | OK 201, 203 | complete / material success | S40 / compatible | pass | pass | pass | 20,480 / 3 / 2 |
| Equal-weight fusion r8 / q6 | critical | 35 / 62 / 97 | OK 201, 202 | complete / material success | S40 / compatible | pass | pass | pass | 20,480 / 3 / 2 |
CAPPED means explicit RETRIEVAL_LIMIT with one inspected eligible ID. FAILURE means UNAVAILABLE at the 150 ms deadline with no delivered IDs. Neither is a complete empty answer; these queries have eligible documents. Weight-2 fusion has no manufactured request rows.
Selected frozen query answers
q5: Diagnose a live cutover that missed a replacement or a deletion; require incident evidence.
Eligible IDs 201, 202, 203; judged relevant eligible IDs 201, 202. These are the C40/S40 reference.
- Lexical baseline: selected 201, 202; judged recall 1 (100%); precision 1 (100%) .
- Dense candidates: selected 203, 202; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 201, 203; judged recall 1/2 (50%); precision 1/2 (50%) .
- Lexical-weight-2 fusion (post-hoc): selected 201, 202; judged recall 1 (100%); precision 1 (100%) .
q6: Diagnose a deleted result reappearing after cache-generation rerouting; require incident evidence.
Eligible IDs 201, 202, 203; judged relevant eligible IDs 202, 203. These are the C40/S40 reference.
- Lexical baseline: selected 202, 203; judged recall 1 (100%); precision 1 (100%) .
- Dense candidates: selected 201, 203; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 201, 202; judged recall 1/2 (50%); precision 1/2 (50%) .
- Lexical-weight-2 fusion (post-hoc): selected 202, 203; judged recall 1 (100%); precision 1 (100%) .
Required slices — All scheduled attempts
Which required slice prevents the proposed all-slice launch? Selected population: All scheduled attempts .
Overall proposal evidence: Critical incident recall falls from 1 to 1/2; common recall rises from 5/8 to 1.
Reject the proposed all-slice equal-weight launch under the stipulated no-regression requirement. A common-only scope still needs independent evidence and a separately declared rollout.
All scheduled attempts: 6 distinct information needs; planned-workload share 1 (100%). 8 scheduled attempts; request-schedule share 1 (100%). These are separate denominators.
Frozen C40/S40 reference audit
Binary judged recall@2 measures relevant eligible hits over all judged-relevant eligible IDs. Precision@2 uses requested k=2; it equals recall in these six full quality rows. Dense fidelity measures agreement with the exact eligible dense top-2, separately from relevance. A selected-slice quality pass cannot approve an all-slice proposal.
Lexical baseline
- Frozen judged recall@2
- 53/80 (66.25%)
- Dense exact-neighbor fidelity
- not applicable; distinct answer function
- All-required-slice quality gate
- pass
- Material-success coverage
- 8 /8; 1 (100%) — pass against 99%
- Mean bytes per nominal OK
- 9216 bytes; pass against the soft per-response budget
Continue testing. The fixture cannot establish a real deployment approval.
Dense candidates
- Frozen judged recall@2
- 1/2 (50%)
- Dense exact-neighbor fidelity
- 1 (100%)
- All-required-slice quality gate
- FAIL
- Material-success coverage
- 6 /8; 3/4 (75%) — FAIL against 99%
- Mean bytes per nominal OK
- 4096 bytes; pass against the soft per-response budget
Reject the proposed all-slice launch. The fixture cannot establish a real deployment approval.
Equal-weight fusion
- Frozen judged recall@2
- 19/20 (95%)
- Dense exact-neighbor fidelity
- not applicable; distinct answer function
- All-required-slice quality gate
- FAIL
- Material-success coverage
- 8 /8; 1 (100%) — pass against 99%
- Mean bytes per nominal OK
- 14336 bytes; FAIL against the soft per-response budget
Reject the proposed all-slice launch. The fixture cannot establish a real deployment approval.
Lexical-weight-2 fusion (post-hoc)
- Frozen judged recall@2
- 1 (100%)
- Dense exact-neighbor fidelity
- not applicable; distinct answer function
- All-required-slice quality gate
- pass
- Operating evidence
- unknown — post-hoc weights have no operating schedule
Continue testing. Quality was repaired after inspecting this fixture; freeze a new hypothesis and test fresh held-out needs.
Material success means the full response passes compatible-reference, snapshot eligibility, freshness and current disclosure requirements. Its denominator includes every scheduled attempt. Nominal OK and complete count can remain unchanged when a gate fails. Durations above are synthetic elapsed durations, not measured latency or p99.
Every selected request, including cap/failure and gate violations
| Plan / request / query | Slice | Scheduled / duration / completion ms | Outcome / IDs sent | Complete / material | S / reference | Snapshot eligibility | Freshness | Current disclosure | Bytes / requests / rounds |
|---|---|---|---|---|---|---|---|---|---|
| Lexical baseline r1 / q1 | common | 0 / 70 / 70 | OK 101, 103 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r2 / q2 | common | 5 / 75 / 80 | OK 101, 102 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r3 / q3 | common | 10 / 80 / 90 | OK 102, 101 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r4 / q4 | common | 15 / 85 / 100 | OK 103, 104 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r5 / q1 | common | 20 / 75 / 95 | OK 101, 103 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r6 / q2 | common | 25 / 80 / 105 | OK 101, 102 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r7 / q5 | critical | 30 / 100 / 130 | OK 201, 202 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 2 / 2 |
| Lexical baseline r8 / q6 | critical | 35 / 110 / 145 | OK 202, 203 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 2 / 2 |
| Dense candidates r1 / q1 | common | 0 / 20 / 20 | OK 102, 104 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r2 / q2 | common | 5 / 25 / 30 | OK 103, 104 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r3 / q3 | common | 10 / 100 / 110 | RETRIEVAL_LIMIT 104 | incomplete / NOT material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r4 / q4 | common | 15 / 150 / 165 | UNAVAILABLE none | incomplete / NOT material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r5 / q1 | common | 20 / 22 / 42 | OK 102, 104 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r6 / q2 | common | 25 / 26 / 51 | OK 103, 104 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r7 / q5 | critical | 30 / 30 / 60 | OK 203, 202 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r8 / q6 | critical | 35 / 32 / 67 | OK 201, 203 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Equal-weight fusion r1 / q1 | common | 0 / 40 / 40 | OK 101, 102 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r2 / q2 | common | 5 / 42 / 47 | OK 101, 103 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r3 / q3 | common | 10 / 44 / 54 | OK 102, 104 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r4 / q4 | common | 15 / 46 / 61 | OK 104, 103 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r5 / q1 | common | 20 / 42 / 62 | OK 101, 102 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r6 / q2 | common | 25 / 45 / 70 | OK 101, 103 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r7 / q5 | critical | 30 / 60 / 90 | OK 201, 203 | complete / material success | S40 / compatible | pass | pass | pass | 20,480 / 3 / 2 |
| Equal-weight fusion r8 / q6 | critical | 35 / 62 / 97 | OK 201, 202 | complete / material success | S40 / compatible | pass | pass | pass | 20,480 / 3 / 2 |
CAPPED means explicit RETRIEVAL_LIMIT with one inspected eligible ID. FAILURE means UNAVAILABLE at the 150 ms deadline with no delivered IDs. Neither is a complete empty answer; these queries have eligible documents. Weight-2 fusion has no manufactured request rows.
Selected frozen query answers
q1: Understand safe tenant-scoped reads immediately after an acknowledged update.
Eligible IDs 101, 102, 103, 104; judged relevant eligible IDs 101, 102. These are the C40/S40 reference.
- Lexical baseline: selected 101, 103; judged recall 1/2 (50%); precision 1/2 (50%) .
- Dense candidates: selected 102, 104; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 101, 102; judged recall 1 (100%); precision 1 (100%) .
- Lexical-weight-2 fusion (post-hoc): selected 101, 102; judged recall 1 (100%); precision 1 (100%) .
q2: Route a tenant request safely when its warm cache node is replaced.
Eligible IDs 101, 102, 103, 104; judged relevant eligible IDs 101, 103. These are the C40/S40 reference.
- Lexical baseline: selected 101, 102; judged recall 1/2 (50%); precision 1/2 (50%) .
- Dense candidates: selected 103, 104; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 101, 103; judged recall 1 (100%); precision 1 (100%) .
- Lexical-weight-2 fusion (post-hoc): selected 101, 103; judged recall 1 (100%); precision 1 (100%) .
q3: Preserve fresh results when ingestion backlog competes with finite query budgets.
Eligible IDs 101, 102, 103, 104; judged relevant eligible IDs 102, 104. These are the C40/S40 reference.
- Lexical baseline: selected 102, 101; judged recall 1/2 (50%); precision 1/2 (50%) .
- Dense candidates: selected 104, 103; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 102, 104; judged recall 1 (100%); precision 1 (100%) .
- Lexical-weight-2 fusion (post-hoc): selected 102, 104; judged recall 1 (100%); precision 1 (100%) .
q4: Plan admission during a cold-cache replacement and traffic burst.
Eligible IDs 101, 102, 103, 104; judged relevant eligible IDs 103, 104. These are the C40/S40 reference.
- Lexical baseline: selected 103, 104; judged recall 1 (100%); precision 1 (100%) .
- Dense candidates: selected 104, 102; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 104, 103; judged recall 1 (100%); precision 1 (100%) .
- Lexical-weight-2 fusion (post-hoc): selected 103, 104; judged recall 1 (100%); precision 1 (100%) .
q5: Diagnose a live cutover that missed a replacement or a deletion; require incident evidence.
Eligible IDs 201, 202, 203; judged relevant eligible IDs 201, 202. These are the C40/S40 reference.
- Lexical baseline: selected 201, 202; judged recall 1 (100%); precision 1 (100%) .
- Dense candidates: selected 203, 202; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 201, 203; judged recall 1/2 (50%); precision 1/2 (50%) .
- Lexical-weight-2 fusion (post-hoc): selected 201, 202; judged recall 1 (100%); precision 1 (100%) .
q6: Diagnose a deleted result reappearing after cache-generation rerouting; require incident evidence.
Eligible IDs 201, 202, 203; judged relevant eligible IDs 202, 203. These are the C40/S40 reference.
- Lexical baseline: selected 202, 203; judged recall 1 (100%); precision 1 (100%) .
- Dense candidates: selected 201, 203; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 201, 202; judged recall 1/2 (50%); precision 1/2 (50%) .
- Lexical-weight-2 fusion (post-hoc): selected 202, 203; judged recall 1 (100%); precision 1 (100%) .
Required slices — Common tenant-A queries
Which required slice prevents the proposed all-slice launch? Selected population: Common tenant-A queries .
Overall proposal evidence: Critical incident recall falls from 1 to 1/2; common recall rises from 5/8 to 1.
Reject the proposed all-slice equal-weight launch under the stipulated no-regression requirement. A common-only scope still needs independent evidence and a separately declared rollout.
Common tenant-A queries: 4 distinct information needs; planned-workload share 9/10 (90%). 6 scheduled attempts; request-schedule share 3/4 (75%). These are separate denominators.
Frozen C40/S40 reference audit
Binary judged recall@2 measures relevant eligible hits over all judged-relevant eligible IDs. Precision@2 uses requested k=2; it equals recall in these six full quality rows. Dense fidelity measures agreement with the exact eligible dense top-2, separately from relevance. A selected-slice quality pass cannot approve an all-slice proposal.
Lexical baseline
- Frozen judged recall@2
- 5/8 (62.5%)
- Dense exact-neighbor fidelity
- not applicable; distinct answer function
- All-required-slice quality gate
- pass
- Material-success coverage
- 6 /6; 1 (100%) — pass against 99%
- Mean bytes per nominal OK
- 8192 bytes; pass against the soft per-response budget
Continue testing. The fixture cannot establish a real deployment approval.
Dense candidates
- Frozen judged recall@2
- 1/2 (50%)
- Dense exact-neighbor fidelity
- 1 (100%)
- All-required-slice quality gate
- FAIL
- Material-success coverage
- 4 /6; 2/3 (66.67%) — FAIL against 99%
- Mean bytes per nominal OK
- 4096 bytes; pass against the soft per-response budget
Reject the proposed all-slice launch. The fixture cannot establish a real deployment approval.
Equal-weight fusion
- Frozen judged recall@2
- 1 (100%)
- Dense exact-neighbor fidelity
- not applicable; distinct answer function
- All-required-slice quality gate
- FAIL
- Material-success coverage
- 6 /6; 1 (100%) — pass against 99%
- Mean bytes per nominal OK
- 12288 bytes; pass against the soft per-response budget
Reject the proposed all-slice launch. The fixture cannot establish a real deployment approval.
Lexical-weight-2 fusion (post-hoc)
- Frozen judged recall@2
- 1 (100%)
- Dense exact-neighbor fidelity
- not applicable; distinct answer function
- All-required-slice quality gate
- pass
- Operating evidence
- unknown — post-hoc weights have no operating schedule
Continue testing. Quality was repaired after inspecting this fixture; freeze a new hypothesis and test fresh held-out needs.
Material success means the full response passes compatible-reference, snapshot eligibility, freshness and current disclosure requirements. Its denominator includes every scheduled attempt. Nominal OK and complete count can remain unchanged when a gate fails. Durations above are synthetic elapsed durations, not measured latency or p99.
Every selected request, including cap/failure and gate violations
| Plan / request / query | Slice | Scheduled / duration / completion ms | Outcome / IDs sent | Complete / material | S / reference | Snapshot eligibility | Freshness | Current disclosure | Bytes / requests / rounds |
|---|---|---|---|---|---|---|---|---|---|
| Lexical baseline r1 / q1 | common | 0 / 70 / 70 | OK 101, 103 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r2 / q2 | common | 5 / 75 / 80 | OK 101, 102 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r3 / q3 | common | 10 / 80 / 90 | OK 102, 101 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r4 / q4 | common | 15 / 85 / 100 | OK 103, 104 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r5 / q1 | common | 20 / 75 / 95 | OK 101, 103 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r6 / q2 | common | 25 / 80 / 105 | OK 101, 102 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Dense candidates r1 / q1 | common | 0 / 20 / 20 | OK 102, 104 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r2 / q2 | common | 5 / 25 / 30 | OK 103, 104 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r3 / q3 | common | 10 / 100 / 110 | RETRIEVAL_LIMIT 104 | incomplete / NOT material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r4 / q4 | common | 15 / 150 / 165 | UNAVAILABLE none | incomplete / NOT material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r5 / q1 | common | 20 / 22 / 42 | OK 102, 104 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r6 / q2 | common | 25 / 26 / 51 | OK 103, 104 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Equal-weight fusion r1 / q1 | common | 0 / 40 / 40 | OK 101, 102 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r2 / q2 | common | 5 / 42 / 47 | OK 101, 103 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r3 / q3 | common | 10 / 44 / 54 | OK 102, 104 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r4 / q4 | common | 15 / 46 / 61 | OK 104, 103 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r5 / q1 | common | 20 / 42 / 62 | OK 101, 102 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r6 / q2 | common | 25 / 45 / 70 | OK 101, 103 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
CAPPED means explicit RETRIEVAL_LIMIT with one inspected eligible ID. FAILURE means UNAVAILABLE at the 150 ms deadline with no delivered IDs. Neither is a complete empty answer; these queries have eligible documents. Weight-2 fusion has no manufactured request rows.
Selected frozen query answers
q1: Understand safe tenant-scoped reads immediately after an acknowledged update.
Eligible IDs 101, 102, 103, 104; judged relevant eligible IDs 101, 102. These are the C40/S40 reference.
- Lexical baseline: selected 101, 103; judged recall 1/2 (50%); precision 1/2 (50%) .
- Dense candidates: selected 102, 104; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 101, 102; judged recall 1 (100%); precision 1 (100%) .
- Lexical-weight-2 fusion (post-hoc): selected 101, 102; judged recall 1 (100%); precision 1 (100%) .
q2: Route a tenant request safely when its warm cache node is replaced.
Eligible IDs 101, 102, 103, 104; judged relevant eligible IDs 101, 103. These are the C40/S40 reference.
- Lexical baseline: selected 101, 102; judged recall 1/2 (50%); precision 1/2 (50%) .
- Dense candidates: selected 103, 104; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 101, 103; judged recall 1 (100%); precision 1 (100%) .
- Lexical-weight-2 fusion (post-hoc): selected 101, 103; judged recall 1 (100%); precision 1 (100%) .
q3: Preserve fresh results when ingestion backlog competes with finite query budgets.
Eligible IDs 101, 102, 103, 104; judged relevant eligible IDs 102, 104. These are the C40/S40 reference.
- Lexical baseline: selected 102, 101; judged recall 1/2 (50%); precision 1/2 (50%) .
- Dense candidates: selected 104, 103; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 102, 104; judged recall 1 (100%); precision 1 (100%) .
- Lexical-weight-2 fusion (post-hoc): selected 102, 104; judged recall 1 (100%); precision 1 (100%) .
q4: Plan admission during a cold-cache replacement and traffic burst.
Eligible IDs 101, 102, 103, 104; judged relevant eligible IDs 103, 104. These are the C40/S40 reference.
- Lexical baseline: selected 103, 104; judged recall 1 (100%); precision 1 (100%) .
- Dense candidates: selected 104, 102; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 104, 103; judged recall 1 (100%); precision 1 (100%) .
- Lexical-weight-2 fusion (post-hoc): selected 103, 104; judged recall 1 (100%); precision 1 (100%) .
Required slices — Required tenant-B incident queries
Which required slice prevents the proposed all-slice launch? Selected population: Required tenant-B incident queries .
Overall proposal evidence: Critical incident recall falls from 1 to 1/2; common recall rises from 5/8 to 1.
Reject the proposed all-slice equal-weight launch under the stipulated no-regression requirement. A common-only scope still needs independent evidence and a separately declared rollout.
Required tenant-B incident queries: 2 distinct information needs; planned-workload share 1/10 (10%). 2 scheduled attempts; request-schedule share 1/4 (25%). These are separate denominators.
Frozen C40/S40 reference audit
Binary judged recall@2 measures relevant eligible hits over all judged-relevant eligible IDs. Precision@2 uses requested k=2; it equals recall in these six full quality rows. Dense fidelity measures agreement with the exact eligible dense top-2, separately from relevance. A selected-slice quality pass cannot approve an all-slice proposal.
Lexical baseline
- Frozen judged recall@2
- 1 (100%)
- Dense exact-neighbor fidelity
- not applicable; distinct answer function
- All-required-slice quality gate
- pass
- Material-success coverage
- 2 /2; 1 (100%) — pass against 99%
- Mean bytes per nominal OK
- 12288 bytes; pass against the soft per-response budget
Continue testing. The fixture cannot establish a real deployment approval.
Dense candidates
- Frozen judged recall@2
- 1/2 (50%)
- Dense exact-neighbor fidelity
- 1 (100%)
- All-required-slice quality gate
- FAIL
- Material-success coverage
- 2 /2; 1 (100%) — pass against 99%
- Mean bytes per nominal OK
- 4096 bytes; pass against the soft per-response budget
Reject the proposed all-slice launch. The fixture cannot establish a real deployment approval.
Equal-weight fusion
- Frozen judged recall@2
- 1/2 (50%)
- Dense exact-neighbor fidelity
- not applicable; distinct answer function
- All-required-slice quality gate
- FAIL
- Material-success coverage
- 2 /2; 1 (100%) — pass against 99%
- Mean bytes per nominal OK
- 20480 bytes; FAIL against the soft per-response budget
Reject the proposed all-slice launch. The fixture cannot establish a real deployment approval.
Lexical-weight-2 fusion (post-hoc)
- Frozen judged recall@2
- 1 (100%)
- Dense exact-neighbor fidelity
- not applicable; distinct answer function
- All-required-slice quality gate
- pass
- Operating evidence
- unknown — post-hoc weights have no operating schedule
Continue testing. Quality was repaired after inspecting this fixture; freeze a new hypothesis and test fresh held-out needs.
Material success means the full response passes compatible-reference, snapshot eligibility, freshness and current disclosure requirements. Its denominator includes every scheduled attempt. Nominal OK and complete count can remain unchanged when a gate fails. Durations above are synthetic elapsed durations, not measured latency or p99.
Every selected request, including cap/failure and gate violations
| Plan / request / query | Slice | Scheduled / duration / completion ms | Outcome / IDs sent | Complete / material | S / reference | Snapshot eligibility | Freshness | Current disclosure | Bytes / requests / rounds |
|---|---|---|---|---|---|---|---|---|---|
| Lexical baseline r7 / q5 | critical | 30 / 100 / 130 | OK 201, 202 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 2 / 2 |
| Lexical baseline r8 / q6 | critical | 35 / 110 / 145 | OK 202, 203 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 2 / 2 |
| Dense candidates r7 / q5 | critical | 30 / 30 / 60 | OK 203, 202 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r8 / q6 | critical | 35 / 32 / 67 | OK 201, 203 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Equal-weight fusion r7 / q5 | critical | 30 / 60 / 90 | OK 201, 203 | complete / material success | S40 / compatible | pass | pass | pass | 20,480 / 3 / 2 |
| Equal-weight fusion r8 / q6 | critical | 35 / 62 / 97 | OK 201, 202 | complete / material success | S40 / compatible | pass | pass | pass | 20,480 / 3 / 2 |
CAPPED means explicit RETRIEVAL_LIMIT with one inspected eligible ID. FAILURE means UNAVAILABLE at the 150 ms deadline with no delivered IDs. Neither is a complete empty answer; these queries have eligible documents. Weight-2 fusion has no manufactured request rows.
Selected frozen query answers
q5: Diagnose a live cutover that missed a replacement or a deletion; require incident evidence.
Eligible IDs 201, 202, 203; judged relevant eligible IDs 201, 202. These are the C40/S40 reference.
- Lexical baseline: selected 201, 202; judged recall 1 (100%); precision 1 (100%) .
- Dense candidates: selected 203, 202; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 201, 203; judged recall 1/2 (50%); precision 1/2 (50%) .
- Lexical-weight-2 fusion (post-hoc): selected 201, 202; judged recall 1 (100%); precision 1 (100%) .
q6: Diagnose a deleted result reappearing after cache-generation rerouting; require incident evidence.
Eligible IDs 201, 202, 203; judged relevant eligible IDs 202, 203. These are the C40/S40 reference.
- Lexical baseline: selected 202, 203; judged recall 1 (100%); precision 1 (100%) .
- Dense candidates: selected 201, 203; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 201, 202; judged recall 1/2 (50%); precision 1/2 (50%) .
- Lexical-weight-2 fusion (post-hoc): selected 202, 203; judged recall 1 (100%); precision 1 (100%) .
Completion coverage — All scheduled attempts
Can the dense schedule's 32 ms successful-only maximum establish the service goal? Selected population: All scheduled attempts .
Overall proposal evidence: Dense delivers 6 complete material successes, 1 capped response and 1 failure across 8 scheduled attempts: 75% coverage against a stipulated 99% goal. The all-attempt maximum is 150 ms.
The durations are synthetic. A conditional duration excludes unsuccessful work and does not replace completion coverage. No measured latency or p99 is available.
All scheduled attempts: 6 distinct information needs; planned-workload share 1 (100%). 8 scheduled attempts; request-schedule share 1 (100%). These are separate denominators.
Frozen C40/S40 reference audit
Binary judged recall@2 measures relevant eligible hits over all judged-relevant eligible IDs. Precision@2 uses requested k=2; it equals recall in these six full quality rows. Dense fidelity measures agreement with the exact eligible dense top-2, separately from relevance. A selected-slice quality pass cannot approve an all-slice proposal.
Lexical baseline
- Frozen judged recall@2
- 53/80 (66.25%)
- Dense exact-neighbor fidelity
- not applicable; distinct answer function
- All-required-slice quality gate
- pass
- Material-success coverage
- 8 /8; 1 (100%) — pass against 99%
Continue testing. The fixture cannot establish a real deployment approval.
Dense candidates
- Frozen judged recall@2
- 1/2 (50%)
- Dense exact-neighbor fidelity
- 1 (100%)
- All-required-slice quality gate
- FAIL
- Material-success coverage
- 6 /8; 3/4 (75%) — FAIL against 99%
Reject the proposed all-slice launch. The fixture cannot establish a real deployment approval.
Equal-weight fusion
- Frozen judged recall@2
- 19/20 (95%)
- Dense exact-neighbor fidelity
- not applicable; distinct answer function
- All-required-slice quality gate
- FAIL
- Material-success coverage
- 8 /8; 1 (100%) — pass against 99%
Reject the proposed all-slice launch. The fixture cannot establish a real deployment approval.
Lexical-weight-2 fusion (post-hoc)
- Frozen judged recall@2
- 1 (100%)
- Dense exact-neighbor fidelity
- not applicable; distinct answer function
- All-required-slice quality gate
- pass
- Operating evidence
- unknown — post-hoc weights have no operating schedule
Continue testing. Quality was repaired after inspecting this fixture; freeze a new hypothesis and test fresh held-out needs.
Material success means the full response passes compatible-reference, snapshot eligibility, freshness and current disclosure requirements. Its denominator includes every scheduled attempt. Nominal OK and complete count can remain unchanged when a gate fails. Durations above are synthetic elapsed durations, not measured latency or p99.
Every selected request, including cap/failure and gate violations
| Plan / request / query | Slice | Scheduled / duration / completion ms | Outcome / IDs sent | Complete / material | S / reference | Snapshot eligibility | Freshness | Current disclosure | Bytes / requests / rounds |
|---|---|---|---|---|---|---|---|---|---|
| Lexical baseline r1 / q1 | common | 0 / 70 / 70 | OK 101, 103 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r2 / q2 | common | 5 / 75 / 80 | OK 101, 102 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r3 / q3 | common | 10 / 80 / 90 | OK 102, 101 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r4 / q4 | common | 15 / 85 / 100 | OK 103, 104 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r5 / q1 | common | 20 / 75 / 95 | OK 101, 103 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r6 / q2 | common | 25 / 80 / 105 | OK 101, 102 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r7 / q5 | critical | 30 / 100 / 130 | OK 201, 202 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 2 / 2 |
| Lexical baseline r8 / q6 | critical | 35 / 110 / 145 | OK 202, 203 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 2 / 2 |
| Dense candidates r1 / q1 | common | 0 / 20 / 20 | OK 102, 104 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r2 / q2 | common | 5 / 25 / 30 | OK 103, 104 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r3 / q3 | common | 10 / 100 / 110 | RETRIEVAL_LIMIT 104 | incomplete / NOT material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r4 / q4 | common | 15 / 150 / 165 | UNAVAILABLE none | incomplete / NOT material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r5 / q1 | common | 20 / 22 / 42 | OK 102, 104 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r6 / q2 | common | 25 / 26 / 51 | OK 103, 104 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r7 / q5 | critical | 30 / 30 / 60 | OK 203, 202 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r8 / q6 | critical | 35 / 32 / 67 | OK 201, 203 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Equal-weight fusion r1 / q1 | common | 0 / 40 / 40 | OK 101, 102 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r2 / q2 | common | 5 / 42 / 47 | OK 101, 103 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r3 / q3 | common | 10 / 44 / 54 | OK 102, 104 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r4 / q4 | common | 15 / 46 / 61 | OK 104, 103 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r5 / q1 | common | 20 / 42 / 62 | OK 101, 102 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r6 / q2 | common | 25 / 45 / 70 | OK 101, 103 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r7 / q5 | critical | 30 / 60 / 90 | OK 201, 203 | complete / material success | S40 / compatible | pass | pass | pass | 20,480 / 3 / 2 |
| Equal-weight fusion r8 / q6 | critical | 35 / 62 / 97 | OK 201, 202 | complete / material success | S40 / compatible | pass | pass | pass | 20,480 / 3 / 2 |
CAPPED means explicit RETRIEVAL_LIMIT with one inspected eligible ID. FAILURE means UNAVAILABLE at the 150 ms deadline with no delivered IDs. Neither is a complete empty answer; these queries have eligible documents. Weight-2 fusion has no manufactured request rows.
Selected frozen query answers
q1: Understand safe tenant-scoped reads immediately after an acknowledged update.
Eligible IDs 101, 102, 103, 104; judged relevant eligible IDs 101, 102. These are the C40/S40 reference.
- Lexical baseline: selected 101, 103; judged recall 1/2 (50%); precision 1/2 (50%) .
- Dense candidates: selected 102, 104; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 101, 102; judged recall 1 (100%); precision 1 (100%) .
- Lexical-weight-2 fusion (post-hoc): selected 101, 102; judged recall 1 (100%); precision 1 (100%) .
q2: Route a tenant request safely when its warm cache node is replaced.
Eligible IDs 101, 102, 103, 104; judged relevant eligible IDs 101, 103. These are the C40/S40 reference.
- Lexical baseline: selected 101, 102; judged recall 1/2 (50%); precision 1/2 (50%) .
- Dense candidates: selected 103, 104; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 101, 103; judged recall 1 (100%); precision 1 (100%) .
- Lexical-weight-2 fusion (post-hoc): selected 101, 103; judged recall 1 (100%); precision 1 (100%) .
q3: Preserve fresh results when ingestion backlog competes with finite query budgets.
Eligible IDs 101, 102, 103, 104; judged relevant eligible IDs 102, 104. These are the C40/S40 reference.
- Lexical baseline: selected 102, 101; judged recall 1/2 (50%); precision 1/2 (50%) .
- Dense candidates: selected 104, 103; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 102, 104; judged recall 1 (100%); precision 1 (100%) .
- Lexical-weight-2 fusion (post-hoc): selected 102, 104; judged recall 1 (100%); precision 1 (100%) .
q4: Plan admission during a cold-cache replacement and traffic burst.
Eligible IDs 101, 102, 103, 104; judged relevant eligible IDs 103, 104. These are the C40/S40 reference.
- Lexical baseline: selected 103, 104; judged recall 1 (100%); precision 1 (100%) .
- Dense candidates: selected 104, 102; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 104, 103; judged recall 1 (100%); precision 1 (100%) .
- Lexical-weight-2 fusion (post-hoc): selected 103, 104; judged recall 1 (100%); precision 1 (100%) .
q5: Diagnose a live cutover that missed a replacement or a deletion; require incident evidence.
Eligible IDs 201, 202, 203; judged relevant eligible IDs 201, 202. These are the C40/S40 reference.
- Lexical baseline: selected 201, 202; judged recall 1 (100%); precision 1 (100%) .
- Dense candidates: selected 203, 202; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 201, 203; judged recall 1/2 (50%); precision 1/2 (50%) .
- Lexical-weight-2 fusion (post-hoc): selected 201, 202; judged recall 1 (100%); precision 1 (100%) .
q6: Diagnose a deleted result reappearing after cache-generation rerouting; require incident evidence.
Eligible IDs 201, 202, 203; judged relevant eligible IDs 202, 203. These are the C40/S40 reference.
- Lexical baseline: selected 202, 203; judged recall 1 (100%); precision 1 (100%) .
- Dense candidates: selected 201, 203; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 201, 202; judged recall 1/2 (50%); precision 1/2 (50%) .
- Lexical-weight-2 fusion (post-hoc): selected 202, 203; judged recall 1 (100%); precision 1 (100%) .
Completion coverage — Common tenant-A queries
Can the dense schedule's 32 ms successful-only maximum establish the service goal? Selected population: Common tenant-A queries .
Overall proposal evidence: Dense delivers 6 complete material successes, 1 capped response and 1 failure across 8 scheduled attempts: 75% coverage against a stipulated 99% goal. The all-attempt maximum is 150 ms.
The durations are synthetic. A conditional duration excludes unsuccessful work and does not replace completion coverage. No measured latency or p99 is available.
Common tenant-A queries: 4 distinct information needs; planned-workload share 9/10 (90%). 6 scheduled attempts; request-schedule share 3/4 (75%). These are separate denominators.
Frozen C40/S40 reference audit
Binary judged recall@2 measures relevant eligible hits over all judged-relevant eligible IDs. Precision@2 uses requested k=2; it equals recall in these six full quality rows. Dense fidelity measures agreement with the exact eligible dense top-2, separately from relevance. A selected-slice quality pass cannot approve an all-slice proposal.
Lexical baseline
- Frozen judged recall@2
- 5/8 (62.5%)
- Dense exact-neighbor fidelity
- not applicable; distinct answer function
- All-required-slice quality gate
- pass
- Material-success coverage
- 6 /6; 1 (100%) — pass against 99%
Continue testing. The fixture cannot establish a real deployment approval.
Dense candidates
- Frozen judged recall@2
- 1/2 (50%)
- Dense exact-neighbor fidelity
- 1 (100%)
- All-required-slice quality gate
- FAIL
- Material-success coverage
- 4 /6; 2/3 (66.67%) — FAIL against 99%
Reject the proposed all-slice launch. The fixture cannot establish a real deployment approval.
Equal-weight fusion
- Frozen judged recall@2
- 1 (100%)
- Dense exact-neighbor fidelity
- not applicable; distinct answer function
- All-required-slice quality gate
- FAIL
- Material-success coverage
- 6 /6; 1 (100%) — pass against 99%
Reject the proposed all-slice launch. The fixture cannot establish a real deployment approval.
Lexical-weight-2 fusion (post-hoc)
- Frozen judged recall@2
- 1 (100%)
- Dense exact-neighbor fidelity
- not applicable; distinct answer function
- All-required-slice quality gate
- pass
- Operating evidence
- unknown — post-hoc weights have no operating schedule
Continue testing. Quality was repaired after inspecting this fixture; freeze a new hypothesis and test fresh held-out needs.
Material success means the full response passes compatible-reference, snapshot eligibility, freshness and current disclosure requirements. Its denominator includes every scheduled attempt. Nominal OK and complete count can remain unchanged when a gate fails. Durations above are synthetic elapsed durations, not measured latency or p99.
Every selected request, including cap/failure and gate violations
| Plan / request / query | Slice | Scheduled / duration / completion ms | Outcome / IDs sent | Complete / material | S / reference | Snapshot eligibility | Freshness | Current disclosure | Bytes / requests / rounds |
|---|---|---|---|---|---|---|---|---|---|
| Lexical baseline r1 / q1 | common | 0 / 70 / 70 | OK 101, 103 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r2 / q2 | common | 5 / 75 / 80 | OK 101, 102 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r3 / q3 | common | 10 / 80 / 90 | OK 102, 101 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r4 / q4 | common | 15 / 85 / 100 | OK 103, 104 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r5 / q1 | common | 20 / 75 / 95 | OK 101, 103 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r6 / q2 | common | 25 / 80 / 105 | OK 101, 102 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Dense candidates r1 / q1 | common | 0 / 20 / 20 | OK 102, 104 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r2 / q2 | common | 5 / 25 / 30 | OK 103, 104 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r3 / q3 | common | 10 / 100 / 110 | RETRIEVAL_LIMIT 104 | incomplete / NOT material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r4 / q4 | common | 15 / 150 / 165 | UNAVAILABLE none | incomplete / NOT material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r5 / q1 | common | 20 / 22 / 42 | OK 102, 104 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r6 / q2 | common | 25 / 26 / 51 | OK 103, 104 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Equal-weight fusion r1 / q1 | common | 0 / 40 / 40 | OK 101, 102 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r2 / q2 | common | 5 / 42 / 47 | OK 101, 103 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r3 / q3 | common | 10 / 44 / 54 | OK 102, 104 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r4 / q4 | common | 15 / 46 / 61 | OK 104, 103 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r5 / q1 | common | 20 / 42 / 62 | OK 101, 102 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r6 / q2 | common | 25 / 45 / 70 | OK 101, 103 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
CAPPED means explicit RETRIEVAL_LIMIT with one inspected eligible ID. FAILURE means UNAVAILABLE at the 150 ms deadline with no delivered IDs. Neither is a complete empty answer; these queries have eligible documents. Weight-2 fusion has no manufactured request rows.
Selected frozen query answers
q1: Understand safe tenant-scoped reads immediately after an acknowledged update.
Eligible IDs 101, 102, 103, 104; judged relevant eligible IDs 101, 102. These are the C40/S40 reference.
- Lexical baseline: selected 101, 103; judged recall 1/2 (50%); precision 1/2 (50%) .
- Dense candidates: selected 102, 104; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 101, 102; judged recall 1 (100%); precision 1 (100%) .
- Lexical-weight-2 fusion (post-hoc): selected 101, 102; judged recall 1 (100%); precision 1 (100%) .
q2: Route a tenant request safely when its warm cache node is replaced.
Eligible IDs 101, 102, 103, 104; judged relevant eligible IDs 101, 103. These are the C40/S40 reference.
- Lexical baseline: selected 101, 102; judged recall 1/2 (50%); precision 1/2 (50%) .
- Dense candidates: selected 103, 104; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 101, 103; judged recall 1 (100%); precision 1 (100%) .
- Lexical-weight-2 fusion (post-hoc): selected 101, 103; judged recall 1 (100%); precision 1 (100%) .
q3: Preserve fresh results when ingestion backlog competes with finite query budgets.
Eligible IDs 101, 102, 103, 104; judged relevant eligible IDs 102, 104. These are the C40/S40 reference.
- Lexical baseline: selected 102, 101; judged recall 1/2 (50%); precision 1/2 (50%) .
- Dense candidates: selected 104, 103; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 102, 104; judged recall 1 (100%); precision 1 (100%) .
- Lexical-weight-2 fusion (post-hoc): selected 102, 104; judged recall 1 (100%); precision 1 (100%) .
q4: Plan admission during a cold-cache replacement and traffic burst.
Eligible IDs 101, 102, 103, 104; judged relevant eligible IDs 103, 104. These are the C40/S40 reference.
- Lexical baseline: selected 103, 104; judged recall 1 (100%); precision 1 (100%) .
- Dense candidates: selected 104, 102; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 104, 103; judged recall 1 (100%); precision 1 (100%) .
- Lexical-weight-2 fusion (post-hoc): selected 103, 104; judged recall 1 (100%); precision 1 (100%) .
Completion coverage — Required tenant-B incident queries
Can the dense schedule's 32 ms successful-only maximum establish the service goal? Selected population: Required tenant-B incident queries .
Overall proposal evidence: Dense delivers 6 complete material successes, 1 capped response and 1 failure across 8 scheduled attempts: 75% coverage against a stipulated 99% goal. The all-attempt maximum is 150 ms.
The durations are synthetic. A conditional duration excludes unsuccessful work and does not replace completion coverage. No measured latency or p99 is available.
Required tenant-B incident queries: 2 distinct information needs; planned-workload share 1/10 (10%). 2 scheduled attempts; request-schedule share 1/4 (25%). These are separate denominators.
Frozen C40/S40 reference audit
Binary judged recall@2 measures relevant eligible hits over all judged-relevant eligible IDs. Precision@2 uses requested k=2; it equals recall in these six full quality rows. Dense fidelity measures agreement with the exact eligible dense top-2, separately from relevance. A selected-slice quality pass cannot approve an all-slice proposal.
Lexical baseline
- Frozen judged recall@2
- 1 (100%)
- Dense exact-neighbor fidelity
- not applicable; distinct answer function
- All-required-slice quality gate
- pass
- Material-success coverage
- 2 /2; 1 (100%) — pass against 99%
Continue testing. The fixture cannot establish a real deployment approval.
Dense candidates
- Frozen judged recall@2
- 1/2 (50%)
- Dense exact-neighbor fidelity
- 1 (100%)
- All-required-slice quality gate
- FAIL
- Material-success coverage
- 2 /2; 1 (100%) — pass against 99%
Reject the proposed all-slice launch. The fixture cannot establish a real deployment approval.
Equal-weight fusion
- Frozen judged recall@2
- 1/2 (50%)
- Dense exact-neighbor fidelity
- not applicable; distinct answer function
- All-required-slice quality gate
- FAIL
- Material-success coverage
- 2 /2; 1 (100%) — pass against 99%
Reject the proposed all-slice launch. The fixture cannot establish a real deployment approval.
Lexical-weight-2 fusion (post-hoc)
- Frozen judged recall@2
- 1 (100%)
- Dense exact-neighbor fidelity
- not applicable; distinct answer function
- All-required-slice quality gate
- pass
- Operating evidence
- unknown — post-hoc weights have no operating schedule
Continue testing. Quality was repaired after inspecting this fixture; freeze a new hypothesis and test fresh held-out needs.
Material success means the full response passes compatible-reference, snapshot eligibility, freshness and current disclosure requirements. Its denominator includes every scheduled attempt. Nominal OK and complete count can remain unchanged when a gate fails. Durations above are synthetic elapsed durations, not measured latency or p99.
Every selected request, including cap/failure and gate violations
| Plan / request / query | Slice | Scheduled / duration / completion ms | Outcome / IDs sent | Complete / material | S / reference | Snapshot eligibility | Freshness | Current disclosure | Bytes / requests / rounds |
|---|---|---|---|---|---|---|---|---|---|
| Lexical baseline r7 / q5 | critical | 30 / 100 / 130 | OK 201, 202 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 2 / 2 |
| Lexical baseline r8 / q6 | critical | 35 / 110 / 145 | OK 202, 203 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 2 / 2 |
| Dense candidates r7 / q5 | critical | 30 / 30 / 60 | OK 203, 202 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r8 / q6 | critical | 35 / 32 / 67 | OK 201, 203 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Equal-weight fusion r7 / q5 | critical | 30 / 60 / 90 | OK 201, 203 | complete / material success | S40 / compatible | pass | pass | pass | 20,480 / 3 / 2 |
| Equal-weight fusion r8 / q6 | critical | 35 / 62 / 97 | OK 201, 202 | complete / material success | S40 / compatible | pass | pass | pass | 20,480 / 3 / 2 |
CAPPED means explicit RETRIEVAL_LIMIT with one inspected eligible ID. FAILURE means UNAVAILABLE at the 150 ms deadline with no delivered IDs. Neither is a complete empty answer; these queries have eligible documents. Weight-2 fusion has no manufactured request rows.
Selected frozen query answers
q5: Diagnose a live cutover that missed a replacement or a deletion; require incident evidence.
Eligible IDs 201, 202, 203; judged relevant eligible IDs 201, 202. These are the C40/S40 reference.
- Lexical baseline: selected 201, 202; judged recall 1 (100%); precision 1 (100%) .
- Dense candidates: selected 203, 202; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 201, 203; judged recall 1/2 (50%); precision 1/2 (50%) .
- Lexical-weight-2 fusion (post-hoc): selected 201, 202; judged recall 1 (100%); precision 1 (100%) .
q6: Diagnose a deleted result reappearing after cache-generation rerouting; require incident evidence.
Eligible IDs 201, 202, 203; judged relevant eligible IDs 202, 203. These are the C40/S40 reference.
- Lexical baseline: selected 202, 203; judged recall 1 (100%); precision 1 (100%) .
- Dense candidates: selected 201, 203; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 201, 202; judged recall 1/2 (50%); precision 1/2 (50%) .
- Lexical-weight-2 fusion (post-hoc): selected 202, 203; judged recall 1 (100%); precision 1 (100%) .
Change lexical weight — All scheduled attempts
If lexical weight changes from 1 to 2, may we approve a broad launch? Selected population: All scheduled attempts .
Overall proposal evidence: All six fixture queries now have recall 1. This is a post-hoc hypothesis; operational evidence for the changed variant is unknown.
Continue testing on a fresh frozen holdout and an independently scheduled arrival experiment. Reusing this fixture to claim an unbiased gain would tune to the test evidence.
All scheduled attempts: 6 distinct information needs; planned-workload share 1 (100%). 8 scheduled attempts; request-schedule share 1 (100%). These are separate denominators.
Frozen C40/S40 reference audit
Binary judged recall@2 measures relevant eligible hits over all judged-relevant eligible IDs. Precision@2 uses requested k=2; it equals recall in these six full quality rows. Dense fidelity measures agreement with the exact eligible dense top-2, separately from relevance. A selected-slice quality pass cannot approve an all-slice proposal.
Lexical baseline
- Frozen judged recall@2
- 53/80 (66.25%)
- Dense exact-neighbor fidelity
- not applicable; distinct answer function
- All-required-slice quality gate
- pass
- Material-success coverage
- 8 /8; 1 (100%) — pass against 99%
Continue testing. The fixture cannot establish a real deployment approval.
Dense candidates
- Frozen judged recall@2
- 1/2 (50%)
- Dense exact-neighbor fidelity
- 1 (100%)
- All-required-slice quality gate
- FAIL
- Material-success coverage
- 6 /8; 3/4 (75%) — FAIL against 99%
Reject the proposed all-slice launch. The fixture cannot establish a real deployment approval.
Equal-weight fusion
- Frozen judged recall@2
- 19/20 (95%)
- Dense exact-neighbor fidelity
- not applicable; distinct answer function
- All-required-slice quality gate
- FAIL
- Material-success coverage
- 8 /8; 1 (100%) — pass against 99%
Reject the proposed all-slice launch. The fixture cannot establish a real deployment approval.
Lexical-weight-2 fusion (post-hoc)
- Frozen judged recall@2
- 1 (100%)
- Dense exact-neighbor fidelity
- not applicable; distinct answer function
- All-required-slice quality gate
- pass
- Operating evidence
- unknown — post-hoc weights have no operating schedule
Continue testing. Quality was repaired after inspecting this fixture; freeze a new hypothesis and test fresh held-out needs.
Material success means the full response passes compatible-reference, snapshot eligibility, freshness and current disclosure requirements. Its denominator includes every scheduled attempt. Nominal OK and complete count can remain unchanged when a gate fails. Durations above are synthetic elapsed durations, not measured latency or p99.
Every selected request, including cap/failure and gate violations
| Plan / request / query | Slice | Scheduled / duration / completion ms | Outcome / IDs sent | Complete / material | S / reference | Snapshot eligibility | Freshness | Current disclosure | Bytes / requests / rounds |
|---|---|---|---|---|---|---|---|---|---|
| Lexical baseline r1 / q1 | common | 0 / 70 / 70 | OK 101, 103 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r2 / q2 | common | 5 / 75 / 80 | OK 101, 102 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r3 / q3 | common | 10 / 80 / 90 | OK 102, 101 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r4 / q4 | common | 15 / 85 / 100 | OK 103, 104 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r5 / q1 | common | 20 / 75 / 95 | OK 101, 103 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r6 / q2 | common | 25 / 80 / 105 | OK 101, 102 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r7 / q5 | critical | 30 / 100 / 130 | OK 201, 202 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 2 / 2 |
| Lexical baseline r8 / q6 | critical | 35 / 110 / 145 | OK 202, 203 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 2 / 2 |
| Dense candidates r1 / q1 | common | 0 / 20 / 20 | OK 102, 104 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r2 / q2 | common | 5 / 25 / 30 | OK 103, 104 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r3 / q3 | common | 10 / 100 / 110 | RETRIEVAL_LIMIT 104 | incomplete / NOT material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r4 / q4 | common | 15 / 150 / 165 | UNAVAILABLE none | incomplete / NOT material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r5 / q1 | common | 20 / 22 / 42 | OK 102, 104 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r6 / q2 | common | 25 / 26 / 51 | OK 103, 104 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r7 / q5 | critical | 30 / 30 / 60 | OK 203, 202 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r8 / q6 | critical | 35 / 32 / 67 | OK 201, 203 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Equal-weight fusion r1 / q1 | common | 0 / 40 / 40 | OK 101, 102 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r2 / q2 | common | 5 / 42 / 47 | OK 101, 103 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r3 / q3 | common | 10 / 44 / 54 | OK 102, 104 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r4 / q4 | common | 15 / 46 / 61 | OK 104, 103 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r5 / q1 | common | 20 / 42 / 62 | OK 101, 102 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r6 / q2 | common | 25 / 45 / 70 | OK 101, 103 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r7 / q5 | critical | 30 / 60 / 90 | OK 201, 203 | complete / material success | S40 / compatible | pass | pass | pass | 20,480 / 3 / 2 |
| Equal-weight fusion r8 / q6 | critical | 35 / 62 / 97 | OK 201, 202 | complete / material success | S40 / compatible | pass | pass | pass | 20,480 / 3 / 2 |
CAPPED means explicit RETRIEVAL_LIMIT with one inspected eligible ID. FAILURE means UNAVAILABLE at the 150 ms deadline with no delivered IDs. Neither is a complete empty answer; these queries have eligible documents. Weight-2 fusion has no manufactured request rows.
Selected frozen query answers
q1: Understand safe tenant-scoped reads immediately after an acknowledged update.
Eligible IDs 101, 102, 103, 104; judged relevant eligible IDs 101, 102. These are the C40/S40 reference.
- Lexical baseline: selected 101, 103; judged recall 1/2 (50%); precision 1/2 (50%) .
- Dense candidates: selected 102, 104; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 101, 102; judged recall 1 (100%); precision 1 (100%) .
- Lexical-weight-2 fusion (post-hoc): selected 101, 102; judged recall 1 (100%); precision 1 (100%) .
q2: Route a tenant request safely when its warm cache node is replaced.
Eligible IDs 101, 102, 103, 104; judged relevant eligible IDs 101, 103. These are the C40/S40 reference.
- Lexical baseline: selected 101, 102; judged recall 1/2 (50%); precision 1/2 (50%) .
- Dense candidates: selected 103, 104; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 101, 103; judged recall 1 (100%); precision 1 (100%) .
- Lexical-weight-2 fusion (post-hoc): selected 101, 103; judged recall 1 (100%); precision 1 (100%) .
q3: Preserve fresh results when ingestion backlog competes with finite query budgets.
Eligible IDs 101, 102, 103, 104; judged relevant eligible IDs 102, 104. These are the C40/S40 reference.
- Lexical baseline: selected 102, 101; judged recall 1/2 (50%); precision 1/2 (50%) .
- Dense candidates: selected 104, 103; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 102, 104; judged recall 1 (100%); precision 1 (100%) .
- Lexical-weight-2 fusion (post-hoc): selected 102, 104; judged recall 1 (100%); precision 1 (100%) .
q4: Plan admission during a cold-cache replacement and traffic burst.
Eligible IDs 101, 102, 103, 104; judged relevant eligible IDs 103, 104. These are the C40/S40 reference.
- Lexical baseline: selected 103, 104; judged recall 1 (100%); precision 1 (100%) .
- Dense candidates: selected 104, 102; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 104, 103; judged recall 1 (100%); precision 1 (100%) .
- Lexical-weight-2 fusion (post-hoc): selected 103, 104; judged recall 1 (100%); precision 1 (100%) .
q5: Diagnose a live cutover that missed a replacement or a deletion; require incident evidence.
Eligible IDs 201, 202, 203; judged relevant eligible IDs 201, 202. These are the C40/S40 reference.
- Lexical baseline: selected 201, 202; judged recall 1 (100%); precision 1 (100%) .
- Dense candidates: selected 203, 202; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 201, 203; judged recall 1/2 (50%); precision 1/2 (50%) .
- Lexical-weight-2 fusion (post-hoc): selected 201, 202; judged recall 1 (100%); precision 1 (100%) .
q6: Diagnose a deleted result reappearing after cache-generation rerouting; require incident evidence.
Eligible IDs 201, 202, 203; judged relevant eligible IDs 202, 203. These are the C40/S40 reference.
- Lexical baseline: selected 202, 203; judged recall 1 (100%); precision 1 (100%) .
- Dense candidates: selected 201, 203; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 201, 202; judged recall 1/2 (50%); precision 1/2 (50%) .
- Lexical-weight-2 fusion (post-hoc): selected 202, 203; judged recall 1 (100%); precision 1 (100%) .
Change lexical weight — Common tenant-A queries
If lexical weight changes from 1 to 2, may we approve a broad launch? Selected population: Common tenant-A queries .
Overall proposal evidence: All six fixture queries now have recall 1. This is a post-hoc hypothesis; operational evidence for the changed variant is unknown.
Continue testing on a fresh frozen holdout and an independently scheduled arrival experiment. Reusing this fixture to claim an unbiased gain would tune to the test evidence.
Common tenant-A queries: 4 distinct information needs; planned-workload share 9/10 (90%). 6 scheduled attempts; request-schedule share 3/4 (75%). These are separate denominators.
Frozen C40/S40 reference audit
Binary judged recall@2 measures relevant eligible hits over all judged-relevant eligible IDs. Precision@2 uses requested k=2; it equals recall in these six full quality rows. Dense fidelity measures agreement with the exact eligible dense top-2, separately from relevance. A selected-slice quality pass cannot approve an all-slice proposal.
Lexical baseline
- Frozen judged recall@2
- 5/8 (62.5%)
- Dense exact-neighbor fidelity
- not applicable; distinct answer function
- All-required-slice quality gate
- pass
- Material-success coverage
- 6 /6; 1 (100%) — pass against 99%
Continue testing. The fixture cannot establish a real deployment approval.
Dense candidates
- Frozen judged recall@2
- 1/2 (50%)
- Dense exact-neighbor fidelity
- 1 (100%)
- All-required-slice quality gate
- FAIL
- Material-success coverage
- 4 /6; 2/3 (66.67%) — FAIL against 99%
Reject the proposed all-slice launch. The fixture cannot establish a real deployment approval.
Equal-weight fusion
- Frozen judged recall@2
- 1 (100%)
- Dense exact-neighbor fidelity
- not applicable; distinct answer function
- All-required-slice quality gate
- FAIL
- Material-success coverage
- 6 /6; 1 (100%) — pass against 99%
Reject the proposed all-slice launch. The fixture cannot establish a real deployment approval.
Lexical-weight-2 fusion (post-hoc)
- Frozen judged recall@2
- 1 (100%)
- Dense exact-neighbor fidelity
- not applicable; distinct answer function
- All-required-slice quality gate
- pass
- Operating evidence
- unknown — post-hoc weights have no operating schedule
Continue testing. Quality was repaired after inspecting this fixture; freeze a new hypothesis and test fresh held-out needs.
Material success means the full response passes compatible-reference, snapshot eligibility, freshness and current disclosure requirements. Its denominator includes every scheduled attempt. Nominal OK and complete count can remain unchanged when a gate fails. Durations above are synthetic elapsed durations, not measured latency or p99.
Every selected request, including cap/failure and gate violations
| Plan / request / query | Slice | Scheduled / duration / completion ms | Outcome / IDs sent | Complete / material | S / reference | Snapshot eligibility | Freshness | Current disclosure | Bytes / requests / rounds |
|---|---|---|---|---|---|---|---|---|---|
| Lexical baseline r1 / q1 | common | 0 / 70 / 70 | OK 101, 103 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r2 / q2 | common | 5 / 75 / 80 | OK 101, 102 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r3 / q3 | common | 10 / 80 / 90 | OK 102, 101 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r4 / q4 | common | 15 / 85 / 100 | OK 103, 104 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r5 / q1 | common | 20 / 75 / 95 | OK 101, 103 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r6 / q2 | common | 25 / 80 / 105 | OK 101, 102 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Dense candidates r1 / q1 | common | 0 / 20 / 20 | OK 102, 104 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r2 / q2 | common | 5 / 25 / 30 | OK 103, 104 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r3 / q3 | common | 10 / 100 / 110 | RETRIEVAL_LIMIT 104 | incomplete / NOT material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r4 / q4 | common | 15 / 150 / 165 | UNAVAILABLE none | incomplete / NOT material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r5 / q1 | common | 20 / 22 / 42 | OK 102, 104 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r6 / q2 | common | 25 / 26 / 51 | OK 103, 104 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Equal-weight fusion r1 / q1 | common | 0 / 40 / 40 | OK 101, 102 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r2 / q2 | common | 5 / 42 / 47 | OK 101, 103 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r3 / q3 | common | 10 / 44 / 54 | OK 102, 104 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r4 / q4 | common | 15 / 46 / 61 | OK 104, 103 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r5 / q1 | common | 20 / 42 / 62 | OK 101, 102 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r6 / q2 | common | 25 / 45 / 70 | OK 101, 103 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
CAPPED means explicit RETRIEVAL_LIMIT with one inspected eligible ID. FAILURE means UNAVAILABLE at the 150 ms deadline with no delivered IDs. Neither is a complete empty answer; these queries have eligible documents. Weight-2 fusion has no manufactured request rows.
Selected frozen query answers
q1: Understand safe tenant-scoped reads immediately after an acknowledged update.
Eligible IDs 101, 102, 103, 104; judged relevant eligible IDs 101, 102. These are the C40/S40 reference.
- Lexical baseline: selected 101, 103; judged recall 1/2 (50%); precision 1/2 (50%) .
- Dense candidates: selected 102, 104; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 101, 102; judged recall 1 (100%); precision 1 (100%) .
- Lexical-weight-2 fusion (post-hoc): selected 101, 102; judged recall 1 (100%); precision 1 (100%) .
q2: Route a tenant request safely when its warm cache node is replaced.
Eligible IDs 101, 102, 103, 104; judged relevant eligible IDs 101, 103. These are the C40/S40 reference.
- Lexical baseline: selected 101, 102; judged recall 1/2 (50%); precision 1/2 (50%) .
- Dense candidates: selected 103, 104; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 101, 103; judged recall 1 (100%); precision 1 (100%) .
- Lexical-weight-2 fusion (post-hoc): selected 101, 103; judged recall 1 (100%); precision 1 (100%) .
q3: Preserve fresh results when ingestion backlog competes with finite query budgets.
Eligible IDs 101, 102, 103, 104; judged relevant eligible IDs 102, 104. These are the C40/S40 reference.
- Lexical baseline: selected 102, 101; judged recall 1/2 (50%); precision 1/2 (50%) .
- Dense candidates: selected 104, 103; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 102, 104; judged recall 1 (100%); precision 1 (100%) .
- Lexical-weight-2 fusion (post-hoc): selected 102, 104; judged recall 1 (100%); precision 1 (100%) .
q4: Plan admission during a cold-cache replacement and traffic burst.
Eligible IDs 101, 102, 103, 104; judged relevant eligible IDs 103, 104. These are the C40/S40 reference.
- Lexical baseline: selected 103, 104; judged recall 1 (100%); precision 1 (100%) .
- Dense candidates: selected 104, 102; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 104, 103; judged recall 1 (100%); precision 1 (100%) .
- Lexical-weight-2 fusion (post-hoc): selected 103, 104; judged recall 1 (100%); precision 1 (100%) .
Change lexical weight — Required tenant-B incident queries
If lexical weight changes from 1 to 2, may we approve a broad launch? Selected population: Required tenant-B incident queries .
Overall proposal evidence: All six fixture queries now have recall 1. This is a post-hoc hypothesis; operational evidence for the changed variant is unknown.
Continue testing on a fresh frozen holdout and an independently scheduled arrival experiment. Reusing this fixture to claim an unbiased gain would tune to the test evidence.
Required tenant-B incident queries: 2 distinct information needs; planned-workload share 1/10 (10%). 2 scheduled attempts; request-schedule share 1/4 (25%). These are separate denominators.
Frozen C40/S40 reference audit
Binary judged recall@2 measures relevant eligible hits over all judged-relevant eligible IDs. Precision@2 uses requested k=2; it equals recall in these six full quality rows. Dense fidelity measures agreement with the exact eligible dense top-2, separately from relevance. A selected-slice quality pass cannot approve an all-slice proposal.
Lexical baseline
- Frozen judged recall@2
- 1 (100%)
- Dense exact-neighbor fidelity
- not applicable; distinct answer function
- All-required-slice quality gate
- pass
- Material-success coverage
- 2 /2; 1 (100%) — pass against 99%
Continue testing. The fixture cannot establish a real deployment approval.
Dense candidates
- Frozen judged recall@2
- 1/2 (50%)
- Dense exact-neighbor fidelity
- 1 (100%)
- All-required-slice quality gate
- FAIL
- Material-success coverage
- 2 /2; 1 (100%) — pass against 99%
Reject the proposed all-slice launch. The fixture cannot establish a real deployment approval.
Equal-weight fusion
- Frozen judged recall@2
- 1/2 (50%)
- Dense exact-neighbor fidelity
- not applicable; distinct answer function
- All-required-slice quality gate
- FAIL
- Material-success coverage
- 2 /2; 1 (100%) — pass against 99%
Reject the proposed all-slice launch. The fixture cannot establish a real deployment approval.
Lexical-weight-2 fusion (post-hoc)
- Frozen judged recall@2
- 1 (100%)
- Dense exact-neighbor fidelity
- not applicable; distinct answer function
- All-required-slice quality gate
- pass
- Operating evidence
- unknown — post-hoc weights have no operating schedule
Continue testing. Quality was repaired after inspecting this fixture; freeze a new hypothesis and test fresh held-out needs.
Material success means the full response passes compatible-reference, snapshot eligibility, freshness and current disclosure requirements. Its denominator includes every scheduled attempt. Nominal OK and complete count can remain unchanged when a gate fails. Durations above are synthetic elapsed durations, not measured latency or p99.
Every selected request, including cap/failure and gate violations
| Plan / request / query | Slice | Scheduled / duration / completion ms | Outcome / IDs sent | Complete / material | S / reference | Snapshot eligibility | Freshness | Current disclosure | Bytes / requests / rounds |
|---|---|---|---|---|---|---|---|---|---|
| Lexical baseline r7 / q5 | critical | 30 / 100 / 130 | OK 201, 202 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 2 / 2 |
| Lexical baseline r8 / q6 | critical | 35 / 110 / 145 | OK 202, 203 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 2 / 2 |
| Dense candidates r7 / q5 | critical | 30 / 30 / 60 | OK 203, 202 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r8 / q6 | critical | 35 / 32 / 67 | OK 201, 203 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Equal-weight fusion r7 / q5 | critical | 30 / 60 / 90 | OK 201, 203 | complete / material success | S40 / compatible | pass | pass | pass | 20,480 / 3 / 2 |
| Equal-weight fusion r8 / q6 | critical | 35 / 62 / 97 | OK 201, 202 | complete / material success | S40 / compatible | pass | pass | pass | 20,480 / 3 / 2 |
CAPPED means explicit RETRIEVAL_LIMIT with one inspected eligible ID. FAILURE means UNAVAILABLE at the 150 ms deadline with no delivered IDs. Neither is a complete empty answer; these queries have eligible documents. Weight-2 fusion has no manufactured request rows.
Selected frozen query answers
q5: Diagnose a live cutover that missed a replacement or a deletion; require incident evidence.
Eligible IDs 201, 202, 203; judged relevant eligible IDs 201, 202. These are the C40/S40 reference.
- Lexical baseline: selected 201, 202; judged recall 1 (100%); precision 1 (100%) .
- Dense candidates: selected 203, 202; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 201, 203; judged recall 1/2 (50%); precision 1/2 (50%) .
- Lexical-weight-2 fusion (post-hoc): selected 201, 202; judged recall 1 (100%); precision 1 (100%) .
q6: Diagnose a deleted result reappearing after cache-generation rerouting; require incident evidence.
Eligible IDs 201, 202, 203; judged relevant eligible IDs 202, 203. These are the C40/S40 reference.
- Lexical baseline: selected 202, 203; judged recall 1 (100%); precision 1 (100%) .
- Dense candidates: selected 201, 203; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 201, 202; judged recall 1/2 (50%); precision 1/2 (50%) .
- Lexical-weight-2 fusion (post-hoc): selected 202, 203; judged recall 1 (100%); precision 1 (100%) .
Change observed snapshot — All scheduled attempts
What changes when r7 observes S39 but requires commit 40? Selected population: All scheduled attempts .
Overall proposal evidence: Freshness fails; S40 judgments cannot score the S39 response as a compatible comparison. Material-success coverage becomes 7/8 (87.5%). Snapshot eligibility is unknown because no S39 corpus is provided.
Keep freshness, reference compatibility and unknown eligibility separate from relevance. A nominal OK response can fail material-success requirements.
All scheduled attempts: 6 distinct information needs; planned-workload share 1 (100%). 8 scheduled attempts; request-schedule share 1 (100%). These are separate denominators.
Frozen C40/S40 reference audit; these values do not score the changed S39 response
Binary judged recall@2 measures relevant eligible hits over all judged-relevant eligible IDs. Precision@2 uses requested k=2; it equals recall in these six full quality rows. Dense fidelity measures agreement with the exact eligible dense top-2, separately from relevance. A selected-slice quality pass cannot approve an all-slice proposal.
Lexical baseline
- Frozen judged recall@2
- 53/80 (66.25%)
- Dense exact-neighbor fidelity
- not applicable; distinct answer function
- All-required-slice quality gate
- pass
- Material-success coverage
- 8 /8; 1 (100%) — pass against 99%
Continue testing. The fixture cannot establish a real deployment approval.
Dense candidates
- Frozen judged recall@2
- 1/2 (50%)
- Dense exact-neighbor fidelity
- 1 (100%)
- All-required-slice quality gate
- FAIL
- Material-success coverage
- 6 /8; 3/4 (75%) — FAIL against 99%
Reject the proposed all-slice launch. The fixture cannot establish a real deployment approval.
Equal-weight fusion
- Frozen judged recall@2
- 19/20 (95%)
- Dense exact-neighbor fidelity
- not applicable; distinct answer function
- All-required-slice quality gate
- FAIL
- Material-success coverage
- 7 /8; 7/8 (87.5%) — FAIL against 99%
Reject the proposed all-slice launch. The fixture cannot establish a real deployment approval.
Lexical-weight-2 fusion (post-hoc)
- Frozen judged recall@2
- 1 (100%)
- Dense exact-neighbor fidelity
- not applicable; distinct answer function
- All-required-slice quality gate
- pass
- Operating evidence
- unknown — post-hoc weights have no operating schedule
Continue testing. Quality was repaired after inspecting this fixture; freeze a new hypothesis and test fresh held-out needs.
Material success means the full response passes compatible-reference, snapshot eligibility, freshness and current disclosure requirements. Its denominator includes every scheduled attempt. Nominal OK and complete count can remain unchanged when a gate fails. Durations above are synthetic elapsed durations, not measured latency or p99.
Every selected request, including cap/failure and gate violations
| Plan / request / query | Slice | Scheduled / duration / completion ms | Outcome / IDs sent | Complete / material | S / reference | Snapshot eligibility | Freshness | Current disclosure | Bytes / requests / rounds |
|---|---|---|---|---|---|---|---|---|---|
| Lexical baseline r1 / q1 | common | 0 / 70 / 70 | OK 101, 103 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r2 / q2 | common | 5 / 75 / 80 | OK 101, 102 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r3 / q3 | common | 10 / 80 / 90 | OK 102, 101 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r4 / q4 | common | 15 / 85 / 100 | OK 103, 104 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r5 / q1 | common | 20 / 75 / 95 | OK 101, 103 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r6 / q2 | common | 25 / 80 / 105 | OK 101, 102 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r7 / q5 | critical | 30 / 100 / 130 | OK 201, 202 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 2 / 2 |
| Lexical baseline r8 / q6 | critical | 35 / 110 / 145 | OK 202, 203 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 2 / 2 |
| Dense candidates r1 / q1 | common | 0 / 20 / 20 | OK 102, 104 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r2 / q2 | common | 5 / 25 / 30 | OK 103, 104 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r3 / q3 | common | 10 / 100 / 110 | RETRIEVAL_LIMIT 104 | incomplete / NOT material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r4 / q4 | common | 15 / 150 / 165 | UNAVAILABLE none | incomplete / NOT material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r5 / q1 | common | 20 / 22 / 42 | OK 102, 104 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r6 / q2 | common | 25 / 26 / 51 | OK 103, 104 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r7 / q5 | critical | 30 / 30 / 60 | OK 203, 202 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r8 / q6 | critical | 35 / 32 / 67 | OK 201, 203 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Equal-weight fusion r1 / q1 | common | 0 / 40 / 40 | OK 101, 102 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r2 / q2 | common | 5 / 42 / 47 | OK 101, 103 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r3 / q3 | common | 10 / 44 / 54 | OK 102, 104 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r4 / q4 | common | 15 / 46 / 61 | OK 104, 103 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r5 / q1 | common | 20 / 42 / 62 | OK 101, 102 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r6 / q2 | common | 25 / 45 / 70 | OK 101, 103 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r7 / q5 | critical | 30 / 60 / 90 | OK 201, 203 | complete / NOT material success | S39 / INCOMPATIBLE | unknown | FAIL | pass | 20,480 / 3 / 2 |
| Equal-weight fusion r8 / q6 | critical | 35 / 62 / 97 | OK 201, 202 | complete / material success | S40 / compatible | pass | pass | pass | 20,480 / 3 / 2 |
CAPPED means explicit RETRIEVAL_LIMIT with one inspected eligible ID. FAILURE means UNAVAILABLE at the 150 ms deadline with no delivered IDs. Neither is a complete empty answer; these queries have eligible documents. Weight-2 fusion has no manufactured request rows.
Selected frozen query answers
q1: Understand safe tenant-scoped reads immediately after an acknowledged update.
Eligible IDs 101, 102, 103, 104; judged relevant eligible IDs 101, 102. These are the C40/S40 reference.
- Lexical baseline: selected 101, 103; judged recall 1/2 (50%); precision 1/2 (50%) .
- Dense candidates: selected 102, 104; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 101, 102; judged recall 1 (100%); precision 1 (100%) .
- Lexical-weight-2 fusion (post-hoc): selected 101, 102; judged recall 1 (100%); precision 1 (100%) .
q2: Route a tenant request safely when its warm cache node is replaced.
Eligible IDs 101, 102, 103, 104; judged relevant eligible IDs 101, 103. These are the C40/S40 reference.
- Lexical baseline: selected 101, 102; judged recall 1/2 (50%); precision 1/2 (50%) .
- Dense candidates: selected 103, 104; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 101, 103; judged recall 1 (100%); precision 1 (100%) .
- Lexical-weight-2 fusion (post-hoc): selected 101, 103; judged recall 1 (100%); precision 1 (100%) .
q3: Preserve fresh results when ingestion backlog competes with finite query budgets.
Eligible IDs 101, 102, 103, 104; judged relevant eligible IDs 102, 104. These are the C40/S40 reference.
- Lexical baseline: selected 102, 101; judged recall 1/2 (50%); precision 1/2 (50%) .
- Dense candidates: selected 104, 103; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 102, 104; judged recall 1 (100%); precision 1 (100%) .
- Lexical-weight-2 fusion (post-hoc): selected 102, 104; judged recall 1 (100%); precision 1 (100%) .
q4: Plan admission during a cold-cache replacement and traffic burst.
Eligible IDs 101, 102, 103, 104; judged relevant eligible IDs 103, 104. These are the C40/S40 reference.
- Lexical baseline: selected 103, 104; judged recall 1 (100%); precision 1 (100%) .
- Dense candidates: selected 104, 102; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 104, 103; judged recall 1 (100%); precision 1 (100%) .
- Lexical-weight-2 fusion (post-hoc): selected 103, 104; judged recall 1 (100%); precision 1 (100%) .
q5: Diagnose a live cutover that missed a replacement or a deletion; require incident evidence.
Eligible IDs 201, 202, 203; judged relevant eligible IDs 201, 202. These are the C40/S40 reference.
- Lexical baseline: selected 201, 202; judged recall 1 (100%); precision 1 (100%) .
- Dense candidates: selected 203, 202; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 201, 203; judged recall 1/2 (50%); precision 1/2 (50%) .
- Lexical-weight-2 fusion (post-hoc): selected 201, 202; judged recall 1 (100%); precision 1 (100%) .
q6: Diagnose a deleted result reappearing after cache-generation rerouting; require incident evidence.
Eligible IDs 201, 202, 203; judged relevant eligible IDs 202, 203. These are the C40/S40 reference.
- Lexical baseline: selected 202, 203; judged recall 1 (100%); precision 1 (100%) .
- Dense candidates: selected 201, 203; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 201, 202; judged recall 1/2 (50%); precision 1/2 (50%) .
- Lexical-weight-2 fusion (post-hoc): selected 202, 203; judged recall 1 (100%); precision 1 (100%) .
All-slice launch gate: FAIL. Changed r7 lacks a fresh compatible reference; current-response relevance is unavailable. A common-only display does not change the proposed exposure scope.
Change observed snapshot — Common tenant-A queries
What changes when r7 observes S39 but requires commit 40? Selected population: Common tenant-A queries .
Overall proposal evidence: Freshness fails; S40 judgments cannot score the S39 response as a compatible comparison. Material-success coverage becomes 7/8 (87.5%). Snapshot eligibility is unknown because no S39 corpus is provided.
Keep freshness, reference compatibility and unknown eligibility separate from relevance. A nominal OK response can fail material-success requirements.
Common tenant-A queries: 4 distinct information needs; planned-workload share 9/10 (90%). 6 scheduled attempts; request-schedule share 3/4 (75%). These are separate denominators.
Frozen C40/S40 reference audit; these values do not score the changed S39 response
Binary judged recall@2 measures relevant eligible hits over all judged-relevant eligible IDs. Precision@2 uses requested k=2; it equals recall in these six full quality rows. Dense fidelity measures agreement with the exact eligible dense top-2, separately from relevance. A selected-slice quality pass cannot approve an all-slice proposal.
Lexical baseline
- Frozen judged recall@2
- 5/8 (62.5%)
- Dense exact-neighbor fidelity
- not applicable; distinct answer function
- All-required-slice quality gate
- pass
- Material-success coverage
- 6 /6; 1 (100%) — pass against 99%
Continue testing. The fixture cannot establish a real deployment approval.
Dense candidates
- Frozen judged recall@2
- 1/2 (50%)
- Dense exact-neighbor fidelity
- 1 (100%)
- All-required-slice quality gate
- FAIL
- Material-success coverage
- 4 /6; 2/3 (66.67%) — FAIL against 99%
Reject the proposed all-slice launch. The fixture cannot establish a real deployment approval.
Equal-weight fusion
- Frozen judged recall@2
- 1 (100%)
- Dense exact-neighbor fidelity
- not applicable; distinct answer function
- All-required-slice quality gate
- FAIL
- Material-success coverage
- 6 /6; 1 (100%) — pass against 99%
Reject the proposed all-slice launch. The fixture cannot establish a real deployment approval.
Lexical-weight-2 fusion (post-hoc)
- Frozen judged recall@2
- 1 (100%)
- Dense exact-neighbor fidelity
- not applicable; distinct answer function
- All-required-slice quality gate
- pass
- Operating evidence
- unknown — post-hoc weights have no operating schedule
Continue testing. Quality was repaired after inspecting this fixture; freeze a new hypothesis and test fresh held-out needs.
Material success means the full response passes compatible-reference, snapshot eligibility, freshness and current disclosure requirements. Its denominator includes every scheduled attempt. Nominal OK and complete count can remain unchanged when a gate fails. Durations above are synthetic elapsed durations, not measured latency or p99.
Every selected request, including cap/failure and gate violations
| Plan / request / query | Slice | Scheduled / duration / completion ms | Outcome / IDs sent | Complete / material | S / reference | Snapshot eligibility | Freshness | Current disclosure | Bytes / requests / rounds |
|---|---|---|---|---|---|---|---|---|---|
| Lexical baseline r1 / q1 | common | 0 / 70 / 70 | OK 101, 103 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r2 / q2 | common | 5 / 75 / 80 | OK 101, 102 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r3 / q3 | common | 10 / 80 / 90 | OK 102, 101 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r4 / q4 | common | 15 / 85 / 100 | OK 103, 104 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r5 / q1 | common | 20 / 75 / 95 | OK 101, 103 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r6 / q2 | common | 25 / 80 / 105 | OK 101, 102 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Dense candidates r1 / q1 | common | 0 / 20 / 20 | OK 102, 104 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r2 / q2 | common | 5 / 25 / 30 | OK 103, 104 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r3 / q3 | common | 10 / 100 / 110 | RETRIEVAL_LIMIT 104 | incomplete / NOT material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r4 / q4 | common | 15 / 150 / 165 | UNAVAILABLE none | incomplete / NOT material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r5 / q1 | common | 20 / 22 / 42 | OK 102, 104 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r6 / q2 | common | 25 / 26 / 51 | OK 103, 104 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Equal-weight fusion r1 / q1 | common | 0 / 40 / 40 | OK 101, 102 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r2 / q2 | common | 5 / 42 / 47 | OK 101, 103 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r3 / q3 | common | 10 / 44 / 54 | OK 102, 104 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r4 / q4 | common | 15 / 46 / 61 | OK 104, 103 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r5 / q1 | common | 20 / 42 / 62 | OK 101, 102 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r6 / q2 | common | 25 / 45 / 70 | OK 101, 103 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
CAPPED means explicit RETRIEVAL_LIMIT with one inspected eligible ID. FAILURE means UNAVAILABLE at the 150 ms deadline with no delivered IDs. Neither is a complete empty answer; these queries have eligible documents. Weight-2 fusion has no manufactured request rows.
Selected frozen query answers
q1: Understand safe tenant-scoped reads immediately after an acknowledged update.
Eligible IDs 101, 102, 103, 104; judged relevant eligible IDs 101, 102. These are the C40/S40 reference.
- Lexical baseline: selected 101, 103; judged recall 1/2 (50%); precision 1/2 (50%) .
- Dense candidates: selected 102, 104; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 101, 102; judged recall 1 (100%); precision 1 (100%) .
- Lexical-weight-2 fusion (post-hoc): selected 101, 102; judged recall 1 (100%); precision 1 (100%) .
q2: Route a tenant request safely when its warm cache node is replaced.
Eligible IDs 101, 102, 103, 104; judged relevant eligible IDs 101, 103. These are the C40/S40 reference.
- Lexical baseline: selected 101, 102; judged recall 1/2 (50%); precision 1/2 (50%) .
- Dense candidates: selected 103, 104; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 101, 103; judged recall 1 (100%); precision 1 (100%) .
- Lexical-weight-2 fusion (post-hoc): selected 101, 103; judged recall 1 (100%); precision 1 (100%) .
q3: Preserve fresh results when ingestion backlog competes with finite query budgets.
Eligible IDs 101, 102, 103, 104; judged relevant eligible IDs 102, 104. These are the C40/S40 reference.
- Lexical baseline: selected 102, 101; judged recall 1/2 (50%); precision 1/2 (50%) .
- Dense candidates: selected 104, 103; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 102, 104; judged recall 1 (100%); precision 1 (100%) .
- Lexical-weight-2 fusion (post-hoc): selected 102, 104; judged recall 1 (100%); precision 1 (100%) .
q4: Plan admission during a cold-cache replacement and traffic burst.
Eligible IDs 101, 102, 103, 104; judged relevant eligible IDs 103, 104. These are the C40/S40 reference.
- Lexical baseline: selected 103, 104; judged recall 1 (100%); precision 1 (100%) .
- Dense candidates: selected 104, 102; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 104, 103; judged recall 1 (100%); precision 1 (100%) .
- Lexical-weight-2 fusion (post-hoc): selected 103, 104; judged recall 1 (100%); precision 1 (100%) .
All-slice launch gate: FAIL. Changed r7 lacks a fresh compatible reference; current-response relevance is unavailable. A common-only display does not change the proposed exposure scope.
Change observed snapshot — Required tenant-B incident queries
What changes when r7 observes S39 but requires commit 40? Selected population: Required tenant-B incident queries .
Overall proposal evidence: Freshness fails; S40 judgments cannot score the S39 response as a compatible comparison. Material-success coverage becomes 7/8 (87.5%). Snapshot eligibility is unknown because no S39 corpus is provided.
Keep freshness, reference compatibility and unknown eligibility separate from relevance. A nominal OK response can fail material-success requirements.
Required tenant-B incident queries: 2 distinct information needs; planned-workload share 1/10 (10%). 2 scheduled attempts; request-schedule share 1/4 (25%). These are separate denominators.
Frozen C40/S40 reference audit; these values do not score the changed S39 response
Binary judged recall@2 measures relevant eligible hits over all judged-relevant eligible IDs. Precision@2 uses requested k=2; it equals recall in these six full quality rows. Dense fidelity measures agreement with the exact eligible dense top-2, separately from relevance. A selected-slice quality pass cannot approve an all-slice proposal.
Lexical baseline
- Frozen judged recall@2
- 1 (100%)
- Dense exact-neighbor fidelity
- not applicable; distinct answer function
- All-required-slice quality gate
- pass
- Material-success coverage
- 2 /2; 1 (100%) — pass against 99%
Continue testing. The fixture cannot establish a real deployment approval.
Dense candidates
- Frozen judged recall@2
- 1/2 (50%)
- Dense exact-neighbor fidelity
- 1 (100%)
- All-required-slice quality gate
- FAIL
- Material-success coverage
- 2 /2; 1 (100%) — pass against 99%
Reject the proposed all-slice launch. The fixture cannot establish a real deployment approval.
Equal-weight fusion
- Frozen judged recall@2
- 1/2 (50%)
- Dense exact-neighbor fidelity
- not applicable; distinct answer function
- All-required-slice quality gate
- FAIL
- Material-success coverage
- 1 /2; 1/2 (50%) — FAIL against 99%
Reject the proposed all-slice launch. The fixture cannot establish a real deployment approval.
Lexical-weight-2 fusion (post-hoc)
- Frozen judged recall@2
- 1 (100%)
- Dense exact-neighbor fidelity
- not applicable; distinct answer function
- All-required-slice quality gate
- pass
- Operating evidence
- unknown — post-hoc weights have no operating schedule
Continue testing. Quality was repaired after inspecting this fixture; freeze a new hypothesis and test fresh held-out needs.
Material success means the full response passes compatible-reference, snapshot eligibility, freshness and current disclosure requirements. Its denominator includes every scheduled attempt. Nominal OK and complete count can remain unchanged when a gate fails. Durations above are synthetic elapsed durations, not measured latency or p99.
Every selected request, including cap/failure and gate violations
| Plan / request / query | Slice | Scheduled / duration / completion ms | Outcome / IDs sent | Complete / material | S / reference | Snapshot eligibility | Freshness | Current disclosure | Bytes / requests / rounds |
|---|---|---|---|---|---|---|---|---|---|
| Lexical baseline r7 / q5 | critical | 30 / 100 / 130 | OK 201, 202 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 2 / 2 |
| Lexical baseline r8 / q6 | critical | 35 / 110 / 145 | OK 202, 203 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 2 / 2 |
| Dense candidates r7 / q5 | critical | 30 / 30 / 60 | OK 203, 202 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r8 / q6 | critical | 35 / 32 / 67 | OK 201, 203 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Equal-weight fusion r7 / q5 | critical | 30 / 60 / 90 | OK 201, 203 | complete / NOT material success | S39 / INCOMPATIBLE | unknown | FAIL | pass | 20,480 / 3 / 2 |
| Equal-weight fusion r8 / q6 | critical | 35 / 62 / 97 | OK 201, 202 | complete / material success | S40 / compatible | pass | pass | pass | 20,480 / 3 / 2 |
CAPPED means explicit RETRIEVAL_LIMIT with one inspected eligible ID. FAILURE means UNAVAILABLE at the 150 ms deadline with no delivered IDs. Neither is a complete empty answer; these queries have eligible documents. Weight-2 fusion has no manufactured request rows.
Selected frozen query answers
q5: Diagnose a live cutover that missed a replacement or a deletion; require incident evidence.
Eligible IDs 201, 202, 203; judged relevant eligible IDs 201, 202. These are the C40/S40 reference.
- Lexical baseline: selected 201, 202; judged recall 1 (100%); precision 1 (100%) .
- Dense candidates: selected 203, 202; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 201, 203; judged recall 1/2 (50%); precision 1/2 (50%) .
- Lexical-weight-2 fusion (post-hoc): selected 201, 202; judged recall 1 (100%); precision 1 (100%) .
q6: Diagnose a deleted result reappearing after cache-generation rerouting; require incident evidence.
Eligible IDs 201, 202, 203; judged relevant eligible IDs 202, 203. These are the C40/S40 reference.
- Lexical baseline: selected 202, 203; judged recall 1 (100%); precision 1 (100%) .
- Dense candidates: selected 201, 203; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 201, 202; judged recall 1/2 (50%); precision 1/2 (50%) .
- Lexical-weight-2 fusion (post-hoc): selected 202, 203; judged recall 1 (100%); precision 1 (100%) .
All-slice launch gate: FAIL. Changed r7 lacks a fresh compatible reference; current-response relevance is unavailable. A common-only display does not change the proposed exposure scope.
Change current policy — All scheduled attempts
Is a ranked and snapshot-eligible document safe to disclose after current policy denies it? Selected population: All scheduled attempts .
Overall proposal evidence: r7 still discloses document 201 despite the current deny decision. Disclosure fails; material-success coverage becomes 7/8 (87.5%).
Current disclosure is a separate hard gate. Ranking quality and snapshot eligibility cannot excuse this payload release.
All scheduled attempts: 6 distinct information needs; planned-workload share 1 (100%). 8 scheduled attempts; request-schedule share 1 (100%). These are separate denominators.
Frozen C40/S40 reference audit
Binary judged recall@2 measures relevant eligible hits over all judged-relevant eligible IDs. Precision@2 uses requested k=2; it equals recall in these six full quality rows. Dense fidelity measures agreement with the exact eligible dense top-2, separately from relevance. A selected-slice quality pass cannot approve an all-slice proposal.
Lexical baseline
- Frozen judged recall@2
- 53/80 (66.25%)
- Dense exact-neighbor fidelity
- not applicable; distinct answer function
- All-required-slice quality gate
- pass
- Material-success coverage
- 8 /8; 1 (100%) — pass against 99%
Continue testing. The fixture cannot establish a real deployment approval.
Dense candidates
- Frozen judged recall@2
- 1/2 (50%)
- Dense exact-neighbor fidelity
- 1 (100%)
- All-required-slice quality gate
- FAIL
- Material-success coverage
- 6 /8; 3/4 (75%) — FAIL against 99%
Reject the proposed all-slice launch. The fixture cannot establish a real deployment approval.
Equal-weight fusion
- Frozen judged recall@2
- 19/20 (95%)
- Dense exact-neighbor fidelity
- not applicable; distinct answer function
- All-required-slice quality gate
- FAIL
- Material-success coverage
- 7 /8; 7/8 (87.5%) — FAIL against 99%
Reject the proposed all-slice launch. The fixture cannot establish a real deployment approval.
Lexical-weight-2 fusion (post-hoc)
- Frozen judged recall@2
- 1 (100%)
- Dense exact-neighbor fidelity
- not applicable; distinct answer function
- All-required-slice quality gate
- pass
- Operating evidence
- unknown — post-hoc weights have no operating schedule
Continue testing. Quality was repaired after inspecting this fixture; freeze a new hypothesis and test fresh held-out needs.
Material success means the full response passes compatible-reference, snapshot eligibility, freshness and current disclosure requirements. Its denominator includes every scheduled attempt. Nominal OK and complete count can remain unchanged when a gate fails. Durations above are synthetic elapsed durations, not measured latency or p99.
Every selected request, including cap/failure and gate violations
| Plan / request / query | Slice | Scheduled / duration / completion ms | Outcome / IDs sent | Complete / material | S / reference | Snapshot eligibility | Freshness | Current disclosure | Bytes / requests / rounds |
|---|---|---|---|---|---|---|---|---|---|
| Lexical baseline r1 / q1 | common | 0 / 70 / 70 | OK 101, 103 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r2 / q2 | common | 5 / 75 / 80 | OK 101, 102 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r3 / q3 | common | 10 / 80 / 90 | OK 102, 101 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r4 / q4 | common | 15 / 85 / 100 | OK 103, 104 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r5 / q1 | common | 20 / 75 / 95 | OK 101, 103 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r6 / q2 | common | 25 / 80 / 105 | OK 101, 102 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r7 / q5 | critical | 30 / 100 / 130 | OK 201, 202 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 2 / 2 |
| Lexical baseline r8 / q6 | critical | 35 / 110 / 145 | OK 202, 203 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 2 / 2 |
| Dense candidates r1 / q1 | common | 0 / 20 / 20 | OK 102, 104 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r2 / q2 | common | 5 / 25 / 30 | OK 103, 104 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r3 / q3 | common | 10 / 100 / 110 | RETRIEVAL_LIMIT 104 | incomplete / NOT material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r4 / q4 | common | 15 / 150 / 165 | UNAVAILABLE none | incomplete / NOT material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r5 / q1 | common | 20 / 22 / 42 | OK 102, 104 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r6 / q2 | common | 25 / 26 / 51 | OK 103, 104 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r7 / q5 | critical | 30 / 30 / 60 | OK 203, 202 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r8 / q6 | critical | 35 / 32 / 67 | OK 201, 203 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Equal-weight fusion r1 / q1 | common | 0 / 40 / 40 | OK 101, 102 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r2 / q2 | common | 5 / 42 / 47 | OK 101, 103 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r3 / q3 | common | 10 / 44 / 54 | OK 102, 104 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r4 / q4 | common | 15 / 46 / 61 | OK 104, 103 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r5 / q1 | common | 20 / 42 / 62 | OK 101, 102 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r6 / q2 | common | 25 / 45 / 70 | OK 101, 103 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r7 / q5 | critical | 30 / 60 / 90 | OK 201, 203 | complete / NOT material success | S40 / compatible | pass | pass | FAIL ; denied 201 still sent | 20,480 / 3 / 2 |
| Equal-weight fusion r8 / q6 | critical | 35 / 62 / 97 | OK 201, 202 | complete / material success | S40 / compatible | pass | pass | pass | 20,480 / 3 / 2 |
CAPPED means explicit RETRIEVAL_LIMIT with one inspected eligible ID. FAILURE means UNAVAILABLE at the 150 ms deadline with no delivered IDs. Neither is a complete empty answer; these queries have eligible documents. Weight-2 fusion has no manufactured request rows.
Selected frozen query answers
q1: Understand safe tenant-scoped reads immediately after an acknowledged update.
Eligible IDs 101, 102, 103, 104; judged relevant eligible IDs 101, 102. These are the C40/S40 reference.
- Lexical baseline: selected 101, 103; judged recall 1/2 (50%); precision 1/2 (50%) .
- Dense candidates: selected 102, 104; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 101, 102; judged recall 1 (100%); precision 1 (100%) .
- Lexical-weight-2 fusion (post-hoc): selected 101, 102; judged recall 1 (100%); precision 1 (100%) .
q2: Route a tenant request safely when its warm cache node is replaced.
Eligible IDs 101, 102, 103, 104; judged relevant eligible IDs 101, 103. These are the C40/S40 reference.
- Lexical baseline: selected 101, 102; judged recall 1/2 (50%); precision 1/2 (50%) .
- Dense candidates: selected 103, 104; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 101, 103; judged recall 1 (100%); precision 1 (100%) .
- Lexical-weight-2 fusion (post-hoc): selected 101, 103; judged recall 1 (100%); precision 1 (100%) .
q3: Preserve fresh results when ingestion backlog competes with finite query budgets.
Eligible IDs 101, 102, 103, 104; judged relevant eligible IDs 102, 104. These are the C40/S40 reference.
- Lexical baseline: selected 102, 101; judged recall 1/2 (50%); precision 1/2 (50%) .
- Dense candidates: selected 104, 103; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 102, 104; judged recall 1 (100%); precision 1 (100%) .
- Lexical-weight-2 fusion (post-hoc): selected 102, 104; judged recall 1 (100%); precision 1 (100%) .
q4: Plan admission during a cold-cache replacement and traffic burst.
Eligible IDs 101, 102, 103, 104; judged relevant eligible IDs 103, 104. These are the C40/S40 reference.
- Lexical baseline: selected 103, 104; judged recall 1 (100%); precision 1 (100%) .
- Dense candidates: selected 104, 102; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 104, 103; judged recall 1 (100%); precision 1 (100%) .
- Lexical-weight-2 fusion (post-hoc): selected 103, 104; judged recall 1 (100%); precision 1 (100%) .
q5: Diagnose a live cutover that missed a replacement or a deletion; require incident evidence.
Eligible IDs 201, 202, 203; judged relevant eligible IDs 201, 202. These are the C40/S40 reference.
- Lexical baseline: selected 201, 202; judged recall 1 (100%); precision 1 (100%) .
- Dense candidates: selected 203, 202; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 201, 203; judged recall 1/2 (50%); precision 1/2 (50%) .
- Lexical-weight-2 fusion (post-hoc): selected 201, 202; judged recall 1 (100%); precision 1 (100%) .
q6: Diagnose a deleted result reappearing after cache-generation rerouting; require incident evidence.
Eligible IDs 201, 202, 203; judged relevant eligible IDs 202, 203. These are the C40/S40 reference.
- Lexical baseline: selected 202, 203; judged recall 1 (100%); precision 1 (100%) .
- Dense candidates: selected 201, 203; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 201, 202; judged recall 1/2 (50%); precision 1/2 (50%) .
- Lexical-weight-2 fusion (post-hoc): selected 202, 203; judged recall 1 (100%); precision 1 (100%) .
All-slice launch gate: FAIL. Changed r7 discloses a currently denied payload. A common-only display does not change the proposed exposure scope.
Change current policy — Common tenant-A queries
Is a ranked and snapshot-eligible document safe to disclose after current policy denies it? Selected population: Common tenant-A queries .
Overall proposal evidence: r7 still discloses document 201 despite the current deny decision. Disclosure fails; material-success coverage becomes 7/8 (87.5%).
Current disclosure is a separate hard gate. Ranking quality and snapshot eligibility cannot excuse this payload release.
Common tenant-A queries: 4 distinct information needs; planned-workload share 9/10 (90%). 6 scheduled attempts; request-schedule share 3/4 (75%). These are separate denominators.
Frozen C40/S40 reference audit
Binary judged recall@2 measures relevant eligible hits over all judged-relevant eligible IDs. Precision@2 uses requested k=2; it equals recall in these six full quality rows. Dense fidelity measures agreement with the exact eligible dense top-2, separately from relevance. A selected-slice quality pass cannot approve an all-slice proposal.
Lexical baseline
- Frozen judged recall@2
- 5/8 (62.5%)
- Dense exact-neighbor fidelity
- not applicable; distinct answer function
- All-required-slice quality gate
- pass
- Material-success coverage
- 6 /6; 1 (100%) — pass against 99%
Continue testing. The fixture cannot establish a real deployment approval.
Dense candidates
- Frozen judged recall@2
- 1/2 (50%)
- Dense exact-neighbor fidelity
- 1 (100%)
- All-required-slice quality gate
- FAIL
- Material-success coverage
- 4 /6; 2/3 (66.67%) — FAIL against 99%
Reject the proposed all-slice launch. The fixture cannot establish a real deployment approval.
Equal-weight fusion
- Frozen judged recall@2
- 1 (100%)
- Dense exact-neighbor fidelity
- not applicable; distinct answer function
- All-required-slice quality gate
- FAIL
- Material-success coverage
- 6 /6; 1 (100%) — pass against 99%
Reject the proposed all-slice launch. The fixture cannot establish a real deployment approval.
Lexical-weight-2 fusion (post-hoc)
- Frozen judged recall@2
- 1 (100%)
- Dense exact-neighbor fidelity
- not applicable; distinct answer function
- All-required-slice quality gate
- pass
- Operating evidence
- unknown — post-hoc weights have no operating schedule
Continue testing. Quality was repaired after inspecting this fixture; freeze a new hypothesis and test fresh held-out needs.
Material success means the full response passes compatible-reference, snapshot eligibility, freshness and current disclosure requirements. Its denominator includes every scheduled attempt. Nominal OK and complete count can remain unchanged when a gate fails. Durations above are synthetic elapsed durations, not measured latency or p99.
Every selected request, including cap/failure and gate violations
| Plan / request / query | Slice | Scheduled / duration / completion ms | Outcome / IDs sent | Complete / material | S / reference | Snapshot eligibility | Freshness | Current disclosure | Bytes / requests / rounds |
|---|---|---|---|---|---|---|---|---|---|
| Lexical baseline r1 / q1 | common | 0 / 70 / 70 | OK 101, 103 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r2 / q2 | common | 5 / 75 / 80 | OK 101, 102 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r3 / q3 | common | 10 / 80 / 90 | OK 102, 101 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r4 / q4 | common | 15 / 85 / 100 | OK 103, 104 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r5 / q1 | common | 20 / 75 / 95 | OK 101, 103 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Lexical baseline r6 / q2 | common | 25 / 80 / 105 | OK 101, 102 | complete / material success | S40 / compatible | pass | pass | pass | 8,192 / 2 / 2 |
| Dense candidates r1 / q1 | common | 0 / 20 / 20 | OK 102, 104 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r2 / q2 | common | 5 / 25 / 30 | OK 103, 104 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r3 / q3 | common | 10 / 100 / 110 | RETRIEVAL_LIMIT 104 | incomplete / NOT material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r4 / q4 | common | 15 / 150 / 165 | UNAVAILABLE none | incomplete / NOT material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r5 / q1 | common | 20 / 22 / 42 | OK 102, 104 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r6 / q2 | common | 25 / 26 / 51 | OK 103, 104 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Equal-weight fusion r1 / q1 | common | 0 / 40 / 40 | OK 101, 102 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r2 / q2 | common | 5 / 42 / 47 | OK 101, 103 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r3 / q3 | common | 10 / 44 / 54 | OK 102, 104 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r4 / q4 | common | 15 / 46 / 61 | OK 104, 103 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r5 / q1 | common | 20 / 42 / 62 | OK 101, 102 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
| Equal-weight fusion r6 / q2 | common | 25 / 45 / 70 | OK 101, 103 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 3 / 2 |
CAPPED means explicit RETRIEVAL_LIMIT with one inspected eligible ID. FAILURE means UNAVAILABLE at the 150 ms deadline with no delivered IDs. Neither is a complete empty answer; these queries have eligible documents. Weight-2 fusion has no manufactured request rows.
Selected frozen query answers
q1: Understand safe tenant-scoped reads immediately after an acknowledged update.
Eligible IDs 101, 102, 103, 104; judged relevant eligible IDs 101, 102. These are the C40/S40 reference.
- Lexical baseline: selected 101, 103; judged recall 1/2 (50%); precision 1/2 (50%) .
- Dense candidates: selected 102, 104; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 101, 102; judged recall 1 (100%); precision 1 (100%) .
- Lexical-weight-2 fusion (post-hoc): selected 101, 102; judged recall 1 (100%); precision 1 (100%) .
q2: Route a tenant request safely when its warm cache node is replaced.
Eligible IDs 101, 102, 103, 104; judged relevant eligible IDs 101, 103. These are the C40/S40 reference.
- Lexical baseline: selected 101, 102; judged recall 1/2 (50%); precision 1/2 (50%) .
- Dense candidates: selected 103, 104; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 101, 103; judged recall 1 (100%); precision 1 (100%) .
- Lexical-weight-2 fusion (post-hoc): selected 101, 103; judged recall 1 (100%); precision 1 (100%) .
q3: Preserve fresh results when ingestion backlog competes with finite query budgets.
Eligible IDs 101, 102, 103, 104; judged relevant eligible IDs 102, 104. These are the C40/S40 reference.
- Lexical baseline: selected 102, 101; judged recall 1/2 (50%); precision 1/2 (50%) .
- Dense candidates: selected 104, 103; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 102, 104; judged recall 1 (100%); precision 1 (100%) .
- Lexical-weight-2 fusion (post-hoc): selected 102, 104; judged recall 1 (100%); precision 1 (100%) .
q4: Plan admission during a cold-cache replacement and traffic burst.
Eligible IDs 101, 102, 103, 104; judged relevant eligible IDs 103, 104. These are the C40/S40 reference.
- Lexical baseline: selected 103, 104; judged recall 1 (100%); precision 1 (100%) .
- Dense candidates: selected 104, 102; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 104, 103; judged recall 1 (100%); precision 1 (100%) .
- Lexical-weight-2 fusion (post-hoc): selected 103, 104; judged recall 1 (100%); precision 1 (100%) .
All-slice launch gate: FAIL. Changed r7 discloses a currently denied payload. A common-only display does not change the proposed exposure scope.
Change current policy — Required tenant-B incident queries
Is a ranked and snapshot-eligible document safe to disclose after current policy denies it? Selected population: Required tenant-B incident queries .
Overall proposal evidence: r7 still discloses document 201 despite the current deny decision. Disclosure fails; material-success coverage becomes 7/8 (87.5%).
Current disclosure is a separate hard gate. Ranking quality and snapshot eligibility cannot excuse this payload release.
Required tenant-B incident queries: 2 distinct information needs; planned-workload share 1/10 (10%). 2 scheduled attempts; request-schedule share 1/4 (25%). These are separate denominators.
Frozen C40/S40 reference audit
Binary judged recall@2 measures relevant eligible hits over all judged-relevant eligible IDs. Precision@2 uses requested k=2; it equals recall in these six full quality rows. Dense fidelity measures agreement with the exact eligible dense top-2, separately from relevance. A selected-slice quality pass cannot approve an all-slice proposal.
Lexical baseline
- Frozen judged recall@2
- 1 (100%)
- Dense exact-neighbor fidelity
- not applicable; distinct answer function
- All-required-slice quality gate
- pass
- Material-success coverage
- 2 /2; 1 (100%) — pass against 99%
Continue testing. The fixture cannot establish a real deployment approval.
Dense candidates
- Frozen judged recall@2
- 1/2 (50%)
- Dense exact-neighbor fidelity
- 1 (100%)
- All-required-slice quality gate
- FAIL
- Material-success coverage
- 2 /2; 1 (100%) — pass against 99%
Reject the proposed all-slice launch. The fixture cannot establish a real deployment approval.
Equal-weight fusion
- Frozen judged recall@2
- 1/2 (50%)
- Dense exact-neighbor fidelity
- not applicable; distinct answer function
- All-required-slice quality gate
- FAIL
- Material-success coverage
- 1 /2; 1/2 (50%) — FAIL against 99%
Reject the proposed all-slice launch. The fixture cannot establish a real deployment approval.
Lexical-weight-2 fusion (post-hoc)
- Frozen judged recall@2
- 1 (100%)
- Dense exact-neighbor fidelity
- not applicable; distinct answer function
- All-required-slice quality gate
- pass
- Operating evidence
- unknown — post-hoc weights have no operating schedule
Continue testing. Quality was repaired after inspecting this fixture; freeze a new hypothesis and test fresh held-out needs.
Material success means the full response passes compatible-reference, snapshot eligibility, freshness and current disclosure requirements. Its denominator includes every scheduled attempt. Nominal OK and complete count can remain unchanged when a gate fails. Durations above are synthetic elapsed durations, not measured latency or p99.
Every selected request, including cap/failure and gate violations
| Plan / request / query | Slice | Scheduled / duration / completion ms | Outcome / IDs sent | Complete / material | S / reference | Snapshot eligibility | Freshness | Current disclosure | Bytes / requests / rounds |
|---|---|---|---|---|---|---|---|---|---|
| Lexical baseline r7 / q5 | critical | 30 / 100 / 130 | OK 201, 202 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 2 / 2 |
| Lexical baseline r8 / q6 | critical | 35 / 110 / 145 | OK 202, 203 | complete / material success | S40 / compatible | pass | pass | pass | 12,288 / 2 / 2 |
| Dense candidates r7 / q5 | critical | 30 / 30 / 60 | OK 203, 202 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Dense candidates r8 / q6 | critical | 35 / 32 / 67 | OK 201, 203 | complete / material success | S40 / compatible | pass | pass | pass | 4,096 / 1 / 1 |
| Equal-weight fusion r7 / q5 | critical | 30 / 60 / 90 | OK 201, 203 | complete / NOT material success | S40 / compatible | pass | pass | FAIL ; denied 201 still sent | 20,480 / 3 / 2 |
| Equal-weight fusion r8 / q6 | critical | 35 / 62 / 97 | OK 201, 202 | complete / material success | S40 / compatible | pass | pass | pass | 20,480 / 3 / 2 |
CAPPED means explicit RETRIEVAL_LIMIT with one inspected eligible ID. FAILURE means UNAVAILABLE at the 150 ms deadline with no delivered IDs. Neither is a complete empty answer; these queries have eligible documents. Weight-2 fusion has no manufactured request rows.
Selected frozen query answers
q5: Diagnose a live cutover that missed a replacement or a deletion; require incident evidence.
Eligible IDs 201, 202, 203; judged relevant eligible IDs 201, 202. These are the C40/S40 reference.
- Lexical baseline: selected 201, 202; judged recall 1 (100%); precision 1 (100%) .
- Dense candidates: selected 203, 202; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 201, 203; judged recall 1/2 (50%); precision 1/2 (50%) .
- Lexical-weight-2 fusion (post-hoc): selected 201, 202; judged recall 1 (100%); precision 1 (100%) .
q6: Diagnose a deleted result reappearing after cache-generation rerouting; require incident evidence.
Eligible IDs 201, 202, 203; judged relevant eligible IDs 202, 203. These are the C40/S40 reference.
- Lexical baseline: selected 202, 203; judged recall 1 (100%); precision 1 (100%) .
- Dense candidates: selected 201, 203; judged recall 1/2 (50%); precision 1/2 (50%) ; dense exact fidelity 1 (100%) .
- Equal-weight fusion: selected 201, 202; judged recall 1/2 (50%); precision 1/2 (50%) .
- Lexical-weight-2 fusion (post-hoc): selected 202, 203; judged recall 1 (100%); precision 1 (100%) .
All-slice launch gate: FAIL. Changed r7 discloses a currently denied payload. A common-only display does not change the proposed exposure scope.
Exact fixture inputs, full rankings and relevance reference
Every row is live version 1 in C40/S40. Document vectors are eight-dimensional unit basis vectors. Query vectors are explicit. Lexical scores and binary judgments are authored; the numeric distance calculation follows squared Euclidean distance and ID tie order. No actual tokenizer, embedding model or ANN traversal ran.
| ID / version | Tenant / kind | Title | Body vector |
|---|---|---|---|
| 101 v1 | A / guide | Trusted identity and safe tenant access | (1, 0, 0, 0, 0, 0, 0, 0) |
| 102 v1 | A / guide | Index lag, write visibility and ingest backlog | (0, 1, 0, 0, 0, 0, 0, 0) |
| 103 v1 | A / guide | Cache node replacement and compatible routing | (0, 0, 1, 0, 0, 0, 0, 0) |
| 104 v1 | A / guide | Admission, contention and finite request budgets | (0, 0, 0, 1, 0, 0, 0, 0) |
| 201 v1 | B / incident | Missing concurrent replacement during index cutover | (0, 0, 0, 0, 1, 0, 0, 0) |
| 202 v1 | B / incident | Delete barrier and stale live-version resurrection | (0, 0, 0, 0, 0, 1, 0, 0) |
| 203 v1 | B / incident | Cache generation routing after a replacement | (0, 0, 0, 0, 0, 0, 1, 0) |
| 204 v1 | B / guide | General search-system orientation | (0, 0, 0, 0, 0, 0, 0, 1) |
q1 — common; information need, vector, scores and all eight judgments
Understand safe tenant-scoped reads immediately after an acknowledged update. Query text: “tenant fresh acknowledged update”. Predicate: tenant A AND kind unrestricted. Planned query weight 9/40. Query vector (2, 4, 1, 3, 0, 0, 0, 0).
Eligible IDs 101, 102, 103, 104; saved dense candidate window 102, 104, 101; exact dense top-2 102, 104.
| Document ID | Eligible at S40 | Authored relevance | Lexical score |
|---|---|---|---|
| 101 | yes | 1 | 10 |
| 102 | yes | 1 | 6 |
| 103 | yes | 0 | 8 |
| 104 | yes | 0 | 4 |
| 201 | no | 0 | not eligible |
| 202 | no | 0 | not eligible |
| 203 | no | 0 | not eligible |
| 204 | no | 0 | not eligible |
| Position | Lexical ID / stipulated score | Exact dense ID / squared distance |
|---|---|---|
| 1 | 101 / 10 | 102 / 23 |
| 2 | 103 / 8 | 104 / 25 |
| 3 | 102 / 6 | 101 / 27 |
| 4 | 104 / 4 | 103 / 29 |
| ID | Lexical rank / contribution | Dense rank / contribution | Exact score |
|---|---|---|---|
| 101 | 1 / 1/61 | 3 / 1/63 | 124/3843 |
| 102 | 3 / 1/63 | 1 / 1/61 | 124/3843 |
| 103 | 2 / 1/62 | absent / 0 | 1/62 |
| 104 | absent / 0 | 2 / 1/62 | 1/62 |
| ID | Lexical rank / contribution | Dense rank / contribution | Exact score |
|---|---|---|---|
| 101 | 1 / 2/61 | 3 / 1/63 | 187/3843 |
| 102 | 3 / 2/63 | 1 / 1/61 | 185/3843 |
| 103 | 2 / 1/31 | absent / 0 | 1/31 |
| 104 | absent / 0 | 2 / 1/62 | 1/62 |
q2 — common; information need, vector, scores and all eight judgments
Route a tenant request safely when its warm cache node is replaced. Query text: “tenant cache node routing”. Predicate: tenant A AND kind unrestricted. Planned query weight 9/40. Query vector (2, 1, 4, 3, 0, 0, 0, 0).
Eligible IDs 101, 102, 103, 104; saved dense candidate window 103, 104, 101; exact dense top-2 103, 104.
| Document ID | Eligible at S40 | Authored relevance | Lexical score |
|---|---|---|---|
| 101 | yes | 1 | 10 |
| 102 | yes | 0 | 8 |
| 103 | yes | 1 | 6 |
| 104 | yes | 0 | 4 |
| 201 | no | 0 | not eligible |
| 202 | no | 0 | not eligible |
| 203 | no | 0 | not eligible |
| 204 | no | 0 | not eligible |
| Position | Lexical ID / stipulated score | Exact dense ID / squared distance |
|---|---|---|
| 1 | 101 / 10 | 103 / 23 |
| 2 | 102 / 8 | 104 / 25 |
| 3 | 103 / 6 | 101 / 27 |
| 4 | 104 / 4 | 102 / 29 |
| ID | Lexical rank / contribution | Dense rank / contribution | Exact score |
|---|---|---|---|
| 101 | 1 / 1/61 | 3 / 1/63 | 124/3843 |
| 103 | 3 / 1/63 | 1 / 1/61 | 124/3843 |
| 102 | 2 / 1/62 | absent / 0 | 1/62 |
| 104 | absent / 0 | 2 / 1/62 | 1/62 |
| ID | Lexical rank / contribution | Dense rank / contribution | Exact score |
|---|---|---|---|
| 101 | 1 / 2/61 | 3 / 1/63 | 187/3843 |
| 103 | 3 / 2/63 | 1 / 1/61 | 185/3843 |
| 102 | 2 / 1/31 | absent / 0 | 1/31 |
| 104 | absent / 0 | 2 / 1/62 | 1/62 |
q3 — common; information need, vector, scores and all eight judgments
Preserve fresh results when ingestion backlog competes with finite query budgets. Query text: “index lag admission budget”. Predicate: tenant A AND kind unrestricted. Planned query weight 9/40. Query vector (1, 2, 3, 4, 0, 0, 0, 0).
Eligible IDs 101, 102, 103, 104; saved dense candidate window 104, 103, 102; exact dense top-2 104, 103.
| Document ID | Eligible at S40 | Authored relevance | Lexical score |
|---|---|---|---|
| 101 | yes | 0 | 8 |
| 102 | yes | 1 | 10 |
| 103 | yes | 0 | 4 |
| 104 | yes | 1 | 6 |
| 201 | no | 0 | not eligible |
| 202 | no | 0 | not eligible |
| 203 | no | 0 | not eligible |
| 204 | no | 0 | not eligible |
| Position | Lexical ID / stipulated score | Exact dense ID / squared distance |
|---|---|---|
| 1 | 102 / 10 | 104 / 23 |
| 2 | 101 / 8 | 103 / 25 |
| 3 | 104 / 6 | 102 / 27 |
| 4 | 103 / 4 | 101 / 29 |
| ID | Lexical rank / contribution | Dense rank / contribution | Exact score |
|---|---|---|---|
| 102 | 1 / 1/61 | 3 / 1/63 | 124/3843 |
| 104 | 3 / 1/63 | 1 / 1/61 | 124/3843 |
| 101 | 2 / 1/62 | absent / 0 | 1/62 |
| 103 | absent / 0 | 2 / 1/62 | 1/62 |
| ID | Lexical rank / contribution | Dense rank / contribution | Exact score |
|---|---|---|---|
| 102 | 1 / 2/61 | 3 / 1/63 | 187/3843 |
| 104 | 3 / 2/63 | 1 / 1/61 | 185/3843 |
| 101 | 2 / 1/31 | absent / 0 | 1/31 |
| 103 | absent / 0 | 2 / 1/62 | 1/62 |
q4 — common; information need, vector, scores and all eight judgments
Plan admission during a cold-cache replacement and traffic burst. Query text: “cold cache traffic burst”. Predicate: tenant A AND kind unrestricted. Planned query weight 9/40. Query vector (1, 3, 2, 4, 0, 0, 0, 0).
Eligible IDs 101, 102, 103, 104; saved dense candidate window 104, 102, 103; exact dense top-2 104, 102.
| Document ID | Eligible at S40 | Authored relevance | Lexical score |
|---|---|---|---|
| 101 | yes | 0 | 4 |
| 102 | yes | 0 | 6 |
| 103 | yes | 1 | 10 |
| 104 | yes | 1 | 8 |
| 201 | no | 0 | not eligible |
| 202 | no | 0 | not eligible |
| 203 | no | 0 | not eligible |
| 204 | no | 0 | not eligible |
| Position | Lexical ID / stipulated score | Exact dense ID / squared distance |
|---|---|---|
| 1 | 103 / 10 | 104 / 23 |
| 2 | 104 / 8 | 102 / 25 |
| 3 | 102 / 6 | 103 / 27 |
| 4 | 101 / 4 | 101 / 29 |
| ID | Lexical rank / contribution | Dense rank / contribution | Exact score |
|---|---|---|---|
| 104 | 2 / 1/62 | 1 / 1/61 | 123/3782 |
| 103 | 1 / 1/61 | 3 / 1/63 | 124/3843 |
| 102 | 3 / 1/63 | 2 / 1/62 | 125/3906 |
| ID | Lexical rank / contribution | Dense rank / contribution | Exact score |
|---|---|---|---|
| 103 | 1 / 2/61 | 3 / 1/63 | 187/3843 |
| 104 | 2 / 1/31 | 1 / 1/61 | 92/1891 |
| 102 | 3 / 2/63 | 2 / 1/62 | 187/3906 |
q5 — critical; information need, vector, scores and all eight judgments
Diagnose a live cutover that missed a replacement or a deletion; require incident evidence. Query text: “cutover missing update delete incident”. Predicate: tenant B AND kind incident. Planned query weight 1/20. Query vector (0, 0, 0, 0, 1, 2, 3, 0).
Eligible IDs 201, 202, 203; saved dense candidate window 203, 202, 201; exact dense top-2 203, 202.
| Document ID | Eligible at S40 | Authored relevance | Lexical score |
|---|---|---|---|
| 101 | no | 0 | not eligible |
| 102 | no | 0 | not eligible |
| 103 | no | 0 | not eligible |
| 104 | no | 0 | not eligible |
| 201 | yes | 1 | 10 |
| 202 | yes | 1 | 8 |
| 203 | yes | 0 | 6 |
| 204 | no | 0 | not eligible |
| Position | Lexical ID / stipulated score | Exact dense ID / squared distance |
|---|---|---|
| 1 | 201 / 10 | 203 / 9 |
| 2 | 202 / 8 | 202 / 11 |
| 3 | 203 / 6 | 201 / 13 |
| ID | Lexical rank / contribution | Dense rank / contribution | Exact score |
|---|---|---|---|
| 201 | 1 / 1/61 | 3 / 1/63 | 124/3843 |
| 203 | 3 / 1/63 | 1 / 1/61 | 124/3843 |
| 202 | 2 / 1/62 | 2 / 1/62 | 1/31 |
| ID | Lexical rank / contribution | Dense rank / contribution | Exact score |
|---|---|---|---|
| 201 | 1 / 2/61 | 3 / 1/63 | 187/3843 |
| 202 | 2 / 1/31 | 2 / 1/62 | 3/62 |
| 203 | 3 / 2/63 | 1 / 1/61 | 185/3843 |
q6 — critical; information need, vector, scores and all eight judgments
Diagnose a deleted result reappearing after cache-generation rerouting; require incident evidence. Query text: “deleted result cache generation incident”. Predicate: tenant B AND kind incident. Planned query weight 1/20. Query vector (0, 0, 0, 0, 3, 1, 2, 0).
Eligible IDs 201, 202, 203; saved dense candidate window 201, 203, 202; exact dense top-2 201, 203.
| Document ID | Eligible at S40 | Authored relevance | Lexical score |
|---|---|---|---|
| 101 | no | 0 | not eligible |
| 102 | no | 0 | not eligible |
| 103 | no | 0 | not eligible |
| 104 | no | 0 | not eligible |
| 201 | yes | 0 | 6 |
| 202 | yes | 1 | 10 |
| 203 | yes | 1 | 8 |
| 204 | no | 0 | not eligible |
| Position | Lexical ID / stipulated score | Exact dense ID / squared distance |
|---|---|---|
| 1 | 202 / 10 | 201 / 9 |
| 2 | 203 / 8 | 203 / 11 |
| 3 | 201 / 6 | 202 / 13 |
| ID | Lexical rank / contribution | Dense rank / contribution | Exact score |
|---|---|---|---|
| 201 | 3 / 1/63 | 1 / 1/61 | 124/3843 |
| 202 | 1 / 1/61 | 3 / 1/63 | 124/3843 |
| 203 | 2 / 1/62 | 2 / 1/62 | 1/31 |
| ID | Lexical rank / contribution | Dense rank / contribution | Exact score |
|---|---|---|---|
| 202 | 1 / 2/61 | 3 / 1/63 | 187/3843 |
| 203 | 2 / 1/31 | 2 / 1/62 | 3/62 |
| 201 | 3 / 2/63 | 1 / 1/61 | 185/3843 |
Fusion takes the union of the first three lexical and dense-window IDs. Input positions are unique, one-based after their declared total order. Missing-list contribution is zero. Constant 60; exact-rational score descending, then document ID ascending. Equal weights 1/1 were frozen; lexical weight 2 is post-hoc. For zero eligible IDs return empty and mark precision/fidelity N/A; zero judged-relevant eligibility makes recall N/A. Neither boundary population occurs in these six queries. Oracle audit work is separate from serving-plan costs.
Read cost at the same scope as the decision
The fixture also stipulates remote-read accounting:
| Plan | Common bytes per attempt | Incident bytes per attempt | Remote requests | Dependency rounds |
|---|---|---|---|---|
| Lexical | 8,192 | 12,288 | 2 | 2 |
| Dense | 4,096 | 4,096 | 1 | 1 |
| Equal fusion | 12,288 | 20,480 | 3 | 2 |
Capped and failed dense work retains its 4,096-byte cost. The exact-oracle audit is separate from normal serving-plan costs. CPU, cache behavior, storage amplification and actual stage timing remain unknown.
The soft goal is at most 16,384 remote bytes per successful request. Equal fusion’s eight-attempt mean is 14,336 bytes, yet each incident success uses 20,480 bytes. The aggregate fits while a required slice exceeds the preference. We may renegotiate a soft cost target explicitly. We cannot exchange it for permission to disclose a denied payload or silently waive a required relevance condition.
Guided change: double lexical weight
Change only lexical weight from 1 to 2; keep the same windows, judgments and references. Predict whether q5 now includes 202. Then decide whether its result permits a broad launch.
Check the repaired fixture and the remaining evidence gap
q5 scores are 187/3843 for 201, 3/62 for 202 and 185/3843 for 203. The order becomes 201, 202, 203, so its top-2 are both relevant. q6 becomes 202, 203, 201. All six fixture queries now have judged recall 1; the planned-workload score is1.
This choice was made after inspecting the fixture. It is a post-hoc hypothesis, not a new unbiased effectiveness estimate. The development/test passage explains why tuning weights to maximize a test collection overstates expected performance. Freeze weight 2 on development evidence, then test fresh held-out information needs. The variant has no request schedule: its operating coverage, durations, costs and gates are unknown. Copying equal-fusion timing into its columns would manufacture evidence.
Continue testing. Even a scoped common-only proposal needs separately declared exposure, critical-user protection and evidence for that scope.
Change freshness, then change disclosure
First reset to equal fusion and change only r7’s observed snapshot from S40 to S39. The request requires commit 40 and still nominally returns 201, 203. Then reset to S40 and change only the trusted current-policy decision: deny 201, while the response still sends it. Predict material-success coverage and the status of the quality comparison in each case.
Check two independent hard gates
At S39, freshness fails because 39 is below required commit 40. The response also lacks a compatible reference: no S39 corpus/live-version/judgment set is provided. Changed-response judged relevance is unavailable, and snapshot eligibility at S39 is unknown. Do not show the frozen 95% beside S39 as its current-response score. Any retained C40/S40 quality table is explicitly a frozen offline reference audit.
At S40 with current deny 201, snapshot eligibility still passes and the S40 relevance reference is compatible, but current disclosure fails: 201 is still released. Relevance and predicate membership do not authorize that payload. The fixture stipulates the trusted policy decision point; it does not establish a real permission-race/fencing protocol.
Each changed schedule has eight nominally complete OK responses but only 7/8 material successes. Common remains 6/6; the incident slice becomes 1/2. Reject the proposed all-slice launch under either hard-gate violation. Viewing the unaffected common slice does not change the proposal’s scope.
Replace authored evidence with a distinguishing experiment
A real comparison begins by freezing the answer and the test: document/live versions, snapshot boundary, tenant/filter/query slices, information needs, independent relevance judgments, lexical/model/index configurations, ANN budgets, fusion windows/weights/ties, k/projection, runtime versions and environment. Save raw rankings and exact dense references as well as summary scripts. Separate development tuning from untouched held-out needs.
Use domain experts to formulate needs appropriate to the corpus and predicted usage. Preserve assessor disagreements and evaluate sensitivity to them. The assessing-relevance passage distinguishes exhaustive tiny collections from larger pooled judgments. At scale, record unjudged pairs and pool bias; an unjudged item is not automatically irrelevant. If the full relevant set is unknown, do not claim exhaustive true relevance recall.
Compare methods on paired identical needs and compatible snapshots. Report query-level differences, required-slice sample counts and uncertainty for paired differences, with a documented resampling unit/seed and treatment of correlated repeated queries or tenants. Predeclare non-regression margins/confidence rules and multiple-comparison treatment. Determine adequate sampling from the decision precision, variance and rare-slice requirements; a historical query-count rule of thumb is insufficient. No confidence interval follows from the authored labels here.
For operational evidence, schedule arrivals independently of completion under declared steady, burst, cold/replacement and update/delete/maintenance load. Record scheduled arrival, actual dispatch, admission/queue/service-stage timestamps and client completion. Track generator delay/saturation so it cannot appear to improve capacity by postponing offered work. Keep scheduled-to-complete, dispatch-to-complete, queue wait, service and network timing distinct.
Count every scheduled logical request and preserve attempt/retry relationships. Report complete material successes, caps, failures, admission rejections, cancellations and timeouts separately, including their costs. Report successful-only distributions alongside coverage and all-outcome time to terminal response. A real tail estimate needs adequate measured samples, an estimator, resolution and uncertainty; tail adequacy can remain unknown. Separate bytes, requests, dependency rounds, local CPU, memory/cache behavior, storage amplification and maintenance lag.
Use acknowledged replacements/deletes to audit the required visibility boundary, chosen snapshot and winning live versions. Audit current disclosure at the declared trusted decision point. A mixed/stale reference is a gate failure or an evidence gap, not an omitted quality row. Request-stage evidence should distinguish a queue/admission problem, candidate work, serial remote discovery, transfer/projection or maintenance deficit; the next controlled input change should falsify one explanation.
Before exposure, name the hard gates, negotiable preferences, owner and stop/rollback conditions. A shadow comparison can collect compatible answer evidence; it does not alone establish serving capacity or current disclosure behavior. A canary needs limited exposure and audited actual outcomes. Preserve and verify the prior plan/state needed for rollback before widening. The next lesson’s live-index decision addresses derived-generation coverage and retained state; a scoring/model change still needs its own effectiveness evidence.
Write the decision that could change
Keep one page with the proposal’s exposure scope, required needs/slices and goals; the alternatives and named metric denominators; the decision and decisive evidence; consequential unknowns; guarded exposure/rollback with an owner; and the next experiment whose outcome could reverse the choice.
For this case, decide among rejecting equal fusion, continuing with a fresh weight 2 hypothesis, or approving the broad launch. State what would have to become true before approval is justified.
Compare a defensible decision record
Reject the proposed all-slice equal-weight launch: the required incident slice falls from 1 to 1/2 despite aggregate 66.25%→95%. Retain lexical as the comparator. Dense also fails required-slice relevance and its synthetic material-success coverage goal, despite perfect dense fidelity and smaller successful-only durations.
Continue testing the post-hoc weight 2 hypothesis on fresh frozen held-out needs and an arrival-driven operational/safety experiment. Require adequate rare-slice and assessor evidence, compatible exact/reference audits, material-success coverage and freshness/current-disclosure checks. Evaluate the soft cost preference explicitly; do not hide the incident bytes in the mean.
Approve only a declared exposure scope that the collected real evidence supports, with uncertainty rules and guarded rollback. This fixture establishes no such approval. The distinguishing next experiment pairs the frozen variant against lexical on untouched incident/common needs while measuring all scheduled outcomes under cold cache and concurrent update/delete load. That experiment could expose an unseen relevance loss or operational gate failure and change the proposal before users receive it.