Every place where arms were compared or a candidate was measured, and a decision followed. Several ended by deleting what had been built.
The pattern that recurred
Almost every result that looked strong on a small sample shrank on a larger one. The project stopped trusting a candidate until it had been run on a population large enough to disappoint it.
Agreement with gold, small sample against larger sample
Small sampleLarger sample
Three of the arms tried for breaking ties between concepts that share a name. The arm that shipped, a cross-encoder reading sentence and candidate together, held its level at scale.
Judges
Every hard negative is audited by a model judge before it is published, so the judge's errors decide what a retriever learns.
7 Sep 2026judge-reasoning-trial
Reasoning on or off
Question
Was a judge running with reasoning disabled discarding real planted flaws as "safe"?
Setup
DeepSeek V4 Flash, same prompt, reasoning enabled with a 16,384-token budget, on 104 negatives: 52 it had discarded and 52 it had approved.
Result
Of the 52 discards, 50 carried the planted flaw. With reasoning on, 39 were approved with the flaw named and 13 stayed "safe"; 49 of 52 approvals held. Median latency rose from about 3 s to 21.3 s, at about $0.0009 per call.
Decision
Reasoning re-enabled for the judge, with its own output budget. Before the trial, 22% of 4,142 negatives were being discarded, 26% of contraindication traps.
10 Sep 2026judge-arms-mercury-ling
A faster judge with the same error rate
Question
Can a faster model replace the judge on contraindication traps?
Ling and the incumbent agreed on only one of their five discards each. Ling called flaws "consistent with current standard of care" that the incumbent named exactly.
Decision
Judge not swapped. An equal error rate is not an equal judge when the errors fall on different items.
17 Aug 2026judge-batching
One negative per request, or four
Question
Can one judge request carry four negatives without corrupting verdicts?
Setup
Replay of about 234 already-judged negatives, a serial control against a batched arm, same seed.
Result
Agreement with the ledger was 0.9657 serial and 0.9060 batched. Dangerous flips (discarded before, approved now) were 3 serial and 5 batched. The control mattered: at temperature 0 the judge disagrees with its own earlier verdict 3.4% of the time.
Decision
Batches of four adopted, cutting worst-case requests per unit from 26 to 11. A reply is attributed by the index the model states, never by position; an unattributable reply is re-asked one negative at a time.
6 Sep 2026judge-provider-trial
The same model through different providers
Question
Can the judge move to a cheaper route without losing quality?
Setup
12 real four-negative batches, production verdicts as reference, several builds of one model pinned to individual providers and unpinned.
Result
Every variant agreed with production 92–100%. The differences were operational: malformed JSON and rate limits, 9 of 12 valid unpinned against 12 of 12 for the best pin. The whole corpus costs roughly $20–40 of judge either way.
Decision
Judge choice is a price and reliability question here; the generator dominates cost.
11 Sep 2026judge-deepseek-flash-lr
The best judge, then it vanished
Question
Does spreading one model across providers explain differences in judge quality?
Result
A single-provider build called 0.10–0.15 of flawed negatives safe; the multi-provider build of the same model, 0.25–0.30. About five hours into a build, the single-provider id disappeared from the router's catalogue and judge circuits opened across shards. Because the judge fails closed, the damage was quarantine, not approval: 583 quarantined, 2,113 accepted, 1,097 rejected.
Decision
The judge is now a pool: the single-provider id first, the multi-provider id as failover. A multi-day run is never pinned to one route with no fallback.
6 Sep 2026judge-retry-budget
Errors hidden inside successful responses
Question
Why were negatives going unverified?
Result
34 negatives in five minutes: 28 were "operation aborted" error envelopes inside HTTP 200, read as a permanent failure because they carried no status; 6 were rate limits.
Decision
A separate retry budget for the judge, and an envelope that says the request was interrupted is retried. A genuine refusal still is not.
Generators
5 Sep 2026ling-generation-trial
A cheaper generator against two incumbents
Question
Can a cheaper, faster model write the training data?
Setup
One clinical handbook, 106 units, one judge for every arm.
Result
Generator
Accepted
Negatives approved
Off-topic negatives
Wall time
Ling 3.0 Flash
61%
64%
25%
64 min
GPT-5.6 Luna
77%
79%
11%
194 min
MiniMax M3
74%
81%
10%
116 min
The fast model's approved negatives had median similarity 0.87 to their positive, and 31% opened with the positive's own text.
Decision
Not a replacement, at most a last fallback. Any wider use would need the audit to refuse a negative that opens by restating its positive.
Vision readers
30 Aug 2026candidate-gateway-probe
A chart with a known answer
Question
Can a candidate model read a figure, or only return a well-formed answer?
Setup
A drawn logarithmic axis from 102 to 107, sent to candidates on five gateways by the project's probe script.
Result
One candidate read the scale correctly apart from one sign. Two frontier models reached through one gateway answered confidently and wrongly, "0 to 6000, linear" and "0 to 100, linear", having passed every text contract. Two others never received the image intact. Every reply came back HTTP 200.
Decision
No replacement. A reader that invents axis values is worse than one that refuses, and a vision candidate is judged on what it says about a known chart.
25 Aug 2026medical-mcq-vision-recovery
Answer keys from page images
Question
Can question forms printed as images be recovered without a second reader on every item?
Setup
13 image-only PDFs, 1,364 pages, read by MiniMax M3; a sampled audit by a different model on a different provider.
Result
761 of 854 keyed items recovered (89.1%); 604 checked against the author's printed list of wrong options with no disagreement; the audit agreed on 79 of 80. Keyed on question number alone, one four-section form recovered 0 of 50; keyed by section and number, 199.
Decision
One reader plus a sampled audit, because a second reader on every item would take 2.5 hours against 15 minutes. The audit keeps agreement a measured rate. A strict JSON-schema mode hung for 16 minutes on one page that answered in 9.5 seconds when JSON was requested in the prompt.
5 Sep 2026vision-ladder-pixel-path
The model list that became a model name
Result
The first run to reach a raster figure refused every one: 105 assets. The pixel reader passed the whole list of fallback models as a single model name and got HTTP 404. Figures with native tables took another path and were fine.
Decision
One owner for model routing, read by both readers.
Retrieval and reranking
31 Aug 2026retrieval-comparison
Does removing bibliography pages change what is retrieved?
Setup
Two collections, 20,437 cleaned and 22,003 uncleaned passages, asked the same 50 clinical questions with one embedding per question; top-10 slots counted by hand, 100 per arm.
Result
Reference-list text fell from 16 to 6 of 100 slots (Fisher p = 0.040). Distinct sources per top-10 rose from 4.70 to 5.16. Mean cosine was higher for the uncleaned collection, because reference lists are dense in exactly the terms a clinical query carries.
Decision
The hand count is the number. Two automated classifiers were rejected: one reached precision 1.00 at recall 0.50, and loosening it to recall 0.88 dropped precision to 0.64.
31 Aug 2026reranker-cannot-detect-bibliography
Can a reranker spot a retrieved bibliography?
Setup
22 hand-labelled bibliography passages against 178 clinical ones, each scored against its own query.
Result
Separation was +0.016 for Voyage rerank 2.5 lite, +0.076 for Qwen3 reranker 8B and −0.047 for Cohere rerank 3.5, which preferred the bibliography. A three-document probe had shown 8× separation, but its citation was off-topic, which a real retrieved hit never is.
Decision
The reranker arm became a safety check; the hand count stayed the headline.
9 Sep 2026cross-document-grounding
A score floor for "supported"
Setup
16 searched claims, 24 reranked candidates each, over 63,785 passages.
Result
A case finding scored 0.7526; the safety rule "alpha-blockade must precede any beta-blocker" scored 0.0061. A floor of 0.005 kept 14 of 16 claims and 93 passages; 0.05 kept 9 and 19. The ranking was right: the rule's top three passages were genuine.
Decision
Absolute floor refuted. Support is admitted relative to the best score for the same claim. See the chart on the results page.
31 Jul 2026vector-live-verification
Embedder and reranker, live
Result
With the selected NVIDIA Nemotron embedder and reranker, 33 queries scored 0.9091 first-stage accuracy, 1.0 recall at 5 and 0.9697 reranked accuracy at 5. Ids and payloads survived a container restart and an address change.
Decision
Shipped as the retrieval stack.
Concept linking
30 Jul – 2 Aug 2026equal-best-concept-ties · cross-encoder-tie-break
Ten ways to break a tie
Question
When several concepts score exactly the same for one span, what should decide?
Setup
Up to 700 held-out MedMentions documents, 1,775 tie surfaces, 7,180 tie occurrences.
Result
Nine arms that score each candidate on its own failed: canonical name 84.6%, attestation 77.8%, frequency 65.0%, a clinical problem list 53.8%, plus the regressions charted above. The best of them, a multi-word canonical-name rule, reached 92.97% on 327 selections. The tenth arm, a cross-encoder reading sentence and candidate together, reached 0.9636 on 165 single-word picks, never disagreed with the name rule, and together they covered 12.1% of ties at 0.9411 agreement, at no cost.
Decision
Ties fail closed by default; the cross-encoder breaks single-word ties at a winner-to-runner-up ratio of 6. Checked once on an external word-sense set against a floor fixed in advance: 0.9499 agreement, Wilson lower bound 0.9440 against a 0.9008 floor.
2 Aug 2026local-reranker-substitutes-rejected
Six local rerankers
Result
The best, bge-reranker-v2-m3, reached 0.8554 at the coverage where the remote model reached 0.9636. The biomedical model MedCPT was anti-calibrated: its agreement fell as its confidence rose. Plain top-1 agreement was nearly identical across models, so the gap was calibration, not ranking.
Decision
All six rejected and deleted. The same public weights as the remote reranker now run locally and reproduce its curve.
Aug 2026tie-degeneracy
Ties that are one concept spelled twice
Result
7.1% of tie groups on textbook text and 8.8% on MedMentions are one referent under two identifiers. Over 1,040 passages, collapsing them offline admitted 1,741 concepts against 782 for the paid cross-encoder.
Decision
A free synonym collapse now runs before the tie-break.
Document reading
10 Sep 2026body-type-size-bistable
What size is the book set in?
Setup
All 325 prepared documents, four ways of estimating body type size.
Result
The median line size left 24 documents reading more than half their text as headings. The size carrying the most characters left none; it improved 31 documents and worsened 1. Two books went from 63.7% and 82.1% heading share to 2.3% and 1.9%.
Decision
The character-weighted size shipped. An earlier claim that it was not the fix had been measured on a copy that was not in the corpus.
One broken font family shifted text on 459 of 1,825 pages, putting 3,693 garbled tokens into 157 accepted units. Across the corpus, 72,560 private-use characters appeared in 67 documents; 72% are now read and 49 documents cleared. For ligatures, demanding two proving words everywhere cut coverage from 84% to 69%, because a per-figure font sets one word.
Decision
Prove on the document, repair per span, leave unproven text as printed. Illustrated on How it works.
4 Sep 2026dehyphenation-document-vocabulary
When is a line-end hyphen one word?
Result
Of 1,230,210 hyphenated pairs, 40.0% join because the book prints the joined word more often, 35,284 stay hyphenated because the compound is printed more, and 702,739 are never printed whole and stay as printed. Every sampled join was a broken word and every sampled keep a real compound.
Decision
The document's own vocabulary decides; a tie or silence leaves the hyphen.
5 Aug 2026pdf-and-research-tools-review
Replace the PDF parser?
Setup
Five hand-labelled textbook slices, 16 pages, identical bytes to PyMuPDF, pdf-inspector and Docling.
Result
pdf-inspector was 22× faster and found 8 of 15 table regions; Docling was 20× slower and found 13, with one exact table-grid match and mean overlap 0.58.
Decision
No switch. Each challenger has a narrow role as a second opinion on difficult pages.
Aug 2026axis-run-discriminator · page-sized-raster-population
Axes, and pages that are pictures
Result
Even spacing cannot tell an axis from a numbered table column, because the table is more regular than the axis. Recovering the axis's facing edge from its box would have deleted 13 of 16 real axes. Of 1,445 page-sized rasters, only 10 had under 100 characters of text; the harm was the 177 that carried a caption.
Decision
No axis discriminator built. A page-sized raster carrying one of several captions is refused by name, never bought.
Evaluation methods
22 Jul 2026assertion-accuracy
Self-authored cases against real notes
Result
100% on 44 self-authored cases, 87.7% on 2,376 independently labelled pairs, with two invariant-breaking defects found. After repair, 90.08% strict.
Decision
The held-out harness runs before any assertion change.
27 Aug 2026mcq-explanation-pairing-audit
Does this explanation belong to this question?
Result
Four lexical proxies each failed a hand-read set of 32 in a different direction; one flagged 32 records of which 21 were fine. A model that names both topics before giving its verdict kept 21 of 21 good records and caught 10 of 11 bad ones; an exact rule catches the last.
Decision
The model audit shipped. It reports and never repairs.
17 Aug 2026gateway-throughput-measurement-trap
Is throughput capped per model?
Result
Numbers taken while a build shared the gateway, with the pooled arm given twice the requests in flight, suggested pooling models doubled throughput. With the build stopped and depth held equal, one model reached 1.48 batches per second and the two-model pool 1.41, collapsing to 0.65 at higher concurrency.
Decision
The conclusion was withdrawn before it shipped. Pooling is for failover. Never benchmark a gateway during a build.
Built, measured and deleted
An abbreviation ranker. Wired into production and measured on 199 passages: 8 ambiguous abbreviations before, 8 after. It decided nothing, and 627 lines were deleted.
Six local rerankers. None matched the remote cross-encoder's calibration.
Name-based and context-based tie ranks. Each looked good small and regressed at scale.
Two automated bibliography classifiers. Rejected against hand labels.
An even-spacing axis detector. Refused before it was built, once the measurement showed tables are more regular than axes.