simpledoctor · Lessons

Lessons learned

Each lesson below cost something first. It is written as the incident that taught it and the rule the code now enforces, so the rule outlives the memory of why.

Measuring

Recall is not evidence

A semantic-type "cleanup" dropped one type and silently deleted smoking and alcohol status from the engine. A claim of "no leak, peak 1.07 GB" was measured while the machine was deep in swap; the true figure was 12.4 GB. A measurement of a filter re-applied the same filter to its own output, so it could only report 100%.

RuleEvery number carries the script that regenerates it, and no number is quoted in the operating document, where it would go stale.

A self-authored gold set proves nothing

The assertion rules scored 100% on 44 cases written alongside them and 87.7% on independently labelled clinical notes. The held-out notes exposed a rule that turned a whole negated review of systems positive, and negation that crossed a section heading. Neither could appear in cases written by the rules' author. The held-out set has its own blind spots too: only four family-history rows and no negation-with-duration construction.

RuleThe held-out harness runs before any assertion change is trusted, and constructed adversarial cases cover the shapes it lacks.

A green suite is not a measurement

A review-queue fix shipped wrong with 534 tests passing, because no test asserted the property that broke. A 50-article re-measurement caught it.

RuleA change with a measurable effect is measured, whatever the suite says.

Two independent reviews beat one

A Claude review and a Codex adversarial review ran in parallel over the same commits. The first round found 10 defects behind 608 passing tests; the second found 9 more behind 650, six of them introduced by the first round's fixes. The two reviewers overlapped on only 3 of 10 and 2 of 9.

RuleRun both, and say in writing which shapes a measurement cannot judge.

Never benchmark a shared gateway under load

A finding that pooling two models doubled throughput nearly shipped. It had been measured while a four-shard build used the same gateway, with the pooled arm given twice the requests in flight. Measured fairly, one model reached 1.48 batches per second and the pool 1.41, collapsing to 0.65 at higher concurrency.

RulePooling models is for failover, never throughput. Hold in-flight depth equal across arms, and never measure during a build.

Hold the cheap proxy to the real outcome once

Removing bibliography pages cut a corpus-level contamination count by 71%. Asked 50 clinical questions against both corpora, the retrieval effect was 62%: 16 against 6 of 100 top-10 slots. Two automated classifiers built to count it were rejected against hand labels.

RuleThe hand count is the reported number. The proxy was confirmed good enough to direct work, which is its only job.

Spend a one-shot verdict once

The tie-break was checked on an external word-sense set against a floor fixed in advance and cleared it: 0.9499 agreement, lower bound 0.9440 against 0.9008. Its grid of other thresholds is kept for the record only, because choosing a better-looking ratio after seeing it would spend the external set.

RuleThresholds are selected on development labels; the external check runs once.

Wire it, then measure whether it decides anything

An ontology-based abbreviation ranker was wired into production and measured on 199 passages: 8 ambiguous abbreviations before, 8 after. It abstained because the concepts it needed were themselves tied, and the gate fails closed. It was deleted.

RuleAbbreviations are expanded only from ranked evidence (a definition in the text, then the book's own table, then a lexicon), and ties at the best rank are refused.

Working with models

A model transcribing is not a regex matching

A pattern that does not match returns nothing. A model that cannot read a page returns fluent, wrong medicine. Recovering answer keys from page images therefore checks every item against the author's own statement of which options are wrong, which must partition the options exactly.

RuleA single-reader shortcut is always paired with a sampled audit by an independent model, so agreement is a measured rate rather than an assumption.

Relevance is not support

A reranker scored a case finding 0.7526 and the safety rule "alpha-blockade must precede any beta-blocker" 0.0061. Every absolute floor that kept the first dropped the second. The ranking itself was right.

RuleSupport is admitted relative to the best score for the same claim. The only absolute rule is at exactly zero.

An equal error rate is not an equal judge

A fast judge matched the incumbent's rate of calling flawed negatives safe, 0.125 each, but the two agreed on only one of their five mistakes. The fast judge called flaws "consistent with current standard of care" that the incumbent named exactly.

RuleCompare judges item by item, not by rate.

The rejects are the dataset

The reasoning stage refuses most of what it is offered: 4 of 11 accepted in the prototype. Read as yield, that is failure, and the standing incentive is to loosen the judge. Read as data, each refusal is a labelled negative on the same concept as an accepted item, with the layer that refused it and the judge's reason.

RuleOnly judged items take a side in a pair; an outage is not a clinical flaw. A stricter judge now produces better-labelled data, not less of it.

A successful response can carry a failure

Twenty-eight negatives in five minutes were lost to "operation aborted" errors delivered inside HTTP 200 responses, read as permanent failures because they carried no status code. A vision model that fabricated a chart's axes also answered 200.

RuleAn error envelope that says the request was interrupted is retried; a model is judged on what it says about a known answer, not on its status code.

A tie is usually two concepts, sometimes one

About 7% of linker ties are one referent spelled twice under two identifiers, such as an imaging test and the service that performs it. The rest are genuinely different concepts sharing a name.

RuleSynonymous identifiers are collapsed offline first, for free; real ties go to the cross-encoder or are quarantined.

Identity and accounting

A number is not a key

Answer-key recovery failed four times on the same shape. A page prints two items and the model was asked for "the item". A full-length form runs four sections each numbered 1 to 50, so keyed on the number alone one form recovered 0 of 50 while its siblings got 42 to 50. An audit joined on number alone. A positional label named different files in the audit and the recovery.

RuleIdentity is a composite key owned by one function. Keyed by section and number, that form recovered 199 items and reader disagreements fell from 52 to 2.

Count before you page, and check after

The graph store capped a read at 10,000 rows without saying so, against a 49,276-unit graph. The prune reported "stored units: 10000", a number that looked measured, and never compared anything past the cap.

RuleEvery store read counts first, pages in a total order, and raises if the pages do not reach the count.

An outage is not a verdict about a document

Nine of nineteen document failures in one build were "failed to open file", in two bursts milliseconds apart; every file reads fine today. A source had been briefly unreachable, and the run marked nine books failed.

RuleA source that vanishes stops the run. Failed, skipped and ingested are verdicts about a document, and an outage is not one.

The published file can lag its ledger

A build held 17,981 accepted outcomes and had published 7,072: 60.7% of the accepted work was missing from the training file. One shard published nothing after an error on its way out; two others resumed and never rebuilt the file.

RuleThe verify gate checks that every accepted outcome is published exactly once, as content rather than a timestamp, and one script rebuilds the file from the ledger.

One question, one answer

A supervisor decided whether a shard should retry by searching its report for an open circuit. One shard recorded 79 judge failures without its circuit ever opening, and was printed complete while its own warnings said three of four documents failed. Three scripts held three different answers to that question.

RuleOne script owns the retry verdict, and it refuses rather than defaults: a report with nothing in it is not a report saying nothing failed.

Scale

Tests pass; production runs out of memory

The first real corpus run wrote a 92.9 MB ledger for 30 units behind 1,231 passing tests; 99.7% of it was each concept's ontology neighbourhood, copied from a store that already held it. Its in-memory twin surfaced weeks later: about 17 MB per unit, so a 2,190-unit book reached 37 GB and swap hit 39.2 of 39.9 GB.

RuleA derived value is never stored beside its source; a unit drops the neighbourhood when it is constructed.

Refuse a bad character; do not strip it

One NUL byte stopped a 50,348-unit graph ingest at 48.6%. Stripping it in the transport would have left the graph disagreeing with the ledger and the vector collection, silently.

RuleOne module defines a control character. The reader replaces it with a separator, and both stores refuse a unit still carrying one, by name.

No crash is not thread safety

A run at concurrency 5 used the shared language pipeline and finished without crashing. Review found the hazard anyway: the lock never covered the second cached pipeline, and a lock around the whole admission step held a network call.

RuleThe backend serialises only its own parse, and the network call runs outside the lock.

Reading documents

Furniture is position, not repetition

"Investigations" recurs 301 times in one book as a real heading and never lands at the same height twice. A watermark lands at the same coordinate on 151 of 218 pages. One book's licence footer rode along in 289 of its 370 passages.

RuleA line stamped at one fixed position on most pages is furniture. The threshold is chosen from below, because a false positive deletes a section from every page.

The median line misleads

Body type size was the median line size, and a populous population of small type pulled it down, so body text rose above the heading floor. Across 325 documents, 24 read more than half their text as headings.

RuleThe body size is the size carrying the most characters. Two books went from 63.7% and 82.1% heading share to 2.3% and 1.9%.

Two heading lines can be one header

Treatment boxes in one book are set kind-first: a small "TREATMENT", then the larger subject. The reader treated the subject as closing the kind and discarded it, at a cost of 537 passages in the largest book.

RuleTwo adjacent heading lines with no prose between them, the second larger, are one header.

A heading locates; a shape proves

Chapter-end reference lists begin under the last section's prose, so no page looks like a bibliography. 791 of 12,033 passages under a reference heading were typed as clinical text, asserting concepts from article titles. Later, 93 of one shard's first 309 accepted triplets rested on a reference list, caught by reading one marginal number rather than the aggregate metrics.

RuleA reference list is claimed only when a heading from a measured vocabulary and the shape of the entries agree; generation and the verify gate both refuse one.

Outline titles collide

One book's own bookmark outline has 7,786 entries, and 471 of its 5,581 leaf titles repeat: "Management" 419 times. Letting any matching line replace the heading path filed one chapter's text under another in 29 places.

RuleA re-read is refused for any document whose new read accepts under 75% of its old one, and headings that name a subject are not forced into a section vocabulary.

Rules and vocabularies

Deleting a vocabulary is a change

Consolidating the relative words into one module dropped generic kin terms such as "relatives" and "family members". Every phrasing that named no specific relative then became a patient assertion.

RuleOne module owns who counts as a relative, and two lists are diffed before one defers to the other.

A conditional inside a shared pattern inverts

A regular-expression member meant to exclude "of all" sat in an alternation used both positively and inside negative lookaheads. In the negative position it admitted exactly what it was written to exclude.

RuleNo conditional member in a pattern used in both polarities; a composed pattern is expanded literally before it is trusted.

Sixteen right rows hide the seventeenth

Sixteen of seventeen hand-typed fallback concepts were correct, which hid that pulmonary embolism had been typed as a disease where the ontology says pathologic function.

RuleA test checks every hand-written row against the ontology, and semantic type always comes from the concept.