History
Eight weeks, from a first commit that measured assertion accuracy before adding anything, to a verified dataset build across four shards.
At a glance
| First commit | 22 July 2026 |
| Commits | 546 |
| Lines of Python | 126,235 lines |
| Automated tests collected | 2,516 |
| Architecture decision records | 21 |
| Tracked issues, closed of total | 284 of 391 |
Phases
Measure the assertion layer before building on it
The first day built a labelled assertion set, then a held-out harness on real clinical notes. The constructed set said 100%; real text said 87.7%. Negation, experiencer and family history were repaired first, and the three-verdict adjudication vocabulary was defined, because a negated or relative's finding must never become a patient assertion.
One owner per rule
The pipeline was refactored so each rule has one owning module: admission, passages, the page reader, corpus recipes, reference matter and abbreviations. Decisions 0001 to 0007 were written: model output is a proposal, native tables are read from their text, and a concurrency-safe ledger replaces an orchestration framework.
Stores, linking and the first external measurements
The linker was rebuilt from the filtered ontology and measured against published systems. The cross-encoder tie-break became the first positive result after nine rejected arms. The graph and vector stores went live with assertion encoded as the edge type.
Gates and resumability
The mechanical read is written down before anything is paid for; runs stop where a gate can read the artifact; verdicts are bound to the run and the file they approved. A question-bank corpus was added, including answer keys recovered from page images.
Cleaner reading
Bibliography and index pages are removed before reading. Broken fonts, private-use glyphs, ligatures and hyphenated wraps are repaired on proof from each document. A curation stall at 37 GB was traced to a derived value carried on every unit, and a corpus five times larger was indexed.
Reasoning items and preference pairs
Reasoning items are generated from a concept subtree and proved in three layers. Refused items became the rejected halves of preference pairs. The judge's reasoning trial and the grounding-floor refutation both landed this week.
Rebuild and verify
The corpus was re-read on the corrected tree, curation was costed and completed (109 documents, 26,894 units, 3.41 hours), and a four-shard verification build was launched, carrying forward 17,261 outcomes already paid for.
Milestones
Held-out assertion harness
The self-authored set said 100% while real notes scored 87.7%.
Assertion repaired on real text
Strict 90.08%, lenient 97.38% on 2,376 pairs; negation F1 0.939.
Relatives stop leaking into the patient
Experiencer, relation, lineage and degree modelled.
One admission module
Every candidate that does not clear is now counted rather than dropped.
One PDF reader
A single walk yields prose, visuals and apparatus.
Ontology subset measured
792,516 concepts from 8.5 GB of raw UMLS.
Assertion becomes the graph edge type
A careless query can no longer return a denial.
Vector store live
0.9697 reranked accuracy at 5 on 33 queries.
Production link precision
0.496 on annotated spans of 100 held-out MedMentions documents, against the corpus baseline of 0.471.
Cross-encoder tie-break
First positive result after nine rejected arms.
A plotted number must be printed
A value read off a chart is refused unless the chart prints it.
Stage gates
Runs stop where a gate can read the output.
Answer keys from page images
761 of 854 recovered; independent audit agreed on 79 of 80.
Cleanup measured on retrieval
Reference-list text in top-10 slots fell from 16 to 6 of 100.
Five-times-larger corpus indexed
325 documents, 50,348 units, 49,278 accepted points.
Font, glyph and ligature repair
Proved per document, repaired per span.
Curation memory stall fixed
One 2,190-unit book had reached 37 GB.
Reasoning items
Generated from a concept subtree; the judge sees only the cited passages.
Refusals become preference data
23 labelled pairs from the chapter's first 30 judged items.
Judge reasoning trial
39 of 52 wrongly discarded negatives recovered with reasoning on.
Support floor refuted
A 123-fold score gap between a case fact and a safety rule.
Heading re-read
Prose characters up 5.7% across 314 documents.
Verification build launched
About 43,000 outcomes targeted across four shards.
Architecture decisions
Twenty-one recorded decisions. Number 14 was never used.
| No. | Date | Decision |
|---|---|---|
| 0001 | 22 Jul | An analyzer's attestation is verified, not trusted: it proposes a mention and a concept, and the engine checks the concept and derives its type. |
| 0002 | 22 Jul | scikit-learn stays below 1.5, where newer versions silently load the linker without IDF weighting. The weights are repaired at load or the load raises. |
| 0003 | 22 Jul | Repair the weak assertion layer before building adjudication on top of it. |
| 0004 | 23 Jul | Document, passage and figure ids are content-addressed; paths and detection order are not part of them. |
| 0005 | 23 Jul | No orchestration framework: plain Python, strict contracts, a bounded worker pool and a concurrency-safe ledger. |
| 0006 | 24 Jul | Vision analysis is a proposal, verified by admission before it reaches any output. |
| 0007 | 27 Jul | A native table is read from its extraction and never rasterized. |
| 0008 | 29 Jul | An ungrounded negative is not an unverified one: quarantine only what could not be checked. |
| 0009 | 29 Jul | The linker is built from the filtered ontology, not from a stock artifact. |
| 0010 | 30 Jul | A reviewer's decision is a proposal with a scope, and cannot overrule the ontology. |
| 0011 | 30 Jul | An assertion is an edge type, not an edge property. |
| 0012 | 6 Aug | A plotted quantity is refused unless its numbers are printed. |
| 0013 | 8 Aug | An authored question is not an authored answer; unanswered questions never become passages. |
| 0015 | 9 Aug | An absent edge is not an absent fact: the graph gates candidates and never retrieves them. |
| 0016 | 9 Aug | A wrapped table cell is a rendering fact; a table flattens only when that is its sole defect. |
| 0017 | 23 Aug | A curated unit is keyed by what curated it, so a model swap does not throw away curation. |
| 0018 | 25 Aug | Raw UMLS stays local; UMLS identifiers in the pipeline's output are releasable with attribution. |
| 0019 | 3 Sep | Retrievable is not generation-worthy: two questions, two flags. |
| 0020 | 3 Sep | Prevention is a section, and a swapped-number negative is judgeable. |
| 0021 | 6 Sep | A reasoning item is generated from a concept subtree, not a passage, and proved in three layers. |
| 0022 | 7 Sep | A refused reasoning item is the rejected half of a preference pair. |