simpledoctor · History

History

Eight weeks, from a first commit that measured assertion accuracy before adding anything, to a verified dataset build across four shards.

At a glance

First commit22 July 2026
Commits546
Lines of Python126,235 lines
Automated tests collected2,516
Architecture decision records21
Tracked issues, closed of total284 of 391
Commits per week
Commits per week starting 20 July: 131, then 90, 48, 8, 64, 67, 93, 44, and 1 in the week of 14 September. 100 50 0 131 93 Jul 20 Jul 27 Aug 3 Aug 10 Aug 17 Aug 24 Aug 31 Sep 7 Sep 14 Week starting, 2026
Commits on the working branch. From mid-September the work was running and measuring builds rather than changing code.

Phases

22–23 July

Measure the assertion layer before building on it

The first day built a labelled assertion set, then a held-out harness on real clinical notes. The constructed set said 100%; real text said 87.7%. Negation, experiencer and family history were repaired first, and the three-verdict adjudication vocabulary was defined, because a negated or relative's finding must never become a patient assertion.

23–27 July

One owner per rule

The pipeline was refactored so each rule has one owning module: admission, passages, the page reader, corpus recipes, reference matter and abbreviations. Decisions 0001 to 0007 were written: model output is a proposal, native tables are read from their text, and a concurrency-safe ledger replaces an orchestration framework.

28 July – 10 August

Stores, linking and the first external measurements

The linker was rebuilt from the filtered ontology and measured against published systems. The cross-encoder tie-break became the first positive result after nine rejected arms. The graph and vector stores went live with assertion encoded as the edge type.

17–29 August

Gates and resumability

The mechanical read is written down before anything is paid for; runs stop where a gate can read the artifact; verdicts are bound to the run and the file they approved. A question-bank corpus was added, including answer keys recovered from page images.

31 August – 4 September

Cleaner reading

Bibliography and index pages are removed before reading. Broken fonts, private-use glyphs, ligatures and hyphenated wraps are repaired on proof from each document. A curation stall at 37 GB was traced to a derived value carried on every unit, and a corpus five times larger was indexed.

3–10 September

Reasoning items and preference pairs

Reasoning items are generated from a concept subtree and proved in three layers. Refused items became the rejected halves of preference pairs. The judge's reasoning trial and the grounding-floor refutation both landed this week.

10–13 September

Rebuild and verify

The corpus was re-read on the corrected tree, curation was costed and completed (109 documents, 26,894 units, 3.41 hours), and a four-shard verification build was launched, carrying forward 17,261 outcomes already paid for.

Milestones

  1. Held-out assertion harness

    The self-authored set said 100% while real notes scored 87.7%.

  2. Assertion repaired on real text

    Strict 90.08%, lenient 97.38% on 2,376 pairs; negation F1 0.939.

  3. Relatives stop leaking into the patient

    Experiencer, relation, lineage and degree modelled.

  4. One admission module

    Every candidate that does not clear is now counted rather than dropped.

  5. One PDF reader

    A single walk yields prose, visuals and apparatus.

  6. Ontology subset measured

    792,516 concepts from 8.5 GB of raw UMLS.

  7. Assertion becomes the graph edge type

    A careless query can no longer return a denial.

  8. Vector store live

    0.9697 reranked accuracy at 5 on 33 queries.

  9. Production link precision

    0.496 on annotated spans of 100 held-out MedMentions documents, against the corpus baseline of 0.471.

  10. Cross-encoder tie-break

    First positive result after nine rejected arms.

  11. A plotted number must be printed

    A value read off a chart is refused unless the chart prints it.

  12. Stage gates

    Runs stop where a gate can read the output.

  13. Answer keys from page images

    761 of 854 recovered; independent audit agreed on 79 of 80.

  14. Cleanup measured on retrieval

    Reference-list text in top-10 slots fell from 16 to 6 of 100.

  15. Five-times-larger corpus indexed

    325 documents, 50,348 units, 49,278 accepted points.

  16. Font, glyph and ligature repair

    Proved per document, repaired per span.

  17. Curation memory stall fixed

    One 2,190-unit book had reached 37 GB.

  18. Reasoning items

    Generated from a concept subtree; the judge sees only the cited passages.

  19. Refusals become preference data

    23 labelled pairs from the chapter's first 30 judged items.

  20. Judge reasoning trial

    39 of 52 wrongly discarded negatives recovered with reasoning on.

  21. Support floor refuted

    A 123-fold score gap between a case fact and a safety rule.

  22. Heading re-read

    Prose characters up 5.7% across 314 documents.

  23. Verification build launched

    About 43,000 outcomes targeted across four shards.

Architecture decisions

Twenty-one recorded decisions. Number 14 was never used.

No.DateDecision
000122 JulAn analyzer's attestation is verified, not trusted: it proposes a mention and a concept, and the engine checks the concept and derives its type.
000222 Julscikit-learn stays below 1.5, where newer versions silently load the linker without IDF weighting. The weights are repaired at load or the load raises.
000322 JulRepair the weak assertion layer before building adjudication on top of it.
000423 JulDocument, passage and figure ids are content-addressed; paths and detection order are not part of them.
000523 JulNo orchestration framework: plain Python, strict contracts, a bounded worker pool and a concurrency-safe ledger.
000624 JulVision analysis is a proposal, verified by admission before it reaches any output.
000727 JulA native table is read from its extraction and never rasterized.
000829 JulAn ungrounded negative is not an unverified one: quarantine only what could not be checked.
000929 JulThe linker is built from the filtered ontology, not from a stock artifact.
001030 JulA reviewer's decision is a proposal with a scope, and cannot overrule the ontology.
001130 JulAn assertion is an edge type, not an edge property.
00126 AugA plotted quantity is refused unless its numbers are printed.
00138 AugAn authored question is not an authored answer; unanswered questions never become passages.
00159 AugAn absent edge is not an absent fact: the graph gates candidates and never retrieves them.
00169 AugA wrapped table cell is a rendering fact; a table flattens only when that is its sole defect.
001723 AugA curated unit is keyed by what curated it, so a model swap does not throw away curation.
001825 AugRaw UMLS stays local; UMLS identifiers in the pipeline's output are releasable with attribution.
00193 SepRetrievable is not generation-worthy: two questions, two flags.
00203 SepPrevention is a section, and a swapped-number negative is judgeable.
00216 SepA reasoning item is generated from a concept subtree, not a passage, and proved in three layers.
00227 SepA refused reasoning item is the rejected half of a preference pair.