simpledoctor
Measured results from a fail-closed clinical knowledge engine that turns medical documents into ontology-verified, assertion-aware records for retrieval and for training data.
Summary
simpledoctor reads clinical documents and emits validated records for a knowledge graph, a vector store and model-training datasets. Its errors would become clinical assertions about patients, so every stage treats a model's output as a proposal and accepts it only against an authority outside the model: the UMLS ontology, the source text, or the author's own answer key. Candidates that cannot be decided are kept and queued for review, never guessed.
This report collects the measurements behind that design. Each figure was produced by a committed harness and, wherever such data exists, against independently labelled data rather than cases written by the engine's author. Three results bear on how frontier models should be used as verifiers. A judge with reasoning disabled passed planted contraindications that the same model caught with reasoning on. Retrieval scores measure topical relevance and cannot be calibrated into a test of whether a passage supports a claim. And items a verifier refuses, labelled by the layer that refused them, form preference data.
| Finding | Result | Sample |
|---|---|---|
| Negation detection on held-out clinical notes | F1 0.939 | 2,376 labelled concept–sentence pairs |
| Assertion status: self-authored cases, then held-out notes | 100% → 90.08% | 44 cases; 2,376 pairs |
| Concept-linking precision, against the corpus baseline of 0.471 | 0.496 | 100 held-out MedMentions documents |
| Planted contraindications the judge caught once reasoning was on | 39 of 52 | 104 re-judged negatives |
| Spread between best support scores of two claims | 123× | 16 searched claims |
| Answer keys recovered from image-only PDFs, with no disagreement from the author's own check | 89.1% | 854 keyed items, 1,364 pages |
| Reference-list text in top-10 retrieval, before and after cleanup | 16% → 6% | 100 slots per arm, p = 0.040 |
| Preference pairs from refused reasoning items | 23 | 30 judged items, one chapter |
How the engine is built
The illustrated walkthrough follows one document through every stage.
Filters 8.5 GB of raw UMLS to 792,516 clinical concepts, once and offline. A concept code is checked against this subset, never against a format pattern.
Links mentions to concepts, reads negation, time and who the finding belongs to, and adjudicates every candidate.
Curates passages, writes retrieval triplets and hard negatives in four clinical failure modes, audits each with a judge, and logs every call to append-only ledgers.
Every candidate ends in exactly one of three verdicts, and the totals must account for all of them. The rule separating them is whether more information would change the answer.
Rules the code enforces
- A negated finding never becomes a positive assertion.
- A relative's finding is never attributed to the patient.
- A concept's semantic type comes from the ontology, never from a caller or a model.
- A concept code is verified against the ontology subset.
C9999999matches the format and names nothing. - Nothing is silently dropped.
- Ambiguity that cannot be resolved is refused or quarantined, never guessed.
How it is measured
Each figure below names the harness that regenerates it. Held-out data outranks self-authored data. The assertion rules scored 100% on 44 cases written alongside them and 87.7% on their first run against independently labelled clinical notes. That run exposed two defects no self-authored case could have shown: a cue-counting rule that turned an entire negated review of systems into positive findings, and negation scope that crossed section headings, so "Complications: none" negated the diagnosis beneath it.
Stage gates refuse to spend on an artifact that has not passed its checks, and each gate's verdict names the file and run it approved. A committed golden set pins what the document reader produces. Runs are keyed by a fingerprint of model, prompt version and settings, so a changed generator cannot silently reuse another run's work.
Negation, time and experiencer
Measured on the NegEx/ConText corpus: concept–sentence pairs from de-identified discharge summaries and radiology, echocardiography and colonoscopy reports, labelled by the corpus authors. It is the only held-out clinical assertion set available without a data-use agreement.
| Measure | Result | Note |
|---|---|---|
| Assertion status, strict / lenient | 90.08% / 97.38% | 2,376 pairs |
| Negation | F1 0.939P 0.905 · R 0.976 | 12 negated findings of 491 leaked through |
| Negation, rules alone | P 0.976R 0.921 | Before the ConText pass sees the text |
| Historical, the weakest class | P 0.746R 0.608 | 39 of 89 misses are a bare concept with no context to read |
| Experiencer | 4 of 4 | Too few family rows to judge; 23 constructed cases hold this rule |
Concept linking
On 100 MedMentions test documents held out of all tuning, the production gate ran on each document's own text. A link counts as correct only when it names the gold concept on an overlapping span.
| System, predicted mentions | Precision | Recall |
|---|---|---|
| simpledoctor, on annotated spans | 0.496 | 0.286in-scope gold |
| TaggerOne, the corpus's own baseline, trained 9 days on a 900 GB machine | 0.471 | 0.436 |
| scispaCy linker | 0.387 | — |
Recall is lower by design: the ontology subset leaves out literature-indexing types, the confidence floor is 0.80, and ties fail closed. Of the 1,183 wrong links, 75.8% have a gold concept outside the clinical subset, which no linking decision could reach; 4.1% are parent or child of the gold concept, 6.0% are siblings, and 14.1% are unrelated. MedMentions is research abstracts, so this is an off-domain lower bound for clinical text.
Ties and scale-free gates
When several concepts carry exactly the name found in a sentence, the linker ties. Nine methods that score each candidate on its own failed on single-word names, because what separates the candidates is how each one fits the sentence. A cross-encoder that reads sentence and candidate together succeeded. Its raw scores ran only 0.001 to 0.014 and shifted from one tie to the next, so no absolute margin carried over; the ratio of winner to runner-up did. Over 4,064 scoreable ties drawn from 53,776 mentions in 700 held-out documents, the single-word arm agreed with the gold concept 0.951 of the time at a ratio of 6, at zero cost. Every call it declines costs coverage, never correctness, because those mentions are quarantined otherwise.
Frontier models as judges
Each generated hard negative is a passage rewritten to be clinically wrong in one of four ways: a contraindication trap, outdated practice, a near-miss entity, or unwarranted specificity. A model judge must confirm the flaw before the negative is published. Across 4,142 negatives, a judge running with reasoning disabled (DeepSeek V4 Flash) discarded 22% as clinically accurate and safe, and 26% of contraindication traps. Reading 52 of those discards found the planted flaw in 50, in sentences absent from the source. One told clinicians no adjustment was needed for a patient on therapeutic warfarin; the judge called each one standard practice.
In this pipeline a wrongly discarded negative costs data and shows up in the counts. The opposite error, a correct passage approved as a flawed negative, looks identical to a correct approval in the ledger. Used as a reward model or safety evaluator, a judge that calls a planted contraindication safe makes the costly error. Evidence was also thin. Only 43.0% of judged negatives, and 50.3% of contraindication traps, had any FDA label text available to the judge; the rest were judged on the model's recall alone.
Relevance is not support
A reasoning item's steps cite passages. When a step's own citations do not support it, the engine searches a 63,785-passage clinical collection and reranks the candidates with a cross-encoder. Sixteen claims actually reach that search. The best score for a case finding about plasma metanephrines was 0.753. For "alpha-blockade must precede any beta-blocker", the safety rule this search exists to protect, it was 0.0061.
The reranker scores how topical a passage is to its own query, and those scores do not compare across queries. Its ordering was sound: the top three passages for the safety rule were genuine pre-operative alpha-blockade guidance, and junk began at rank six. The shipped gate therefore admits support only relative to the best score for the same claim. Four to six of the sixteen searched claims were the case's own invented values, which should never have been searched at all; that defect is filed.
Models as readers
Thirteen question-bank answer PDFs are page images with no text layer. A vision model (MiniMax-M3) transcribed every item, and the author's own printed complement, such as "Incorrect answers: A, B, D and E", checked each key where the page printed one.
| Measure | Result |
|---|---|
| Items with a printed key, across 1,364 pages | 854 |
| Recovered | 76189.1% |
| Checked against the author's complement | 6040 disagreements |
| Independent re-read by a second model on another provider | 79 of 8098.75% agree |
The one disagreement was the auditor answering G on an item whose options run A to E. Every apparent fabrication traced to how items were identified rather than to the model: a page that prints two items, and a full-length form whose four sections each number their items 1 to 50. Throughput turned out to be a property of the gateway, not the model: 1.56 calls per second on one provider against under 0.15 for every model on another.
What a PDF loses before a model sees it
| Measure | Result | Note |
|---|---|---|
| Reference-list text in sampled top-10 retrieval slots, before and after removing bibliography pages | 16 → 6of 100 | Fisher exact p = 0.040; Wilson 95% intervals 10.1–24.4% and 2.8–12.5%. Counted by hand after two automated classifiers failed on measurement. |
| Reference-matter pages claimed in 8 textbooks | 1157,077 entries | 0 false positives |
| Private-use glyphs read from each document's own font program | 72%of 72,560 | 49 of 67 affected documents cleared; unreadable glyphs left as printed |
| Naive sentence cuts that are not sentence ends | 7.4%290 of 3,921 | 7 fell inside a negation's reach, where a cut detaches "no" from its finding |
Reference lists ranked because they are dense in exactly the medical terms a clinical query carries: mean cosine similarity was higher for the uncleaned collection. Similarity is not a quality measure.
Refusals as supervision
A reasoning item is a case or a multiple-choice question with a step-by-step rationale that cites its passages. It passes three layers in order or is refused with the layer named: structure, citation anchoring, and clinical judgement. On one textbook chapter, three generators produced 30 judged items: 6 accepted, 23 rejected and 1 quarantined because the judge was unreachable. Paired on the same concept and item kind, they form 23 preference pairs.
| Refusing layer | Pairs | What the rejected half teaches |
|---|---|---|
| Clinical judgement | 13 | The medicine or its support; each carries the judge's statement of what was unsupported |
| Citation anchoring | 7 | A claim that may be true but is not in the passage it cites |
| Structure | 3 | Form |
A quarantined item takes no side, because an outage is not a clinical flaw. A rejected item enters at most one pair. Re-judging 12 items changed one verdict, and an item judged both ways is refused as a pair rather than arbitrated.
Limitations
- One investigator built the engine and hand-counted the samples reported here.
- The held-out assertion corpus has only four family-history rows and no negation-with-duration construction; constructed adversarial cases cover those shapes.
- Link precision is measured on research abstracts, off-domain for clinical text. Gold clinical linking corpora require a data-use agreement.
- The judge trial is one model on 104 items; the grounding result is 16 claims from one chapter; the preference result is 23 pairs from one chapter.
- The retrieval comparison has 100 slots per arm and overlapping intervals. Its reranker arm has not yet run.
Reproducibility and data
Every figure names its harness. Each harness reads the pipeline's own ledgers or a public corpus and writes its raw output beside the result. The engine was built between July and September 2026: 546 commits, about 126,000 lines of Python, 2,516 automated tests and 21 recorded architecture decisions.
The repository is private because its build depends on the licensed UMLS Metathesaurus and on reference texts that may not be redistributed. The engine records a redistribution status for every corpus and treats an unverified status as prohibited. No patient data is used. Code, harnesses and raw outputs are available to reviewers on request.
This work uses the UMLS Metathesaurus under the NLM UMLS licence. No UMLS content is reproduced in this report.