Technical report · Version 1 · 29 September 2026

simpledoctor

Measured results from a fail-closed clinical knowledge engine that turns medical documents into ontology-verified, assertion-aware records for retrieval and for training data.

Summary

simpledoctor reads clinical documents and emits validated records for a knowledge graph, a vector store and model-training datasets. Its errors would become clinical assertions about patients, so every stage treats a model's output as a proposal and accepts it only against an authority outside the model: the UMLS ontology, the source text, or the author's own answer key. Candidates that cannot be decided are kept and queued for review, never guessed.

This report collects the measurements behind that design. Each figure was produced by a committed harness and, wherever such data exists, against independently labelled data rather than cases written by the engine's author. Three results bear on how frontier models should be used as verifiers. A judge with reasoning disabled passed planted contraindications that the same model caught with reasoning on. Retrieval scores measure topical relevance and cannot be calibrated into a test of whether a passage supports a claim. And items a verifier refuses, labelled by the layer that refused them, form preference data.

Results at a glance
FindingResultSample
Negation detection on held-out clinical notesF1 0.9392,376 labelled concept–sentence pairs
Assertion status: self-authored cases, then held-out notes100% → 90.08%44 cases; 2,376 pairs
Concept-linking precision, against the corpus baseline of 0.4710.496100 held-out MedMentions documents
Planted contraindications the judge caught once reasoning was on39 of 52104 re-judged negatives
Spread between best support scores of two claims123×16 searched claims
Answer keys recovered from image-only PDFs, with no disagreement from the author's own check89.1%854 keyed items, 1,364 pages
Reference-list text in top-10 retrieval, before and after cleanup16% → 6%100 slots per arm, p = 0.040
Preference pairs from refused reasoning items2330 judged items, one chapter

How the engine is built

The illustrated walkthrough follows one document through every stage.

Layer 1 · Ontology

Filters 8.5 GB of raw UMLS to 792,516 clinical concepts, once and offline. A concept code is checked against this subset, never against a format pattern.

Layer 2 · Clinical reading

Links mentions to concepts, reads negation, time and who the finding belongs to, and adjudicates every candidate.

Layer 3 · Dataset building

Curates passages, writes retrieval triplets and hard negatives in four clinical failure modes, audits each with a judge, and logs every call to append-only ledgers.

Every candidate ends in exactly one of three verdicts, and the totals must account for all of them. The rule separating them is whether more information would change the answer.

AdmittedCleared every gate.
RejectedWrong by definition: a malformed code, an out-of-scope type, a superseded span. A reviewer could not overturn it.
QuarantinedThe engine could not decide and a person could. Kept in full, excluded from output, and queued for review.

Rules the code enforces

  1. A negated finding never becomes a positive assertion.
  2. A relative's finding is never attributed to the patient.
  3. A concept's semantic type comes from the ontology, never from a caller or a model.
  4. A concept code is verified against the ontology subset. C9999999 matches the format and names nothing.
  5. Nothing is silently dropped.
  6. Ambiguity that cannot be resolved is refused or quarantined, never guessed.

How it is measured

Each figure below names the harness that regenerates it. Held-out data outranks self-authored data. The assertion rules scored 100% on 44 cases written alongside them and 87.7% on their first run against independently labelled clinical notes. That run exposed two defects no self-authored case could have shown: a cue-counting rule that turned an entire negated review of systems into positive findings, and negation scope that crossed section headings, so "Complications: none" negated the diagnosis beneath it.

Stage gates refuse to spend on an artifact that has not passed its checks, and each gate's verdict names the file and run it approved. A committed golden set pins what the document reader produces. Runs are keyed by a fingerprint of model, prompt version and settings, so a changed generator cannot silently reuse another run's work.

Negation, time and experiencer

Measured on the NegEx/ConText corpus: concept–sentence pairs from de-identified discharge summaries and radiology, echocardiography and colonoscopy reports, labelled by the corpus authors. It is the only held-out clinical assertion set available without a data-use agreement.

MeasureResultNote
Assertion status, strict / lenient90.08% / 97.38%2,376 pairs
NegationF1 0.939P 0.905 · R 0.97612 negated findings of 491 leaked through
Negation, rules aloneP 0.976R 0.921Before the ConText pass sees the text
Historical, the weakest classP 0.746R 0.60839 of 89 misses are a bare concept with no context to read
Experiencer4 of 4Too few family rows to judge; 23 constructed cases hold this rule
Reproducescripts/measure_assertion_heldout.py --fetch

Concept linking

On 100 MedMentions test documents held out of all tuning, the production gate ran on each document's own text. A link counts as correct only when it names the gold concept on an overlapping span.

System, predicted mentionsPrecisionRecall
simpledoctor, on annotated spans0.4960.286in-scope gold
TaggerOne, the corpus's own baseline, trained 9 days on a 900 GB machine0.4710.436
scispaCy linker0.387—

Recall is lower by design: the ontology subset leaves out literature-indexing types, the confidence floor is 0.80, and ties fail closed. Of the 1,183 wrong links, 75.8% have a gold concept outside the clinical subset, which no linking decision could reach; 4.1% are parent or child of the gold concept, 6.0% are siblings, and 14.1% are unrelated. MedMentions is research abstracts, so this is an off-domain lower bound for clinical text.

Reproducescripts/measure_link_precision.py --skip 200 --limit 100Measured 1 Aug 2026

Ties and scale-free gates

When several concepts carry exactly the name found in a sentence, the linker ties. Nine methods that score each candidate on its own failed on single-word names, because what separates the candidates is how each one fits the sentence. A cross-encoder that reads sentence and candidate together succeeded. Its raw scores ran only 0.001 to 0.014 and shifted from one tie to the next, so no absolute margin carried over; the ratio of winner to runner-up did. Over 4,064 scoreable ties drawn from 53,776 mentions in 700 held-out documents, the single-word arm agreed with the gold concept 0.951 of the time at a ratio of 6, at zero cost. Every call it declines costs coverage, never correctness, because those mentions are quarantined otherwise.

Reproducescripts/measure_rerank_ties.pyMeasured 2 Aug 2026

Frontier models as judges

Each generated hard negative is a passage rewritten to be clinically wrong in one of four ways: a contraindication trap, outdated practice, a near-miss entity, or unwarranted specificity. A model judge must confirm the flaw before the negative is published. Across 4,142 negatives, a judge running with reasoning disabled (DeepSeek V4 Flash) discarded 22% as clinically accurate and safe, and 26% of contraindication traps. Reading 52 of those discards found the planted flaw in 50, in sentences absent from the source. One told clinicians no adjustment was needed for a patient on therapeutic warfarin; the judge called each one standard practice.

The same judge, reasoning on, re-reading 104 negatives
Of 52 negatives the judge called safe with reasoning off, 39 had their flaw named with reasoning on and 13 were still called safe. Of 52 it approved with reasoning off, 49 stayed approved and 3 were called safe. Called safe, reasoning off n = 52 Approved, reasoning off n = 52 39 flaw named 13 still safe 49 still approved 3 now safe
Same model and same single-passage prompt; only reasoning changed. Reasoning cost about seven times the latency (median 21.3 s against about 3 s) at about $0.0009 per call. Measured 7 Sep 2026.

In this pipeline a wrongly discarded negative costs data and shows up in the counts. The opposite error, a correct passage approved as a flawed negative, looks identical to a correct approval in the ledger. Used as a reward model or safety evaluator, a judge that calls a planted contraindication safe makes the costly error. Evidence was also thin. Only 43.0% of judged negatives, and 50.3% of contraindication traps, had any FDA label text available to the judge; the rest were judged on the model's recall alone.

Relevance is not support

A reasoning item's steps cite passages. When a step's own citations do not support it, the engine searches a 63,785-passage clinical collection and reranks the candidates with a cross-encoder. Sixteen claims actually reach that search. The best score for a case finding about plasma metanephrines was 0.753. For "alpha-blockade must precede any beta-blocker", the safety rule this search exists to protect, it was 0.0061.

Claims that keep any support, by absolute score floor
Claims retaining support fall from 14 of 16 at a floor of 0.005 to 12 at 0.01, 9 at 0.05 and 4 at 0.1. The safety rule's best score is 0.0061, so any floor above it drops the rule; the case finding scores 0.753. Floors that keep the rule 16 12 8 4 0 0.001 0.01 0.1 1 Absolute floor on reranker score (log scale) Safety rule's best: 0.0061 Case finding's best: 0.753 4 of 16 12 of 16
Every floor high enough to keep the evidence tight drops the safety rule; a floor low enough to keep it admits 14 of 16 claims and 93 passages. Measured 9 Sep 2026.

The reranker scores how topical a passage is to its own query, and those scores do not compare across queries. Its ordering was sound: the top three passages for the safety rule were genuine pre-operative alpha-blockade guidance, and junk began at rank six. The shipped gate therefore admits support only relative to the best score for the same claim. Four to six of the sixteen searched claims were the case's own invented values, which should never have been searched at all; that defect is filed.

Reproducescripts/measure_cross_document_grounding.py

Models as readers

Thirteen question-bank answer PDFs are page images with no text layer. A vision model (MiniMax-M3) transcribed every item, and the author's own printed complement, such as "Incorrect answers: A, B, D and E", checked each key where the page printed one.

MeasureResult
Items with a printed key, across 1,364 pages854
Recovered76189.1%
Checked against the author's complement6040 disagreements
Independent re-read by a second model on another provider79 of 8098.75% agree

The one disagreement was the auditor answering G on an item whose options run A to E. Every apparent fabrication traced to how items were identified rather than to the model: a page that prints two items, and a full-length form whose four sections each number their items 1 to 50. Throughput turned out to be a property of the gateway, not the model: 1.56 calls per second on one provider against under 0.15 for every model on another.

Reproducescripts/audit_vision_recovery.pyMeasured 25 Aug 2026

What a PDF loses before a model sees it

MeasureResultNote
Reference-list text in sampled top-10 retrieval slots, before and after removing bibliography pages16 → 6of 100Fisher exact p = 0.040; Wilson 95% intervals 10.1–24.4% and 2.8–12.5%. Counted by hand after two automated classifiers failed on measurement.
Reference-matter pages claimed in 8 textbooks1157,077 entries0 false positives
Private-use glyphs read from each document's own font program72%of 72,56049 of 67 affected documents cleared; unreadable glyphs left as printed
Naive sentence cuts that are not sentence ends7.4%290 of 3,9217 fell inside a negation's reach, where a cut detaches "no" from its finding

Reference lists ranked because they are dense in exactly the medical terms a clinical query carries: mean cosine similarity was higher for the uncleaned collection. Similarity is not a quality measure.

Refusals as supervision

A reasoning item is a case or a multiple-choice question with a step-by-step rationale that cites its passages. It passes three layers in order or is refused with the layer named: structure, citation anchoring, and clinical judgement. On one textbook chapter, three generators produced 30 judged items: 6 accepted, 23 rejected and 1 quarantined because the judge was unreachable. Paired on the same concept and item kind, they form 23 preference pairs.

Refusing layerPairsWhat the rejected half teaches
Clinical judgement13The medicine or its support; each carries the judge's statement of what was unsupported
Citation anchoring7A claim that may be true but is not in the passage it cites
Structure3Form

A quarantined item takes no side, because an outage is not a clinical flaw. A rejected item enters at most one pair. Re-judging 12 items changed one verdict, and an item judged both ways is refused as a pair rather than arbitrated.

Reproducescripts/build_preference_pairs.pyMeasured 7 Sep 2026

Limitations

Reproducibility and data

Every figure names its harness. Each harness reads the pipeline's own ledgers or a public corpus and writes its raw output beside the result. The engine was built between July and September 2026: 546 commits, about 126,000 lines of Python, 2,516 automated tests and 21 recorded architecture decisions.

The repository is private because its build depends on the licensed UMLS Metathesaurus and on reference texts that may not be redistributed. The engine records a redistribution status for every corpus and treats an unverified status as prohibited. No patient data is used. Code, harnesses and raw outputs are available to reviewers on request.

This work uses the UMLS Metathesaurus under the NLM UMLS licence. No UMLS content is reproduced in this report.