simpledoctor · Testing

Testing and verification

Each layer of checking proves something the layer before it cannot. Most of the defects that mattered were caught by a layer above the test suite.

Seven layers of evidence

  1. Tests

    2,516 automated tests

    Prove the code does what its author meant: 1,335 for the clinical reader, 1,181 for the dataset builder and its scripts. They need no network, container or licensed data, so heavy optional dependencies are absent by design.

    pytestruffmypy
  2. Invariants

    Rules that make a violating edit unimportable

    The graph vocabulary is verified when the module is imported, so an edit that lets an edge carry a denial as a property cannot load. Tests also pin that no definitional adjudication reason can reach a reviewer and no uncertain one can be dropped.

    graph._verify_vocabularytest_adjudicationtest_identity
  3. Golden sets

    Committed fixtures and what they produce

    The reader's golden set pins 26 passages; curation's pins 40 units from three article fixtures, with the linker off so no licensed artifact is needed. A lost row, an invented row and a changed row are three findings, and order is part of the expectation.

    gate_stage.py --golden
  4. Stage gates

    Real artifacts, checked before money is spent on them

    Structural rules block always. A verdict names the run, the file and the file's hash, and a run that requires gates re-checks all three.

    gates.pystage-gates-v1.json
  5. Held-out data

    57 measurement harnesses

    External labelled corpora and re-runnable measurements. This layer turned 100% into 87.7% on day one.

    scripts/measure_*.py
  6. Review

    Two independent AI reviewers

    A Claude review and a Codex adversarial review found 10 defects behind 608 passing tests, then 9 more behind 650.

  7. Reading

    A person reading samples

    Hand counts are the reported number wherever an automated classifier failed against hand labels. A targeted read of one marginal number found 93 of 309 accepted triplets resting on reference lists.

Before every commit

.venv/bin/python -m pytest -q
.venv/bin/ruff check . && .venv/bin/ruff format --check .
.venv/bin/mypy clinical_preprocessor clinical_dataset_builder scripts

A commit is checked in a separate worktree of its own tree, so a concurrent session's uncommitted work can neither pass nor fail it.

Stage gates

Three stages are gated: the read, curation and verification. A gate reads the stage's artifact, sets aside a seeded sample for a person to read, and writes its verdict beside the artifact. Exit code 0 means passed, 1 failed, 2 could not run. A refusal is recorded too, because "gated and refused" and "never gated" ask the operator for different things.

Structural ruleWhat it catches
chunk_ids_derive_from_their_own_contentA record that joins to nothing while every count stays right
every_read_still_hashes_to_what_it_wroteAn artifact changed after it was written
the_accounting_reconcilesWork the ledger says was done that the artifact does not hold
the_stage_produced_somethingAn empty artifact, which satisfies every other rule vacuously
Hard negatives are exactly the approved auditA surplus negative or a lost approval, asked in both directions
No negative restates its own positiveA "negative" that is the right answer again
The published file holds every accepted outcome onceA run that died writing the dataset, or a resume that never rebuilt it

Thresholds are different: they are claims about a corpus, and they ship unset. Unset means "not asserted", never "no limit", and an unset threshold still reports what it observed. An unmeasured threshold that blocks teaches an operator to bypass the gate.

Measurement harnesses

AreaHarnesses
Assertionmeasure_assertion_heldout on 2,376 externally labelled pairs; measure_assertion on the self-authored set, kept as a demonstration
Linking and tiesmeasure_link_precision, measure_rerank_ties, measure_tie_degeneracy, measure_tie_break_yield, and measure_msh_wsd_verdict, the one-shot external check
Reading pagesmeasure_body_type_size, measure_heading_attribution, measure_outline_collisions, measure_reference_matter, measure_sentence_boundaries, measure_pua_glyphs, measure_axis_runs
Tables and figuresmeasure_flatten_refusals, measure_table_bounds, measure_visual_yield, measure_vision_analysis, measure_plot_vocabulary
Questionsmeasure_mcq_recovery against question banks and prose controls
Judges and generationmeasure_judge_arms, measure_judge_batching, measure_generator_candidates, measure_label_coverage, measure_model_candidates
Retrievalmeasure_retrieval_arms, measure_embedders, measure_query_gates, measure_cross_document_grounding
Cost and capacitymeasure_generation_concurrency, measure_service_capacity, measure_split_cost, measure_read_vs_curate

What the suite missed

Every item below passed the full test suite at the time. The layer that caught it is named.

DefectCaught by
Enumerated negation flipped to positive; negation crossing a section headingHeld-out clinical notes
19 defects across two rounds, 6 of them introduced by the first round's fixesIndependent review
A 92.9 MB ledger for 30 units, and a 37 GB curation stallThe first real corpus runs
A store read silently capped at 10,000 of 49,276 unitsScale
One NUL byte stopping a graph ingest at 48.6%Scale
60.7% of accepted work missing from the published fileA gate rule that reads the published file itself
93 of 309 accepted triplets resting on reference listsReading one marginal number
Explanations attached to the wrong question in 3 of 4 records of one fileA hand-read set, then a model audit

The lessons behind each of these are written up on the Lessons page.