Testing and verification
Each layer of checking proves something the layer before it cannot. Most of the defects that mattered were caught by a layer above the test suite.
Seven layers of evidence
- Tests
2,516 automated tests
Prove the code does what its author meant: 1,335 for the clinical reader, 1,181 for the dataset builder and its scripts. They need no network, container or licensed data, so heavy optional dependencies are absent by design.
- Invariants
Rules that make a violating edit unimportable
The graph vocabulary is verified when the module is imported, so an edit that lets an edge carry a denial as a property cannot load. Tests also pin that no definitional adjudication reason can reach a reviewer and no uncertain one can be dropped.
- Golden sets
Committed fixtures and what they produce
The reader's golden set pins 26 passages; curation's pins 40 units from three article fixtures, with the linker off so no licensed artifact is needed. A lost row, an invented row and a changed row are three findings, and order is part of the expectation.
- Stage gates
Real artifacts, checked before money is spent on them
Structural rules block always. A verdict names the run, the file and the file's hash, and a run that requires gates re-checks all three.
- Held-out data
57 measurement harnesses
External labelled corpora and re-runnable measurements. This layer turned 100% into 87.7% on day one.
- Review
Two independent AI reviewers
A Claude review and a Codex adversarial review found 10 defects behind 608 passing tests, then 9 more behind 650.
- Reading
A person reading samples
Hand counts are the reported number wherever an automated classifier failed against hand labels. A targeted read of one marginal number found 93 of 309 accepted triplets resting on reference lists.
Before every commit
.venv/bin/ruff check . && .venv/bin/ruff format --check .
.venv/bin/mypy clinical_preprocessor clinical_dataset_builder scripts
A commit is checked in a separate worktree of its own tree, so a concurrent session's uncommitted work can neither pass nor fail it.
Stage gates
Three stages are gated: the read, curation and verification. A gate reads the stage's artifact, sets aside a seeded sample for a person to read, and writes its verdict beside the artifact. Exit code 0 means passed, 1 failed, 2 could not run. A refusal is recorded too, because "gated and refused" and "never gated" ask the operator for different things.
| Structural rule | What it catches |
|---|---|
chunk_ids_derive_from_their_own_content | A record that joins to nothing while every count stays right |
every_read_still_hashes_to_what_it_wrote | An artifact changed after it was written |
the_accounting_reconciles | Work the ledger says was done that the artifact does not hold |
the_stage_produced_something | An empty artifact, which satisfies every other rule vacuously |
| Hard negatives are exactly the approved audit | A surplus negative or a lost approval, asked in both directions |
| No negative restates its own positive | A "negative" that is the right answer again |
| The published file holds every accepted outcome once | A run that died writing the dataset, or a resume that never rebuilt it |
Thresholds are different: they are claims about a corpus, and they ship unset. Unset means "not asserted", never "no limit", and an unset threshold still reports what it observed. An unmeasured threshold that blocks teaches an operator to bypass the gate.
Measurement harnesses
| Area | Harnesses |
|---|---|
| Assertion | measure_assertion_heldout on 2,376 externally labelled pairs; measure_assertion on the self-authored set, kept as a demonstration |
| Linking and ties | measure_link_precision, measure_rerank_ties, measure_tie_degeneracy, measure_tie_break_yield, and measure_msh_wsd_verdict, the one-shot external check |
| Reading pages | measure_body_type_size, measure_heading_attribution, measure_outline_collisions, measure_reference_matter, measure_sentence_boundaries, measure_pua_glyphs, measure_axis_runs |
| Tables and figures | measure_flatten_refusals, measure_table_bounds, measure_visual_yield, measure_vision_analysis, measure_plot_vocabulary |
| Questions | measure_mcq_recovery against question banks and prose controls |
| Judges and generation | measure_judge_arms, measure_judge_batching, measure_generator_candidates, measure_label_coverage, measure_model_candidates |
| Retrieval | measure_retrieval_arms, measure_embedders, measure_query_gates, measure_cross_document_grounding |
| Cost and capacity | measure_generation_concurrency, measure_service_capacity, measure_split_cost, measure_read_vs_curate |
What the suite missed
Every item below passed the full test suite at the time. The layer that caught it is named.
| Defect | Caught by |
|---|---|
| Enumerated negation flipped to positive; negation crossing a section heading | Held-out clinical notes |
| 19 defects across two rounds, 6 of them introduced by the first round's fixes | Independent review |
| A 92.9 MB ledger for 30 units, and a 37 GB curation stall | The first real corpus runs |
| A store read silently capped at 10,000 of 49,276 units | Scale |
| One NUL byte stopping a graph ingest at 48.6% | Scale |
| 60.7% of accepted work missing from the published file | A gate rule that reads the published file itself |
| 93 of 309 accepted triplets resting on reference lists | Reading one marginal number |
| Explanations attached to the wrong question in 3 of 4 records of one file | A hand-read set, then a model audit |
The lessons behind each of these are written up on the Lessons page.