From a written test plan to patients, with evidence.
You declare the situations your software must handle. The engine constructs a population that contains them, checks independently which ones it achieved, and records what is in the data, what is not, and what nobody tested.
The seven steps
- 1
Declare
The situations to test are written as a specification in YAML: each scenario, how many patients need it, and the checkable conditions that define it.
You getA versioned test plan
- 2
Plan
Before building anything, the engine works out how many patients go to each scenario, to each combination of scenarios, and to background fill.
You getA preview of the run
- 3
Construct
Patients are built to satisfy the scenarios: demographics, encounters, conditions, lab results, medications and allergies, linked with dates in a valid order. Values come from stated distributions using a recorded random seed.
You getA synthetic population
- 4
Verify
Separate code checks each scenario against the finished population. An automated test forbids it from importing the construction code. Each scenario is reported as met, short or not verifiable.
You getA coverage report
- 5
Measure
Structure, internal consistency, clinical rule checks, subgroups and privacy are measured. Each measurement states its method, its assumptions and what it cannot detect. There is no combined score.
You getAn evaluation report
- 6
Evidence
Everything is written into a Data Passport and a manifest that records every input, including its hash.
You getA Passport and a reproducible record
- 7
Use
The population is exported as FHIR R4: one bundle for the population, and one file per patient as a test fixture. It can be sent to a FHIR endpoint you name.
You getTest data your system can load
Also available at library level
These work in the engine today, outside the standard seven steps.
- Controlled mutations
- Make one recorded change to a patient (shift a date, change a lab value, drop a field, duplicate an entry) and get the before-and-after pair.
- Failure minimisation
- When a sequence of changes breaks your system, reduce it to the smallest remaining set that still reproduces the failure.
- Data-quality modes
- Build a population that should pass validation (valid), one made only of deliberate defects (adversarial), or both (mixed).
- Declared missing data
- Remove values at a stated rate: completely at random, dependent on another value, or missing together with another field.
- Timeline view
- See any patient’s record as snapshots over time.
- Constructed vs incidental
- For every scenario, how many patients were built for it and how many turned up by chance.
What you receive
- FHIR R4 data
- Patient, Encounter, Condition, Observation, MedicationRequest and AllergyIntolerance in a population bundle. Structure, cardinality, value formats, required codes and references are checked against the pinned R4 specification (not a full HL7 validator).
- Test fixtures
- One bundle per patient, as files your test suite can load.
- Coverage report
- Each declared scenario marked met, short or not verifiable, with how many patients were built for it and how many turned up by chance.
- Data Passport
- Eleven sections covering what is in the data, what is not, and what nobody tested. Its “Not tested” section is never empty.
- Reproducible manifest
- Every input to the run, with its hash. Rebuilding from the manifest produces the same population.
# Data Passport
## 1. Identity and intended use
- Generator: hde.compose.Composer/0.2-construct-fill-repair EVAL-0.1
- Run: RUN-745029d75136
- Specification: medication_safety v1
- Seed: 7
- Population: 3000 patients
## 9. Privacy measurements
- privacy.singling_out_rate = 0.0267 proportion of records unique on the stated quasi-identifiers
- assumptions: those three fields are what an adversary would know in advance.
That choice is a judgement and a different choice gives a different number
- privacy.exact_match_rate = NOT MEASURED
- NaN here means NOT MEASURED, not zero risk and not a low value.- Identity and intended use
- Scope and limits
- Fitness statement: validated for / not validated for / not tested
- Coverage report
- Downstream utility
- Statistical fidelity
- Structural validity
- Clinical consistency
- Privacy measurements
- Not tested (never empty, cannot be switched off)
- Manifest (every input, with hashes)
Structural and reproducibility evidence
From the 3,000-patient run of 27 Sep 2026 unless stated.
- FHIR R4 validation issues, valid mode
- 0 at 3,000, 10,000 and 30,000 patients
- FHIR R4 validation issues, mixed mode
- 60, caused on purpose by the declared data-integrity scenarios
- Events in valid time order
- 100%
- Identifier uniqueness
- 100%
- References that resolve
- 99.46% in mixed mode; the rest are deliberate broken-reference scenarios
- Events attached to an encounter
- 100%
- Same inputs, run twice
- Identical run ID and byte-identical population
- Different seed
- Different run ID
- Rebuild from manifest
- Inputs match; results match
Sentences every Passport carries, word for word
They show how we write about what the data does and does not establish.
This assessment measures coverage of a declared specification and the properties listed. It does not establish that the data is anonymous, that its use is lawful, or that any system tested against it is clinically safe.
This rule set is incomplete and covers only the named domain. A dataset can satisfy every rule applied here and still be clinically wrong in ways these rules do not check.
Not tested: attacks using auxiliary external datasets; attacks with access to generator parameters; disclosure through attributes not declared as quasi-identifiers; longitudinal structure beyond the tested columns; any risk arising after the data leaves the environment where it was assessed.
Scenario coverage from a real run is on the scenario library page.
Bring the list of situations your software must handle.
We will turn it into a scenario specification and show you what a pilot would deliver. Or email us at hello@supermerco.com.