From a written test plan to patients, with evidence.

You declare the situations your software must handle. The engine constructs a population that contains them, checks independently which ones it achieved, and records what is in the data, what is not, and what nobody tested.

The seven steps

  1. 1

    Declare

    The situations to test are written as a specification in YAML: each scenario, how many patients need it, and the checkable conditions that define it.

    You getA versioned test plan

  2. 2

    Plan

    Before building anything, the engine works out how many patients go to each scenario, to each combination of scenarios, and to background fill.

    You getA preview of the run

  3. 3

    Construct

    Patients are built to satisfy the scenarios: demographics, encounters, conditions, lab results, medications and allergies, linked with dates in a valid order. Values come from stated distributions using a recorded random seed.

    You getA synthetic population

  4. 4

    Verify

    Separate code checks each scenario against the finished population. An automated test forbids it from importing the construction code. Each scenario is reported as met, short or not verifiable.

    You getA coverage report

  5. 5

    Measure

    Structure, internal consistency, clinical rule checks, subgroups and privacy are measured. Each measurement states its method, its assumptions and what it cannot detect. There is no combined score.

    You getAn evaluation report

  6. 6

    Evidence

    Everything is written into a Data Passport and a manifest that records every input, including its hash.

    You getA Passport and a reproducible record

  7. 7

    Use

    The population is exported as FHIR R4: one bundle for the population, and one file per patient as a test fixture. It can be sent to a FHIR endpoint you name.

    You getTest data your system can load

Also available at library level

These work in the engine today, outside the standard seven steps.

Controlled mutations
Make one recorded change to a patient (shift a date, change a lab value, drop a field, duplicate an entry) and get the before-and-after pair.
Failure minimisation
When a sequence of changes breaks your system, reduce it to the smallest remaining set that still reproduces the failure.
Data-quality modes
Build a population that should pass validation (valid), one made only of deliberate defects (adversarial), or both (mixed).
Declared missing data
Remove values at a stated rate: completely at random, dependent on another value, or missing together with another field.
Timeline view
See any patient’s record as snapshots over time.
Constructed vs incidental
For every scenario, how many patients were built for it and how many turned up by chance.

What you receive

FHIR R4 data
Patient, Encounter, Condition, Observation, MedicationRequest and AllergyIntolerance in a population bundle. Structure, cardinality, value formats, required codes and references are checked against the pinned R4 specification (not a full HL7 validator).
Test fixtures
One bundle per patient, as files your test suite can load.
Coverage report
Each declared scenario marked met, short or not verifiable, with how many patients were built for it and how many turned up by chance.
Data Passport
Eleven sections covering what is in the data, what is not, and what nobody tested. Its “Not tested” section is never empty.
Reproducible manifest
Every input to the run, with its hash. Rebuilding from the manifest produces the same population.
Data Passport excerpt, run of 27 Sep 2026
# Data Passport

## 1. Identity and intended use
- Generator: hde.compose.Composer/0.2-construct-fill-repair EVAL-0.1
- Run: RUN-745029d75136
- Specification: medication_safety v1
- Seed: 7
- Population: 3000 patients

## 9. Privacy measurements
- privacy.singling_out_rate = 0.0267 proportion of records unique on the stated quasi-identifiers
  - assumptions: those three fields are what an adversary would know in advance.
    That choice is a judgement and a different choice gives a different number
- privacy.exact_match_rate = NOT MEASURED
  - NaN here means NOT MEASURED, not zero risk and not a low value.

Inside the Data Passport

Eleven sections, in this order, in every Passport.

Explore each section
  1. Identity and intended use
  2. Scope and limits
  3. Fitness statement: validated for / not validated for / not tested
  4. Coverage report
  5. Downstream utility
  6. Statistical fidelity
  7. Structural validity
  8. Clinical consistency
  9. Privacy measurements
  10. Not tested (never empty, cannot be switched off)
  11. Manifest (every input, with hashes)

Structural and reproducibility evidence

From the 3,000-patient run of 27 Sep 2026 unless stated.

FHIR R4 validation issues, valid mode
0 at 3,000, 10,000 and 30,000 patients
FHIR R4 validation issues, mixed mode
60, caused on purpose by the declared data-integrity scenarios
Events in valid time order
100%
Identifier uniqueness
100%
References that resolve
99.46% in mixed mode; the rest are deliberate broken-reference scenarios
Events attached to an encounter
100%
Same inputs, run twice
Identical run ID and byte-identical population
Different seed
Different run ID
Rebuild from manifest
Inputs match; results match

Sentences every Passport carries, word for word

They show how we write about what the data does and does not establish.

This assessment measures coverage of a declared specification and the properties listed. It does not establish that the data is anonymous, that its use is lawful, or that any system tested against it is clinically safe.

This rule set is incomplete and covers only the named domain. A dataset can satisfy every rule applied here and still be clinically wrong in ways these rules do not check.

Not tested: attacks using auxiliary external datasets; attacks with access to generator parameters; disclosure through attributes not declared as quasi-identifiers; longitudinal structure beyond the tested columns; any risk arising after the data leaves the environment where it was assessed.

Scenario coverage from a real run is on the scenario library page.

Where this is going

In development Supermerco is building healthcare scenario infrastructure. You define the situations your software must handle; we construct synthetic patient trajectories with known clinical truth, the record your system is allowed to see, the workflow actions that happened, and any declared errors or delays. Each can be verified independently, used as data on its own, or run against your system and kept as a regression case.

See the roadmap

Bring the list of situations your software must handle.

We will turn it into a scenario specification and show you what a pilot would deliver. Or email us at hello@supermerco.com.

Talk to us about a pilot