Trust, shown rather than claimed
Each rule below is enforced by an automated test, and each number comes from a run we can reproduce. Where we have not measured something, we say so.
Enforced by tests, not by good intentions
- 01
The checker cannot see how the data was made
Verification and evaluation code are forbidden from importing construction code.
- 02
No combined quality score, anywhere
One number would hide the trade we make on purpose: we change the population so it covers your cases, which makes it look less like any source.
- 03
Every measurement declares what it cannot detect
It is a required field with no default.
- 04
A missing clinical fact is “unknown”, never “none”
The engine refuses to count a scenario that depends on an unsigned clinical fact.
- 05
Clinical rules live in versioned files
Each has a named reviewer and is never hidden in code, so a clinician can correct it.
- 06
Nothing leaves your machine
No telemetry and no network calls at runtime.
- 07
Every run is reproducible
From its manifest. The run ID is derived from the inputs, not from a clock.
- 08
Our wording is tested
An automated check fails the build if our output makes any of the legal or safety claims our policy forbids.
Engineering evidence
As of 27 Sep 2026.
- Automated tests
- 1,296 passing
- Current code, in normal and forced-ASCII locales.
- Continuous integration
- Green on Ubuntu, Windows and macOS
- Plus a security job, on the last push referenced (24 Sep 2026).
- Runtime dependencies
- 1
- PyYAML.
- Known vulnerabilities in locked packages
- 0
- OSV database, 18 packages, checked 27 Sep 2026.
- Network calls at runtime
- None by default
- The one exception is the explicit command that sends fixtures to an address you give it.
- Software bill of materials
- 19 components
- Checked against the lockfile in CI.
Your data
- No real patient data enters the system. Every patient is constructed from a specification and stated distributions. The engine does not take in customer patient records, and we do not offer to.
- Nothing leaves the machine. The engine sends no telemetry and makes no network calls while running. The one exception is the command that sends fixtures to an address you give it.
- This website stores what you send through the contact form so we can reply. Please do not include real patient information in it.
Privacy measurements, stated as measurements
From the 3,000-patient run. These are measurements under stated assumptions, not a verdict, and we do not turn them into words like “private” or “safe”.
| Measurement | Value | Note |
|---|---|---|
| Records unique on birth year, sex and condition count | 2.67% | A floor: more attributes can only raise it. |
| Exact-match rate | Not measured | Needs a real source dataset to compare against, and we deliberately use none. |
| Distance to closest record | Not measured | Same reason. |
| Membership inference | Not measured | Same reason. |
Benchmark status
No quotable result There is no valid comparative benchmark result we can publish, so we do not say that Supermerco beats any other tool.
Before an overnight run, we found flaws in the comparison as it was built:
- Vocabulary
- Synthea uses coding standards such as LOINC and RxNorm, while our scenario checker used plain-language names, so real cases were missed.
- Reading length
- A local AI judge could not read most Synthea records whole under its earlier settings.
- Population size
- Population sizes were unequal, which favoured the larger Supermerco arm.
- Against us
- The judge also could not see our deliberate broken-reference scenarios, which biased the other way.
In development The rebuilt comparison uses equal populations and compares specification satisfaction, exact, rare and multi-constraint coverage, temporal requirements, incidental contamination, verification, reproducibility, time to usable test data and downstream failure discovery.
How the numbers were produced
All from the command line on 27 Sep 2026, on a MacBook Pro (Apple M3, 8 GB), with outputs saved at the time. Timings from /usr/bin/time -l.
uv run hde generate --spec scenarios/healthcare/medication_safety.v1.yaml --seed 7 -n 3000
uv run hde generate ... --seed 7 -n 3000 # again: same run ID, same hash
uv run hde generate ... --seed 8 -n 3000 # different run ID
uv run hde verify --population <run>/population.json
uv run hde evaluate --population <run>/population.json
uv run hde reproduce --manifest <run>/manifest.json # inputs MATCH, results MATCH
uv run hde fixtures --population <run>/population.json --out <dir>
uv run hde generate ... -n 3000 | 10000 | 30000 --mode validWhat we do not claim
- The Data Passport does not establish that the data is anonymous, that its use is lawful, or that any system tested against it is clinically safe.
- Nothing has been signed by a clinician yet. Scenarios that depend on an unsigned clinical fact are counted as not verifiable, and 8 scenarios have no named clinical reviewer.
- We do not publish a benchmark comparison. No valid result exists yet.
- Sending fixtures to a FHIR server has been tested only against a stand-in, never against a real FHIR server.
- End-to-end runs are measured up to 30,000 patients. Above that, the FHIR export does not fit in 8 GB of memory.
- We construct declared test trajectories. We never predict what will happen to a real patient.
Questions about any of this? Ask us directly.
Bring the list of situations your software must handle.
We will turn it into a scenario specification and show you what a pilot would deliver. Or email us at hello@supermerco.com.