What we are building next

Supermerco is building healthcare scenario infrastructure. You define the situations your software must handle; we construct synthetic patient trajectories with known clinical truth, the record your system is allowed to see, the workflow actions that happened, and any declared errors or delays. Each can be verified independently, used as data on its own, or run against your system and kept as a regression case.

In development Nothing on this page is shipped unless it is marked Today.

Programmable healthcare test worlds with verifiable ground truth.

We do not aim to win generic synthetic data or generic AI evaluation. We aim to own the healthcare reality those systems are tested against.

The chart is not necessarily the patient

Each synthetic patient will carry three separate realities, and a scenario can make them disagree on purpose.

  • Clinical truth

    What is actually true in the synthetic world: the medicine taken, the allergy, the condition, whether the patient takes the dose.

  • Recorded state

    What the EHR or downstream system shows. It can be incomplete, stale, delayed, duplicated, contradictory, wrongly linked or simply wrong.

  • Workflow state

    What people and systems actually did: ordered, performed, returned, reviewed, changed, handed off, followed up.

Same patient · same seed · identical history until 17:45

09:00blood collected10:00result produced10:03visible in record17:45review time18:10action timeLaterdeclared branch
Clinical truthPotassium 6.2Medicine stoppedStable
Recorded stateResult producedResult visibleChange recorded
WorkflowBlood collectedResult reviewedMedicine changedFollow-up arranged
In developmentIllustration of a counterfactual twin. Exactly one declared divergence; every later difference follows from declared rules, and the pair is verified with a machine-readable diff. A declared branch, not a prediction.

Human and process error, as data

Not random corruption. Each error is a declared object with a cause, a consequence and a place in time.

Documentation errors

  • Allergy omitted
  • Stale medication list
  • Wrong dose recorded
  • Wrong encounter association

Clinical-process failures

  • Monitoring not ordered
  • Result available but never reviewed
  • Missed follow-up
  • Handoff omission

Patient behaviour

  • Medication non-adherence
  • Wrong dose taken
  • Lab not completed
  • Late presentation

Integration and system failures

  • Result delayed
  • Duplicate message
  • Wrong-patient reference
  • Change not propagated

Every error records its type, actor, event time, visible time, affected state, declared cause, declared consequence, detection, correction, provenance and clinical-review status.

Four ways in, without buying everything

Start with data. Add coverage, error trajectories or regression when you need them.

A

Scenario data

Today

For teams that only need data.

  1. requirement
  2. specification
  3. construct
  4. verify
  5. FHIR and fixtures
  6. Data Passport
B

Coverage

In development

For teams with existing test data.

  1. existing suite
  2. coverage analysis
  3. gaps
  4. construct missing cases
  5. merged, verified suite
C

Clinical reality and error trajectories

In development

For advanced teams.

  1. clinical truth
  2. record state
  3. workflow state
  4. controlled divergence
  5. counterfactual twin
D

Regression

In development

For a continuous testing workflow.

  1. scenario
  2. your system
  3. failure
  4. minimisation
  5. regression case
  6. rerun on each release

Milestones

In order. None of them is shipped yet.

  1. Milestone A

    Clinical reality model

    Clinical truth, recorded state and workflow state as separate models, with multiple timestamps, first-class error events and an independent verifier.

  2. Milestone B

    Five deep failure trajectories

    True allergy absent from the chart; medication discontinued but stale in the record; result produced but never reviewed; handoff information loss; medication recorded active while the patient does not take it.

  3. Milestone C

    Counterfactual twins

    For every supported error: a baseline and an error branch with identical history before the divergence, an exact diff, and independently verified lineage.

  4. Milestone D

    An external customer system

    Run trajectories against a real external AI or software product and find a failure the customer did not already have as a regression case.

  5. Milestone E

    Continuous regression

    Keep customer-private scenario suites and rerun them across versions.

  6. Later

    Deeper clinical trajectories

    Complications, recovery branches and multi-year progression, only once clinical-review capacity exists.

The first trajectory families

Medication safety and data integrity first: a few deep, high-value scenarios rather than hundreds of shallow ones.

  1. True allergy missing from the record
  2. Medication prescribed despite a documented allergy
  3. Medication discontinued but still active in the record
  4. Medication recorded active but the patient does not take it
  5. Monitoring not ordered
  6. Monitoring ordered but not performed
  7. Result available but not reviewed
  8. Abnormal result followed by delayed action
  9. Medication or allergy information lost during a handoff
  10. Two systems disagree about the patient’s current state
  11. Wrong encounter or reference linkage
  12. Duplicate patient, or demographic drift
  13. Stale or contradictory medication history
  14. Delayed documentation versus actual event time

Your eval tools run the test. We supply the healthcare world.

In development Generic agent and evaluation tools can already run scenarios, call a system and grade the answer. We would rather work with them than against them: your tools or agents call Supermerco for the healthcare scenario, its hidden ground truth and the evidence, and run it against your system as they do today.

  1. Your eval tools or AI agents
  2. Supermerco: the healthcare test world
  3. Your system

The workflow we are working toward

Ten steps. 2 exist today, 2 exist only at library level, 1 is in beta, and 5 are in development.

  1. 1

    Define the healthcare situations the system must handle

    Scenario specification in YAML.

    Today
  2. 2

    Measure which situations the existing test population already covers

    Test-coverage scanner.

    In development
  3. 3

    Identify missing and under-covered situations

    Test-coverage scanner.

    In development
  4. 4

    Construct the missing cases

    Today we construct whole populations; gap-only construction is in development.

    In development
  5. 5

    Independently verify the resulting coverage

    Independent verifier.

    Today
  6. 6

    Run the cases against your AI or software

    Fixture sender, tested only against a stand-in.

    Beta
  7. 7

    Identify and attribute failures

    Oracle contract and Test Passport.

    In development
  8. 8

    Use controlled counterfactuals to isolate the cause

    Mutations exist at library level, not yet in the workflow.

    Library level
  9. 9

    Save the failure as a reproducible regression case

    Failure minimisation exists at library level.

    Library level
  10. 10

    Rerun the regression suite on future releases

    Regression suites.

    In development

Today, and next

Area by area.

AreaTodayIn development
Data constructionScenario-driven synthetic patient construction.Coverage-gap construction against your existing test fixtures.
VerificationIndependent scenario verification and structured measurements.A framework for judging system behaviour (oracles).
EvidenceData Passport and manifest.Test Passport with failures and regression lineage.
System executionA fixture sender exists, tested only against a stand-in.Real FHIR endpoint and model integrations.
CounterfactualsControlled mutations exist at library level.Counterfactual twins: one declared divergence, an exact diff and verified lineage.
Clinical realityNot modelled separately yet.Separate clinical truth, recorded state and workflow state, with multiple clocks.
RegressionFailure minimisation exists at library level.Stored regression suites, rerun on each release.
AI modelThe core product does not require an AI model.Possibly a self-hosted model for assisted language tasks.
InterfaceCommand line.Other interfaces only once customer use validates them.
Clinical contentSourced but unsigned, so partly not verifiable.Clinician-reviewed rules with provenance.
BenchmarkNo quotable comparative result.A fair external benchmark, published whichever way it falls.

What we are working on now

The near-term work that changes what this site is allowed to say.

  1. Find a clinician or clinical pharmacist reviewerUnblocks the 5 not-verifiable scenarios, the unsigned drug knowledge and the inactive clinical rules.
  2. Run against a real FHIR serverTurns guided test execution from stand-in-only beta into external integration evidence.
  3. Build one external system-under-test adapterNeeded for a real scenario, failure and regression demonstration.
  4. Wire mutations and failure minimisation into test executionMoves library capability into a usable workflow.
  5. Build the external test-coverage scannerCreates the coverage-guided workflow.
  6. Build the Test Passport and regression artifactCaptures evidence about your system, not only about the data.

Where an AI model might fit, and where it never will

Proposed direction Paid third-party AI services will not be a hard runtime dependency of the product core. Frequent scenario analysis, counterfactual exploration and system testing could become expensive and dependent on third-party costs, quotas, latency and availability. The preferred direction is a strong self-hosted or fine-tuned model for assisted tasks, if it meets measured quality requirements.

Stays deterministic

  • Scenario constraints and generation planning
  • Patient construction
  • Independent verification
  • Coverage counts
  • FHIR structural checks
  • Evidence and manifest hashing
  • Reproducibility
  • Regression case identity

Candidate assisted tasks

  • Turn a plain-language requirement into a draft scenario specification for human review
  • Propose terminology and code mappings, with deterministic validation and reviewer approval
  • Summarise or classify outputs from a system under test
  • Cluster similar failures after deterministic evidence is recorded
  • Suggest counterfactuals or next tests, with the change and its verification kept deterministic
  • Write explanatory text for a Test Passport from structured results

A model is never

  • The sole verifier of whether a patient satisfies the specification
  • A substitute for clinician review of clinical rules
  • The source of a regulatory, compliance or safety conclusion
  • An untracked clinical oracle whose output becomes ground truth
  • A required network service for basic generation or reproduction

Before any model is chosen, candidates are benchmarked on our exact workload: accuracy, structured output, latency, memory and cost.

What we are deliberately not building

Every substantial feature has to make the healthcare test world more valuable, or be required by a real customer.

  • A general-purpose AI evaluation or red-teaming tool
  • A universal digital twin, or a simulator of every disease
  • Predictions of what will happen to real patients
  • Ingestion of production patient data
  • Regulatory sign-off services
  • Arbitrary “realism” scores

Bring the list of situations your software must handle.

We will turn it into a scenario specification and show you what a pilot would deliver. Or email us at hello@supermerco.com.

Talk to us about a pilot