FICTA · MERIDIAN IS A SYNTHETIC INSTITUTION — EVERY NAME, CLIENT, TRADE AND NUMBER IS INVENTED
MeridianCapital Markets Operations

A bank that exists to be measured.

Meridian runs the full operational day of a capital-markets institution — trades booked, allocated, reconciled, margined, settled and released — entirely in fiction. It is the world inside a sealed evaluation environment: coherent enough to work in, synthetic enough to publish.

Operations desk

Trade lifecyclecapture through confirmation, every state transition receipted
Booking & allocationsblock-to-account splits under allocation contracts
Reconciliationledger-to-ledger breaks investigated to closure
Collateral & margincalls issued, disputed, resolved against agreements
Settlement exceptionsfails chased across depots and deadlines
Permissions & releasefour-eyes approvals; nothing self-released
Temporal causalityevery action ordered against what was knowable when
Evidence sufficiencyclaims stand only on retrievable, versioned evidence

The desk, mid-morning

09:04 booking block trade split 4 ways · allocation contract v3 · applied 09:12 reconcile ledger break, 2 legs · evidence attached · escalated to supervisor 09:31 collateral margin call issued · counterparty dispute window open 09:47 settlement fail aged 2 days · buy-in warning drafted · pending approval 10:02 permissions release request · preparer ≠ approver enforced · approved 10:18 lifecycle amendment on confirmed trade · stale version rejected ✓

Every line above is generated fiction — and every line is the kind of decision an AI agent is graded on, inside.

Why Meridian exists

Frontier AI agents are being sent into operational finance. Whether they act correctly — complete the work, respect authorization, refuse shortcuts, touch nothing outside contract — has to be measured somewhere consequences can't reach. Meridian is that somewhere. The Lab holds the instruments.

The instruments behind the institution.

A sealed, preregistered agent-evaluation environment: hash-sealed episodes, contracted tool surfaces, armed reward-hacking gates, and an evidence chain that binds every reported number to run bytes.

Headline results · harness v2, preregistered

27 / 40frontier-model completions in ≤24 turns — unsaturated
39 / 40open-weight arm — the spread is the instrument's resolution
0gate trips: no hidden-field probes, no canary echoes, any arm
243 / 243sealed adversarial mutations rejected with expected errors

The evidence chain

Runs are preregistered; the scripted honest baseline must pass 40/40 before any model result is reported. Configs, transcripts, metrics and reports bind to a manifest by SHA-256 — change a byte, invalidate the report. Independence from any real corpus is measured at release (shingle-overlap gate: zero overlaps), not asserted.

pip install -e .            # github.com/thefazzer/cleanroom-eval
python -m cleanroom_eval.contract verify-set
python -m cleanroom_eval.free_run --policy chat --out runs --run-id my-model

Working with the lab

TierDeliverable
S1Private hash-sealed episode set — fresh worlds, namespaces, canaries, mutation battery; per-customer exclusive
S2Evaluation run of your model on a private set, delivered as a hash-bound evidence pack
S3Custom operational families on the same contract-and-gates machinery
S4Generator licence for sealed-set minting at lab scale

Public sample is MIT. Contact: open an issue at github.com/thefazzer/cleanroom-eval.