How it works

← The ledger

The experiment tests one question: does an explicit model of how the world works add forecasting skill beyond the AI that runs it? Everything on this site exists to make that answerable in public, where it can’t be quietly abandoned.

The world events, releases, filings, incidents Observation live web reads — dated, sourced items per watch area The instruments (sealed) six independent world-models, each reading observations through its own mechanisms A B C D F G Instrument E the composite: sees all six, arbitrates conflicts, covers seams Predictions falsifiable claim · written resolution criteria · resolve-by date instrument probability · mechanism (sealed) No-model baseline same LLM, same claim — no instrument reasoning in context Pre-registration committed to git before resolution — tamper-evident hashes, append-only Resolution & grading after resolve-by: hit / partial / miss against the written criteria — reasoning unsealed on grading Scoreboard instrument vs baseline Brier, per instrument and pooled — hits and misses alike learnings feed back into the instruments

The prediction logic, step by step

1 · Observation

Each instrument has watch areas. On each round, current events in those areas are gathered from the live web as dated, sourced items — the raw material every instrument sees.

2 · The instruments

Six independent world-models (A–D, F, G), each with its own mechanisms, read the observations in isolation — no instrument sees another’s reasoning. A seventh, Instrument E, is the composite: it alone sees all six and their open predictions, and it specializes in three moves — arbitrating head-on disagreements, predicting in the seams none of the six covers, and hedging assumptions they all share. The instruments’ identities and mechanisms are sealed until grading.

3 · Predictions

Each round yields falsifiable claims: a concrete observable, written resolution criteria, a resolve-by date, and the instrument’s probability. Vague trends and near-certain base-rate events are banned by the quality bar — a good prediction is one a smart skeptic might bet against.

4 · The baseline

Every prediction is paired with a no-model baseline: the same underlying LLM, given the same claim and criteria, with none of the instrument’s reasoning in context. If the instruments are just decoration, instrument and baseline scores will match. The delta (Δ) shown on each card is where the instruments stick their necks out.

5 · Pre-registration

Predictions are committed to a git repository before resolution; the commit hashes are printed on the ledger. Entries are append-only — once registered, a prediction can resolve, but it can never be edited or deleted.

6 · Grading

After each resolve-by date, every due claim is graded hit / partial / miss against its written criteria (unresolvable claims are excluded from scoring). Graded predictions move to Past predictions together with what happened and the instrument’s unsealed reasoning — the seal exists precisely so the reveal can be checked against the pre-registered hash.

7 · The scoreboard & the loop

Brier scores (lower is better) are published instrument-vs-baseline, per instrument and pooled — hits and misses alike. Each miss produces a recorded learning that feeds back into the instrument for its next round: the instruments are versioned and expected to improve on the record, not by claim.

Not investment or safety advice. The site predicts observables, never allocations or actions.