Calibration: we score our own predictions in public
What we grade, today
| Ledger | Type | Current state |
|---|---|---|
| 8 Situational Awareness verdicts | Categorical verdicts; 3 of 8 carry a written flip condition, 2 a resolution date, 1 a watch item, 2 a stated blocker (all in data.json) | 2 on track · 2 open · 2 pending · 1 exceeded · 1 wrong → Thesis Tracker 62.5/100 as of 2026-09-06 |
| SunWatch market-call ledger | Dated, falsifiable market calls | 8 scored, 5 hits (63%), 7 pending, ledger as of 2026-09-28. n=8 is small: the Wilson 95% interval is roughly 31–86%, so this is a work-in-progress sample, not proof of skill. |
| Red-team survival odds | Editorial probabilities on open calls | 6 open calls carry a stated probability (52–65%); confidence cuts are published the day counter-evidence lands. |
| Metaculus FutureEval bot | Probabilities on third-party questions, scored by Metaculus; each committed in public (hashed and sealed) when submitted, then timestamped in Bitcoin | 0 forecast(s) committed, as of 2026-09-27 |
| AGI consensus board | Third-party forecasts, not ours | Cross-venue median and spread, recomputable from the published snapshot; it is a reference we quote, not a call we are scored on. |
The number nobody wants to print: Brier-eligible n = 0
A Brier score needs a stated probability and an outcome. Of the 8 scored calls, 0
carried a probability in the ledger's odds field; 4 of them state a scenario percentage in prose,
which we do not Brier-score because a scenario weight is not a calibrated call. The 6 calls that do carry a
structured probability are still open. So the honest reading is: the hit rate above is real, the calibration curve does not
exist yet, and it cannot exist until 20 probability-bearing calls have resolved.
The commitment
When the pool of scored probability calls reaches n≥20, this page publishes a Brier score and a calibration curve (stated probability vs realized frequency), recomputed on every ledger change — the same way the Thesis Tracker recomputes on every verdict change. The commitment is itself pre-registered as a dated line in the public bet ledger of this site's repository; if the pool never gets there, the line settles as insufficient and says so here. The forecast ledger is public and timestamped, so anyone can compute it before we do.
Timestamp proofs: which versions can no longer be backdated
Versions of the records below stamped since 2026-09-26 carry an OpenTimestamps proof: the file's SHA-256 is submitted to public calendars and, once their aggregate transaction confirms, sits in a Bitcoin block. Status today: 9 version(s) listed, 8 with a Bitcoin block attestation, 1 pending (calendar receipt only, not in a block yet). A proof shows that those exact bytes existed no later than the block it names; it says nothing about whether a verdict is right, and versions before 2026-09-26 rest on the public git log only. Index: /ots/manifest.json.
| Record | Versions | Latest stamped | Status | Proof |
|---|---|---|---|---|
| agi-consensus.json | 2 | 2026-09-26 | bitcoin | agi-consensus.json.c327ae396014.ots |
| data.json | 1 | 2026-09-26 | bitcoin | data.json.12d283936aaf.ots |
| independent-grades.json | 2 | 2026-09-27 | pending | independent-grades.json.796c1e3e43a5.ots |
| index-history.json | 1 | 2026-09-26 | bitcoin | index-history.json.89363c1abcda.ots |
| market-board.json | 2 | 2026-09-26 | bitcoin | market-board.json.aec19e4fc7de.ots |
| odds-history.json | 1 | 2026-09-26 | bitcoin | odds-history.json.a5903c50e0c5.ots |
Verify: ots verify <proof> -f <file> (client or web verifier at opentimestamps.org); pick the manifest entry whose sha256 matches the file you downloaded. We issue no coin, hold no coin and pay nothing; the calendars are free public services.
Check them yourself, in this browser
“Don't trust, verify” only builds trust if verifying is cheap. The button below runs the first half of the command-line
check with nothing to install: it downloads the manifest and every record it lists, hashes each
one with your browser's own SHA-256 (Web Crypto), and looks that hash up among the listed versions; for a match it also downloads
the proof and checks that the proof file commits to that same hash. It then recomputes the Thesis Tracker from
/data.json with the published weights (On track, Exceeded, Holding = 1; Open, Pending = 0.5; Wrong = 0; mean ×100) and compares it with the score the same file
publishes. Nothing is fetched until you click. Those downloads are tagged ?utm_source=verify and are not counted as
visits to the files; the run leaves one event with no identifiers (how many records matched, whether the score agreed).
What a match proves: the bytes served to you are the bytes a listed proof was made for, and once that proof has a
Bitcoin attestation, those bytes existed no later than that block. What it does not prove: that any verdict is right,
that a source was read correctly, or that nothing was left out; an agreeing recompute only shows the score is the stated arithmetic
of the verdicts. The button does not walk the proof up to a Bitcoin block header: the verifier at opentimestamps.org does, and so
does ots verify on a machine with access to a Bitcoin node. A record changed since the last daily stamp is reported as
“not among the listed versions yet”, not as a failure. These files are served exactly as committed (checked 2026-09-27: neither the deploy
build nor the edge worker rewrites them), so a stamped version should match byte for byte.
The forecast ledger: sealed before the question closes, opened after
The one place this network states probabilities that someone else scores is its bot in Metaculus's
FutureEval bot tournament. Metaculus keeps what the bot submits;
it does not keep what the bot would have said without this site's own research as a prior, and that counterfactual is the only
test of whether the site's judgement helps. So every submitted forecast is committed in public — and, for binary AI questions where the
prior was used (up to five a run), so is a shadow forecast made without it: a line in data/metaculus/forecasts.jsonl
carries the SHA-256 of the sealed forecast plus a random 256-bit nonce (without the nonce, a probability on a 1–99 grid could be recovered
from its hash in milliseconds — we measured it on the first version and fixed it before the bot ever ran). The ledger only grows. In the run
that writes a new version, that version is submitted to the public OpenTimestamps calendars and its byte length and our runner's clock time
are recorded; a Bitcoin block attests it later, usually within hours, and only that block's time is proof-backed. Every
earlier version must remain an exact prefix of today's file. After a question closes, the exact sealed text is published in
revealed.jsonl: anyone can hash it, find the matching line, and see which proof covers it — the line provably existed before the
close only if that proof's block came before the close; the runner's own stamp time is shown too, labelled as self-reported. The whole check is
one standard-library script, tools/fleet/verify_commitments.py, run daily by the repository's
heartbeat, which turns red on any rewritten line or mismatched reveal.
State as of 2026-09-27: the bot has not submitted a forecast yet; 0 forecast line(s) committed, 0 with a shadow; house-prior comparison on resolved questions: n = 0. A Brier comparison appears here only once questions resolve; until then this section says exactly this.
Why this page exists
Every AI-era answer engine can generate confident takes; almost none can show you a scored history. Being auditable — misses kept on the page next to hits, flip conditions registered before outcomes, probabilities graded against reality, new versions timestamped so they cannot be quietly rewritten — is this network's entire moat. This page is that moat made explicit, including the part where the sample is still too small.
Frequently asked questions
A measure of probability-forecast accuracy: the mean squared difference between stated probabilities and outcomes (0 = perfect, 0.25 = coin-flip guessing on binary events). We publish ours once scored probability calls reach n≥20.
Because the Brier-eligible sample is 0: the 8 scored calls were graded hit/miss without a structured probability, and the 6 probability-bearing calls have not resolved. Publishing a curve from that would be theater. The raw ledger is public and timestamped, so nothing is hidden in the meantime.
Outcomes are graded against pre-registered falsification conditions written before the outcome, with dated multi-source verification, and misses stay published with their lesson. The grading rules are public in the eight-layer method, including the red-team layer.
Versions of agi-consensus.json, data.json, independent-grades.json, index-history.json, market-board.json, odds-history.json stamped since 2026-09-26 carry an OpenTimestamps proof under /ots/ (the manifest lists each proof's status: pending until the calendar's transaction is in a Bitcoin block). The “Verify these records in your browser” button on this page downloads each listed file, hashes it locally and checks the hash against the manifest and the proof file; to follow a proof to its Bitcoin block, run the free client against the entry whose sha256 matches your download. That proves when those bytes existed, not that they are correct; earlier history rests on the public git log.
The live scorecard updates as models ship and verdicts change.
View the live scorecard →