← Live scorecard
Accountability

Calibration: we score our own predictions in public

Last updated: September 28, 2026 · Updated as verdicts change
We score our own predictions in public — and the sample is still small. This page inventories every probability-shaped claim the AGI Scorecard network makes, states how many of them can actually be Brier-scored today (0), and pre-commits to publishing a Brier score and calibration curve once scored probability calls reach n≥20. We would rather show a small honest n than a big fake curve.

What we grade, today

LedgerTypeCurrent state
8 Situational Awareness verdictsCategorical verdicts; 3 of 8 carry a written flip condition, 2 a resolution date, 1 a watch item, 2 a stated blocker (all in data.json)2 on track · 2 open · 2 pending · 1 exceeded · 1 wrong → Thesis Tracker 62.5/100 as of 2026-09-06
SunWatch market-call ledgerDated, falsifiable market calls8 scored, 5 hits (63%), 7 pending, ledger as of 2026-09-28. n=8 is small: the Wilson 95% interval is roughly 31–86%, so this is a work-in-progress sample, not proof of skill.
Red-team survival oddsEditorial probabilities on open calls6 open calls carry a stated probability (52–65%); confidence cuts are published the day counter-evidence lands.
Metaculus FutureEval botProbabilities on third-party questions, scored by Metaculus; each committed in public (hashed and sealed) when submitted, then timestamped in Bitcoin0 forecast(s) committed, as of 2026-09-27
AGI consensus boardThird-party forecasts, not oursCross-venue median and spread, recomputable from the published snapshot; it is a reference we quote, not a call we are scored on.

The number nobody wants to print: Brier-eligible n = 0

A Brier score needs a stated probability and an outcome. Of the 8 scored calls, 0 carried a probability in the ledger's odds field; 4 of them state a scenario percentage in prose, which we do not Brier-score because a scenario weight is not a calibrated call. The 6 calls that do carry a structured probability are still open. So the honest reading is: the hit rate above is real, the calibration curve does not exist yet, and it cannot exist until 20 probability-bearing calls have resolved.

The commitment

When the pool of scored probability calls reaches n≥20, this page publishes a Brier score and a calibration curve (stated probability vs realized frequency), recomputed on every ledger change — the same way the Thesis Tracker recomputes on every verdict change. The commitment is itself pre-registered as a dated line in the public bet ledger of this site's repository; if the pool never gets there, the line settles as insufficient and says so here. The forecast ledger is public and timestamped, so anyone can compute it before we do.

Timestamp proofs: which versions can no longer be backdated

Versions of the records below stamped since 2026-09-26 carry an OpenTimestamps proof: the file's SHA-256 is submitted to public calendars and, once their aggregate transaction confirms, sits in a Bitcoin block. Status today: 9 version(s) listed, 8 with a Bitcoin block attestation, 1 pending (calendar receipt only, not in a block yet). A proof shows that those exact bytes existed no later than the block it names; it says nothing about whether a verdict is right, and versions before 2026-09-26 rest on the public git log only. Index: /ots/manifest.json.

RecordVersionsLatest stampedStatusProof
agi-consensus.json22026-09-26bitcoinagi-consensus.json.c327ae396014.ots
data.json12026-09-26bitcoindata.json.12d283936aaf.ots
independent-grades.json22026-09-27pendingindependent-grades.json.796c1e3e43a5.ots
index-history.json12026-09-26bitcoinindex-history.json.89363c1abcda.ots
market-board.json22026-09-26bitcoinmarket-board.json.aec19e4fc7de.ots
odds-history.json12026-09-26bitcoinodds-history.json.a5903c50e0c5.ots

Verify: ots verify <proof> -f <file> (client or web verifier at opentimestamps.org); pick the manifest entry whose sha256 matches the file you downloaded. We issue no coin, hold no coin and pay nothing; the calendars are free public services.

Check them yourself, in this browser

“Don't trust, verify” only builds trust if verifying is cheap. The button below runs the first half of the command-line check with nothing to install: it downloads the manifest and every record it lists, hashes each one with your browser's own SHA-256 (Web Crypto), and looks that hash up among the listed versions; for a match it also downloads the proof and checks that the proof file commits to that same hash. It then recomputes the Thesis Tracker from /data.json with the published weights (On track, Exceeded, Holding = 1; Open, Pending = 0.5; Wrong = 0; mean ×100) and compares it with the score the same file publishes. Nothing is fetched until you click. Those downloads are tagged ?utm_source=verify and are not counted as visits to the files; the run leaves one event with no identifiers (how many records matched, whether the score agreed).

What a match proves: the bytes served to you are the bytes a listed proof was made for, and once that proof has a Bitcoin attestation, those bytes existed no later than that block. What it does not prove: that any verdict is right, that a source was read correctly, or that nothing was left out; an agreeing recompute only shows the score is the stated arithmetic of the verdicts. The button does not walk the proof up to a Bitcoin block header: the verifier at opentimestamps.org does, and so does ots verify on a machine with access to a Bitcoin node. A record changed since the last daily stamp is reported as “not among the listed versions yet”, not as a failure. These files are served exactly as committed (checked 2026-09-27: neither the deploy build nor the edge worker rewrites them), so a stamped version should match byte for byte.

The forecast ledger: sealed before the question closes, opened after

The one place this network states probabilities that someone else scores is its bot in Metaculus's FutureEval bot tournament. Metaculus keeps what the bot submits; it does not keep what the bot would have said without this site's own research as a prior, and that counterfactual is the only test of whether the site's judgement helps. So every submitted forecast is committed in public — and, for binary AI questions where the prior was used (up to five a run), so is a shadow forecast made without it: a line in data/metaculus/forecasts.jsonl carries the SHA-256 of the sealed forecast plus a random 256-bit nonce (without the nonce, a probability on a 1–99 grid could be recovered from its hash in milliseconds — we measured it on the first version and fixed it before the bot ever ran). The ledger only grows. In the run that writes a new version, that version is submitted to the public OpenTimestamps calendars and its byte length and our runner's clock time are recorded; a Bitcoin block attests it later, usually within hours, and only that block's time is proof-backed. Every earlier version must remain an exact prefix of today's file. After a question closes, the exact sealed text is published in revealed.jsonl: anyone can hash it, find the matching line, and see which proof covers it — the line provably existed before the close only if that proof's block came before the close; the runner's own stamp time is shown too, labelled as self-reported. The whole check is one standard-library script, tools/fleet/verify_commitments.py, run daily by the repository's heartbeat, which turns red on any rewritten line or mismatched reveal.

State as of 2026-09-27: the bot has not submitted a forecast yet; 0 forecast line(s) committed, 0 with a shadow; house-prior comparison on resolved questions: n = 0. A Brier comparison appears here only once questions resolve; until then this section says exactly this.

Why this page exists

Every AI-era answer engine can generate confident takes; almost none can show you a scored history. Being auditable — misses kept on the page next to hits, flip conditions registered before outcomes, probabilities graded against reality, new versions timestamped so they cannot be quietly rewritten — is this network's entire moat. This page is that moat made explicit, including the part where the sample is still too small.

Frequently asked questions

What is a Brier score?

A measure of probability-forecast accuracy: the mean squared difference between stated probabilities and outcomes (0 = perfect, 0.25 = coin-flip guessing on binary events). We publish ours once scored probability calls reach n≥20.

Why not publish a calibration curve now?

Because the Brier-eligible sample is 0: the 8 scored calls were graded hit/miss without a structured probability, and the 6 probability-bearing calls have not resolved. Publishing a curve from that would be theater. The raw ledger is public and timestamped, so nothing is hidden in the meantime.

Who grades the calls?

Outcomes are graded against pre-registered falsification conditions written before the outcome, with dated multi-source verification, and misses stay published with their lesson. The grading rules are public in the eight-layer method, including the red-team layer.

How do I check that a record was not backdated?

Versions of agi-consensus.json, data.json, independent-grades.json, index-history.json, market-board.json, odds-history.json stamped since 2026-09-26 carry an OpenTimestamps proof under /ots/ (the manifest lists each proof's status: pending until the calendar's transaction is in a Bitcoin block). The “Verify these records in your browser” button on this page downloads each listed file, hashes it locally and checks the hash against the manifest and the proof file; to follow a proof to its Bitcoin block, run the free client against the entry whose sha256 matches your download. That proves when those bytes existed, not that they are correct; earlier history rests on the public git log.

The live scorecard updates as models ship and verdicts change.

View the live scorecard →