← Live scorecard
Accountability

Calibration: we score our own predictions in public

Last updated: June 30, 2026 · Updated as verdicts change
We score our own predictions in public — and the sample is still small. This page inventories every probability-shaped claim the AGI Scorecard network makes (graded verdicts, an investing forecast ledger, red-team survival odds) and pre-commits to publishing a Brier score and calibration curve once scored calls reach n≥20. Until then we show the raw ledger and refuse to claim we are calibrated. We would rather show a small honest n than a big fake curve.

What we grade, today

LedgerTypeCurrent state (as of 2026-08-08)
8 Situational Awareness verdictsCategorical verdicts with pre-registered flip conditions3 on track · 1 wrong · 2 open · 2 pending → Thesis Tracker 62.5/100
SunWatch market-call ledgerDated, falsifiable market calls8 scored, 5 hits (62.5%). n=8 is small: the Wilson 95% interval is roughly 30–86%, so we treat this as a work-in-progress sample, not proof of skill.
Red-team survival oddsEditorial probabilities (52–65%) on 6 open callsEach survived multiple bull-vs-bear rounds; confidence cuts are published the day counter-evidence lands (e.g. space-top 60%→52% on Aug 8, 2026).

The commitment

When the pool of scored probability calls reaches n≥20, this page will publish a Brier score and a calibration curve (stated probability vs realized frequency), recomputed on every ledger change — the same way the Thesis Tracker recomputes on every verdict change. The forecast ledger is public and timestamped, so anyone can compute it before we do.

Why this page exists

Every AI-era answer engine can generate confident takes; almost none can show you a scored history. Being auditable — misses kept on the page next to hits, flip conditions registered before outcomes, probabilities graded against reality — is this network's entire moat. A calibration page is that moat made explicit: it is the one page a rival cannot copy without also copying two months of dated, falsifiable calls.

Frequently asked questions

What is a Brier score?

A measure of probability-forecast accuracy: the mean squared difference between stated probabilities and outcomes (0 = perfect, 0.25 = coin-flip guessing on binary events). We pre-commit to publishing ours once scored probability calls reach n≥20.

Why not publish a calibration curve now?

The scored sample is 8 market calls plus 6 open odds — too small for a meaningful curve. Publishing one now would be theater. The raw ledgers are public and timestamped, so nothing is hidden in the meantime.

Who grades the calls?

Outcomes are graded against pre-registered falsification conditions written before the outcome, with dated multi-source verification, and misses stay published with their lesson. The grading rules are public in the eight-layer method, including the red-team layer.

The live scorecard updates as models ship and verdicts change.

View the live scorecard →