Calibration: we score our own predictions in public
What we grade, today
| Ledger | Type | Current state (as of 2026-08-08) |
|---|---|---|
| 8 Situational Awareness verdicts | Categorical verdicts with pre-registered flip conditions | 3 on track · 1 wrong · 2 open · 2 pending → Thesis Tracker 62.5/100 |
| SunWatch market-call ledger | Dated, falsifiable market calls | 8 scored, 5 hits (62.5%). n=8 is small: the Wilson 95% interval is roughly 30–86%, so we treat this as a work-in-progress sample, not proof of skill. |
| Red-team survival odds | Editorial probabilities (52–65%) on 6 open calls | Each survived multiple bull-vs-bear rounds; confidence cuts are published the day counter-evidence lands (e.g. space-top 60%→52% on Aug 8, 2026). |
The commitment
When the pool of scored probability calls reaches n≥20, this page will publish a Brier score and a calibration curve (stated probability vs realized frequency), recomputed on every ledger change — the same way the Thesis Tracker recomputes on every verdict change. The forecast ledger is public and timestamped, so anyone can compute it before we do.
Why this page exists
Every AI-era answer engine can generate confident takes; almost none can show you a scored history. Being auditable — misses kept on the page next to hits, flip conditions registered before outcomes, probabilities graded against reality — is this network's entire moat. A calibration page is that moat made explicit: it is the one page a rival cannot copy without also copying two months of dated, falsifiable calls.
Frequently asked questions
A measure of probability-forecast accuracy: the mean squared difference between stated probabilities and outcomes (0 = perfect, 0.25 = coin-flip guessing on binary events). We pre-commit to publishing ours once scored probability calls reach n≥20.
The scored sample is 8 market calls plus 6 open odds — too small for a meaningful curve. Publishing one now would be theater. The raw ledgers are public and timestamped, so nothing is hidden in the meantime.
Outcomes are graded against pre-registered falsification conditions written before the outcome, with dated multi-source verification, and misses stay published with their lesson. The grading rules are public in the eight-layer method, including the red-team layer.
The live scorecard updates as models ship and verdicts change.
View the live scorecard →