Who else graded Situational Awareness — and where they disagree with us
Our verdicts next to every independent grading we could find
Our verdicts are as of 2026-09-06. Each outside grade is a verbatim sentence from the grader's own post; the label beside it is our mapping of that sentence, which you can check against the quote. Tally over 16 comparisons: 9 exact agreements, 6 same direction, 1 disagreement. 2 of our 8 predictions have no independent grade at all.
| Prediction | Our verdict | Independent grader (verbatim) | Their grade · vs ours |
|---|---|---|---|
| Models outpace college graduates across knowledge work | On track | Jamie Harris (EA Forum, 2026-03-29): “AI capabilities on benchmarks have broadly met his expectations, though the qualitative "shocking leap" hasn't landed as he described.” | On track · agrees |
| Nathan Delisle (LessWrong, 2025-06-23): “Compute, algorithmic efficiency, and unhobbling gains seem to follow Aschenbrenner’s forecasts as well, although with more uncertainty.” | On track · agrees | ||
| Daniel Reeves (AGI Friday (Substack), 2025-10-04): “In terms of homework and exams, absolutely. And I don’t think it was obvious in June 2024 how thoroughly true that would be (the International Math Olympiad gold medal performance being the biggest surprise). But Aschenbrenner meant more than homework and exams or even math competitions so there’s plenty of ambiguity about how accurate this prediction is so far.” | Unresolved · disagrees | ||
| Compute + algorithmic scaling continues at trend | On track | Jamie Harris (EA Forum, 2026-03-29): “Infrastructure investment and algorithmic efficiency are tracking ahead of his predictions.” | Ahead · same direction |
| Nathan Delisle (LessWrong, 2025-06-23): “Overall, Aschenbrenner’s predicted pace of roughly half an order-of-magnitude annual progress is ~roughly supported by available evidence.” | On track · agrees | ||
| Philipp D. Dubach (philippdubach.com, 2026-05-21): “The framework is unfashionably simple. It is also empirically unusually well calibrated.” | On track · agrees | ||
| Massive AI capex acceleration | Exceeded | Jamie Harris (EA Forum, 2026-03-29): “Infrastructure investment and algorithmic efficiency are tracking ahead of his predictions.” | Ahead · agrees |
| Nathan Delisle (LessWrong, 2025-06-23): “Using publicly available data as of June 2025, this audit finds that global AI investment, electricity consumption, and chip production follow Aschenbrenner’s forecasts.” | On track · same direction | ||
| Daniel Reeves (AGI Friday (Substack), 2025-10-04): “And as for the progress toward the trillion-dollar cluster, well, Stargate is happening on schedule. Aschenbrenner doesn’t look particularly wrong with any of his “just trust the trendlines” predictions.” | On track · same direction | ||
| Open source fades; proprietary algorithms create a durable US moat | Wrong | Jamie Harris (EA Forum, 2026-03-29): “But he missed that China would innovate independently and that open-source models would thrive at the frontier.” | Wrong · agrees |
| Philipp D. Dubach (philippdubach.com, 2026-05-21): “Bets on defence-AI consolidation and against open-source diffusion would not have paid.” | Wrong · agrees | ||
| AGI: models do the work of an AI researcher/engineer | Open | Jamie Harris (EA Forum, 2026-03-29): “The AGI-by-2027 timeline is unresolved. It looked implausible six months ago, but recent capability jumps (particularly on software engineering tasks) have made it more credible again.” | Unresolved · agrees |
| Daniel Reeves (AGI Friday (Substack), 2025-10-04): “As for full-fledged AI agents and drop-in remote workers on that timescale, my gut says no, but my gut has said no to a lot of things the past few years that have proceeded to happen.” | Behind · same direction | ||
| Philipp D. Dubach (philippdubach.com, 2026-05-21): “The strongest counter-counter is that the timeline remained open on May 21, 2026.” | Unresolved · agrees | ||
| US government launches formal AGI project | Open | Jamie Harris (EA Forum, 2026-03-29): “Government response is behind what he predicted, though his AGI project timeline (27/28) hasn't elapsed yet.” | Behind · same direction |
| Philipp D. Dubach (philippdubach.com, 2026-05-21): “By May 21, 2026, the full nationalisation that he predicted had not happened.” | Behind · same direction | ||
| Intelligence explosion: a decade of progress in ~1 year | Pending | No independent public grade found — nobody else has graded this one yet. | |
| Superintelligence; decisive strategic advantage | Pending | No independent public grade found — nobody else has graded this one yet. | |
How agreement is counted
Two labels agree exactly when they are the same. They agree in direction when both are supportive (ahead or on track) or both say 'not delivered yet' (behind or unresolved). Anything else — including any pairing with 'wrong' that is not 'wrong' on both sides — is a disagreement. A grader who did not address a prediction has no row for it: silence is never counted as agreement.
What this does and does not show
Four graders is still a small panel, and they graded at different times (2025-06-23 to 2026-05-21), so part of any difference is time, not judgement. What the table does show is the thing a single scorecard cannot show about itself: where an outside reader, looking at the same essay, landed somewhere else. Under the counting rule there is one outright disagreement: on Models outpace college graduates across knowledge work, Daniel Reeves reads the evidence as unresolved where this site says On track. A disagreement stays on this page; it is what the table exists to show. Agreement with four graders who read the same public evidence is weak evidence of being right — it mostly rules out that our verdicts are idiosyncratic.
Graders
- Situational Awareness: A One-Year Retrospective — Nathan Delisle, LessWrong, published 2025-06-23, read 2026-09-27. Scope: Quantitative audit of the 'drivers' (compute, algorithmic efficiency, unhobbling) and 'indicators' (cluster size, investment, chips, revenue, electricity) using public data through June 2025. Does not grade AGI-2027, the government project, open source, intelligence explosion or superintelligence.
- How did Leopold do? Evaluating Situational Awareness's predictions — Jamie Harris, EA Forum, published 2026-03-29, read 2026-09-27. Scope: Qualitative evaluation across infrastructure, benchmarks, government, agents, revenue, China/US, open source, safety and the AGI-2027 timeline.
- Retrospective Book Review: Situational Awareness — Daniel Reeves, AGI Friday (Substack), published 2025-10-04, read 2026-09-27. Scope: Retrospective by the Beeminder co-founder after a group re-read that started a year after publication 'with the goal of assessing its predictions'. Verdicts are explicitly hedged. Grades knowledge work, the trillion-dollar cluster and the drop-in remote worker on the 2027 timescale; his base-model/unhobbling remark is about capability returns, not the ~0.5 OOM/yr pace, so it is not mapped. His agi-2027 sentence ('my gut says no', hedged) could be read as unresolved or behind; it is mapped to behind, the reading that credits this site with less agreement.
- Aschenbrenner's Receipts (Leopold Aschenbrenner Predictions: A Two-Year Scorecard) — Philipp D. Dubach, philippdubach.com, published 2026-05-21, read 2026-09-27. Scope: Personal-site scorecard dated 21 May 2026 (page shows 'Updated 16 August 2026'). Grades the counting-OOMs framework, nationalisation, open-source diffusion and the 2027-28 window; its GPQA and power sentences address claims this site does not track, so they are listed separately. The same post also says 'Compute scaling was alive but contested.'; the compute-scaling row quotes his verdict on the counting-OOMs framework instead.
Graded elsewhere, not one of our eight
- AI revenue — Jamie Harris: “AI revenue isn't quite as high: he predicted $100B run rate by mid-2026, and the best is $60B.”
- AI revenue — Nathan Delisle: “However, recent OpenAI and Anthropic models trail raw-compute trends by about one-third to one-half an order of magnitude, and AI-related revenue growth is several months behind.”
- China/US competition — Jamie Harris: “China/US competition intensified as he predicted, and most of his specifics were right (7nm chips, power infrastructure, espionage, CCP waking up, Middle East).”
- Safety and security — Jamie Harris: “Safety and security remain roughly as inadequate as he described, with his warnings about state-actor exploitation and race-dynamic pressure on safety vindicated.”
- GPQA Diamond saturation — Philipp D. Dubach: “It took roughly 18 months, almost exactly the prediction window.”
- Power as the binding constraint — Philipp D. Dubach: “His power-not-chips framing preceded consensus and later became industry consensus.”
Know of another dated, public grading of Situational Awareness? It goes in this table on the same terms: verbatim sentence, link, date, and a mapping you can argue with. Machine-readable: /grader-consensus.json (sources in /independent-grades.json).
Frequently asked questions
Across 16 comparisons with four independent, dated public gradings, 9 agree exactly, 6 agree in direction and 1 disagrees. 2 of our 8 predictions have no independent grade yet.
Nathan Delisle (LessWrong, 2025-06-23); Jamie Harris (EA Forum, 2026-03-29); Daniel Reeves (AGI Friday (Substack), 2025-10-04); Philipp D. Dubach (philippdubach.com, 2026-05-21). Each grade on this page is a verbatim sentence from their own post, linked.
Under the counting rule there is one outright disagreement: on Models outpace college graduates across knowledge work, Daniel Reeves reads the evidence as unresolved where this site says On track. Differences between ahead and on track (for example on capex), or between behind and unresolved, are counted as same-direction, not as disagreements.
Because a scorecard graded by one editor is only as trustworthy as that editor. Putting every independent grading next to ours, with the quotes, is the cheapest way for a reader to check whether our verdicts are idiosyncratic.
The live scorecard updates as models ship and verdicts change.
View the live scorecard →