← Live scorecard
Analysis

Will AI replace programmers? What the benchmarks actually say

Last updated: September 5, 2026 · Updated as verdicts change
By the AGI Scorecard team · methodology & independence
Partly — and unevenly. Agentic coding is genuinely strong: frontier models score ~80% on SWE-Bench Pro and agents ship real code in production. But “replace” overstates it — no system autonomously owns end-to-end software engineering without human review. On the capability curve this tracks Aschenbrenner’s “outpace college grads” call: On track, not a full replacement.

Status as of 2026-09-05

Start here if the question is about your own job: Should you be worried about AI taking your job?

What AI can already do

The measurable progress is real. Frontier models handle complex, multi-step coding tasks — around 80% on SWE-Bench Pro — and coding agents are deployed in real production workflows, not just demos. On raw capability, AI has cleared the bar of a strong entry-level engineer for many well-scoped tasks. This is the same evidence behind the scorecard's On-track "outpace college graduates" verdict.

Where it still falls short

Benchmarks aren't autonomy. The gap between scoring 80% on a coding test and reliably owning a feature end-to-end — scoping, integrating, debugging in a messy real codebase, and being accountable for it without human cleanup — is exactly where "replace" breaks down. The defining AGI bar Aschenbrenner named, a system that autonomously does AI research (the hardest kind of programming), remains undemonstrated.

The honest verdict

QuestionAnswer (mid-2026)
Can AI write production code?Yes, with review
Does it automate parts of the job?Yes
Can it replace an engineer unsupervised?Not yet

So: AI is automating slices of programming and raising each engineer's output, but it is not a drop-in replacement for a software engineer as of mid-2026. On track on capability; the reliability-and-ownership gap is what to watch.

Five sources, one card

The programmer debate runs on five numbers that sound contradictory. Here they are side by side, each dated and linked.

SourceWhat it measuresReading (dated)Link
BLS Employment Projections 2024–34Projected US employment change by occupation, ten years outComputer programmers −6%; software developers +15% (data scientists +33.5%). Published 2026.bls.gov
NY Fed, “Labor Market for Recent College Graduates”Unemployment of recent graduates by majorAs reported by secondary coverage: 6.1% computer science, 7.5% computer engineering (2026).Secondary reports; not independently verified here
Indeed Hiring LabJob-posting growth by seniority71% of software-development posting growth May 2025→May 2026 was senior-level (2026-07-08).hiringlab.indeed.com
Microsoft Research, “Working with AI”AI applicability: share of an occupation’s work activities Copilot users were observed doing with AIComputer Programmers (SOC 15-1251) 0.31; Software Developers (SOC 15-1252) 0.278 (Jul 2025, CC-BY-4.0).github.com/microsoft
Anthropic Economic IndexShare of occupations’ tasks observed in Claude usage; augmentation vs automation49% of occupations have ≥25% of tasks observed; usage split 55% augmentation vs 42% automation (March 2026).anthropic.com

Why they disagree: exposure ≠ displacement. Microsoft and Anthropic measure how much of the work AI is observed doing; the BLS projects headcount a decade out; Indeed and the NY Fed measure hiring and unemployment now. An occupation can be highly AI-applicable and still grow (developers, +15%) while a neighbouring title shrinks (programmers, −6%), and a junior-hiring slump can coexist with senior growth (71%). Look up your own SOC code on the AI job exposure check.

Five sources, one card — so when does an agent own the whole engineering job?

2025–2620272028–302030s2040+ / never

One tap — see which real forecaster you side with, instantly. No sign-up.

Frequently asked questions

Will AI replace programmers?

Partly, and unevenly. As of mid-2026 agentic coding is strong (~80% on SWE-Bench Pro) and automates parts of the job, but no system reliably owns end-to-end software engineering without human review. It augments programmers rather than replacing them.

How good is AI at coding in 2026?

Frontier models score around 80% on SWE-Bench Pro and ship real code in production for well-scoped tasks — roughly the level Aschenbrenner predicted for 2025/26, graded On track.

What's stopping AI from fully replacing engineers?

The gap between benchmark scores and reliable, accountable, end-to-end ownership in a real codebase. Autonomous software engineering without human cleanup is undemonstrated as of mid-2026.

The live scorecard updates as models ship and verdicts change.

View the live scorecard →

Be first to know when a verdict flips

One email when one of the eight verdicts changes and the AGI-2027 Thesis Tracker moves — the single auditable score no other tracker has. Not a weekly newsletter; nothing arrives until the score actually moves.

Subscribe free →