← Live scorecard
Analysis

Will AI replace programmers? What the benchmarks actually say

Last updated: June 30, 2026 · Updated as verdicts change
By the AGI Scorecard team · methodology & independence
Partly — and unevenly. Agentic coding is genuinely strong: frontier models score ~80% on SWE-Bench Pro and agents ship real code in production. But “replace” overstates it — no system autonomously owns end-to-end software engineering without human review. On the capability curve this tracks Aschenbrenner’s “outpace college grads” call: On track, not a full replacement.

What AI can already do

The measurable progress is real. Frontier models handle complex, multi-step coding tasks — around 80% on SWE-Bench Pro — and coding agents are deployed in real production workflows, not just demos. On raw capability, AI has cleared the bar of a strong entry-level engineer for many well-scoped tasks. This is the same evidence behind the scorecard's On-track "outpace college graduates" verdict.

Where it still falls short

Benchmarks aren't autonomy. The gap between scoring 80% on a coding test and reliably owning a feature end-to-end — scoping, integrating, debugging in a messy real codebase, and being accountable for it without human cleanup — is exactly where "replace" breaks down. The defining AGI bar Aschenbrenner named, a system that autonomously does AI research (the hardest kind of programming), remains undemonstrated.

The honest verdict

QuestionAnswer (mid-2026)
Can AI write production code?Yes, with review
Does it automate parts of the job?Yes
Can it replace an engineer unsupervised?Not yet

So: AI is automating slices of programming and raising each engineer's output, but it is not a drop-in replacement for a software engineer as of mid-2026. On track on capability; the reliability-and-ownership gap is what to watch.

Frequently asked questions

Will AI replace programmers?

Partly, and unevenly. As of mid-2026 agentic coding is strong (~80% on SWE-Bench Pro) and automates parts of the job, but no system reliably owns end-to-end software engineering without human review. It augments programmers rather than replacing them.

How good is AI at coding in 2026?

Frontier models score around 80% on SWE-Bench Pro and ship real code in production for well-scoped tasks — roughly the level Aschenbrenner predicted for 2025/26, graded On track.

What's stopping AI from fully replacing engineers?

The gap between benchmark scores and reliable, accountable, end-to-end ownership in a real codebase. Autonomous software engineering without human cleanup is undemonstrated as of mid-2026.

The live scorecard updates as models ship and verdicts change.

View the live scorecard →

Get the weekly AGI progress briefing

Verdict changes, lab milestones, and what they mean for the 2027 clock. Free — no hype, just signal.

Subscribe free →