Coding agent leaderboard: the scores, with sources.

One table per benchmark, pulled from the official leaderboards, plus the vendor claims nobody has checked yet and what the numbers leave out.

Numbered lanes and the word START painted on a running track
Photo by Irham Setyaki on Unsplashdithered by Cyborb

On the official Terminal-Bench 4.0 coding agent leaderboard, the top score is 58.2%, set by GPT-6 Astra running in OpenAI’s Codex, with Claude Fable 5.1 in Claude Code at 57.9%, a statistical tie. On SWE-Bench Pro’s public leaderboard, Meta’s Muse Spark 1.1 leads at 61.5%, but Scale AI has not yet added the newest Claude, GPT-6 or Gemini models.

Both were checked on September 28, 2026. Anthropic reports 66.4% on Terminal-Bench 4.0 for Claude Opus 5.5, released September 22, but that run is the vendor’s own and has not appeared on the leaderboard yet. Below are the current tables, where each number comes from, and what the scores cannot tell you. If the benchmarks themselves are new to you, start with AI benchmarks explained.

The short version
  • Terminal-Bench 4.0: GPT-6 Astra (58.2%) and Claude Fable 5.1 (57.9%) lead, a tie within the error bars. Claude Opus 5 follows at 53.9%.
  • SWE-Bench Pro: Muse Spark 1.1 leads the public set at 61.5%, but the board has not added a model since July 9, 2026.
  • DeepSWE v1.1: GPT-6 Astra, Gemini 3.8 Flash and Claude Opus 5 tie at 74%. The same Gemini scores 19.1% on Terminal-Bench 4.0.
  • SWE-bench Verified is retired at the frontier. Its official leaderboard has not added an entry since February 2026.
  • A score measures a model, a harness, an effort level and a budget together. Compare like with like, then test on your own tasks.

Which AI model is best at coding right now?

No single model wins every coding benchmark in September 2026. The strongest independently measured results belong to GPT-6 Astra, first on Terminal-Bench 4.0 and tied first on DeepSWE, Claude Fable 5.1, tied with Astra on Terminal-Bench 4.0, and Claude Opus 5, tied first on DeepSWE.

58.2%
top Terminal-Bench 4.0 score, GPT-6 Astra in Codex
tbench.ai, Sep 2026
61.5%
top SWE-Bench Pro public score, Muse Spark 1.1
Scale AI, Jul 2026
74%
top DeepSWE v1.1 score, a three-way tie
Datacurve, Sep 2026
66.4%
Claude Opus 5.5 on Terminal-Bench 4.0, vendor-reported
Anthropic, Sep 2026

The newest releases are not on any independent board yet: Claude Opus 5.5 and GPT-6 Sol and Luna all shipped on September 22. Our dated list of 2026 model releases shows what came out when.

Terminal-Bench 4.0 leaderboard

Terminal-Bench 4.0 tests whether an agent can finish real work in a command-line environment. Each entry covers 66 tasks with five attempts each, 330 attempts in all, and most models run in their maker’s own agent, such as Codex or Claude Code. Checked on September 28, 2026, on the official leaderboard, last updated September 21. We show each model’s best run; cost per attempt is our division of the leaderboard’s total run cost by 330.

ModelAgent (harness)EffortScore, 95% rangeCost per attempt
GPT-6 AstraCodexmax58.2% ± 2.8$9.90
Claude Fable 5.1Claude Codemax57.9% ± 3.8$18.92
Claude Opus 5Claude Codexhigh53.9% ± 3.2$18.44
Claude Fable 5Claude Codemax44.5% ± 3.9$22.02
GLM-5.3Claude Codemax41.8% ± 3.2$8.27
Grok 4.7Grok Buildxhigh37.6% ± 3.5$11.16
GPT-5.6 SolCodexmax37.3% ± 3.8$7.70
Claude Opus 4.8Claude Codemax23.6% ± 3.6$19.64
GPT-5.6 TerraCodexmax21.5% ± 3.2$5.25
Grok 4.6Grok Buildhigh20.3% ± 3.1$10.88
Gemini 3.8 Flashmini-SWE-agenthigh19.1% ± 3.4$5.54
GPT-5.6 LunaCodexmax17.3% ± 2.9$1.05
Grok 4.5Grok Buildhigh12.4% ± 2.6$6.35
Claude Sonnet 5Claude Codemax12.4% ± 3.1$29.10
Gemini 3.7 Flashmini-SWE-agenthigh11.2% ± 2.5$3.82

Three details change how you read this table. Grok 4.7’s run is marked partial on the leaderboard, with 324 of 330 attempts finished. Google’s models ran in the generic mini-SWE-agent harness, while the others ran in their makers’ own agents. And the gap between first and second place, 0.3 points, is far smaller than either margin.

Effort matters as much as the model. GPT-6 Astra scores 50.6% at low effort and 58.2% at max, and the full run’s cost rises from $1,557 to $3,267. Claude Fable 5.1 goes from 43.3% to 57.9% as its run cost rises from $2,359 to $6,244.

SWE-Bench Pro scores

SWE-Bench Pro asks an agent to fix real issues in actively maintained repositories, and Scale AI keeps most tasks private to limit leakage. The first six rows come from Scale’s public leaderboard, checked on September 28, 2026. The rest are the model makers’ own reports, read on September 24 and 25, which nobody has reproduced on the board.

ModelScoreMeasured byDateSetup
Muse Spark 1.161.5% ± 3.1Scale AIJul 9, 2026mini-swe-agent
GPT-5.4 (xHigh)59.1% ± 3.6Scale AIApr 8, 2026mini-swe-agent
Muse Spark55.0% ± 3.6Scale AIApr 8, 2026mini-swe-agent
Claude Opus 4.6 (thinking)51.9% ± 3.6Scale AIApr 8, 2026mini-swe-agent
Gemini 3.1 Pro (thinking)46.1% ± 3.6Scale AIApr 8, 2026mini-swe-agent
Claude Opus 4.545.9% ± 3.6Scale AIDec 11, 2025Scale’s default setup
Claude Mythos Preview77.8%AnthropicApr 7, 2026Invitation-only model
Qwen3.8-Max67.7%QwenAug 3, 2026Claude Code harness, Qwen’s corrected task set
GPT-5.6 Sol64.6%OpenAIJul 9, 2026OpenAI’s own setup
GPT-5.558.6%OpenAIApr 23, 2026Public set, with a memorization warning
DeepSeek V4-Pro (Max)55.4%DeepSeekApr 24, 2026DeepSeek’s own setup

Scale also scores a private set of 276 tasks from startup codebases that are not public. There, Muse Spark 1.1 leads at 51.5%, ten points below its public score, and Claude Opus 4.6 follows at 47.1%.

Two vendor notes deserve weight. OpenAI marks its SWE-Bench Pro figure with the warning that labs have seen evidence of memorization on the public set. Qwen says it corrected problematic tasks and reran every model it compares on its refined version, so its 67.7% is not directly comparable to Scale’s board. Anthropic’s launch posts for Fable 5.1, Opus 5 and Opus 5.5 headline other benchmarks instead.

DeepSWE v1.1 leaderboard

DeepSWE, built by Datacurve, uses 113 tasks written from scratch across 91 repositories in five languages, which it says keeps the solutions out of any model’s training data. Every model runs in the same mini-swe-agent harness. Checked on September 28, 2026, on the official leaderboard, updated September 22.

ModelEffortPass rateAverage cost per task
GPT-6 Astraxhigh74% ± 3$4.43
Gemini 3.8 Flashhigh74% ± 1$2.36
Claude Opus 5max74% ± 4$11.84
GPT-5.6 Solmax73% ± 3$6.46
Claude Fable 5xhigh70% ± 3$13.41
GLM-5.3max69% ± 3$3.99
Kimi K3max69% ± 5$4.65
Grok 4.6medium67% ± 2$3.45
GPT-5.6 Lunamax67% ± 4$0.61
GPT-5.5xhigh67% ± 6$7.23

Open-weight models do well here. GLM-5.3 and Kimi K3 finish five points behind the leaders, at under 40% of Claude Opus 5’s cost per task.

Is SWE-bench Verified still worth quoting?

Not for comparing frontier models. On February 23, 2026, OpenAI said it would stop evaluating on SWE-bench Verified, after an audit found flawed tests that reject correct fixes and signs that models had trained on the solutions. The official SWE-bench leaderboards have not added an entry since February 26, 2026. On the bash-only view, where every model runs in the same minimal agent, the top score is 76.8%, Claude Opus 4.5 at high effort.

Vendors still print Verified scores. Mistral reported 77.6% for Mistral Medium 3.5 in May, DeepSeek 80.6% for V4-Pro in April, and Anthropic 93.9% for the invitation-only Mythos Preview. Each ran its own harness, so these numbers compare poorly with each other and not at all with the frozen leaderboard. Stanford’s 2026 AI Index sums up the problem: scores on the benchmark rose from 60% to near 100% in a single year.

Which benchmark versions are current?

Version numbers matter, because the same model can score 89% on one Terminal-Bench and 19% on another. As of September 28, 2026:

BenchmarkCurrent versionTasksLeaderboard run by
Terminal-Bench4.0, released August 28, 202666The Terminal-Bench team (Stanford, Laude Institute, Harbor)
SWE-Bench ProPublic and private sets731 public, 276 private, 858 held outScale AI
DeepSWEv1.1113Datacurve
SWE-bench VerifiedNo new entries since February 2026500SWE-bench team

Older versions live on in launch posts. Google’s Gemini 3.8 Flash model card shows 89.4% on Terminal-Bench 2.1 next to 19.1% on version 4.0, and OpenAI’s GPT-5.6 post reports 88.8% for GPT-5.6 Sol on version 2.1. Terminal-Bench 3.0, released July 30, was replaced a month later.

Launch posts also lean on newer tests such as CursorBench 4.0 and FrontierCode 1.1. Unless a public leaderboard lists the run, treat the number as the vendor’s own measurement.

What the scores do and do not tell you

A leaderboard score measures a whole setup, not a model in isolation. Every example below comes from the tables above.

What changes the scoreReal example
The benchmarkGemini 3.8 Flash ties for first on DeepSWE (74%) and scores 19.1% on Terminal-Bench 4.0
Who ran itAnthropic reports 55.8% for Fable 5.1 on Terminal-Bench 4.0; the leaderboard’s own max-effort run scored 57.9%
EffortGPT-6 Astra: 50.6% at low effort, 58.2% at max
BudgetFable 5.1 and GPT-6 Astra score within 0.3 points, but Fable 5.1’s run cost nearly twice as much
HarnessTerminal-Bench runs most models in their makers’ agents; DeepSWE runs all of them in mini-swe-agent
FreshnessSWE-Bench Pro’s newest entry is from July 9, before most current models shipped

The practical rule: a gap inside the error bars is a tie, and a vendor score is a claim until an independent board reproduces it. “Effort” is how long a model may reason before it answers, which our guide to reasoning models explains.

What no public leaderboard can tell you is how an agent does on your code. Twenty tasks from last month’s work, graded the same way for every agent, will settle more than any table here. Our guide to evals for beginners shows how to build that set, and our comparison of the best AI coding agents covers the tools, prices and plans behind these models.

FAQ

What is the best AI model for coding in 2026?

On independent leaderboards checked September 28, 2026, GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5 hold the top results, and none wins everywhere. The best choice for you depends on your language, your codebase and your budget, so test two or three on your own tasks.

What is the highest Terminal-Bench 4.0 score?

58.2%, by GPT-6 Astra at max effort in Codex, on the official leaderboard updated September 21, 2026. Claude Fable 5.1 in Claude Code scored 57.9%, inside the same error range. Anthropic reports 66.4% for Claude Opus 5.5, but that run has not been reproduced on the leaderboard.

Why are Claude Opus 5.5 and GPT-6 Sol missing from the leaderboards?

Both shipped on September 22, 2026. Independent boards take days to weeks to run a new model: Terminal-Bench needs 330 attempts per entry. Until then, only the makers’ own numbers exist.

What is the difference between SWE-bench Verified and SWE-Bench Pro?

SWE-bench Verified is 500 human-checked Python issues from public projects, and top models have nearly saturated it. SWE-Bench Pro, from Scale AI, has 1,865 harder tasks in several languages, and most of them are kept private to limit memorization.

Are vendor-reported benchmark scores reliable?

They are real measurements, but of the vendor’s chosen setup. Harness, effort level, task fixes and retries all vary. Use them to shortlist, and trust independent leaderboards and your own tests to decide.

Key takeaways
  • GPT-6 Astra and Claude Fable 5.1 are tied at the top of Terminal-Bench 4.0, at about 58%.
  • SWE-Bench Pro’s public leaderboard is months behind; most frontier scores on it now come from vendors.
  • DeepSWE ranks the same models differently, which is the best argument against trusting any single board.
  • SWE-bench Verified is frozen and saturated. Quote it only for history.
  • Compare scores only when benchmark version, harness, effort and budget match.

Read next: AI benchmarks explained for what each test measures, or Claude Code vs Codex vs Cursor to pick between the agents behind these scores.

Sources
  1. Terminal-Bench 4.0 leaderboard, Terminal-Bench, updated September 2026
  2. Terminal-Bench 4.0 announcement, Terminal-Bench, August 2026
  3. Terminal-Bench 3.0, Terminal-Bench, July 2026
  4. SWE-Bench Pro public and private leaderboards, Scale AI, accessed September 2026
  5. DeepSWE leaderboard, Datacurve, updated September 2026
  6. SWE-bench leaderboards, SWE-bench, accessed September 2026
  7. Why SWE-bench Verified no longer measures frontier coding capabilities, OpenAI, February 2026
  8. Introducing Claude Opus 5.5, Anthropic, September 2026
  9. Introducing Claude Fable 5.1 and Claude Mythos 5.1, Anthropic, September 2026
  10. Project Glasswing, Anthropic, April 2026
  11. GPT-6 Astra, OpenAI, September 2026
  12. GPT-5.6, OpenAI, July 2026
  13. Introducing GPT-5.5, OpenAI, April 2026
  14. Gemini 3.8 Flash model card, Google DeepMind, September 2026
  15. Qwen3.8-Max, Qwen, August 2026
  16. DeepSeek-V4-Pro model card, DeepSeek, April 2026
  17. Remote agents in Vibe, powered by Mistral Medium 3.5, Mistral AI, May 2026
  18. 2026 AI Index Report, Stanford HAI, April 2026
cyborb.ai

Stop reading about it. Build it.

Describe what you want in plain words. Cyborb plans the work, writes and runs the code, makes the assets, and puts the result online.

Download Cyborb

Free to start. No card required.