The frontier, plotted
A season of
machine minds.
Fifteen consequential releases. One comparable software-engineering benchmark. Four months in which the field moved fast.
major releases
24 April—26 August 2026
Fig. 01
Release date × DeepSWE score
Pan the plot on smaller screens
-
DeepSeek V4 Flash
Efficient V4 tier makes million-token context practical.
53% Pass@1 · max effort · ±4%
-
DeepSeek V4 Pro
1.6T open-weight flagship with a one-million-token context window.
63% Pass@1 · max effort · ±6%
-
Claude Opus 4.8
Reliability-focused Opus upgrade for long-running agent work.
59% Pass@1 · max effort · ±2%
-
Claude Fable 5
New generally available capability frontier with additional safeguards.
70% Pass@1 · max effort · ±4%
-
Claude Sonnet 5
Agentic Sonnet release that narrows the gap to the Opus tier.
54% Pass@1 · max effort · ±4%
-
GPT-5.6 Luna
Low-cost GPT-5.6 tier broadens access to strong coding performance.
67% Pass@1 · max effort · ±4%
-
GPT-5.6 Sol
Flagship GPT-5.6 tier sets the family capability ceiling.
73% Pass@1 · max effort · ±3%
-
Kimi K3
First open 3T-class multimodal model for long-horizon work.
69% Pass@1 · max effort · ±5%
-
Gemini 3.6 Flash
Efficiency-focused workhorse for production agents.
47% Pass@1 · high effort · ±4%
-
Claude Opus 5
Generational Opus upgrade for long-running agents.
74% Pass@1 · max effort · ±4%
-
Qwen3.8-Max
2.4T multimodal flagship for autonomous and long-horizon work.
57% Pass@1 · xhigh effort · ±3%
-
Grok 4.6
Frontier upgrade aimed at long-running and visual agents.
67% Pass@1 · xhigh effort · ±2%
-
Gemini 3.7 Flash
Rapid Flash upgrade focused on coding and agent workflows.
65% Pass@1 · high effort · ±2%
-
GLM-5.3
Open-weight coding leader improved entirely through post-training.
69% Pass@1 · max effort · ±3%
-
GLM-5.3-Flash
Native-multimodal GLM with frontier performance at flash cost.
63% Pass@1 · max effort · ±4%
Score bands
- 0–49
- 50–59
- 60–67
- 68–71
- 72–100
Circle size is fixed. Higher placement means a higher score.
Reading note
What the bubbles mean
Each bubble is one release. Its position and visible number show the same benchmark score; the bubble opens the official release evidence.
DeepSWE v1.1 Pass@1 — current leaderboard observation, captured 27 August 2026; leaderboard updated 26 August 2026.
Scores use the source’s best shown model-and-effort configuration. They are not necessarily release-day scores, and effort budgets differ. The displayed ± value is reproduced as published, without relabelling it a confidence interval.
Releases through 26 August 2026. Editorial selection, not an exhaustive model catalog.
View the DeepSWE leaderboard