The frontier, plotted

A season of
machine minds.

Fifteen consequential releases. One comparable software-engineering benchmark. Four months in which the field moved fast.

major releases
24 April—26 August 2026

Fig. 01

Release date × DeepSWE score

Pan the plot on smaller screens

  1. DeepSeek V4 Flash

    Efficient V4 tier makes million-token context practical.

    53% Pass@1 · max effort · ±4%

  2. DeepSeek V4 Pro

    1.6T open-weight flagship with a one-million-token context window.

    63% Pass@1 · max effort · ±6%

  3. Claude Opus 4.8

    Reliability-focused Opus upgrade for long-running agent work.

    59% Pass@1 · max effort · ±2%

  4. Claude Fable 5

    New generally available capability frontier with additional safeguards.

    70% Pass@1 · max effort · ±4%

  5. Claude Sonnet 5

    Agentic Sonnet release that narrows the gap to the Opus tier.

    54% Pass@1 · max effort · ±4%

  6. GPT-5.6 Luna

    Low-cost GPT-5.6 tier broadens access to strong coding performance.

    67% Pass@1 · max effort · ±4%

  7. GPT-5.6 Sol

    Flagship GPT-5.6 tier sets the family capability ceiling.

    73% Pass@1 · max effort · ±3%

  8. Kimi K3

    First open 3T-class multimodal model for long-horizon work.

    69% Pass@1 · max effort · ±5%

  9. Gemini 3.6 Flash

    Efficiency-focused workhorse for production agents.

    47% Pass@1 · high effort · ±4%

  10. Claude Opus 5

    Generational Opus upgrade for long-running agents.

    74% Pass@1 · max effort · ±4%

  11. Qwen3.8-Max

    2.4T multimodal flagship for autonomous and long-horizon work.

    57% Pass@1 · xhigh effort · ±3%

  12. Grok 4.6

    Frontier upgrade aimed at long-running and visual agents.

    67% Pass@1 · xhigh effort · ±2%

  13. Gemini 3.7 Flash

    Rapid Flash upgrade focused on coding and agent workflows.

    65% Pass@1 · high effort · ±2%

  14. GLM-5.3

    Open-weight coding leader improved entirely through post-training.

    69% Pass@1 · max effort · ±3%

  15. GLM-5.3-Flash

    Native-multimodal GLM with frontier performance at flash cost.

    63% Pass@1 · max effort · ±4%

Score bands

  • 0–49
  • 50–59
  • 60–67
  • 68–71
  • 72–100

Circle size is fixed. Higher placement means a higher score.

Reading note

What the bubbles mean

Each bubble is one release. Its position and visible number show the same benchmark score; the bubble opens the official release evidence.

DeepSWE v1.1 Pass@1 — current leaderboard observation, captured 27 August 2026; leaderboard updated 26 August 2026.

Scores use the source’s best shown model-and-effort configuration. They are not necessarily release-day scores, and effort budgets differ. The displayed ± value is reproduced as published, without relabelling it a confidence interval.

Releases through 26 August 2026. Editorial selection, not an exhaustive model catalog.

View the DeepSWE leaderboard