The frontier, remembered

Five years of
machine minds.

Forty consequential releases, from instruction following to long-running agents. Exact DeepSWE scores appear only where the captured leaderboard has a direct model row.

major releases
2022–2026

Fig. 01

A continuous model history

Scroll the years · 2022 → 2026

2022

  1. InstructGPT

    Not evaluated · N/A

    Put RLHF-trained instruction following into the default OpenAI API model line.

    No exact model row in the captured DeepSWE v1.1 leaderboard.

  2. PaLM

    Not evaluated · N/A

    Demonstrated a 540B dense Pathways model with strong few-shot reasoning and code generation.

    No exact model row in the captured DeepSWE v1.1 leaderboard.

  3. OPT-175B

    Not evaluated · N/A

    Opened research access to model weights, code, and a training logbook at GPT-3 scale.

    No exact model row in the captured DeepSWE v1.1 leaderboard.

  4. BLOOM

    Not evaluated · N/A

    Delivered a transparently trained 176B open multilingual model through a global research collaboration.

    No exact model row in the captured DeepSWE v1.1 leaderboard.

  5. ChatGPT

    Not evaluated · N/A

    Turned an RLHF-tuned GPT-3.5 dialogue model into the breakout consumer interface for generative AI.

    No exact model row in the captured DeepSWE v1.1 leaderboard.

2023

  1. LLaMA

    Not evaluated · N/A

    Showed that smaller, data-efficient foundation models could rival much larger systems.

    No exact model row in the captured DeepSWE v1.1 leaderboard.

  2. Claude

    Not evaluated · N/A

    Introduced Anthropic's steerable assistant and API as a durable second frontier-model line.

    No exact model row in the captured DeepSWE v1.1 leaderboard.

  3. GPT-4

    Not evaluated · N/A

    Established a new closed-model frontier with image input and strong professional-benchmark performance.

    No exact model row in the captured DeepSWE v1.1 leaderboard.

  4. Llama 2

    Not evaluated · N/A

    Moved the Llama family to broadly available weights licensed for research and commercial use.

    No exact model row in the captured DeepSWE v1.1 leaderboard.

  5. Mistral 7B

    Not evaluated · N/A

    Reset expectations for compact open models with strong results and an Apache 2.0 release.

    No exact model row in the captured DeepSWE v1.1 leaderboard.

  6. Gemini 1.0

    Not evaluated · N/A

    Launched Google's natively multimodal Ultra, Pro, and Nano foundation-model family.

    No exact model row in the captured DeepSWE v1.1 leaderboard.

2024

  1. Gemini 1.5

    Not evaluated · N/A

    Made a one-million-token experimental context window the new long-context benchmark.

    No exact model row in the captured DeepSWE v1.1 leaderboard.

  2. Claude 3

    Not evaluated · N/A

    Established the Haiku, Sonnet, and Opus capability-cost tiers still used by Anthropic.

    No exact model row in the captured DeepSWE v1.1 leaderboard.

  3. Llama 3

    Not evaluated · N/A

    Brought materially stronger 8B and 70B open-weight models to a broad platform ecosystem.

    No exact model row in the captured DeepSWE v1.1 leaderboard.

  4. GPT-4o

    Not evaluated · N/A

    Unified text, audio, image, and video interaction in a real-time flagship model.

    No exact model row in the captured DeepSWE v1.1 leaderboard.

  5. Claude 3.5 Sonnet

    Not evaluated · N/A

    Made a mid-tier model the coding and visual-reasoning frontier while introducing Artifacts.

    No exact model row in the captured DeepSWE v1.1 leaderboard.

  6. OpenAI o1-preview

    Not evaluated · N/A

    Made test-time reasoning a distinct product and model-scaling dimension.

    No exact model row in the captured DeepSWE v1.1 leaderboard.

  7. DeepSeek-V3

    Not evaluated · N/A

    Pushed open mixture-of-experts capability and training efficiency into the frontier conversation.

    No exact model row in the captured DeepSWE v1.1 leaderboard.

2025

  1. DeepSeek-R1

    Not evaluated · N/A

    Released a strong reasoning model and distilled variants under an MIT license.

    No exact model row in the captured DeepSWE v1.1 leaderboard.

  2. Claude 3.7 Sonnet

    Not evaluated · N/A

    Combined near-instant and extended reasoning in one model and launched alongside Claude Code.

    No exact model row in the captured DeepSWE v1.1 leaderboard.

  3. Gemini 2.5 Pro

    Not evaluated · N/A

    Made reasoning native across Google's next Gemini generation with strong coding performance.

    No exact model row in the captured DeepSWE v1.1 leaderboard.

  4. Llama 4

    Not evaluated · N/A

    Moved Meta's open-weight family to native multimodality and mixture-of-experts designs.

    No exact model row in the captured DeepSWE v1.1 leaderboard.

  5. Qwen3

    Not evaluated · N/A

    Brought switchable thinking and non-thinking modes to a broad open-weight model family.

    No exact model row in the captured DeepSWE v1.1 leaderboard.

  6. Claude 4

    Not evaluated · N/A

    Established Opus 4 and Sonnet 4 as long-running coding and agent-workflow models.

    No exact model row in the captured DeepSWE v1.1 leaderboard.

  7. GPT-5

    Not evaluated · N/A

    Unified fast responses, deeper reasoning, and routing in OpenAI's next flagship system.

    No exact model row in the captured DeepSWE v1.1 leaderboard.

2026 · through 26 Aug

  1. DeepSeek V4 Flash

    DeepSWE 53%

    Efficient V4 tier makes million-token context practical.

    53% Pass@1 · max effort · ±4%

  2. DeepSeek V4 Pro

    DeepSWE 63%

    1.6T open-weight flagship with a one-million-token context window.

    63% Pass@1 · max effort · ±6%

  3. Claude Opus 4.8

    DeepSWE 59%

    Reliability-focused Opus upgrade for long-running agent work.

    59% Pass@1 · max effort · ±2%

  4. Claude Fable 5

    DeepSWE 70%

    New generally available capability frontier with additional safeguards.

    70% Pass@1 · max effort · ±4%

  5. Claude Sonnet 5

    DeepSWE 54%

    Agentic Sonnet release that narrows the gap to the Opus tier.

    54% Pass@1 · max effort · ±4%

  6. GPT-5.6 Luna

    DeepSWE 67%

    Low-cost GPT-5.6 tier broadens access to strong coding performance.

    67% Pass@1 · max effort · ±4%

  7. GPT-5.6 Sol

    DeepSWE 73%

    Flagship GPT-5.6 tier sets the family capability ceiling.

    73% Pass@1 · max effort · ±3%

  8. Kimi K3

    DeepSWE 69%

    First open 3T-class multimodal model for long-horizon work.

    69% Pass@1 · max effort · ±5%

  9. Gemini 3.6 Flash

    DeepSWE 47%

    Efficiency-focused workhorse for production agents.

    47% Pass@1 · high effort · ±4%

  10. Claude Opus 5

    DeepSWE 74%

    Generational Opus upgrade for long-running agents.

    74% Pass@1 · max effort · ±4%

  11. Qwen3.8-Max

    DeepSWE 57%

    2.4T multimodal flagship for autonomous and long-horizon work.

    57% Pass@1 · xhigh effort · ±3%

  12. Grok 4.6

    DeepSWE 67%

    Frontier upgrade aimed at long-running and visual agents.

    67% Pass@1 · xhigh effort · ±2%

  13. Gemini 3.7 Flash

    DeepSWE 65%

    Rapid Flash upgrade focused on coding and agent workflows.

    65% Pass@1 · high effort · ±2%

  14. GLM-5.3

    DeepSWE 69%

    Open-weight coding leader improved entirely through post-training.

    69% Pass@1 · max effort · ±3%

  15. GLM-5.3-Flash

    DeepSWE 63%

    Native-multimodal GLM with frontier performance at flash cost.

    63% Pass@1 · max effort · ±4%

How to read the marks

  • 63%Filled bubble · exact current DeepSWE score
  • Hollow marker · not evaluated in the captured leaderboard

Circle size is fixed. Higher scored placement means a higher Pass@1 value; hollow markers live on a separate N/A rail.

Reading note

What the marks mean

Each mark is one release and opens its official evidence. A filled bubble is scored; a hollow marker is historical context without a fabricated number.

DeepSWE v1.1 Pass@1 — current leaderboard observation, captured 27 August 2026; leaderboard updated 26 August 2026.

“Not evaluated” means no exact model row in the committed current DeepSWE v1.1 source. It does not mean no coding ability, and it is never treated as a zero.

Scores use the source’s best shown model-and-effort configuration. They are not necessarily release-day scores, and effort budgets differ. The displayed ± value is reproduced as published.

40 major AI model releases from 2022–2026, through 26 August 2026. Editorial selection, not an exhaustive model catalog.

View the DeepSWE leaderboard