The frontier, remembered
Five years of
machine minds.
Forty consequential releases, from instruction following to long-running agents. Exact DeepSWE scores appear only where the captured leaderboard has a direct model row.
major releases
2022–2026
Fig. 01
A continuous model history
Scroll the years · 2022 → 2026
2022
-
InstructGPT
Not evaluated · N/A
Put RLHF-trained instruction following into the default OpenAI API model line.
No exact model row in the captured DeepSWE v1.1 leaderboard.
-
PaLM
Not evaluated · N/A
Demonstrated a 540B dense Pathways model with strong few-shot reasoning and code generation.
No exact model row in the captured DeepSWE v1.1 leaderboard.
-
OPT-175B
Not evaluated · N/A
Opened research access to model weights, code, and a training logbook at GPT-3 scale.
No exact model row in the captured DeepSWE v1.1 leaderboard.
-
BLOOM
Not evaluated · N/A
Delivered a transparently trained 176B open multilingual model through a global research collaboration.
No exact model row in the captured DeepSWE v1.1 leaderboard.
-
ChatGPT
Not evaluated · N/A
Turned an RLHF-tuned GPT-3.5 dialogue model into the breakout consumer interface for generative AI.
No exact model row in the captured DeepSWE v1.1 leaderboard.
2023
-
LLaMA
Not evaluated · N/A
Showed that smaller, data-efficient foundation models could rival much larger systems.
No exact model row in the captured DeepSWE v1.1 leaderboard.
-
Claude
Not evaluated · N/A
Introduced Anthropic's steerable assistant and API as a durable second frontier-model line.
No exact model row in the captured DeepSWE v1.1 leaderboard.
-
GPT-4
Not evaluated · N/A
Established a new closed-model frontier with image input and strong professional-benchmark performance.
No exact model row in the captured DeepSWE v1.1 leaderboard.
-
Llama 2
Not evaluated · N/A
Moved the Llama family to broadly available weights licensed for research and commercial use.
No exact model row in the captured DeepSWE v1.1 leaderboard.
-
Mistral 7B
Not evaluated · N/A
Reset expectations for compact open models with strong results and an Apache 2.0 release.
No exact model row in the captured DeepSWE v1.1 leaderboard.
-
Gemini 1.0
Not evaluated · N/A
Launched Google's natively multimodal Ultra, Pro, and Nano foundation-model family.
No exact model row in the captured DeepSWE v1.1 leaderboard.
2024
-
Gemini 1.5
Not evaluated · N/A
Made a one-million-token experimental context window the new long-context benchmark.
No exact model row in the captured DeepSWE v1.1 leaderboard.
-
Claude 3
Not evaluated · N/A
Established the Haiku, Sonnet, and Opus capability-cost tiers still used by Anthropic.
No exact model row in the captured DeepSWE v1.1 leaderboard.
-
Llama 3
Not evaluated · N/A
Brought materially stronger 8B and 70B open-weight models to a broad platform ecosystem.
No exact model row in the captured DeepSWE v1.1 leaderboard.
-
GPT-4o
Not evaluated · N/A
Unified text, audio, image, and video interaction in a real-time flagship model.
No exact model row in the captured DeepSWE v1.1 leaderboard.
-
Claude 3.5 Sonnet
Not evaluated · N/A
Made a mid-tier model the coding and visual-reasoning frontier while introducing Artifacts.
No exact model row in the captured DeepSWE v1.1 leaderboard.
-
OpenAI o1-preview
Not evaluated · N/A
Made test-time reasoning a distinct product and model-scaling dimension.
No exact model row in the captured DeepSWE v1.1 leaderboard.
-
DeepSeek-V3
Not evaluated · N/A
Pushed open mixture-of-experts capability and training efficiency into the frontier conversation.
No exact model row in the captured DeepSWE v1.1 leaderboard.
2025
-
DeepSeek-R1
Not evaluated · N/A
Released a strong reasoning model and distilled variants under an MIT license.
No exact model row in the captured DeepSWE v1.1 leaderboard.
-
Claude 3.7 Sonnet
Not evaluated · N/A
Combined near-instant and extended reasoning in one model and launched alongside Claude Code.
No exact model row in the captured DeepSWE v1.1 leaderboard.
-
Gemini 2.5 Pro
Not evaluated · N/A
Made reasoning native across Google's next Gemini generation with strong coding performance.
No exact model row in the captured DeepSWE v1.1 leaderboard.
-
Llama 4
Not evaluated · N/A
Moved Meta's open-weight family to native multimodality and mixture-of-experts designs.
No exact model row in the captured DeepSWE v1.1 leaderboard.
-
Qwen3
Not evaluated · N/A
Brought switchable thinking and non-thinking modes to a broad open-weight model family.
No exact model row in the captured DeepSWE v1.1 leaderboard.
-
Claude 4
Not evaluated · N/A
Established Opus 4 and Sonnet 4 as long-running coding and agent-workflow models.
No exact model row in the captured DeepSWE v1.1 leaderboard.
-
GPT-5
Not evaluated · N/A
Unified fast responses, deeper reasoning, and routing in OpenAI's next flagship system.
No exact model row in the captured DeepSWE v1.1 leaderboard.
2026 · through 26 Aug
-
DeepSeek V4 Flash
DeepSWE 53%
Efficient V4 tier makes million-token context practical.
53% Pass@1 · max effort · ±4%
-
DeepSeek V4 Pro
DeepSWE 63%
1.6T open-weight flagship with a one-million-token context window.
63% Pass@1 · max effort · ±6%
-
Claude Opus 4.8
DeepSWE 59%
Reliability-focused Opus upgrade for long-running agent work.
59% Pass@1 · max effort · ±2%
-
Claude Fable 5
DeepSWE 70%
New generally available capability frontier with additional safeguards.
70% Pass@1 · max effort · ±4%
-
Claude Sonnet 5
DeepSWE 54%
Agentic Sonnet release that narrows the gap to the Opus tier.
54% Pass@1 · max effort · ±4%
-
GPT-5.6 Luna
DeepSWE 67%
Low-cost GPT-5.6 tier broadens access to strong coding performance.
67% Pass@1 · max effort · ±4%
-
GPT-5.6 Sol
DeepSWE 73%
Flagship GPT-5.6 tier sets the family capability ceiling.
73% Pass@1 · max effort · ±3%
-
Kimi K3
DeepSWE 69%
First open 3T-class multimodal model for long-horizon work.
69% Pass@1 · max effort · ±5%
-
Gemini 3.6 Flash
DeepSWE 47%
Efficiency-focused workhorse for production agents.
47% Pass@1 · high effort · ±4%
-
Claude Opus 5
DeepSWE 74%
Generational Opus upgrade for long-running agents.
74% Pass@1 · max effort · ±4%
-
Qwen3.8-Max
DeepSWE 57%
2.4T multimodal flagship for autonomous and long-horizon work.
57% Pass@1 · xhigh effort · ±3%
-
Grok 4.6
DeepSWE 67%
Frontier upgrade aimed at long-running and visual agents.
67% Pass@1 · xhigh effort · ±2%
-
Gemini 3.7 Flash
DeepSWE 65%
Rapid Flash upgrade focused on coding and agent workflows.
65% Pass@1 · high effort · ±2%
-
GLM-5.3
DeepSWE 69%
Open-weight coding leader improved entirely through post-training.
69% Pass@1 · max effort · ±3%
-
GLM-5.3-Flash
DeepSWE 63%
Native-multimodal GLM with frontier performance at flash cost.
63% Pass@1 · max effort · ±4%
How to read the marks
- 63%Filled bubble · exact current DeepSWE score
- Hollow marker · not evaluated in the captured leaderboard
Circle size is fixed. Higher scored placement means a higher Pass@1 value; hollow markers live on a separate N/A rail.
Reading note
What the marks mean
Each mark is one release and opens its official evidence. A filled bubble is scored; a hollow marker is historical context without a fabricated number.
DeepSWE v1.1 Pass@1 — current leaderboard observation, captured 27 August 2026; leaderboard updated 26 August 2026.
“Not evaluated” means no exact model row in the committed current DeepSWE v1.1 source. It does not mean no coding ability, and it is never treated as a zero.
Scores use the source’s best shown model-and-effort configuration. They are not necessarily release-day scores, and effort budgets differ. The displayed ± value is reproduced as published.
40 major AI model releases from 2022–2026, through 26 August 2026. Editorial selection, not an exhaustive model catalog.
View the DeepSWE leaderboard