AI Frontier

DeepSWE: Frontier Coding Agents

Frontier models measured by Datacurve's DeepSWE on original, long-horizon software engineering tasks.

Preparing district view
DeepSWE Yards Embedded district view
Explore full city

Latest ranking

August 4, 2026

DeepSWE pass rate (%)
#1 claude-opus-5 73.6
#2 gpt-5-6-sol 72.7
#3 claude-fable-5 69.9
#4 gpt-5-6-terra 69.6
#5 kimi-k3 68.5
#6 gpt-5-6-luna 67.2
#7 gpt-5-5 67
#8 claude-opus-4-8 59
#9 qwen3-8-max 57.5
#10 claude-sonnet-5 53.8
#10 grok-4-5 53.8
#12 muse-spark-1-1 53.3
#13 gpt-5-4 51.8
#14 gemini-3-6-flash 48.6
#15 glm-5-2 43.8

Timeline

Key moments

  • May 31, 2026

    gpt-5-5 leads the DeepSWE pack

  • July 13, 2026

    gpt-5-6-sol leads the DeepSWE pack

  • August 4, 2026

    claude-opus-5 leads the DeepSWE pack

Notes

Methodology

This adapter reads the public DeepSWE v1.1 live leaderboard published by Datacurve, keeps each frontier model's best run across reasoning efforts on 113 original long-horizon software engineering tasks, ranks them by pass rate, and stores daily StoryOps snapshots so the skyline replays how the coding frontier moves.

Observed 51 DeepSWE v1.1 runs across 19 models at 2026-08-04T01:34:32.908483+00:00.

Reference

Sources