AI Coding Benchmark Leaderboard: Cursor + DeepSWE

Authors
  • avatar
    Name
    Hamza Rahman
Published on
-
20 mins read
-

Every time I picked a model I ended up bouncing between two AI coding benchmarks: CursorBench and DeepSWE. They're both solid, but they rank models differently, and on its own neither one told me what I actually wanted to know.

So I merged them into one AI coding benchmark leaderboard: a single ranked table that weighs "how good is it" against "what does it actually cost me." Short version: Opus 5 occupies the first two combined rows and the high-quality end of the price/performance frontier. Grok 4.6 enters across all four effort tiers with Medium delivering standout mid-range value ($2.37 per task for 67.29% correctness). GPT-5.6 Sol and Luna provide the strongest implementation defaults, with Luna Max completely dominating the sub-65% tier on both accuracy and cost ($0.78 per solve). Update (August 31, 2026): CursorBench 3.2 added Grok 4.6 and Gemini 3.7 Flash while adjusting Sonnet 5 pricing. DeepSWE v1.1 updated its snapshot (August 26), adding Grok 4.6 and Gemini 3.7 Flash while repricing GPT-5.6 Sol and Gemini Flash runs on its live leaderboard. This update merges the refreshed runs across 51 matched configurations.

The Combined Leaderboard

Ranked by mean correctness (higher is better), with mean cost alongside. Use the buttons to toggle models on the chart and table. This snapshot uses CursorBench 3.2 (read August 31, 2026) and the DeepSWE v1.1 live leaderboard as updated August 26, 2026. Cost increases left to right on a logarithmic scale, so the better-value direction is up and left. Every effort tier is plotted. Hover or select a family to draw its effort curve, and toggle Show frontier to draw the Pareto staircase of undominated picks.

Strong value zone: ≥50% and ≤$5$0.1$0.25$0.5$1$2$5$10$2020%30%40%50%60%70%Mean cost per task ($, log scale)Mean correctness (%)↖ better = up & leftOpus 5Fable 5GPT-5.6 SolGrok 4.6GPT-5.6 TerraKimi K3GPT-5.6 LunaGemini 3.7 FlashGPT-5.5Opus 4.8Sonnet 5Gemini 3.6 FlashGLM 5.2Kimi K2.7 Code
Mean correctness vs mean cost (average of CursorBench + DeepSWE). Cost increases left to right on a log scale, and every effort tier is plotted. Hover or focus a family to draw its effort curve, or select a point to pin it. The shaded zone is an explicit editorial threshold: at least 50% correctness and at most $5 mean cost. The frontier control draws the dashed Pareto staircase through the ringed undominated points. Models on only one benchmark have no combined mean, so they're listed in the second table rather than plotted here.
RankModel / effortMean correctnessMean costCost / solvedCursorBenchDeepSWE Pass@1Cursor costDeepSWE cost
1Claude Opus 5 - Max71.82%$10.04$13.9870.00%73.65%$8.23$11.84
2Claude Opus 5 - Extra High71.23%$8.21$11.5369.30%73.15%$7.35$9.07
3Fable 5 - Max70.11%$19.48$27.7870.50%69.72%$17.32$21.63
4GPT-5.6 Sol - Max69.93%$6.08$8.6967.20%72.67%$5.69$6.46
5Claude Opus 5 - High69.76%$5.00$7.1766.70%72.83%$3.91$6.08
6Fable 5 - Extra High69.16%$12.57$18.1868.40%69.91%$11.73$13.41
7Grok 4.6 - Extra High68.77%$4.16$6.0570.80%66.74%$2.81$5.50
8GPT-5.6 Sol - Extra High67.62%$3.74$5.5364.50%70.73%$3.88$3.60
9Fable 5 - High67.55%$8.97$13.2866.50%68.60%$8.77$9.18
10Grok 4.6 - High67.54%$3.36$4.9769.90%65.19%$2.34$4.38
11Grok 4.6 - Medium67.29%$2.37$3.5267.10%67.48%$1.28$3.45
12GPT-5.6 Terra - Max67.26%$3.13$4.6564.90%69.62%$2.31$3.96
13Claude Opus 5 - Medium66.60%$3.29$4.9464.30%68.90%$3.29$3.29
14GPT-5.6 Sol - High66.45%$2.73$4.1163.50%69.40%$2.79$2.66
15Fable 5 - Medium65.28%$6.45$9.8865.20%65.37%$6.80$6.09
16Kimi K3 - Max64.66%$3.68$5.6960.80%68.51%$2.70$4.65
17GPT-5.6 Luna - Max64.14%$0.50$0.7861.10%67.19%$0.39$0.61
18Gemini 3.7 Flash - High63.43%$1.69$2.6661.60%65.27%$1.20$2.18
19GPT-5.5 - Extra High62.72%$5.04$8.0458.40%67.04%$2.85$7.23
20Gemini 3.7 Flash - Medium62.24%$1.49$2.3959.00%65.49%$0.95$2.03
21GPT-5.5 - High61.39%$3.57$5.8258.40%64.38%$2.05$5.10
22Fable 5 - Low60.84%$4.11$6.7662.10%59.58%$4.46$3.76
23Claude Opus 4.8 - Max60.64%$9.50$15.6762.30%58.97%$5.77$13.22
24GPT-5.6 Sol - Medium60.53%$1.69$2.7960.00%61.06%$1.95$1.42
25Claude Opus 5 - Low60.46%$2.10$3.4762.80%58.13%$2.55$1.66
26GPT-5.6 Terra - Extra High59.69%$1.42$2.3859.20%60.18%$1.15$1.70
27Claude Sonnet 5 - Max57.67%$15.35$26.6261.50%53.85%$4.30$26.40
28GPT-5.6 Luna - Extra High57.28%$0.27$0.4757.70%56.86%$0.23$0.31
29Claude Opus 4.8 - Extra High56.88%$6.25$10.9959.40%54.36%$4.50$8.01
30Claude Opus 4.8 - High54.88%$3.71$6.7658.00%51.77%$3.15$4.28
31Claude Sonnet 5 - Extra High54.18%$7.33$13.5358.70%49.67%$2.77$11.89
32GPT-5.6 Terra - High53.98%$0.81$1.5054.20%53.76%$0.71$0.91
33GPT-5.5 - Medium53.89%$2.13$3.9553.80%53.98%$1.51$2.75
34Gemini 3.7 Flash - Low53.78%$1.29$2.4053.80%53.76%$0.74$1.83
35Claude Sonnet 5 - High52.57%$4.78$9.0956.90%48.23%$2.13$7.43
36Claude Opus 4.8 - Medium52.39%$3.12$5.9656.10%48.67%$2.81$3.44
37Grok 4.6 - Low51.32%$0.87$1.7061.00%41.65%$0.70$1.04
38GPT-5.6 Luna - High50.52%$0.16$0.3256.80%44.25%$0.16$0.16
39Gemini 3.6 Flash - High50.09%$1.89$3.7753.50%46.68%$1.56$2.21
40GLM 5.2 - Max49.39%$2.84$5.7555.00%43.78%$1.76$3.92
41GPT-5.6 Sol - Low48.98%$0.92$1.8852.60%45.35%$1.01$0.82
42Claude Opus 4.8 - Low46.95%$2.16$4.6053.10%40.80%$2.02$2.29
43Claude Sonnet 5 - Medium46.09%$2.76$5.9952.40%39.78%$1.44$4.08
44GLM 5.2 - High43.89%$2.01$4.5851.50%36.28%$1.19$2.84
45GPT-5.6 Terra - Medium42.71%$0.48$1.1250.30%35.11%$0.49$0.47
46Kimi K2.7 Code40.12%$2.12$5.2849.70%30.53%$1.43$2.82
47Claude Sonnet 5 - Low39.11%$1.53$3.9147.70%30.51%$0.87$2.19
48GPT-5.5 - Low36.80%$1.09$2.9646.60%26.99%$0.98$1.20
49GPT-5.6 Terra - Low35.48%$0.38$1.0746.90%24.05%$0.42$0.34
50GPT-5.6 Luna - Medium29.49%$0.06$0.2047.70%11.28%$0.08$0.04
51GPT-5.6 Luna - Low19.57%$0.02$0.1037.60%1.55%$0.03$0.01

On one benchmark only (no combined mean)

Not plotted above and not ranked: a single score isn't comparable to a mean, and the benchmarks differ in difficulty. These rows are useful references, but they don't answer the combined CursorBench + DeepSWE question.

ModelBenchmarkScoreCostMean
Kimi K3 - HighCursorBench59.70%$1.89n/a
Composer 2.5CursorBench56.10%$0.44n/a
Gemini 3.6 Flash - MediumCursorBench51.20%$1.48n/a
Kimi K3 - LowCursorBench50.50%$0.99n/a
Gemini 3.6 Flash - LowCursorBench47.40%$1.13n/a
GLM 5.3 - MaxDeepSWE Pass@168.96%$3.99n/a
GLM 5.3 Flash - MaxDeepSWE Pass@163.39%$0.24n/a
DeepSeek V4 Pro - MaxDeepSWE Pass@162.83%$1.67n/a
Qwen 3.8 Max - Extra HighDeepSWE Pass@157.46%$3.73n/a
Muse Spark 1.2 - Extra HighDeepSWE Pass@154.87%$3.70n/a
Grok 4.5 - HighDeepSWE Pass@153.76%$2.42n/a
DeepSeek V4 Flash - MaxDeepSWE Pass@153.32%$0.46n/a
Muse Spark 1.1 - Extra HighDeepSWE Pass@153.32%$2.36n/a
GPT-5.4 - Extra HighDeepSWE Pass@151.77%$5.65n/a
Gemini 3.5 Flash - HighDeepSWE Pass@136.06%$3.45n/a
Claude Sonnet 4.6 - HighDeepSWE Pass@129.93%$5.52n/a
Gemini 3.1 Pro Preview - HighDeepSWE Pass@111.73%$2.14n/a

Why Two Benchmarks?

Single benchmarks saturate. Once every frontier model scores in the high 90s, the leaderboard stops telling you anything. CursorBench is closer to real in-editor work: short, underspecified prompts from actual Cursor sessions. DeepSWE is harder and more controlled: 113 hand-written long-horizon tasks across 91 repos, scored with behavioral verifiers.

One is real-world and messy, the other is original and hard. Put together, they land closer to day-to-day usefulness than either does alone.

How I Combined Them

I kept the math deliberately simple so anyone can reproduce it:

  • Mean correctness = average of (CursorBench correctness, DeepSWE Pass@1)
  • Mean cost = average of (CursorBench cost, DeepSWE cost per task)
  • Cost / solved task = mean cost divided by mean correctness as a probability

No weighting, no normalization tricks. Both benchmarks already report cost in dollars per task, so the averages stay in real units. Then I ranked by mean correctness.

One caveat up front: a flat average treats both benchmarks as equally important, which is a choice, not a law. Averaging costs also mixes two different task sizes: DeepSWE tasks are larger, so its cost-per-task is not the same unit as CursorBench's cost-per-task. I still include cost / solved task because it's easier to reason about as a buyer: under this combined view, how much do I pay for one expected successful solve?

If your work looks more like quick in-editor edits, lean on the CursorBench column. If it looks like large multi-file features, lean on DeepSWE.

A few configurations only show up on one benchmark: Kimi K3 High/Low, Composer 2.5, and Gemini 3.6 Flash Medium/Low on CursorBench; GLM 5.3 Max, GLM 5.3 Flash Max, DeepSeek V4 Pro Max, Qwen 3.8 Max Extra High, Muse Spark 1.2 Extra High, Grok 4.5 High, DeepSeek V4 Flash Max, Muse Spark 1.1 Extra High, GPT-5.4 Extra High, Gemini 3.5 Flash High, Claude Sonnet 4.6 High, and Gemini 3.1 Pro Preview High on DeepSWE. Averaging a single score with a missing one isn't a mean, so I keep them out of the ranked table and off the chart.

What Jumps Out

Opus 5 occupies the first two rows. Max reaches 71.82% mean correctness at $10.04 mean cost, while Extra High reaches 71.23% at $8.21. Fable 5 Max is third at 70.11% and $19.48, followed by GPT-5.6 Sol Max at 69.93% and $6.08. Small score gaps still overlap benchmark uncertainty, but Opus 5 has the highest combined score in this snapshot.

Opus 5 High is the standout upper-end value tier at rank five (69.76% mean correctness, $5.00 per task, and $7.17 per expected solve). Medium ranks 13th at 66.60% and $3.29. Max buys the highest score, but Extra High and High make the stronger price/performance cases.

Grok 4.6 is the biggest new addition, entering all four effort tiers on both benchmarks. Extra High reaches rank 7 (68.77% correctness at $4.16 mean cost, $6.05 per solve). Grok 4.6 High is rank 10 (67.54% at $3.36, $4.97 per solve), and Grok 4.6 Medium is rank 11 (67.29% at $2.37, $3.52 per solve). Grok 4.6 Medium stands out at $2.37 per task, beating Terra Max's correctness (67.29% vs 67.26%) for 76 cents less per task and displacing Terra Max from the Pareto frontier.

DeepSWE's latest update repriced GPT-5.6 Sol downward. Sol Max is now $6.08 per task ($8.69 per solve, down from $10.07). Sol Extra High is $3.74 ($5.53 per solve), and Sol High is $2.73 ($4.11 per solve). Together with Luna Max at $0.50 per task ($0.78 per solve), GPT-5.6 remains the most economical stack for everyday code generation.

Gemini 3.7 Flash also gains a combined row across three tiers (High at rank 18 with 63.43% for $1.69; Medium at rank 20 with 62.24% for $1.49). While cheap on paper, Gemini 3.7 Flash is strictly dominated by GPT-5.6 Luna Max: Luna Max is both more accurate (64.14% vs 63.43%) and over three times cheaper per solve ($0.78 vs $2.39–$2.66), keeping Luna Max firmly on the Pareto frontier.

Kimi K3 Max enters at rank 16 with 64.66% mean correctness, $3.68 mean cost, and $5.69 per expected solve. That is a solid combined result, though Grok 4.6 Medium and Luna Max offer better cost-to-correctness ratios.

The low-effort GPT-5.6 rows show why cost alone is dangerous. Luna Low is only $0.02 per task, but its 19.57% mean correctness makes it a poor default. Luna Medium is $0.06 at only 29.49%. The useful frontier starts around Luna High at 50.52% and $0.16, then improves quickly as effort rises.

Do not over-read tiny gaps among the highest rows. DeepSWE reports uncertainty, benchmark tasks differ, and adjacent configurations often overlap. The meaningful differences between Opus 5 Max, Extra High, Fable Max, Sol Max, and Opus 5 High are cost and workflow fit, not just fractions of a percentage point.

Cost Per Solved Task, Ranked

Mean cost tells you what one attempt costs. What you actually pay for is a success. Cost per solved task folds the failures into the price: if a model costs $1 per task but only solves half of them, each solve really costs you $2. That's the number I care about most as a buyer, and it gets a simpler chart than the scatter: a ranked bar list, cheapest solve first.

The chart starts with a 50% correctness floor, because without one the ranking is misleading. Luna Low looks unbeatable at $0.10 per solve, but it only solves 19.57% of tasks, so most of your attempts would be wasted. Luna Medium is similarly cheap at $0.20 per solve but only reaches 29.49%. Pick All models in the dropdown if you want to see those rows anyway, or raise the floor to match your own quality bar.

Opus 5GPT-5.6 SolGPT-5.6 TerraGPT-5.6 LunaGPT-5.5Opus 4.8Sonnet 5GLM 5.2Gemini 3.7 FlashGemini 3.6 FlashKimi K3Kimi K2.7 CodeGrok 4.6Fable 5
Cost per solved task ($), lower is better$0$5$10$15$20$25$30GPT-5.6 Luna - High · 50.52% correct · $0.32 per solved taskGPT-5.6 Luna - High · 51%$0.32GPT-5.6 Luna - Extra High · 57.28% correct · $0.47 per solved taskGPT-5.6 Luna - Extra High · 57%$0.47GPT-5.6 Luna - Max · 64.14% correct · $0.78 per solved taskGPT-5.6 Luna - Max · 64%$0.78GPT-5.6 Terra - High · 53.98% correct · $1.50 per solved taskGPT-5.6 Terra - High · 54%$1.50Grok 4.6 - Low · 51.32% correct · $1.70 per solved taskGrok 4.6 - Low · 51%$1.70GPT-5.6 Terra - Extra High · 59.69% correct · $2.38 per solved taskGPT-5.6 Terra - Extra High · 60%$2.38Gemini 3.7 Flash - Medium · 62.24% correct · $2.39 per solved taskGemini 3.7 Flash - Medium · 62%$2.39Gemini 3.7 Flash - Low · 53.78% correct · $2.40 per solved taskGemini 3.7 Flash - Low · 54%$2.40Gemini 3.7 Flash - High · 63.43% correct · $2.66 per solved taskGemini 3.7 Flash - High · 63%$2.66GPT-5.6 Sol - Medium · 60.53% correct · $2.79 per solved taskGPT-5.6 Sol - Medium · 61%$2.79Claude Opus 5 - Low · 60.46% correct · $3.47 per solved taskOpus 5 - Low · 60%$3.47Grok 4.6 - Medium · 67.29% correct · $3.52 per solved taskGrok 4.6 - Medium · 67%$3.52Gemini 3.6 Flash - High · 50.09% correct · $3.77 per solved taskGemini 3.6 Flash - High · 50%$3.77GPT-5.5 - Medium · 53.89% correct · $3.95 per solved taskGPT-5.5 - Medium · 54%$3.95GPT-5.6 Sol - High · 66.45% correct · $4.11 per solved taskGPT-5.6 Sol - High · 66%$4.11GPT-5.6 Terra - Max · 67.26% correct · $4.65 per solved taskGPT-5.6 Terra - Max · 67%$4.65Claude Opus 5 - Medium · 66.60% correct · $4.94 per solved taskOpus 5 - Medium · 67%$4.94Grok 4.6 - High · 67.54% correct · $4.97 per solved taskGrok 4.6 - High · 68%$4.97GPT-5.6 Sol - Extra High · 67.62% correct · $5.53 per solved taskGPT-5.6 Sol - Extra High · 68%$5.53Kimi K3 - Max · 64.66% correct · $5.69 per solved taskKimi K3 - Max · 65%$5.69GPT-5.5 - High · 61.39% correct · $5.82 per solved taskGPT-5.5 - High · 61%$5.82Claude Opus 4.8 - Medium · 52.39% correct · $5.96 per solved taskOpus 4.8 - Medium · 52%$5.96Grok 4.6 - Extra High · 68.77% correct · $6.05 per solved taskGrok 4.6 - Extra High · 69%$6.05Fable 5 - Low · 60.84% correct · $6.76 per solved taskFable 5 - Low · 61%$6.76Claude Opus 4.8 - High · 54.88% correct · $6.76 per solved taskOpus 4.8 - High · 55%$6.76Claude Opus 5 - High · 69.76% correct · $7.17 per solved taskOpus 5 - High · 70%$7.17GPT-5.5 - Extra High · 62.72% correct · $8.04 per solved taskGPT-5.5 - Extra High · 63%$8.04GPT-5.6 Sol - Max · 69.93% correct · $8.69 per solved taskGPT-5.6 Sol - Max · 70%$8.69Claude Sonnet 5 - High · 52.57% correct · $9.09 per solved taskSonnet 5 - High · 53%$9.09Fable 5 - Medium · 65.28% correct · $9.88 per solved taskFable 5 - Medium · 65%$9.88Claude Opus 4.8 - Extra High · 56.88% correct · $10.99 per solved taskOpus 4.8 - Extra High · 57%$10.99Claude Opus 5 - Extra High · 71.23% correct · $11.53 per solved taskOpus 5 - Extra High · 71%$11.53Fable 5 - High · 67.55% correct · $13.28 per solved taskFable 5 - High · 68%$13.28Claude Sonnet 5 - Extra High · 54.18% correct · $13.53 per solved taskSonnet 5 - Extra High · 54%$13.53Claude Opus 5 - Max · 71.82% correct · $13.98 per solved taskOpus 5 - Max · 72%$13.98Claude Opus 4.8 - Max · 60.64% correct · $15.67 per solved taskOpus 4.8 - Max · 61%$15.67Fable 5 - Extra High · 69.16% correct · $18.18 per solved taskFable 5 - Extra High · 69%$18.18Claude Sonnet 5 - Max · 57.67% correct · $26.62 per solved taskSonnet 5 - Max · 58%$26.62Fable 5 - Max · 70.11% correct · $27.78 per solved taskFable 5 - Max · 70%$27.78
Cost per solved task = mean cost divided by mean correctness as a probability, so failed attempts are amortized into the price of a success. The chart starts at a 50% correctness floor to keep misleadingly cheap low-effort tiers out of view; every row carries its mean correctness. Use the dropdowns to change the floor or cap the price per solve (lower is better). The sort order never changes; filtering removes rows, or fades them if you switch to dim mode.

What this view adds to the scatter:

  • GPT-5.6 supplies the useful cheap end. With a 50% quality floor, Luna High costs $0.32 per solve, Luna Extra High $0.47, and Luna Max $0.78.
  • Grok 4.6 Medium anchors the mid-tier step-up. Grok 4.6 Medium reaches 67.29% for $3.52 per solve, offering the sharpest jump in accuracy once you step past Luna Max.
  • GPT-5.6 Sol provides efficient backend step-ups. Sol High reaches 66.45% for $4.11 per solve, while Sol Extra High reaches 67.62% for $5.53 per solve.
  • Opus 5 covers the quality end. Medium reaches 66.60% for $4.94 per solve, High reaches 69.76% for $7.17, Extra High reaches 71.23% for $11.53, and Max reaches 71.82% for $13.98. High is the value tier; Max is for when the last point matters.
  • Sonnet 5 improved on pricing, but still lags in solve value. CursorBench lowered Sonnet 5 costs, but its heavy output on DeepSWE keeps Sonnet 5 Max at $15.35 mean cost and $26.62 per solve for 57.67% correctness.

Best Value Picks

If you don't want to stare at the table, here's how I'd choose:

PickWhy
GPT-5.6 Luna MaxBest budget & implementation balance: 64.14% at $0.78 per solved task
Grok 4.6 MediumStandout mid-range value: 67.29% at $3.52 per solved task
GPT-5.6 Sol HighEfficient backend step-up: 66.45% at $4.11 per solved task
Claude Opus 5 HighUpper-end value: 69.76% at $7.17 per solved task
Claude Opus 5 Extra HighNear-best correctness: 71.23% at $11.53 per solved task

For most routine implementation tickets, Luna Max is the obvious default. Grok 4.6 Medium and Sol High provide the cleanest step-ups when tasks require heavier backend logic. Sol High gives up only 0.81 correctness points to Terra Max at $4.11 vs $4.65 per solve while using far fewer thinking tokens and steps in these runs.

The important catch is the pricing window. The restored Fable rollout gave some paid subscribers included usage that was originally set to end July 7, and Anthropic has since extended it. That keeps Fable inside normal coding-plan economics for now, but the window is still temporary: once it closes, Fable moves back to separate paid usage or usage credits, and the math changes. The whole point of a coding plan is that you can spend expensive reasoning on planning without worrying about every token.

Opus 5, GPT-5.6, and Kimi K3 availability can still vary by product and account. The benchmark result tells you what the models did under CursorBench and DeepSWE, not whether your editor or API plan exposes every family and effort tier today.

How I Actually Use These (My Take)

Benchmarks rank models on synthetic tasks. They don't tell you what a model feels like when you're staring at an editor for six hours. This is the setup I've settled into based on real-world friction and economics:

The Daily Drivers

  • High-frequency edits & terminal agents → GLM 5.3 Flash (Max). Back when it was tested anonymously under stealth handles like "Ox Alpha" (often called "0x Alpha" across developer forums), it immediately stood out for speed and agentic competence. On DeepSWE, it scores a massive 63.39% Pass@1 at just $0.24 per task ($0.38 per expected solve). It's insanely cheap, responds almost instantly, and handles routine repo edits without burning cash.
  • Daily drafting & large-context dumps → Gemini 3.7. Pure benchmark $/solve ranks tell only part of the story. Google bundles 2TB of Google Drive/Photos cloud storage into its Google One AI Premium subscription ($19.99/mo). If you already pay for cloud storage, the effective marginal cost of Gemini Advanced is basically zero. Combined with the 1M+ token context window, it's my go-to for dumping entire codebases, documentation sets, and rough prototypes.

Planning & Difficult Logic

  • Architecture, planning, & complex backend tasks → GPT-5.6 Sol High. When a problem is genuinely tricky or needs multi-step architectural planning, Sol High is my default. It hits 66.45% combined correctness ($4.11 per solve, reaching 69.40% on DeepSWE) while keeping thinking tokens concise. It plans cleanly without hallucinating fluff or stalling on long inference pauses.

The Frontier Experience: Opus 5 vs Fable 5

  • Opus 5 (Great coder, but not fun to work with). Statistically, Opus 5 is a beast: rank one (71.82% Max) and rank two (71.23% Extra High). But subjectively, it's just not fun to collaborate with. It can feel stiff, pedantic, and overly literal—technically correct on syntax, but tedious when iterating on nuanced designs or complex refactors. I keep Opus 5 Max reserved strictly as an emergency rescue tier for bugs nothing else could solve.
  • Fable 5 (The model you actually want when available). When Anthropic's usage window is open or limits allow, Fable 5 is the model I prefer working with. It has an unmistakable "big model" feel: deeper intuitive judgment, taste in code structure, and a collaborative cadence that makes hard problem-solving feel effortless. Even at higher per-task costs ($19.48 Max, $8.97 High), the developer experience is in a different league.

If you're wiring these models into something beyond the editor, I used the same GPT models in my walkthrough on building a RAG system with Node.js and OpenAI.

Caveats

A few things to keep in mind before you treat this as gospel:

  • It's an unweighted average of two benchmarks. That's a judgment call, so weight the columns toward your real workload.
  • Cost is per task and provider-dependent. Your effective cost shifts with caching, context size, and how you drive the agent. Cost / solved task is easier to read, but it still comes from the same averaged benchmark costs.
  • DeepSWE's live page has repriced Sol, Terra, Luna, and Gemini runs, while its downloadable JSON artifact generated August 26, 2026 contains exact Pass@1 data. This snapshot uses the artifact for exact Pass@1 and the live page for displayed costs.
  • Fable 5 availability depends on plan and platform. Anthropic includes Fable usage for up to 50% of weekly limits on some paid plans, a window originally ending July 7 that has since been extended; usage credits apply once it closes. Cloud providers are being re-enabled separately.
  • Fable 5 has stricter safety routing. Anthropic says blocked requests are sent to Opus 4.8, and the new classifier may flag some normal coding/debugging requests.
  • Opus 5, GPT-5.6, Grok 4.6, Gemini 3.7 Flash, and Kimi K3 availability varies by product and account even when matching runs appear on both source benchmarks.
  • Grok 4.6 now replaces Grok 4.5 across all effort tiers on both benchmarks without earlier training-snapshot caveats.
  • Benchmarks measure benchmark tasks. They're a strong signal, not a substitute for trying the model on your own repo.
  • These numbers are a snapshot. Both sources can change between snapshots, so re-run the merge when either leaderboard or its pricing changes.

What You Get From This

One ranked view of AI coding agents that weighs skill against cost, built from two benchmarks instead of one. The takeaways:

  1. Opus 5 Max has the highest combined score at 71.82%, with Extra High second at 71.23%.
  2. Grok 4.6 enters the combined board with strong numbers across all tiers; Grok 4.6 Medium (67.29% at $2.37) is one of the best price/performance options available.
  3. GPT-5.6 Sol and Luna supply top-tier implementation economics after DeepSWE repricing; Luna Max averages $0.50 per task and $0.78 per expected solve.
  4. Luna Max completely dominates Gemini 3.7 Flash on both accuracy (64.14% vs 63.43%) and cost per solve ($0.78 vs $2.39–$2.66).
  5. My workflow split is Opus 5 for UI and agentic work, GPT-5.6 and Grok 4.6 for implementation and backend work.
  6. Sonnet 5 is repriced lower on CursorBench, but remains expensive per solve due to high token counts on DeepSWE.
  7. Low effort can be too weak even when it is cheap. Apply a quality floor before reading cost-per-solve ranks.

If you want a version tuned to your own work, swap the flat average for a weighted one and re-rank.

References

  1. CursorBench current leaderboard: CursorBench 3.2, read August 31, 2026
  2. DeepSWE v1.1 live leaderboard and current displayed costs
  3. DeepSWE v1.1 artifact for exact Pass@1, generated August 26, 2026
  4. Anthropic: Redeploying Fable 5
  5. Anthropic: statement on the US government directive suspending Fable 5 and Mythos 5 access

Article Changelog

  • August 31, 2026 — Grok 4.6, Gemini 3.7 Flash, and Sol repricing: Added all four Grok 4.6 tiers and three Gemini 3.7 Flash tiers, integrated DeepSWE's August 26 run updates and Sol repricing, factored in CursorBench's Sonnet 5 price reduction, and updated the frontier and Best Value table.
  • August 1, 2026 — Luna and Terra pricing refresh: Updated both benchmarks' GPT-5.6 Luna and Terra costs, recomputed every affected mean and cost-per-solve value, and revised the frontier, Best Value picks, and workflow guidance.
  • July 29, 2026 — Opus 5, Kimi K3, and Gemini 3.6 Flash: Added the configurations with results on both benchmarks to the combined ranking and kept unmatched effort tiers in the one-benchmark-only table.