AI Coding Benchmark Leaderboard: Cursor + DeepSWE
- Authors
- Name
- Hamza Rahman
- Published on
- -20 mins read-
Every time I picked a model I ended up bouncing between two AI coding benchmarks: CursorBench and DeepSWE. They're both solid, but they rank models differently, and on its own neither one told me what I actually wanted to know.
So I merged them into one AI coding benchmark leaderboard: a single ranked table that weighs "how good is it" against "what does it actually cost me." Short version: Opus 5 occupies the first two combined rows and the high-quality end of the price/performance frontier. GPT-5.6 supplies the strongest low-cost tradeoffs after both leaderboards repriced Luna and Terra, and remains my default implementation family, while Kimi K3 Max joins the combined table as a credible alternative.
Update (August 1, 2026): CursorBench 3.2's July 30 changelog says GPT-5.6 Terra and Luna costs were updated. The DeepSWE live leaderboard now shows the same pricing correction across its Terra and Luna rows while keeping the July 25 run snapshot and scores. DeepSWE's downloadable JSON still carries the older cost fields, so I use its exact Pass@1 values and the repriced costs displayed on the live page. Together, the two corrections cut every Terra and Luna combined cost without changing correctness ranks.
The Combined Leaderboard
Ranked by mean correctness (higher is better), with mean cost alongside. Use the buttons to toggle models on the chart and table. This snapshot uses CursorBench 3.2 and the DeepSWE v1.1 live leaderboard as displayed August 1, 2026; DeepSWE's underlying runs are still the July 25 snapshot. Cost increases left to right on a logarithmic scale, so the better-value direction is up and left. Every effort tier is plotted. Hover or select a family to draw its effort curve, and toggle Show frontier to draw the Pareto staircase of undominated picks.
| Rank | Model / effort | Mean correctness | Mean cost | Cost / solved | CursorBench | DeepSWE Pass@1 | Cursor cost | DeepSWE cost |
|---|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 - Max | 71.82% | $10.03 | $13.97 | 70.00% | 73.65% | $8.23 | $11.84 |
| 2 | Claude Opus 5 - Extra High | 71.23% | $8.21 | $11.53 | 69.30% | 73.15% | $7.35 | $9.07 |
| 3 | Fable 5 - Max | 70.11% | $19.48 | $27.78 | 70.50% | 69.72% | $17.32 | $21.63 |
| 4 | GPT-5.6 Sol - Max | 69.93% | $7.04 | $10.07 | 67.20% | 72.67% | $5.69 | $8.39 |
| 5 | Claude Opus 5 - High | 69.76% | $4.99 | $7.15 | 66.70% | 72.83% | $3.91 | $6.08 |
| 6 | Fable 5 - Extra High | 69.16% | $12.57 | $18.18 | 68.40% | 69.91% | $11.73 | $13.41 |
| 7 | GPT-5.6 Sol - Extra High | 67.62% | $4.29 | $6.34 | 64.50% | 70.73% | $3.88 | $4.70 |
| 8 | Fable 5 - High | 67.55% | $8.97 | $13.28 | 66.50% | 68.60% | $8.77 | $9.18 |
| 9 | GPT-5.6 Terra - Max | 67.26% | $3.14 | $4.67 | 64.90% | 69.62% | $2.31 | $3.96 |
| 10 | Claude Opus 5 - Medium | 66.60% | $3.29 | $4.94 | 64.30% | 68.90% | $3.29 | $3.29 |
| 11 | GPT-5.6 Sol - High | 66.45% | $3.13 | $4.71 | 63.50% | 69.40% | $2.79 | $3.47 |
| 12 | Fable 5 - Medium | 65.28% | $6.44 | $9.87 | 65.20% | 65.37% | $6.80 | $6.09 |
| 13 | Kimi K3 - Max | 64.66% | $3.68 | $5.69 | 60.80% | 68.51% | $2.70 | $4.65 |
| 14 | GPT-5.6 Luna - Max | 64.14% | $0.50 | $0.78 | 61.10% | 67.19% | $0.39 | $0.61 |
| 15 | GPT-5.5 - Extra High | 62.72% | $5.04 | $8.04 | 58.40% | 67.04% | $2.85 | $7.23 |
| 16 | GPT-5.5 - High | 61.39% | $3.58 | $5.83 | 58.40% | 64.38% | $2.05 | $5.10 |
| 17 | Fable 5 - Low | 60.84% | $4.11 | $6.76 | 62.10% | 59.58% | $4.46 | $3.76 |
| 18 | Claude Opus 4.8 - Max | 60.64% | $9.50 | $15.67 | 62.30% | 58.97% | $5.77 | $13.22 |
| 19 | GPT-5.6 Sol - Medium | 60.53% | $1.91 | $3.16 | 60.00% | 61.06% | $1.95 | $1.86 |
| 20 | Claude Opus 5 - Low | 60.46% | $2.11 | $3.49 | 62.80% | 58.13% | $2.55 | $1.66 |
| 21 | Grok 4.5 - High* | 60.23% | $1.96 | $3.25 | 66.70% | 53.76% | $1.51 | $2.42 |
| 22 | GPT-5.6 Terra - Extra High | 59.69% | $1.43 | $2.40 | 59.20% | 60.18% | $1.15 | $1.70 |
| 23 | Claude Sonnet 5 - Max | 57.67% | $16.42 | $28.47 | 61.50% | 53.85% | $6.45 | $26.40 |
| 24 | GPT-5.6 Luna - Extra High | 57.28% | $0.27 | $0.47 | 57.70% | 56.86% | $0.23 | $0.31 |
| 25 | Claude Opus 4.8 - Extra High | 56.88% | $6.25 | $10.99 | 59.40% | 54.36% | $4.50 | $8.01 |
| 26 | Claude Opus 4.8 - High | 54.88% | $3.72 | $6.78 | 58.00% | 51.77% | $3.15 | $4.28 |
| 27 | Claude Sonnet 5 - Extra High | 54.18% | $8.03 | $14.82 | 58.70% | 49.67% | $4.16 | $11.89 |
| 28 | GPT-5.6 Terra - High | 53.98% | $0.81 | $1.50 | 54.20% | 53.76% | $0.71 | $0.91 |
| 29 | GPT-5.5 - Medium | 53.89% | $2.13 | $3.95 | 53.80% | 53.98% | $1.51 | $2.75 |
| 30 | Claude Sonnet 5 - High | 52.57% | $5.31 | $10.10 | 56.90% | 48.23% | $3.19 | $7.43 |
| 31 | Claude Opus 4.8 - Medium | 52.39% | $3.13 | $5.97 | 56.10% | 48.67% | $2.81 | $3.44 |
| 32 | Gemini 3.6 Flash - High | 51.03% | $2.55 | $5.00 | 53.50% | 48.56% | $1.56 | $3.53 |
| 33 | GPT-5.6 Luna - High | 50.52% | $0.16 | $0.32 | 56.80% | 44.25% | $0.16 | $0.16 |
| 34 | GLM 5.2 - Max | 49.39% | $2.84 | $5.75 | 55.00% | 43.78% | $1.76 | $3.92 |
| 35 | GPT-5.6 Sol - Low | 48.98% | $1.04 | $2.12 | 52.60% | 45.35% | $1.01 | $1.07 |
| 36 | Claude Opus 4.8 - Low | 46.95% | $2.16 | $4.60 | 53.10% | 40.80% | $2.02 | $2.29 |
| 37 | Claude Sonnet 5 - Medium | 46.09% | $3.12 | $6.77 | 52.40% | 39.78% | $2.16 | $4.08 |
| 38 | GLM 5.2 - High | 43.89% | $2.01 | $4.58 | 51.50% | 36.28% | $1.19 | $2.84 |
| 39 | Gemini 3.5 Flash - Default/Medium | 43.09% | $4.77 | $11.07 | 48.80% | 37.39% | $2.20 | $7.34 |
| 40 | GPT-5.6 Terra - Medium | 42.71% | $0.48 | $1.12 | 50.30% | 35.11% | $0.49 | $0.47 |
| 41 | Kimi K2.7 Code | 40.12% | $2.12 | $5.28 | 49.70% | 30.53% | $1.43 | $2.82 |
| 42 | Claude Sonnet 5 - Low | 39.11% | $1.74 | $4.45 | 47.70% | 30.51% | $1.30 | $2.19 |
| 43 | GPT-5.5 - Low | 36.80% | $1.09 | $2.96 | 46.60% | 26.99% | $0.98 | $1.20 |
| 44 | GPT-5.6 Terra - Low | 35.48% | $0.38 | $1.07 | 46.90% | 24.05% | $0.42 | $0.34 |
| 45 | GPT-5.6 Luna - Medium | 29.49% | $0.06 | $0.20 | 47.70% | 11.28% | $0.08 | $0.04 |
| 46 | GPT-5.6 Luna - Low | 19.57% | $0.02 | $0.10 | 37.60% | 1.55% | $0.03 | $0.01 |
On one benchmark only (no combined mean)
Not plotted above and not ranked: a single score isn't comparable to a mean, and the benchmarks differ in difficulty. These rows are useful references, but they don't answer the combined CursorBench + DeepSWE question.
| Model | Benchmark | Score | Cost | Mean |
|---|---|---|---|---|
| Grok 4.5 - Medium* | CursorBench | 65.40% | $1.54 | n/a |
| Grok 4.5 - Low* | CursorBench | 63.50% | $1.22 | n/a |
| Kimi K3 - High | CursorBench | 59.70% | $1.89 | n/a |
| Composer 2.5 | CursorBench | 56.10% | $0.44 | n/a |
| Gemini 3.6 Flash - Medium | CursorBench | 51.20% | $1.48 | n/a |
| Kimi K3 - Low | CursorBench | 50.50% | $0.99 | n/a |
| Gemini 3.6 Flash - Low | CursorBench | 47.40% | $1.13 | n/a |
| Muse Spark 1.1 - Extra High | DeepSWE Pass@1 | 53.32% | $2.36 | n/a |
| GPT-5.4 - Extra High | DeepSWE Pass@1 | 51.77% | $5.65 | n/a |
| Claude Sonnet 4.6 - High | DeepSWE Pass@1 | 29.93% | $5.52 | n/a |
| Gemini 3.1 Pro Preview - High | DeepSWE Pass@1 | 11.75% | $9.48 | n/a |
Why Two Benchmarks?
Single benchmarks saturate. Once every frontier model scores in the high 90s, the leaderboard stops telling you anything. CursorBench is closer to real in-editor work: short, underspecified prompts from actual Cursor sessions. DeepSWE is harder and more controlled: 113 hand-written long-horizon tasks across 91 repos, scored with behavioral verifiers.
One is real-world and messy, the other is original and hard. Put together, they land closer to day-to-day usefulness than either does alone.
How I Combined Them
I kept the math deliberately simple so anyone can reproduce it:
- Mean correctness = average of (CursorBench correctness, DeepSWE Pass@1)
- Mean cost = average of (CursorBench cost, DeepSWE cost per task)
- Cost / solved task = mean cost divided by mean correctness as a probability
No weighting, no normalization tricks. Both benchmarks already report cost in dollars per task, so the averages stay in real units. Then I ranked by mean correctness.
One caveat up front: a flat average treats both benchmarks as equally important, which is a choice, not a law. Averaging costs also mixes two different task sizes: DeepSWE tasks are larger, so its cost-per-task is not the same unit as CursorBench's cost-per-task. I still include cost / solved task because it's easier to reason about as a buyer: under this combined view, how much do I pay for one expected successful solve?
If your work looks more like quick in-editor edits, lean on the CursorBench column. If it looks like large multi-file features, lean on DeepSWE.
A few configurations only show up on one benchmark: Grok 4.5 Low/Medium, Kimi K3 High/Low, Gemini 3.6 Flash Medium/Low, and Composer 2.5 on CursorBench; Muse Spark 1.1 Extra High, GPT-5.4 Extra High, Claude Sonnet 4.6 High, and Gemini 3.1 Pro Preview High on DeepSWE. Averaging a single score with a missing one isn't a mean, so I keep them out of the ranked table and off the chart. Grok High has both scores, but its combined result still inherits Cursor's training-snapshot caveat.
What Jumps Out
Opus 5 occupies the first two rows. Max reaches 71.82% mean correctness at $10.03 mean cost, while Extra High reaches 71.23% at $8.21. Fable 5 Max is third at 70.11% and $19.48, followed by GPT-5.6 Sol Max at 69.93% and $7.04. Small score gaps still overlap benchmark uncertainty, but Opus 5 has the highest combined score in this snapshot.
Opus 5 High is the standout upper-end value tier: rank five at 69.76% mean correctness, $4.99 per task, and $7.15 per expected solve. Medium ranks tenth at 66.60% and $3.29, narrowly beating Sol High's 66.45% at $3.13. Max buys the highest score, but Extra High and High make the stronger price/performance cases.
Kimi K3 Max now appears on both benchmarks and enters at rank 13 with 64.66% mean correctness, $3.68 mean cost, and $5.69 per expected solve. That is a solid first combined result, although Opus 5 Medium is both more correct and cheaper. Kimi K3 High and Low still lack DeepSWE matches, so they stay in the separate table.
The updated prices put GPT-5.6 across the budget-to-midrange frontier. Luna supplies the cheapest useful steps and now reaches 64.14% at only $0.50 per task. Terra Max reaches 67.26% at $3.14, then Sol and Opus 5 cover the upper-quality steps. The frontier is a cost/correctness view, not a verdict on an entire model family.
GPT-5.6 now has the clearest implementation-value range. The two pricing corrections put Luna Max at 64.14% for $0.50 per task and $0.78 per expected solve. Terra Max reaches 67.26% at $3.14 and $4.67 per solve, narrowly improving on Sol High's 66.45% at $3.13 and $4.71 per solve. Sol Max remains close to the highest scores at 69.93% for $7.04 per task, well below Opus 5 Max's cost.
Gemini 3.6 Flash High also gains a combined row, but it enters at rank 32 with 51.03% mean correctness and $2.55 mean cost. Its Medium and Low tiers remain CursorBench-only. Grok 4.5 High moves to rank 21; it is still caveated because its CursorBench half uses the training-exposed snapshot.
The low-effort GPT-5.6 rows show why cost alone is dangerous. Luna Low is only $0.02 per task, but its 19.57% mean correctness makes it a poor default. Luna Medium is $0.06 at only 29.49%. The useful frontier starts around Luna High at 50.52% and $0.16, then improves quickly as effort rises.
Do not over-read tiny gaps among the highest rows. DeepSWE reports uncertainty, benchmark tasks differ, and adjacent configurations often overlap. The meaningful differences between Opus 5 Max, Extra High, Fable Max, Sol Max, and Opus 5 High are cost and workflow fit, not just fractions of a percentage point.
Cost Per Solved Task, Ranked
Mean cost tells you what one attempt costs. What you actually pay for is a success. Cost per solved task folds the failures into the price: if a model costs $1 per task but only solves half of them, each solve really costs you $2. That's the number I care about most as a buyer, and it gets a simpler chart than the scatter: a ranked bar list, cheapest solve first.
The chart starts with a 50% correctness floor, because without one the ranking is misleading. Luna Low looks unbeatable at $0.10 per solve, but it only solves 19.57% of tasks, so most of your attempts would be wasted. Luna Medium is similarly cheap at $0.20 per solve but only reaches 29.49%. Pick All models in the dropdown if you want to see those rows anyway, or raise the floor to match your own quality bar.
What this view adds to the scatter:
- GPT-5.6 supplies the useful cheap end. With a 50% quality floor, Luna High costs $0.32 per solve, Luna Extra High $0.47, and Luna Max $0.78. Terra's lower tiers cannot beat those Luna rows; Terra Max is the useful jump to 67.26% at $4.67 per solve.
- Opus 5 covers the quality end. Medium reaches 66.60% for $4.94 per solve, High reaches 69.76% for $7.15, Extra High reaches 71.23% for $11.53, and Max reaches 71.82% for $13.97. High is the value tier; Max is for when the last point matters.
- Kimi K3 looks competitive, not dominant. Max reaches 64.66% for $5.69 per solve. That is good enough to test, but Opus 5 Medium is more correct and cheaper in this combined view.
- Sonnet 5 is still the value loser. Sonnet 5 Max costs $28.47 per solve for 57.67% correctness. Sol Medium scores higher at $3.16 per solve. At current benchmark costs, no Sonnet 5 tier earns a value spot.
Best Value Picks
If you don't want to stare at the table, here's how I'd choose:
| Pick | Why |
|---|---|
| GPT-5.6 Luna High | Cheapest tier above 50%: 50.52% mean correctness at $0.32 per solved task |
| GPT-5.6 Luna Extra High | Strong budget step-up: 57.28% at $0.47 per solved task |
| GPT-5.6 Luna Max | Best implementation balance: 64.14% at $0.78 per solved task |
| GPT-5.6 Terra Max | Strong upper implementation tier: 67.26% at $4.67 per solved task |
| Claude Opus 5 High | Upper-end value: 69.76% at $7.15 per solved task |
| Claude Opus 5 Extra High | Near-best correctness: 71.23% at $11.53 per solved task |
My take: Luna Max is my default for most implementation work, while Sol High is the step-up for harder backend changes. Sol High gives up only 0.81 correctness points to Terra Max at essentially the same cost, but uses far fewer thinking tokens and steps in these benchmark runs, so it should usually finish faster. I use Opus 5 High for UI and agentic work, then Opus 5 Max only when I need the strongest rescue or final review.
The important catch is the pricing window. The restored Fable rollout gave some paid subscribers included usage that was originally set to end July 7, and Anthropic has since extended it. That keeps Fable inside normal coding-plan economics for now, but the window is still temporary: once it closes, Fable moves back to separate paid usage or usage credits, and the math changes. The whole point of a coding plan is that you can spend expensive reasoning on planning without worrying about every token.
Opus 5, GPT-5.6, and Kimi K3 availability can still vary by product and account. The benchmark result tells you what the models did under CursorBench and DeepSWE, not whether your editor or API plan exposes every family and effort tier today.
How I Actually Use These (My Take)
Benchmarks rank models. They don't tell you how to actually use them. This is the flow I've settled into.
One subjective caveat: Opus 5 is very capable and well priced, but in my use it does not have the same intelligence or "big model" feel as Fable 5. Fable still feels deeper when a task needs broad judgment and reasoning, even though Opus 5 is the more practical default on capability and price.
- UI and agentic work → Opus 5 High. I use it for visual implementation, interface refinement, unfamiliar codebases, ops, and tasks where the agent has to explore and make judgment calls across several steps. High is the practical default.
- Most implementation work → GPT-5.6 Luna Max. At 64.14% correctness and $0.78 per expected solve, it is the best balance of capability and cost for executing a clear plan or handling routine repo work.
- Harder implementation and backend work → GPT-5.6 Sol High. It reaches 66.45% at $4.71 per solve. Terra Max scores 0.81 points higher at nearly the same price, but Sol High uses far fewer thinking tokens and steps across these runs, which makes it the smarter practical default when speed matters.
- Planning and architecture → Opus 5 High or Sol High. Opus is my preference when the plan depends on product/UI judgment or broad exploration. Sol High is the default when the work is mostly backend architecture and concrete implementation detail.
- Maximum-quality rescue and final review → Opus 5 Max. This is the expensive escalation tier for work the default stack could not crack, or when the cost of missing something matters more than model spend.
- Promising, but not tested yet → Kimi K3 Max. Its new combined result is 64.66% at $3.68 per task. That looks good enough to test, but I have not used it on my own repos yet, so it is not part of my default workflow.
- Use Terra Max selectively; skip the lower tiers. At 67.26% for $3.14, Terra Max beats Sol High by 0.81 points for one cent more. I would only choose it when latency does not matter and squeezing out that small benchmark advantage is worth the extra thinking overhead.
If you're wiring these models into something beyond the editor, I used the same GPT models in my walkthrough on building a RAG system with Node.js and OpenAI.
Caveats
A few things to keep in mind before you treat this as gospel:
- It's an unweighted average of two benchmarks. That's a judgment call, so weight the columns toward your real workload.
- Cost is per task and provider-dependent. Your effective cost shifts with caching, context size, and how you drive the agent. Cost / solved task is easier to read, but it still comes from the same averaged benchmark costs.
- DeepSWE's live page has repriced Terra and Luna, but its downloadable July 25 JSON still contains the older cost fields. This snapshot uses the artifact for exact Pass@1 and the live page for those displayed costs.
- Fable 5 availability depends on plan and platform. Anthropic includes Fable usage for up to 50% of weekly limits on some paid plans, a window originally ending July 7 that has since been extended; usage credits apply once it closes. Cloud providers are being re-enabled separately.
- Fable 5 has stricter safety routing. Anthropic says blocked requests are sent to Opus 4.8, and the new classifier may flag some normal coding/debugging requests.
- Opus 5, GPT-5.6, and Kimi K3 availability varies by product and account even when matching runs appear on both source benchmarks.
- Grok 4.5 High now appears on both benchmarks, but its CursorBench score carries the training-snapshot caveat. Low and Medium remain CursorBench-only.
- Benchmarks measure benchmark tasks. They're a strong signal, not a substitute for trying the model on your own repo.
- These numbers are a snapshot. Both sources can change between snapshots, so re-run the merge when either leaderboard or its pricing changes.
What You Get From This
One ranked view of AI coding agents that weighs skill against cost, built from two benchmarks instead of one. The takeaways:
- Opus 5 Max has the highest combined score at 71.82%, with Extra High second at 71.23%.
- GPT-5.6 supplies the strongest low-cost tradeoffs after both sources repriced Luna and Terra; Luna Max now averages $0.50 per task and $0.78 per expected solve.
- Kimi K3 Max now has data on both benchmarks and enters at rank 13 with 64.66% mean correctness.
- My workflow split is Opus 5 for UI and agentic work, GPT-5.6 for most implementation and backend work.
- Gemini 3.6 Flash High gains a combined row at rank 32; Medium and Low remain CursorBench-only.
- Sonnet 5 is still the value loser at current benchmark costs: its Max tier is the most expensive solve on the board.
- Low effort can be too weak even when it is cheap. Apply a quality floor before reading cost-per-solve ranks.
If you want a version tuned to your own work, swap the flat average for a weighted one and re-rank.
References
- CursorBench current leaderboard: CursorBench 3.2, read August 1, 2026
- DeepSWE v1.1 live leaderboard and current displayed costs
- DeepSWE v1.1 artifact for exact Pass@1, generated July 25, 2026
- Anthropic: Redeploying Fable 5
- Anthropic: statement on the US government directive suspending Fable 5 and Mythos 5 access
Article Changelog
- August 1, 2026 — Luna and Terra pricing refresh: Updated both benchmarks' GPT-5.6 Luna and Terra costs, recomputed every affected mean and cost-per-solve value, and revised the frontier, Best Value picks, and workflow guidance.
- July 29, 2026 — Opus 5, Kimi K3, and Gemini 3.6 Flash: Added the configurations with results on both benchmarks to the combined ranking and kept unmatched effort tiers in the one-benchmark-only table.
- July 9, 2026 — CursorBench 3.2: Recomputed the leaderboard against CursorBench 3.2 and added the Grok 4.5 training-snapshot caveat.
Related articles
Force an LLM to return JSON in JavaScript
Reliably get JSON from an LLM in JavaScript with OpenAI structured outputs and a Zod schema, instead of prompting for JSON and parsing fragile model text yourself.
LangChain JS agent with a custom tool
Create a small LangChain JavaScript agent with one custom tool using createAgent, tool, and a Zod schema, with a runnable end-to-end example.
OpenAI function calling JavaScript example
A minimal OpenAI function calling example in JavaScript using the Responses API and a local tool, including the two-step call-the-tool then answer loop.

