AI Coding Benchmark Leaderboard: Cursor + DeepSWE
- Authors
- Name
- Hamza Rahman
- Published on
- -20 mins read-
Every time I picked a model I ended up bouncing between two AI coding benchmarks: CursorBench and DeepSWE. They're both solid, but they rank models differently, and on its own neither one told me what I actually wanted to know.
So I merged them into one AI coding benchmark leaderboard: a single ranked table that weighs "how good is it" against "what does it actually cost me." Short version: Opus 5 occupies the first two combined rows and the high-quality end of the price/performance frontier. Grok 4.6 enters across all four effort tiers with Medium delivering standout mid-range value ($2.37 per task for 67.29% correctness). GPT-5.6 Sol and Luna provide the strongest implementation defaults, with Luna Max completely dominating the sub-65% tier on both accuracy and cost ($0.78 per solve). Update (August 31, 2026): CursorBench 3.2 added Grok 4.6 and Gemini 3.7 Flash while adjusting Sonnet 5 pricing. DeepSWE v1.1 updated its snapshot (August 26), adding Grok 4.6 and Gemini 3.7 Flash while repricing GPT-5.6 Sol and Gemini Flash runs on its live leaderboard. This update merges the refreshed runs across 51 matched configurations.
The Combined Leaderboard
Ranked by mean correctness (higher is better), with mean cost alongside. Use the buttons to toggle models on the chart and table. This snapshot uses CursorBench 3.2 (read August 31, 2026) and the DeepSWE v1.1 live leaderboard as updated August 26, 2026. Cost increases left to right on a logarithmic scale, so the better-value direction is up and left. Every effort tier is plotted. Hover or select a family to draw its effort curve, and toggle Show frontier to draw the Pareto staircase of undominated picks.
| Rank | Model / effort | Mean correctness | Mean cost | Cost / solved | CursorBench | DeepSWE Pass@1 | Cursor cost | DeepSWE cost |
|---|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 - Max | 71.82% | $10.04 | $13.98 | 70.00% | 73.65% | $8.23 | $11.84 |
| 2 | Claude Opus 5 - Extra High | 71.23% | $8.21 | $11.53 | 69.30% | 73.15% | $7.35 | $9.07 |
| 3 | Fable 5 - Max | 70.11% | $19.48 | $27.78 | 70.50% | 69.72% | $17.32 | $21.63 |
| 4 | GPT-5.6 Sol - Max | 69.93% | $6.08 | $8.69 | 67.20% | 72.67% | $5.69 | $6.46 |
| 5 | Claude Opus 5 - High | 69.76% | $5.00 | $7.17 | 66.70% | 72.83% | $3.91 | $6.08 |
| 6 | Fable 5 - Extra High | 69.16% | $12.57 | $18.18 | 68.40% | 69.91% | $11.73 | $13.41 |
| 7 | Grok 4.6 - Extra High | 68.77% | $4.16 | $6.05 | 70.80% | 66.74% | $2.81 | $5.50 |
| 8 | GPT-5.6 Sol - Extra High | 67.62% | $3.74 | $5.53 | 64.50% | 70.73% | $3.88 | $3.60 |
| 9 | Fable 5 - High | 67.55% | $8.97 | $13.28 | 66.50% | 68.60% | $8.77 | $9.18 |
| 10 | Grok 4.6 - High | 67.54% | $3.36 | $4.97 | 69.90% | 65.19% | $2.34 | $4.38 |
| 11 | Grok 4.6 - Medium | 67.29% | $2.37 | $3.52 | 67.10% | 67.48% | $1.28 | $3.45 |
| 12 | GPT-5.6 Terra - Max | 67.26% | $3.13 | $4.65 | 64.90% | 69.62% | $2.31 | $3.96 |
| 13 | Claude Opus 5 - Medium | 66.60% | $3.29 | $4.94 | 64.30% | 68.90% | $3.29 | $3.29 |
| 14 | GPT-5.6 Sol - High | 66.45% | $2.73 | $4.11 | 63.50% | 69.40% | $2.79 | $2.66 |
| 15 | Fable 5 - Medium | 65.28% | $6.45 | $9.88 | 65.20% | 65.37% | $6.80 | $6.09 |
| 16 | Kimi K3 - Max | 64.66% | $3.68 | $5.69 | 60.80% | 68.51% | $2.70 | $4.65 |
| 17 | GPT-5.6 Luna - Max | 64.14% | $0.50 | $0.78 | 61.10% | 67.19% | $0.39 | $0.61 |
| 18 | Gemini 3.7 Flash - High | 63.43% | $1.69 | $2.66 | 61.60% | 65.27% | $1.20 | $2.18 |
| 19 | GPT-5.5 - Extra High | 62.72% | $5.04 | $8.04 | 58.40% | 67.04% | $2.85 | $7.23 |
| 20 | Gemini 3.7 Flash - Medium | 62.24% | $1.49 | $2.39 | 59.00% | 65.49% | $0.95 | $2.03 |
| 21 | GPT-5.5 - High | 61.39% | $3.57 | $5.82 | 58.40% | 64.38% | $2.05 | $5.10 |
| 22 | Fable 5 - Low | 60.84% | $4.11 | $6.76 | 62.10% | 59.58% | $4.46 | $3.76 |
| 23 | Claude Opus 4.8 - Max | 60.64% | $9.50 | $15.67 | 62.30% | 58.97% | $5.77 | $13.22 |
| 24 | GPT-5.6 Sol - Medium | 60.53% | $1.69 | $2.79 | 60.00% | 61.06% | $1.95 | $1.42 |
| 25 | Claude Opus 5 - Low | 60.46% | $2.10 | $3.47 | 62.80% | 58.13% | $2.55 | $1.66 |
| 26 | GPT-5.6 Terra - Extra High | 59.69% | $1.42 | $2.38 | 59.20% | 60.18% | $1.15 | $1.70 |
| 27 | Claude Sonnet 5 - Max | 57.67% | $15.35 | $26.62 | 61.50% | 53.85% | $4.30 | $26.40 |
| 28 | GPT-5.6 Luna - Extra High | 57.28% | $0.27 | $0.47 | 57.70% | 56.86% | $0.23 | $0.31 |
| 29 | Claude Opus 4.8 - Extra High | 56.88% | $6.25 | $10.99 | 59.40% | 54.36% | $4.50 | $8.01 |
| 30 | Claude Opus 4.8 - High | 54.88% | $3.71 | $6.76 | 58.00% | 51.77% | $3.15 | $4.28 |
| 31 | Claude Sonnet 5 - Extra High | 54.18% | $7.33 | $13.53 | 58.70% | 49.67% | $2.77 | $11.89 |
| 32 | GPT-5.6 Terra - High | 53.98% | $0.81 | $1.50 | 54.20% | 53.76% | $0.71 | $0.91 |
| 33 | GPT-5.5 - Medium | 53.89% | $2.13 | $3.95 | 53.80% | 53.98% | $1.51 | $2.75 |
| 34 | Gemini 3.7 Flash - Low | 53.78% | $1.29 | $2.40 | 53.80% | 53.76% | $0.74 | $1.83 |
| 35 | Claude Sonnet 5 - High | 52.57% | $4.78 | $9.09 | 56.90% | 48.23% | $2.13 | $7.43 |
| 36 | Claude Opus 4.8 - Medium | 52.39% | $3.12 | $5.96 | 56.10% | 48.67% | $2.81 | $3.44 |
| 37 | Grok 4.6 - Low | 51.32% | $0.87 | $1.70 | 61.00% | 41.65% | $0.70 | $1.04 |
| 38 | GPT-5.6 Luna - High | 50.52% | $0.16 | $0.32 | 56.80% | 44.25% | $0.16 | $0.16 |
| 39 | Gemini 3.6 Flash - High | 50.09% | $1.89 | $3.77 | 53.50% | 46.68% | $1.56 | $2.21 |
| 40 | GLM 5.2 - Max | 49.39% | $2.84 | $5.75 | 55.00% | 43.78% | $1.76 | $3.92 |
| 41 | GPT-5.6 Sol - Low | 48.98% | $0.92 | $1.88 | 52.60% | 45.35% | $1.01 | $0.82 |
| 42 | Claude Opus 4.8 - Low | 46.95% | $2.16 | $4.60 | 53.10% | 40.80% | $2.02 | $2.29 |
| 43 | Claude Sonnet 5 - Medium | 46.09% | $2.76 | $5.99 | 52.40% | 39.78% | $1.44 | $4.08 |
| 44 | GLM 5.2 - High | 43.89% | $2.01 | $4.58 | 51.50% | 36.28% | $1.19 | $2.84 |
| 45 | GPT-5.6 Terra - Medium | 42.71% | $0.48 | $1.12 | 50.30% | 35.11% | $0.49 | $0.47 |
| 46 | Kimi K2.7 Code | 40.12% | $2.12 | $5.28 | 49.70% | 30.53% | $1.43 | $2.82 |
| 47 | Claude Sonnet 5 - Low | 39.11% | $1.53 | $3.91 | 47.70% | 30.51% | $0.87 | $2.19 |
| 48 | GPT-5.5 - Low | 36.80% | $1.09 | $2.96 | 46.60% | 26.99% | $0.98 | $1.20 |
| 49 | GPT-5.6 Terra - Low | 35.48% | $0.38 | $1.07 | 46.90% | 24.05% | $0.42 | $0.34 |
| 50 | GPT-5.6 Luna - Medium | 29.49% | $0.06 | $0.20 | 47.70% | 11.28% | $0.08 | $0.04 |
| 51 | GPT-5.6 Luna - Low | 19.57% | $0.02 | $0.10 | 37.60% | 1.55% | $0.03 | $0.01 |
On one benchmark only (no combined mean)
Not plotted above and not ranked: a single score isn't comparable to a mean, and the benchmarks differ in difficulty. These rows are useful references, but they don't answer the combined CursorBench + DeepSWE question.
| Model | Benchmark | Score | Cost | Mean |
|---|---|---|---|---|
| Kimi K3 - High | CursorBench | 59.70% | $1.89 | n/a |
| Composer 2.5 | CursorBench | 56.10% | $0.44 | n/a |
| Gemini 3.6 Flash - Medium | CursorBench | 51.20% | $1.48 | n/a |
| Kimi K3 - Low | CursorBench | 50.50% | $0.99 | n/a |
| Gemini 3.6 Flash - Low | CursorBench | 47.40% | $1.13 | n/a |
| GLM 5.3 - Max | DeepSWE Pass@1 | 68.96% | $3.99 | n/a |
| GLM 5.3 Flash - Max | DeepSWE Pass@1 | 63.39% | $0.24 | n/a |
| DeepSeek V4 Pro - Max | DeepSWE Pass@1 | 62.83% | $1.67 | n/a |
| Qwen 3.8 Max - Extra High | DeepSWE Pass@1 | 57.46% | $3.73 | n/a |
| Muse Spark 1.2 - Extra High | DeepSWE Pass@1 | 54.87% | $3.70 | n/a |
| Grok 4.5 - High | DeepSWE Pass@1 | 53.76% | $2.42 | n/a |
| DeepSeek V4 Flash - Max | DeepSWE Pass@1 | 53.32% | $0.46 | n/a |
| Muse Spark 1.1 - Extra High | DeepSWE Pass@1 | 53.32% | $2.36 | n/a |
| GPT-5.4 - Extra High | DeepSWE Pass@1 | 51.77% | $5.65 | n/a |
| Gemini 3.5 Flash - High | DeepSWE Pass@1 | 36.06% | $3.45 | n/a |
| Claude Sonnet 4.6 - High | DeepSWE Pass@1 | 29.93% | $5.52 | n/a |
| Gemini 3.1 Pro Preview - High | DeepSWE Pass@1 | 11.73% | $2.14 | n/a |
Why Two Benchmarks?
Single benchmarks saturate. Once every frontier model scores in the high 90s, the leaderboard stops telling you anything. CursorBench is closer to real in-editor work: short, underspecified prompts from actual Cursor sessions. DeepSWE is harder and more controlled: 113 hand-written long-horizon tasks across 91 repos, scored with behavioral verifiers.
One is real-world and messy, the other is original and hard. Put together, they land closer to day-to-day usefulness than either does alone.
How I Combined Them
I kept the math deliberately simple so anyone can reproduce it:
- Mean correctness = average of (CursorBench correctness, DeepSWE Pass@1)
- Mean cost = average of (CursorBench cost, DeepSWE cost per task)
- Cost / solved task = mean cost divided by mean correctness as a probability
No weighting, no normalization tricks. Both benchmarks already report cost in dollars per task, so the averages stay in real units. Then I ranked by mean correctness.
One caveat up front: a flat average treats both benchmarks as equally important, which is a choice, not a law. Averaging costs also mixes two different task sizes: DeepSWE tasks are larger, so its cost-per-task is not the same unit as CursorBench's cost-per-task. I still include cost / solved task because it's easier to reason about as a buyer: under this combined view, how much do I pay for one expected successful solve?
If your work looks more like quick in-editor edits, lean on the CursorBench column. If it looks like large multi-file features, lean on DeepSWE.
A few configurations only show up on one benchmark: Kimi K3 High/Low, Composer 2.5, and Gemini 3.6 Flash Medium/Low on CursorBench; GLM 5.3 Max, GLM 5.3 Flash Max, DeepSeek V4 Pro Max, Qwen 3.8 Max Extra High, Muse Spark 1.2 Extra High, Grok 4.5 High, DeepSeek V4 Flash Max, Muse Spark 1.1 Extra High, GPT-5.4 Extra High, Gemini 3.5 Flash High, Claude Sonnet 4.6 High, and Gemini 3.1 Pro Preview High on DeepSWE. Averaging a single score with a missing one isn't a mean, so I keep them out of the ranked table and off the chart.
What Jumps Out
Opus 5 occupies the first two rows. Max reaches 71.82% mean correctness at $10.04 mean cost, while Extra High reaches 71.23% at $8.21. Fable 5 Max is third at 70.11% and $19.48, followed by GPT-5.6 Sol Max at 69.93% and $6.08. Small score gaps still overlap benchmark uncertainty, but Opus 5 has the highest combined score in this snapshot.
Opus 5 High is the standout upper-end value tier at rank five (69.76% mean correctness, $5.00 per task, and $7.17 per expected solve). Medium ranks 13th at 66.60% and $3.29. Max buys the highest score, but Extra High and High make the stronger price/performance cases.
Grok 4.6 is the biggest new addition, entering all four effort tiers on both benchmarks. Extra High reaches rank 7 (68.77% correctness at $4.16 mean cost, $6.05 per solve). Grok 4.6 High is rank 10 (67.54% at $3.36, $4.97 per solve), and Grok 4.6 Medium is rank 11 (67.29% at $2.37, $3.52 per solve). Grok 4.6 Medium stands out at $2.37 per task, beating Terra Max's correctness (67.29% vs 67.26%) for 76 cents less per task and displacing Terra Max from the Pareto frontier.
DeepSWE's latest update repriced GPT-5.6 Sol downward. Sol Max is now $6.08 per task ($8.69 per solve, down from $10.07). Sol Extra High is $3.74 ($5.53 per solve), and Sol High is $2.73 ($4.11 per solve). Together with Luna Max at $0.50 per task ($0.78 per solve), GPT-5.6 remains the most economical stack for everyday code generation.
Gemini 3.7 Flash also gains a combined row across three tiers (High at rank 18 with 63.43% for $1.69; Medium at rank 20 with 62.24% for $1.49). While cheap on paper, Gemini 3.7 Flash is strictly dominated by GPT-5.6 Luna Max: Luna Max is both more accurate (64.14% vs 63.43%) and over three times cheaper per solve ($0.78 vs $2.39–$2.66), keeping Luna Max firmly on the Pareto frontier.
Kimi K3 Max enters at rank 16 with 64.66% mean correctness, $3.68 mean cost, and $5.69 per expected solve. That is a solid combined result, though Grok 4.6 Medium and Luna Max offer better cost-to-correctness ratios.
The low-effort GPT-5.6 rows show why cost alone is dangerous. Luna Low is only $0.02 per task, but its 19.57% mean correctness makes it a poor default. Luna Medium is $0.06 at only 29.49%. The useful frontier starts around Luna High at 50.52% and $0.16, then improves quickly as effort rises.
Do not over-read tiny gaps among the highest rows. DeepSWE reports uncertainty, benchmark tasks differ, and adjacent configurations often overlap. The meaningful differences between Opus 5 Max, Extra High, Fable Max, Sol Max, and Opus 5 High are cost and workflow fit, not just fractions of a percentage point.
Cost Per Solved Task, Ranked
Mean cost tells you what one attempt costs. What you actually pay for is a success. Cost per solved task folds the failures into the price: if a model costs $1 per task but only solves half of them, each solve really costs you $2. That's the number I care about most as a buyer, and it gets a simpler chart than the scatter: a ranked bar list, cheapest solve first.
The chart starts with a 50% correctness floor, because without one the ranking is misleading. Luna Low looks unbeatable at $0.10 per solve, but it only solves 19.57% of tasks, so most of your attempts would be wasted. Luna Medium is similarly cheap at $0.20 per solve but only reaches 29.49%. Pick All models in the dropdown if you want to see those rows anyway, or raise the floor to match your own quality bar.
What this view adds to the scatter:
- GPT-5.6 supplies the useful cheap end. With a 50% quality floor, Luna High costs $0.32 per solve, Luna Extra High $0.47, and Luna Max $0.78.
- Grok 4.6 Medium anchors the mid-tier step-up. Grok 4.6 Medium reaches 67.29% for $3.52 per solve, offering the sharpest jump in accuracy once you step past Luna Max.
- GPT-5.6 Sol provides efficient backend step-ups. Sol High reaches 66.45% for $4.11 per solve, while Sol Extra High reaches 67.62% for $5.53 per solve.
- Opus 5 covers the quality end. Medium reaches 66.60% for $4.94 per solve, High reaches 69.76% for $7.17, Extra High reaches 71.23% for $11.53, and Max reaches 71.82% for $13.98. High is the value tier; Max is for when the last point matters.
- Sonnet 5 improved on pricing, but still lags in solve value. CursorBench lowered Sonnet 5 costs, but its heavy output on DeepSWE keeps Sonnet 5 Max at $15.35 mean cost and $26.62 per solve for 57.67% correctness.
Best Value Picks
If you don't want to stare at the table, here's how I'd choose:
| Pick | Why |
|---|---|
| GPT-5.6 Luna Max | Best budget & implementation balance: 64.14% at $0.78 per solved task |
| Grok 4.6 Medium | Standout mid-range value: 67.29% at $3.52 per solved task |
| GPT-5.6 Sol High | Efficient backend step-up: 66.45% at $4.11 per solved task |
| Claude Opus 5 High | Upper-end value: 69.76% at $7.17 per solved task |
| Claude Opus 5 Extra High | Near-best correctness: 71.23% at $11.53 per solved task |
For most routine implementation tickets, Luna Max is the obvious default. Grok 4.6 Medium and Sol High provide the cleanest step-ups when tasks require heavier backend logic. Sol High gives up only 0.81 correctness points to Terra Max at $4.11 vs $4.65 per solve while using far fewer thinking tokens and steps in these runs.
The important catch is the pricing window. The restored Fable rollout gave some paid subscribers included usage that was originally set to end July 7, and Anthropic has since extended it. That keeps Fable inside normal coding-plan economics for now, but the window is still temporary: once it closes, Fable moves back to separate paid usage or usage credits, and the math changes. The whole point of a coding plan is that you can spend expensive reasoning on planning without worrying about every token.
Opus 5, GPT-5.6, and Kimi K3 availability can still vary by product and account. The benchmark result tells you what the models did under CursorBench and DeepSWE, not whether your editor or API plan exposes every family and effort tier today.
How I Actually Use These (My Take)
Benchmarks rank models on synthetic tasks. They don't tell you what a model feels like when you're staring at an editor for six hours. This is the setup I've settled into based on real-world friction and economics:
The Daily Drivers
- High-frequency edits & terminal agents → GLM 5.3 Flash (Max). Back when it was tested anonymously under stealth handles like "Ox Alpha" (often called "0x Alpha" across developer forums), it immediately stood out for speed and agentic competence. On DeepSWE, it scores a massive 63.39% Pass@1 at just $0.24 per task ($0.38 per expected solve). It's insanely cheap, responds almost instantly, and handles routine repo edits without burning cash.
- Daily drafting & large-context dumps → Gemini 3.7. Pure benchmark $/solve ranks tell only part of the story. Google bundles 2TB of Google Drive/Photos cloud storage into its Google One AI Premium subscription ($19.99/mo). If you already pay for cloud storage, the effective marginal cost of Gemini Advanced is basically zero. Combined with the 1M+ token context window, it's my go-to for dumping entire codebases, documentation sets, and rough prototypes.
Planning & Difficult Logic
- Architecture, planning, & complex backend tasks → GPT-5.6 Sol High. When a problem is genuinely tricky or needs multi-step architectural planning, Sol High is my default. It hits 66.45% combined correctness ($4.11 per solve, reaching 69.40% on DeepSWE) while keeping thinking tokens concise. It plans cleanly without hallucinating fluff or stalling on long inference pauses.
The Frontier Experience: Opus 5 vs Fable 5
- Opus 5 (Great coder, but not fun to work with). Statistically, Opus 5 is a beast: rank one (71.82% Max) and rank two (71.23% Extra High). But subjectively, it's just not fun to collaborate with. It can feel stiff, pedantic, and overly literal—technically correct on syntax, but tedious when iterating on nuanced designs or complex refactors. I keep Opus 5 Max reserved strictly as an emergency rescue tier for bugs nothing else could solve.
- Fable 5 (The model you actually want when available). When Anthropic's usage window is open or limits allow, Fable 5 is the model I prefer working with. It has an unmistakable "big model" feel: deeper intuitive judgment, taste in code structure, and a collaborative cadence that makes hard problem-solving feel effortless. Even at higher per-task costs ($19.48 Max, $8.97 High), the developer experience is in a different league.
If you're wiring these models into something beyond the editor, I used the same GPT models in my walkthrough on building a RAG system with Node.js and OpenAI.
Caveats
A few things to keep in mind before you treat this as gospel:
- It's an unweighted average of two benchmarks. That's a judgment call, so weight the columns toward your real workload.
- Cost is per task and provider-dependent. Your effective cost shifts with caching, context size, and how you drive the agent. Cost / solved task is easier to read, but it still comes from the same averaged benchmark costs.
- DeepSWE's live page has repriced Sol, Terra, Luna, and Gemini runs, while its downloadable JSON artifact generated August 26, 2026 contains exact Pass@1 data. This snapshot uses the artifact for exact Pass@1 and the live page for displayed costs.
- Fable 5 availability depends on plan and platform. Anthropic includes Fable usage for up to 50% of weekly limits on some paid plans, a window originally ending July 7 that has since been extended; usage credits apply once it closes. Cloud providers are being re-enabled separately.
- Fable 5 has stricter safety routing. Anthropic says blocked requests are sent to Opus 4.8, and the new classifier may flag some normal coding/debugging requests.
- Opus 5, GPT-5.6, Grok 4.6, Gemini 3.7 Flash, and Kimi K3 availability varies by product and account even when matching runs appear on both source benchmarks.
- Grok 4.6 now replaces Grok 4.5 across all effort tiers on both benchmarks without earlier training-snapshot caveats.
- Benchmarks measure benchmark tasks. They're a strong signal, not a substitute for trying the model on your own repo.
- These numbers are a snapshot. Both sources can change between snapshots, so re-run the merge when either leaderboard or its pricing changes.
What You Get From This
One ranked view of AI coding agents that weighs skill against cost, built from two benchmarks instead of one. The takeaways:
- Opus 5 Max has the highest combined score at 71.82%, with Extra High second at 71.23%.
- Grok 4.6 enters the combined board with strong numbers across all tiers; Grok 4.6 Medium (67.29% at $2.37) is one of the best price/performance options available.
- GPT-5.6 Sol and Luna supply top-tier implementation economics after DeepSWE repricing; Luna Max averages $0.50 per task and $0.78 per expected solve.
- Luna Max completely dominates Gemini 3.7 Flash on both accuracy (64.14% vs 63.43%) and cost per solve ($0.78 vs $2.39–$2.66).
- My workflow split is Opus 5 for UI and agentic work, GPT-5.6 and Grok 4.6 for implementation and backend work.
- Sonnet 5 is repriced lower on CursorBench, but remains expensive per solve due to high token counts on DeepSWE.
- Low effort can be too weak even when it is cheap. Apply a quality floor before reading cost-per-solve ranks.
If you want a version tuned to your own work, swap the flat average for a weighted one and re-rank.
References
- CursorBench current leaderboard: CursorBench 3.2, read August 31, 2026
- DeepSWE v1.1 live leaderboard and current displayed costs
- DeepSWE v1.1 artifact for exact Pass@1, generated August 26, 2026
- Anthropic: Redeploying Fable 5
- Anthropic: statement on the US government directive suspending Fable 5 and Mythos 5 access
Article Changelog
- August 31, 2026 — Grok 4.6, Gemini 3.7 Flash, and Sol repricing: Added all four Grok 4.6 tiers and three Gemini 3.7 Flash tiers, integrated DeepSWE's August 26 run updates and Sol repricing, factored in CursorBench's Sonnet 5 price reduction, and updated the frontier and Best Value table.
- August 1, 2026 — Luna and Terra pricing refresh: Updated both benchmarks' GPT-5.6 Luna and Terra costs, recomputed every affected mean and cost-per-solve value, and revised the frontier, Best Value picks, and workflow guidance.
- July 29, 2026 — Opus 5, Kimi K3, and Gemini 3.6 Flash: Added the configurations with results on both benchmarks to the combined ranking and kept unmatched effort tiers in the one-benchmark-only table.
Related articles
Force an LLM to return JSON in JavaScript
Reliably get JSON from an LLM in JavaScript with OpenAI structured outputs and a Zod schema, instead of prompting for JSON and parsing fragile model text yourself.
LangChain JS agent with a custom tool
Create a small LangChain JavaScript agent with one custom tool using createAgent, tool, and a Zod schema, with a runnable end-to-end example.
OpenAI function calling JavaScript example
A minimal OpenAI function calling example in JavaScript using the Responses API and a local tool, including the two-step call-the-tool then answer loop.

