CursorBench 3.2

We evaluate agents on ambiguous, multi-file tasks from real Cursor sessions. Higher scores are better.

More about CursorBench
A scatter and line chart comparing Fable 5.1, Fable 5, Opus 5, Opus 4.8, Grok 4.6, GPT-5.6 Sol, GPT-5.6 Terra, GPT-5.6 Luna, GPT-5.5, Sonnet 5, GLM 5.2, Composer 2.5, Gemini 3.8 Flash, Gemini 3.7 Flash, Kimi K3, and Kimi K2.7 Code scores against average cost per task.75%CursorBench 3.2 score70%65%60%55%50%45%$10$5$0Average cost per taskGrok 4.6Fable 5.1Gemini 3.8 FlashOpus 5Kimi K3GPT-5.6 SolSonnet 5Composer 2.5GPT-5.6 TerraGPT-5.6 Luna
Model
1Fable 5.1 Max73.4%$9.6472,06070
2Fable 5.1 Extra High72.8%$6.9651,34955
3Grok 4.6 Extra High70.8%$2.8141,13646
4Fable 5 Max70.5%$17.32103,52572
5Opus 5 Max70.0%$8.2361,83878
6Grok 4.6 High69.9%$2.3432,44939
7Fable 5.1 High69.4%$4.8033,15344
8Opus 5 Extra High69.3%$7.3554,23972
9Gemini 3.8 Flash High69.2%$2.3881,524161
10Fable 5 Extra High68.4%$11.7364,97156
11Fable 5.1 Medium68.0%$3.5323,80136
12GPT-5.6 Sol Max67.2%$5.6928,32048
13Grok 4.6 Medium67.1%$1.2817,94229
14Gemini 3.8 Flash Medium67.0%$1.9361,603136
15Opus 5 High66.7%$3.9127,93248
16Fable 5 High66.5%$8.7743,74748
17Fable 5.1 Low66.2%$2.9019,52231
18Fable 5 Medium65.2%$6.8030,36641
19GPT-5.6 Terra Max64.9%$2.3132,96947
20GPT-5.6 Sol Extra High64.5%$3.8819,69938
21Opus 5 Medium64.3%$3.2923,61244
22GPT-5.6 Sol High63.5%$2.7913,86732
23Opus 5 Low62.8%$2.5518,52937
24Opus 4.8 Max62.3%$5.7771,41144
25Fable 5 Low62.1%$4.4618,18231
26Gemini 3.7 Flash High61.6%$1.2038,44899
27Sonnet 5 Max61.5%$4.3092,88286
28GPT-5.6 Luna Max61.1%$0.3987,97361
29Grok 4.6 Low61.0%$0.7010,65823
30Kimi K3 Max60.8%$2.7038,42857
31GPT-5.6 Sol Medium60.0%$1.959,74727
32Kimi K3 High59.7%$1.8926,84647
33Opus 4.8 Extra High59.4%$4.5051,12140
34GPT-5.6 Terra Extra High59.2%$1.1516,08929
35Gemini 3.7 Flash Medium59.0%$0.9530,95382
36Sonnet 5 Extra High58.7%$2.7752,87167
37GPT-5.5 High58.4%$2.0512,18328
38GPT-5.5 Extra High58.4%$2.8517,53432
39Opus 4.8 High58.0%$3.1533,54833
40GPT-5.6 Luna Extra High57.7%$0.2322,48048
41Sonnet 5 High56.9%$2.1339,48357
42GPT-5.6 Luna High56.8%$0.1615,14140
43Opus 4.8 Medium56.1%$2.8128,38432
44Composer 2.556.1%$0.4414,28633
45GLM 5.2 Max55.0%$1.7635,94658
46GPT-5.6 Terra High54.2%$0.719,46823
47GPT-5.5 Medium53.8%$1.518,52225
48Gemini 3.7 Flash Low53.8%$0.7420,59468
49Opus 4.8 Low53.1%$2.0219,62427
50GPT-5.6 Sol Low52.6%$1.015,10419
51Sonnet 5 Medium52.4%$1.4426,20046
52GLM 5.2 High51.5%$1.1921,82949
53Kimi K3 Low50.5%$0.9913,00733
54GPT-5.6 Terra Medium50.3%$0.496,22220
55Kimi K2.7 Code49.7%$1.4331,24758
56GPT-5.6 Luna Medium47.7%$0.087,09528
57Sonnet 5 Low47.7%$0.8716,26933
58GPT-5.6 Terra Low46.9%$0.425,31219
59GPT-5.5 Low46.6%$0.985,16820
60GPT-5.6 Luna Low37.6%$0.033,20917

Changelog

Reporting

  • Added Gemini 3.8 Flash results. Promoted Gemini 3.8 Flash to Latest; Gemini 3.7 Flash remains as secondary.

Reporting

  • Updated Sonnet 5 results to account for adjusted pricing.

Reporting

  • Updated GPT-5.6 Terra and Luna results to account for adjusted pricing.

Reporting

  • Updated GPT-5.6 Sol, Terra, and Luna results to account for cache write costs.

Tasks

  • CursorBench 3.2
    • Introduced instruction following and advanced tool use problems.

Tasks

  • CursorBench 3.1
    • Introduced problems focused on codebase understanding, bugfinding, planning, and code review.
    • Improved grading criteria for some edit tasks.

Tasks

  • CursorBench 3.0
    • Initial set of tasks focused on edit, refactor, and bugfix problems.

Avg cost / task is computed by applying each model's published per-million-token pricing (input, cache read, cache write, and output) to the tokens it used on each task. Results are subject to variance; small differences in scores may not be statistically meaningful.