Composite index from an independent evaluator, blending many task types into a single capability score.
Coding-agent-specific index, measuring performance inside an agent harness rather than a single model call.
Terminal-Bench 2.1
Not our testEnd-to-end terminal tasks requiring correct multi-step tool use.
Terminal-Bench 3.0
Not our testThe harder revision of Terminal-Bench, where every model still struggles.
- GPT-5.6 SolMax34.6%Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.
- Claude Fable 5Max34.1%Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.
- Grok 4.6High26%Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.
DeepSWE 1.1
Not our testSoftware-engineering tasks grounded in real repositories.
- GPT-5.6 Sol72.7%Reported by OpenAIReported, not independent
- Claude Fable 569.7%Reported by OpenAIReported, not independent
- Grok 4.6High65.9%Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.
Cursor-owned benchmark measuring code-editing work inside Cursor. Not independent for models Cursor helped develop.
- Grok 4.6xhigh70.8%Reported by CursorReported, not independentCursor co-developed the model and owns the benchmark, so this is reported evidence rather than an independent result.
- Claude Fable 5Max70.5%Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.
- Grok 4.6High69.9%Reported by Cursor / SpaceXAIReported, not independentCursor co-developed the model and owns the benchmark, so this is reported evidence rather than an independent result.
- GPT-5.6 SolMax67.2%Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.
FrontierCode
Not our testHigh-difficulty coding tasks included in the Grok 4.6 launch table.
- Grok 4.6High61.3%Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.
APEX-Agents
Not our testAgent benchmark included in the Grok 4.6 launch table.
- Grok 4.6High57.5%Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.
GDPVal-AA v2
Not our testBenchmark weighting the economic value of tasks rather than raw pass rate.
- Grok 4.6High1,753Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.
SWE-Marathon
Not our testLong-horizon engineering tasks that test endurance over many steps.
- Grok 4.6High31.9%Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.
EEBench
Not our testElectrical-engineering benchmark from the Grok 4.6 model card.
CadGen
Not our testCAD-generation benchmark from the Grok 4.6 model card.