Providers
SpaceXAI
Maker of the Grok family. Grok 4.6 was co-developed with Cursor, which makes some launch benchmarks non-independent.
Reported figures
Scores published by the vendor or an outside evaluator. These are not our tests.
Artificial Analysis Intelligence Index
Not our testComposite index from an independent evaluator, blending many task types into a single capability score.
- Grok 4.6High61Reported by Artificial Analysis
CursorBench 3.2
Not our testCursor-owned benchmark measuring code-editing work inside Cursor. Not independent for models Cursor helped develop.
- Grok 4.6xhigh70.8%Reported by CursorReported, not independentCursor co-developed the model and owns the benchmark, so this is reported evidence rather than an independent result.
- Grok 4.6High69.9%Reported by Cursor / SpaceXAIReported, not independentCursor co-developed the model and owns the benchmark, so this is reported evidence rather than an independent result.
DeepSWE 1.1
Not our testSoftware-engineering tasks grounded in real repositories.
- Grok 4.6High65.9%Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.
Terminal-Bench 3.0
Not our testThe harder revision of Terminal-Bench, where every model still struggles.
- Grok 4.6High26%Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.
FrontierCode
Not our testHigh-difficulty coding tasks included in the Grok 4.6 launch table.
- Grok 4.6High61.3%Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.
APEX-Agents
Not our testAgent benchmark included in the Grok 4.6 launch table.
- Grok 4.6High57.5%Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.
GDPVal-AA v2
Not our testBenchmark weighting the economic value of tasks rather than raw pass rate.
- Grok 4.6High1,753Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.
SWE-Marathon
Not our testLong-horizon engineering tasks that test endurance over many steps.
- Grok 4.6High31.9%Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.
EEBench
Not our testElectrical-engineering benchmark from the Grok 4.6 model card.
- Grok 4.6xhigh60Reported by SpaceXAI model cardReported, not independent
CadGen
Not our testCAD-generation benchmark from the Grok 4.6 model card.
- Grok 4.6xhigh40.9Reported by SpaceXAI model cardReported, not independent