Providers

SpaceXAI

Maker of the Grok family. Grok 4.6 was co-developed with Cursor, which makes some launch benchmarks non-independent.

x.ai

Reported figures

Scores published by the vendor or an outside evaluator. These are not our tests.

Composite index from an independent evaluator, blending many task types into a single capability score.

CursorBench 3.2

Not our test

Cursor-owned benchmark measuring code-editing work inside Cursor. Not independent for models Cursor helped develop.

  • Grok 4.6xhigh70.8%Reported by CursorReported, not independentCursor co-developed the model and owns the benchmark, so this is reported evidence rather than an independent result.
  • Grok 4.6High69.9%Reported by Cursor / SpaceXAIReported, not independentCursor co-developed the model and owns the benchmark, so this is reported evidence rather than an independent result.

DeepSWE 1.1

Not our test

Software-engineering tasks grounded in real repositories.

  • Grok 4.6High65.9%Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.

Terminal-Bench 3.0

Not our test

The harder revision of Terminal-Bench, where every model still struggles.

  • Grok 4.6High26%Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.

FrontierCode

Not our test

High-difficulty coding tasks included in the Grok 4.6 launch table.

  • Grok 4.6High61.3%Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.

APEX-Agents

Not our test

Agent benchmark included in the Grok 4.6 launch table.

  • Grok 4.6High57.5%Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.

GDPVal-AA v2

Not our test

Benchmark weighting the economic value of tasks rather than raw pass rate.

  • Grok 4.6High1,753Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.

SWE-Marathon

Not our test

Long-horizon engineering tasks that test endurance over many steps.

  • Grok 4.6High31.9%Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.

EEBench

Not our test

Electrical-engineering benchmark from the Grok 4.6 model card.

CadGen

Not our test

CAD-generation benchmark from the Grok 4.6 model card.