Grok 4.6
SpaceXAI's low-cost frontier model. Ties GPT-5.6 Sol on the Artificial Analysis intelligence index at a far lower price, but is weaker on long-horizon agent benchmarks. Not yet tested by us.
- Input
- $2per million tokens
- Output
- $6per million tokens
- Cached input
- $0.5per million tokens
- Context window
- 500K
- Released
- 2026-08
Above 200,000 tokens the entire request moves to doubled long-context rates: $1 cached input, $4 input, $12 output.
We tested it onOur test
DevShift benchmarks this model took part in.
We haven't put this model through one of our benchmarks yet.
Reported figuresNot our test
Scores published by the vendor or an outside evaluator. These are not our tests.
Artificial Analysis Intelligence Index
Not our testComposite index from an independent evaluator, blending many task types into a single capability score.
- 61Reported by Artificial Analysis
CursorBench 3.2
Not our testCursor-owned benchmark measuring code-editing work inside Cursor. Not independent for models Cursor helped develop.
- 70.8%Reported by CursorReported, not independentCursor co-developed the model and owns the benchmark, so this is reported evidence rather than an independent result.
- 69.9%Reported by Cursor / SpaceXAIReported, not independentCursor co-developed the model and owns the benchmark, so this is reported evidence rather than an independent result.
DeepSWE 1.1
Not our testSoftware-engineering tasks grounded in real repositories.
- 65.9%Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.
Terminal-Bench 3.0
Not our testThe harder revision of Terminal-Bench, where every model still struggles.
- 26%Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.
FrontierCode
Not our testHigh-difficulty coding tasks included in the Grok 4.6 launch table.
- 61.3%Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.
APEX-Agents
Not our testAgent benchmark included in the Grok 4.6 launch table.
- 57.5%Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.
GDPVal-AA v2
Not our testBenchmark weighting the economic value of tasks rather than raw pass rate.
- 1,753Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.
SWE-Marathon
Not our testLong-horizon engineering tasks that test endurance over many steps.
- 31.9%Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.
EEBench
Not our testElectrical-engineering benchmark from the Grok 4.6 model card.
- 60Reported by SpaceXAI model cardReported, not independent
CadGen
Not our testCAD-generation benchmark from the Grok 4.6 model card.
- 40.9Reported by SpaceXAI model cardReported, not independent
Mentioned in articles
- Grok 4.6 Just Broke the Price of Frontier AI · Aug 2026