SpaceXAI

Grok 4.6

SpaceXAI's low-cost frontier model. Ties GPT-5.6 Sol on the Artificial Analysis intelligence index at a far lower price, but is weaker on long-horizon agent benchmarks. Not yet tested by us.

Input
$2per million tokens
Output
$6per million tokens
Cached input
$0.5per million tokens
Context window
500K
Released
2026-08

Above 200,000 tokens the entire request moves to doubled long-context rates: $1 cached input, $4 input, $12 output.

We tested it onOur test

DevShift benchmarks this model took part in.

We haven't put this model through one of our benchmarks yet.

Reported figuresNot our test

Scores published by the vendor or an outside evaluator. These are not our tests.

Composite index from an independent evaluator, blending many task types into a single capability score.

CursorBench 3.2

Not our test

Cursor-owned benchmark measuring code-editing work inside Cursor. Not independent for models Cursor helped develop.

  • 70.8%Reported by CursorReported, not independentCursor co-developed the model and owns the benchmark, so this is reported evidence rather than an independent result.
  • 69.9%Reported by Cursor / SpaceXAIReported, not independentCursor co-developed the model and owns the benchmark, so this is reported evidence rather than an independent result.

DeepSWE 1.1

Not our test

Software-engineering tasks grounded in real repositories.

  • 65.9%Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.

Terminal-Bench 3.0

Not our test

The harder revision of Terminal-Bench, where every model still struggles.

  • 26%Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.

FrontierCode

Not our test

High-difficulty coding tasks included in the Grok 4.6 launch table.

  • 61.3%Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.

APEX-Agents

Not our test

Agent benchmark included in the Grok 4.6 launch table.

  • 57.5%Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.

GDPVal-AA v2

Not our test

Benchmark weighting the economic value of tasks rather than raw pass rate.

  • 1,753Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.

SWE-Marathon

Not our test

Long-horizon engineering tasks that test endurance over many steps.

  • 31.9%Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.

EEBench

Not our test

Electrical-engineering benchmark from the Grok 4.6 model card.

CadGen

Not our test

CAD-generation benchmark from the Grok 4.6 model card.

Mentioned in articles