OpenAI

GPT-5.6 Sol

OpenAI's flagship coding model. In our test it finished in about half the time using ~26% fewer tokens, with a more compact and runtime-efficient implementation.

Input
$5per million tokens
Output
$30per million tokens
Context window
1,050K
Released
2026-07

The model supports 1.05M tokens, but the Codex catalog we tested against exposed only a 272,000-token window.

We tested it onOur test

DevShift benchmarks this model took part in.

July 20, 2026

One-prompt landing page

Codex

Wall-clock time
28m 44s
Total transcript tokens
14.54M
Input tokens
14.48M
Output tokens
57,498

A smaller, more runtime-efficient implementation: bounds are cached, DPR is capped, and rendering stops when motion settles. The tradeoff is a renderer and lifecycle concentrated in one 766-line client component.

View benchmark

Who it ran against

Models benchmarked against this one on the same task.

Reported figuresNot our test

Scores published by the vendor or an outside evaluator. These are not our tests.

Coding-agent-specific index, measuring performance inside an agent harness rather than a single model call.

  • 80Reported by OpenAIReported, not independent

Terminal-Bench 2.1

Not our test

End-to-end terminal tasks requiring correct multi-step tool use.

  • 88.8%Reported by OpenAIReported, not independent

DeepSWE 1.1

Not our test

Software-engineering tasks grounded in real repositories.

  • 72.7%Reported by OpenAIReported, not independent

Composite index from an independent evaluator, blending many task types into a single capability score.

CursorBench 3.2

Not our test

Cursor-owned benchmark measuring code-editing work inside Cursor. Not independent for models Cursor helped develop.

  • 67.2%Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.

Terminal-Bench 3.0

Not our test

The harder revision of Terminal-Bench, where every model still struggles.

  • 34.6%Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.

EEBench

Not our test

Electrical-engineering benchmark from the Grok 4.6 model card.

CadGen

Not our test

CAD-generation benchmark from the Grok 4.6 model card.