Providers

OpenAI

Maker of the GPT family and the Codex agent. Publishes its own benchmark tables alongside each launch.

openai.com

Reported figures

Scores published by the vendor or an outside evaluator. These are not our tests.

Coding-agent-specific index, measuring performance inside an agent harness rather than a single model call.

Terminal-Bench 2.1

Not our test

End-to-end terminal tasks requiring correct multi-step tool use.

DeepSWE 1.1

Not our test

Software-engineering tasks grounded in real repositories.

Composite index from an independent evaluator, blending many task types into a single capability score.

CursorBench 3.2

Not our test

Cursor-owned benchmark measuring code-editing work inside Cursor. Not independent for models Cursor helped develop.

  • GPT-5.6 SolMax67.2%Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.

Terminal-Bench 3.0

Not our test

The harder revision of Terminal-Bench, where every model still struggles.

  • GPT-5.6 SolMax34.6%Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.

EEBench

Not our test

Electrical-engineering benchmark from the Grok 4.6 model card.

CadGen

Not our test

CAD-generation benchmark from the Grok 4.6 model card.