Providers
OpenAI
Maker of the GPT family and the Codex agent. Publishes its own benchmark tables alongside each launch.
Reported figures
Scores published by the vendor or an outside evaluator. These are not our tests.
Artificial Analysis Coding Agent Index
Not our testCoding-agent-specific index, measuring performance inside an agent harness rather than a single model call.
- GPT-5.6 Sol80Reported by OpenAIReported, not independent
Terminal-Bench 2.1
Not our testEnd-to-end terminal tasks requiring correct multi-step tool use.
- GPT-5.6 Sol88.8%Reported by OpenAIReported, not independent
DeepSWE 1.1
Not our testSoftware-engineering tasks grounded in real repositories.
- GPT-5.6 Sol72.7%Reported by OpenAIReported, not independent
Artificial Analysis Intelligence Index
Not our testComposite index from an independent evaluator, blending many task types into a single capability score.
- GPT-5.6 SolMax61Reported by Artificial Analysis
CursorBench 3.2
Not our testCursor-owned benchmark measuring code-editing work inside Cursor. Not independent for models Cursor helped develop.
- GPT-5.6 SolMax67.2%Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.
Terminal-Bench 3.0
Not our testThe harder revision of Terminal-Bench, where every model still struggles.
- GPT-5.6 SolMax34.6%Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.
EEBench
Not our testElectrical-engineering benchmark from the Grok 4.6 model card.
- GPT-5.6 Sol39.4Reported by SpaceXAI model cardReported, not independent
CadGen
Not our testCAD-generation benchmark from the Grok 4.6 model card.
- GPT-5.6 Sol37.1Reported by SpaceXAI model cardReported, not independent