Providers
Anthropic
Maker of the Claude family and the Claude Code agent. Its models carry most of our day-to-day engineering work.
Reported figures
Scores published by the vendor or an outside evaluator. These are not our tests.
Artificial Analysis Coding Agent Index
Not our testCoding-agent-specific index, measuring performance inside an agent harness rather than a single model call.
- Claude Fable 577.2Reported by OpenAIReported, not independent
Terminal-Bench 2.1
Not our testEnd-to-end terminal tasks requiring correct multi-step tool use.
- Claude Fable 583.1%Reported by OpenAIReported, not independent
DeepSWE 1.1
Not our testSoftware-engineering tasks grounded in real repositories.
- Claude Fable 569.7%Reported by OpenAIReported, not independent
Artificial Analysis Intelligence Index
Not our testComposite index from an independent evaluator, blending many task types into a single capability score.
- Claude Fable 5Max62Reported by Artificial Analysis
CursorBench 3.2
Not our testCursor-owned benchmark measuring code-editing work inside Cursor. Not independent for models Cursor helped develop.
- Claude Fable 5Max70.5%Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.
Terminal-Bench 3.0
Not our testThe harder revision of Terminal-Bench, where every model still struggles.
- Claude Fable 5Max34.1%Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.