Providers

Anthropic

Maker of the Claude family and the Claude Code agent. Its models carry most of our day-to-day engineering work.

www.anthropic.com

Reported figures

Scores published by the vendor or an outside evaluator. These are not our tests.

Coding-agent-specific index, measuring performance inside an agent harness rather than a single model call.

Terminal-Bench 2.1

Not our test

End-to-end terminal tasks requiring correct multi-step tool use.

DeepSWE 1.1

Not our test

Software-engineering tasks grounded in real repositories.

Composite index from an independent evaluator, blending many task types into a single capability score.

CursorBench 3.2

Not our test

Cursor-owned benchmark measuring code-editing work inside Cursor. Not independent for models Cursor helped develop.

  • Claude Fable 5Max70.5%Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.

Terminal-Bench 3.0

Not our test

The harder revision of Terminal-Bench, where every model still struggles.

  • Claude Fable 5Max34.1%Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.