Anthropic

Claude Fable 5

Anthropic's flagship coding model. In our test it produced the more modular architecture and the better-looking result - but took twice as long.

Input
$10per million tokens
Output
$50per million tokens
Context window
1,000K
Released
2026-07

We tested it onOur test

DevShift benchmarks this model took part in.

August 23, 2026

Secure multi-tenant app build

Claude Code

Wall-clock time
2h 03m
Agent turns
237
Total tokens
79.4M
Lines of code (ts/tsx)
6,271

Isolation is enforced inside the SQLite engine: code never touches base tables, only principal-scoped views, and a forgotten predicate throws at prepare-time instead of leaking - we verified this with seven bypass attempts against its own fixture, all blocked. It wrote 87 tests nobody asked for, and the UI updates after every change without a reload. The cost: a dense, unfamiliar mechanism concentrated in two maintenance-critical files.

View benchmark

July 20, 2026

One-prompt landing page

Claude Code

Wall-clock time
56m 16s
Total transcript tokens
19.61M
Input tokens
19.42M
Output tokens
185,220

A more modular 3D engine, with geometry, renderer, and lifecycle split apart. No critical or high findings remained after review fixes, but its live render loop is heavier and a direct stage link can leave the HUD one stage behind until scrolling continues.

View benchmark

Who it ran against

Models benchmarked against this one on the same task.

Reported figuresNot our test

Scores published by the vendor or an outside evaluator. These are not our tests.

Coding-agent-specific index, measuring performance inside an agent harness rather than a single model call.

  • 77.2Reported by OpenAIReported, not independent

Terminal-Bench 2.1

Not our test

End-to-end terminal tasks requiring correct multi-step tool use.

  • 83.1%Reported by OpenAIReported, not independent

DeepSWE 1.1

Not our test

Software-engineering tasks grounded in real repositories.

  • 69.7%Reported by OpenAIReported, not independent

Composite index from an independent evaluator, blending many task types into a single capability score.

CursorBench 3.2

Not our test

Cursor-owned benchmark measuring code-editing work inside Cursor. Not independent for models Cursor helped develop.

  • 70.5%Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.

Terminal-Bench 3.0

Not our test

The harder revision of Terminal-Bench, where every model still struggles.

  • 34.1%Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.