Benchmarks

What models actually do on real work

We hand several models the identical task, measure what happened, and publish the prompt, the output, and the numbers - including when the result contradicts what we expected.

Our testAugust 23, 2026

Secure multi-tenant app build

One brief for a four-role incident-intake app with tenant isolation, expiring insurer share links, and an auditor boundary - spec-first, then built, then attacked.

  • Claude Fable 5
  • ox-alpha
View benchmark

Wall-clock time

  1. Claude Fable 5 + Claude Code2h 03m
  2. ox-alpha + Claude Code1h 00m

Total tokens

  1. Claude Fable 5 + Claude Code79.4M
  2. ox-alpha + Claude Code14.2M

Tests written

  1. Claude Fable 5 + Claude Code87 passing
  2. ox-alpha + Claude Code0

Fable clearly won the artifact - fail-closed isolation, structural security, its own test suite, and a polished product that updates itself - but paid for it: double the time, five times the tokens, and $107.83 versus $0. ox-alpha built a correct, readable, far cheaper app, but with discipline-based isolation its own docs admit is weaker, no tests, and no automatic refresh. On a security build, the difference between 'fails closed' and 'leaks silently' is the whole difference.

Benchmarks we track

Vendor and evaluator figures we cite in articles. Useful as context, but none of them is our measurement.

Composite index from an independent evaluator, blending many task types into a single capability score.

Coding-agent-specific index, measuring performance inside an agent harness rather than a single model call.

Terminal-Bench 2.1

Not our test

End-to-end terminal tasks requiring correct multi-step tool use.

Terminal-Bench 3.0

Not our test

The harder revision of Terminal-Bench, where every model still struggles.

  • GPT-5.6 SolMax34.6%Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.
  • Claude Fable 5Max34.1%Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.
  • Grok 4.6High26%Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.

DeepSWE 1.1

Not our test

Software-engineering tasks grounded in real repositories.

  • GPT-5.6 Sol72.7%Reported by OpenAIReported, not independent
  • Claude Fable 569.7%Reported by OpenAIReported, not independent
  • Grok 4.6High65.9%Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.

CursorBench 3.2

Not our test

Cursor-owned benchmark measuring code-editing work inside Cursor. Not independent for models Cursor helped develop.

  • Grok 4.6xhigh70.8%Reported by CursorReported, not independentCursor co-developed the model and owns the benchmark, so this is reported evidence rather than an independent result.
  • Claude Fable 5Max70.5%Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.
  • Grok 4.6High69.9%Reported by Cursor / SpaceXAIReported, not independentCursor co-developed the model and owns the benchmark, so this is reported evidence rather than an independent result.
  • GPT-5.6 SolMax67.2%Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.

FrontierCode

Not our test

High-difficulty coding tasks included in the Grok 4.6 launch table.

  • Grok 4.6High61.3%Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.

APEX-Agents

Not our test

Agent benchmark included in the Grok 4.6 launch table.

  • Grok 4.6High57.5%Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.

GDPVal-AA v2

Not our test

Benchmark weighting the economic value of tasks rather than raw pass rate.

  • Grok 4.6High1,753Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.

SWE-Marathon

Not our test

Long-horizon engineering tasks that test endurance over many steps.

  • Grok 4.6High31.9%Reported by Cursor / SpaceXAIReported, not independentPeer results are the best self-reported or public figures included at launch; effort levels and harnesses are not identical.

EEBench

Not our test

Electrical-engineering benchmark from the Grok 4.6 model card.

CadGen

Not our test

CAD-generation benchmark from the Grok 4.6 model card.

Have a task worth benchmarking?

If you're weighing models for a product you're building, we'd like to hear about your case.

Let's talk