Claude Code

Anthropic#4 of 7 by coverage
Compare harnesses
Best success69%with GPT-6 Astra
Cheapest run$0.25per task with DeepSeek V4 Pro
Fastest run194sper task with GPT-6 Astra
Model support
Open
Graded runs
2 of 2 models
Latest run
2026-09-10

Success rate

% of tasks passed by model · higher is better
GPT-6 Astra
69.0%
DeepSeek V4 Pro
50.0%

Cost per task

USD per task by model · lower is better
DeepSeek V4 Pro
$0.25
GPT-6 Astra
$15.72

Time per task

Seconds per task by model · lower is better
GPT-6 Astra
194s
DeepSeek V4 Pro
414s

Runs

  1. 01GPT-6 Astra
    69%
    $ / task$15.72Time / task194s
  2. 02DeepSeek V4 Pro
    50%
    $ / task$0.25Time / task414s