Compare models.

Every model on the bench, side by side on the same 30 tasks. Toggle entries off to focus the comparison.

By harness

The full matrix. Cell tint scales with the value; the strongest tint wins the row.

ModelClaude CodeCodexOpenCodeHermes AgentPi AgentDeepSeekCommand Code
DeepSeek V4 ProDeepSeek50.0%15/3053.3%16/3050.0%15/3056.7%17/3056.7%17/3053.3%16/3060.0%18/30
GPT-6 AstraOpenAI69.0%20/2972.4%21/2965.5%19/2972.4%21/2965.5%19/2972.4%21/29

Overall

Each column shows that model's selected results on Claude Code, across the same 30 tasks. Choose a harness to compare results across the suite.

0.0%
DeepSeek V4 Pro
0.0%
GPT-6 Astra

Success with retries

Each bar is one model's success on the picked harness, with observed pass@1 and projected pass@2, pass@3 and pass^3. These are expected rates, not measured repeat runs. One scale whichever harness is picked, sorted by pass@3.

Modelpass@3pass^3
  1. GPT-6 Astra97.0%32.8%
  2. DeepSeek V4 Pro55.0%45.0%
0%50%100%
  • pass@1
  • pass@2
  • pass@3
  • pass^3

Projected retries · pass@1 observed · every model on the picked harness · sorted by pass@3

Shape by harness

Success rate on every harness, one polygon per entry.

DeepSeek V4 ProGPT-6 Astra
Success rate radarClaude CodeClaude CodeCodexCodexOpenCodeOpenCodeHermes AgentHermesAgentPi AgentPi AgentDeepSeekDeepSeekCommand CodeCommandCode25%50%75%100%
Claude Code
50%69%
Codex
53%72%
OpenCode
50%66%
Hermes Agent
57%72%
Pi Agent
57%66%
DeepSeek
53%
Command Code
60%72%

Cost analysis

What a task costs each model on Claude Code, cheapest first, with how often it passes beside it.

Cost per task
$0.036
$0.10
$1.00
$10.00
$51.95
  1. DeepSeek V4 Pro
    Success50.0%$ / task$0.25Time414s
  2. GPT-6 Astra
    Success69.0%$ / task$15.72Time194s

Each tick is one task · faint ticks did not pass · the square is the average