Compare models.
Every model on the bench, side by side on the same 30 tasks. Toggle entries off to focus the comparison.
By harness
The full matrix. Cell tint scales with the value; the strongest tint wins the row.
| harness | ||
|---|---|---|
| 50.0%15/30 | 69.0%20/29 | |
| 53.3%16/30 | 72.4%21/29 | |
| 50.0%15/30 | 65.5%19/29 | |
| 56.7%17/30 | 72.4%21/29 | |
| 56.7%17/30 | 65.5%19/29 | |
| 53.3%16/30 | — | |
| 60.0%18/30 | 72.4%21/29 |
| Model | |||||||
|---|---|---|---|---|---|---|---|
| 50.0%15/30 | 53.3%16/30 | 50.0%15/30 | 56.7%17/30 | 56.7%17/30 | 53.3%16/30 | 60.0%18/30 | |
| 69.0%20/29 | 72.4%21/29 | 65.5%19/29 | 72.4%21/29 | 65.5%19/29 | — | 72.4%21/29 |
Overall
Each column shows that model's selected results on Claude Code, across the same 30 tasks. Choose a harness to compare results across the suite.
Success with retries
Each bar is one model's success on the picked harness, with observed pass@1 and projected pass@2, pass@3 and pass^3. These are expected rates, not measured repeat runs. One scale whichever harness is picked, sorted by pass@3.
GPT-6 Astra97.0%32.8%
DeepSeek V4 Pro55.0%45.0%
- pass@1
- pass@2
- pass@3
- pass^3
Projected retries · pass@1 observed · every model on the picked harness · sorted by pass@3
Shape by harness
Success rate on every harness, one polygon per entry.
Cost analysis
What a task costs each model on Claude Code, cheapest first, with how often it passes beside it.
DeepSeek V4 Pro
Success50.0%$ / task$0.25Time414sGPT-6 Astra
Success69.0%$ / task$15.72Time194s
Each tick is one task · faint ticks did not pass · the square is the average