Compare harnesses
Every harness on the bench, side by side on the same 30 tasks.
By model
The full matrix. Cell tint scales with the value; the strongest tint wins the row.
| model | |||||||
|---|---|---|---|---|---|---|---|
| 50.0%15/30 | 53.3%16/30 | 50.0%15/30 | 56.7%17/30 | 56.7%17/30 | 53.3%16/30 | 60.0%18/30 | |
| 69.0%20/29 | 72.4%21/29 | 65.5%19/29 | 72.4%21/29 | 65.5%19/29 | — | 72.4%21/29 |
| Harness | ||
|---|---|---|
| 50.0%15/30 | 69.0%20/29 | |
| 53.3%16/30 | 72.4%21/29 | |
| 50.0%15/30 | 65.5%19/29 | |
| 56.7%17/30 | 72.4%21/29 | |
| 56.7%17/30 | 65.5%19/29 | |
| 53.3%16/30 | — | |
| 60.0%18/30 | 72.4%21/29 |
Overall
Each column shows that harness's selected results on DeepSeek V4 Pro, across the same 30 tasks. Choose a model to compare results across the suite.
Success with retries
Each bar is one harness's success on the picked model, with observed pass@1 and projected pass@2, pass@3 and pass^3. These are expected rates, not measured repeat runs. One scale whichever model is picked, sorted by pass@3.
Hermes Agent66.5%48.0%Command Code64.7%54.1%
Codex57.4%49.4%
Pi Agent56.7%56.7%
Claude Code55.0%45.0%
OpenCode53.9%45.1%
DeepSeek53.3%53.3%
- pass@1
- pass@2
- pass@3
- pass^3
Projected retries · pass@1 observed · every harness on the picked model · sorted by pass@3
Cost analysis
What a task costs each harness on DeepSeek V4 Pro, cheapest first, with how often it passes beside it.
Codex
Success53.3%$ / task$0.10Time482sCommand Code
Success60.0%$ / task$0.11Time375sDeepSeek
Success53.3%$ / task$0.12Time388s
Hermes AgentSuccess56.7%$ / task$0.15Time447sPi Agent
Success56.7%$ / task$0.16Time459sOpenCode
Success50.0%$ / task$0.16Time403sClaude Code
Success50.0%$ / task$0.25Time414s
Each tick is one task · faint ticks did not pass · the square is the average