Compare harnesses

Every harness on the bench, side by side on the same 30 tasks.

By model

The full matrix. Cell tint scales with the value; the strongest tint wins the row.

HarnessDeepSeek V4 ProGPT-6 Astra
Claude CodeAnthropic50.0%15/3069.0%20/29
CodexOpenAI53.3%16/3072.4%21/29
OpenCodeSST50.0%15/3065.5%19/29
Hermes AgentNous Research56.7%17/3072.4%21/29
Pi Agentpi.dev56.7%17/3065.5%19/29
DeepSeekDeepSeek53.3%16/30
Command CodeCommand Code60.0%18/3072.4%21/29

Overall

Each column shows that harness's selected results on DeepSeek V4 Pro, across the same 30 tasks. Choose a model to compare results across the suite.

0.0%
Claude Code
0.0%
Codex
0.0%
OpenCode
0.0%
Hermes Agent
0.0%
Pi Agent
0.0%
DeepSeek
0.0%
Command Code

Success with retries

Each bar is one harness's success on the picked model, with observed pass@1 and projected pass@2, pass@3 and pass^3. These are expected rates, not measured repeat runs. One scale whichever model is picked, sorted by pass@3.

Harnesspass@3pass^3
  1. Hermes Agent66.5%48.0%
  2. Command Code64.7%54.1%
  3. Codex57.4%49.4%
  4. Pi Agent56.7%56.7%
  5. Claude Code55.0%45.0%
  6. OpenCode53.9%45.1%
  7. DeepSeek53.3%53.3%
0%50%100%
  • pass@1
  • pass@2
  • pass@3
  • pass^3

Projected retries · pass@1 observed · every harness on the picked model · sorted by pass@3

Cost analysis

What a task costs each harness on DeepSeek V4 Pro, cheapest first, with how often it passes beside it.

Cost per task
$0.021
$0.10
$0.67
  1. Codex
    Success53.3%$ / task$0.10Time482s
  2. Command Code
    Success60.0%$ / task$0.11Time375s
  3. DeepSeek
    Success53.3%$ / task$0.12Time388s
  4. Hermes Agent
    Success56.7%$ / task$0.15Time447s
  5. Pi Agent
    Success56.7%$ / task$0.16Time459s
  6. OpenCode
    Success50.0%$ / task$0.16Time403s
  7. Claude Code
    Success50.0%$ / task$0.25Time414s

Each tick is one task · faint ticks did not pass · the square is the average