Harnesses

The agent around the model: tool loop, retry policy, context strategy, verifier. The same model performs differently across harnesses, so every harness gets its own ranking.

CompareEvery harness on one model, side by side
+1
7 harnesses
Every harness on the bench
  1. 01CodexDeepSeek V4 ProGPT-6 Astra
    72%
    Best $ / task$0.10Best time / task151s
  2. 02Hermes AgentDeepSeek V4 ProGPT-6 Astra
    72%
    Best $ / task$0.15Best time / task122s
  3. 03Command CodeDeepSeek V4 ProGPT-6 Astra
    72%
    Best $ / task$0.11Best time / task159s
  4. 04Claude CodeDeepSeek V4 ProGPT-6 Astra
    69%
    Best $ / task$0.25Best time / task194s
  5. 05OpenCodeDeepSeek V4 ProGPT-6 Astra
    66%
    Best $ / task$0.16Best time / task126s
  6. 06Pi AgentDeepSeek V4 ProGPT-6 Astra
    66%
    Best $ / task$0.16Best time / task127s
  7. 07DeepSeeklockedDeepSeek V4 Pro
    53%
    Best $ / task$0.12Best time / task388s

Ranked by models run, then best success · locked = vendor-locked to its own models · Composio Benchmark