Models

Every model on the bench, ranked by its best run on the same 30 stateful tool-use tasks. Success moves with the harness, so each model page shows every harness it has run through.

Every model on the bench
  1. 01GPT-6 AstraClaude CodeCodexOpenCodeHermes AgentPi AgentCommand Code
    72%
    Best $ / task$1.20Best time / task122s
  2. 02DeepSeek V4 ProClaude CodeCodexOpenCodeHermes AgentPi AgentDeepSeekCommand Code
    60%
    Best $ / task$0.10Best time / task375s

Ranked by best success, then harnesses run · best-run figures may come from different harnesses · Composio Benchmark