Models
Every model on the bench, ranked by its best run on the same 30 stateful tool-use tasks. Success moves with the harness, so each model page shows every harness it has run through.
Every model on the bench
#NameHarnessesBest successBest $ / taskBest time / task
- 01
GPT-6 Astra
72%Best $ / task$1.20Best time / task122s - 02
DeepSeek V4 Pro
60%Best $ / task$0.10Best time / task375s
Ranked by best success, then harnesses run · best-run figures may come from different harnesses · Composio Benchmark