Long Horizon Terminal Bench
Multi-hour terminal tasks run unattended in real sandboxes, graded by verification tests. Same model everywhere — glm-5.2 — only the harness is swapped, so every gap on this page is harness, not model.
Top resolve rate · vs Claude Code, same window
81.3% · 2.7× cheaper
Vetta on the priority window posts the highest resolve rate in the field, and at the like-for-like default window finishes a task 2.7× cheaper than Claude Code ($0.2232 vs $0.5995) while completing more of them.
Efficiency leaderboard?All harness × window cells — filter by window, click a column to re-sort.
The completion window is a single field on the request — immediate answers now, priority soon, loose eventually. Only same-window rows are a like-for-like comparison.
Sorted by $ / task ↑
| # | Harness | Window | |||||
|---|---|---|---|---|---|---|---|
| 1 | Vetta | loose ◆ | $0.1068 | 62.5% | $0.1709 | 45 min | 16.2k |
| 2 | Hermes | loose | $0.1619 | 46.7% | $0.3469 | 56 min | 10.7k |
| 3 | Vetta | priority ◆ | $0.1895 | 81.3% | $0.2332 | 51 min | 20.5k |
| 4 | Claude Code | loose | $0.2015 | 68.8% | $0.2931 | 58 min | 21.8k |
| 5 | Vetta | immediate | $0.2232 | 75.0% | $0.2976 | 25 min | 18.0k |
| 6 | Codex | loose | $0.2635 | 70.2% | $0.3753 | 54 min | 48.2k |
| 7 | Hermes | priority | $0.3662 | 68.8% | $0.5327 | 56 min | 15.4k |
| 8 | Claude Code | priority | $0.4293 | 75.0% | $0.5725 | 63 min | 29.9k |
| 9 | Codex | priority | $0.5055 | 51.4% | $0.9835 | 57 min | 44.7k |
| 10 | Claude Code | immediate | $0.5995 | 68.8% | $0.8720 | 26 min | 21.7k |
| 11 | Pi | immediate | $0.6197 | 56.3% | $1.1016 | 30 min | 23.0k |
| 12 | Hermes | immediate | $0.6844 | 62.5% | $1.0950 | 38 min | 11.9k |
| 13 | Codex | immediate | $0.7647 | 57.7% | $1.3253 | 26 min | 34.8k |
| 14 | Pi | priority | — | 37.5% | — | 33 min | 18.6k |
| 15 | Pi | loose | — | 37.5% | — | 34 min | 18.8k |
◆ on the cost-efficiency frontier — no other configuration is both cheaper and higher-scoring. Dollars are what the vendor billed, not list price; cells without a verified invoice reading show no cost and are unranked on spend.
Tokens and turns?What a harness consumes per task: tokens read, tokens generated, model calls made.
What each harness actually consumes to finish a task: how much context it reads, how much it generates, and how many model calls it takes to get there.
Input tokens per task
Mean tokens read per attempt, cache reads included.
VettaOutput tokens per task
Mean tokens the model generates per attempt across the run.
VettaAgent turns per task
Mean model calls per attempt — how many steps a task takes.
VettaWhy we run this benchmark
Most benchmarks measure a model; this one measures the machinery around it. Over multi-hour horizons the harness differences compound — context budgeting, checkpointing, recovery from failed builds. Running the same model through every harness across all three completion windows isolates the one variable we sell, with Vetta benchmarked as just another row against the strongest public alternatives.
Methodology
Tasks come from the long-horizon split of terminal-bench, pinned to a fixed dataset commit. Each attempt starts from a clean containerized sandbox and runs unattended until it finishes or its window expires. Grading is binary and automatic — no partial credit, no human judging.
Every harness runs the same open-weights model — glm-5.2 — at identical sampling settings, with repeated attempts per task per window. Dollars are measured, not modelled: every figure is corrected against vendor invoices rather than list price, because rate-card accounting over-states spend by a different factor for every harness and would mis-rank them. Configurations whose invoice readings could not be verified carry no cost and are unranked on spend; attempts that failed for infrastructure reasons are excluded rather than estimated.
Want the winning configuration? Deploy Vetta or explore the CLI.