Terminal-Bench 2.1
Audited Jcode runs over 89 terminal tasks through Harbor on Modal. Cutoff-scored and uncapped results are separated so runtime and capability are not conflated.
GPT-5.6 Sol · high effort · k=5
Earlier Opus 4.8 effort sweep
These older runs used Jcode v0.37.0 and Anthropic pricing. They remain useful for within-model effort and cost comparisons, but are not the current Sol result.
Accuracy against cost
Each point is one run: accuracy over all scored trials against cost per trial. The line follows reasoning effort at k=2. Dashed lines are competitor scores.
The medium cell is the best one: it beats xhigh on accuracy while costing less per trial and finishing tasks sooner. On this benchmark, more thinking stops paying for itself past medium.
Where the failures go
All runs
Click a column to sort. Each run links to its per-task breakdown and trial transcripts.
| Date | Model | Config | k | Scored | Accuracy | $ / trial | Median time | Note |
|---|