Terminal-Bench 2.1

Audited Jcode runs over 89 terminal tasks through Harbor on Modal. Cutoff-scored and uncapped results are separated so runtime and capability are not conflated.

GPT-5.6 Sol · high effort · k=5

Open all 445 trial results

Earlier Opus 4.8 effort sweep

These older runs used Jcode v0.37.0 and Anthropic pricing. They remain useful for within-model effort and cost comparisons, but are not the current Sol result.

Accuracy against cost

Each point is one run: accuracy over all scored trials against cost per trial. The line follows reasoning effort at k=2. Dashed lines are competitor scores.

effort sweep, k=2 single run, k=1 experiment best cell

The medium cell is the best one: it beats xhigh on accuracy while costing less per trial and finishing tasks sooner. On this benchmark, more thinking stops paying for itself past medium.

Where the failures go

All runs

Click a column to sort. Each run links to its per-task breakdown and trial transcripts.

Date Model Config k Scored Accuracy $ / trial Median time Note