DeepSWE
Matched harness comparisons and audited Jcode results with benchmark version, reasoning effort, cutoff policy, exact versions, and denominators shown explicitly.
DeepSWE v1.1 · Claude Opus 4.8 high
| Harness | Untimed | 90m cutoff | Timeout delta | Overtime | Median / max | Infra replacements |
|---|
Download all 452 audited per-task outcomes and provenance
DeepSWE v1.1 · GPT-5.6 Sol max
v1.1 comparison
DataCurve used mini-swe-agent at k=4. Jcode used one attempt per task at k=1, so this is context, not a paired harness experiment.
Cost and token reporting
Dollar totals are shown only when the source result reported them.
| Run | Scored | Total cost | Cost / trial | Cost / pass |
|---|
Download the audited v1.1 summary
DataCurve mini-SWE-agent publication
Official DeepSWE v1.1 GPT-5.6 Sol effort sweep. Each row aggregates four whole-benchmark runs. Provider, verifier, and network errors are excluded from the denominator under DataCurve's published policy.
| Effort | Pass@1 | 95% interval | Passed / scored | Avg cost | Avg output | Avg steps |
|---|
Download the derived per-task data and provenance · Browse DataCurve's official trials
DeepSWE v1 · matched harness comparison
The same 113 DeepSWE v1 tasks, GPT-5.6 Sol, high reasoning effort, k=1. Only the harness changes.
Leaderboard
Official post-hoc scoring at the 90 minute cutoff. Our bar is green.
Per-task outcomes
Agent runtime
Per-task breakdown
Click a column to sort. Filters narrow to the tasks where the harnesses disagreed.
| Task | Outcome | jcode | Codex CLI | jcode time | Transcripts |
|---|
Methodology and provenance
Everything needed to reproduce or audit the comparison. Hashes are SHA-256.