DeepSWE

Matched harness comparisons and audited Jcode results with benchmark version, reasoning effort, cutoff policy, exact versions, and denominators shown explicitly.

DeepSWE v1.1 · Claude Opus 4.8 high

Harness Untimed 90m cutoff Timeout delta Overtime Median / max Infra replacements

Download all 452 audited per-task outcomes and provenance

DeepSWE v1.1 · GPT-5.6 Sol max

v1.1 comparison

DataCurve used mini-swe-agent at k=4. Jcode used one attempt per task at k=1, so this is context, not a paired harness experiment.

Cost and token reporting

Dollar totals are shown only when the source result reported them.

Run Scored Total cost Cost / trial Cost / pass

Download the audited v1.1 summary

DataCurve mini-SWE-agent publication

Official DeepSWE v1.1 GPT-5.6 Sol effort sweep. Each row aggregates four whole-benchmark runs. Provider, verifier, and network errors are excluded from the denominator under DataCurve's published policy.

Effort Pass@1 95% interval Passed / scored Avg cost Avg output Avg steps

Download the derived per-task data and provenance · Browse DataCurve's official trials


DeepSWE v1 · matched harness comparison

The same 113 DeepSWE v1 tasks, GPT-5.6 Sol, high reasoning effort, k=1. Only the harness changes.

Leaderboard

Official post-hoc scoring at the 90 minute cutoff. Our bar is green.

Per-task outcomes

jcode win codex win both pass both fail

Agent runtime

jcode passed jcode failed

Per-task breakdown

Click a column to sort. Filters narrow to the tasks where the harnesses disagreed.

Task Outcome jcode Codex CLI jcode time Transcripts

Methodology and provenance

Everything needed to reproduce or audit the comparison. Hashes are SHA-256.

Version boundary. This is DeepSWE v1. The current DataCurve leaderboard uses DeepSWE v1.1, so those scores are not directly comparable.

Download the audited leaderboard data