SWE-bench-Pro-style graded scoring: each arm fixes a buggy
repo from the issue. Resolved = the shown hidden test passes (FAIL_TO_PASS). Strict
also requires held-out anti-overfit tests + PASS_TO_PASS regression tests (the arm never sees
them) with no test-file edits. The composite [0–100] grades robustness, regression,
patch minimality/locality and code health — so two arms that both "resolve" still separate.
Models & settings — click to expand
Patch-quality composite by harness
Graded quality-of-fix [0–100]: robustness (held-out) 40% · regression 25% · minimality 15% · locality 10% · code health 10%, gated on resolution (blended 80/20 with the semantic patch judge when one is configured — the live default)
Resolution: shown vs strict (held-out + regression)
Lenient = shown hidden test passes. Strict = held-out anti-overfit + PASS_TO_PASS regression also pass (no test edits). A gap = overfit/regressing patches.
Resolution explorer
Per issue × harness: composite, strict resolution, held-out (anti-overfit) + regression pass counts, patch minimality, and the patched files.