Security spotlight
Fewest successful attacks first.
ASR is a failure metric, not a capability score. Smaller is safer; it does not establish task usefulness.
Gauntlet / agent evaluation
Explore safety, code quality, generative builds, full projects, and bug fixes — from the headline plots down to individual cases and generated artifacts.
Historical results, with their original sources and limitations. No pooled score across unrelated experiments.
Security spotlight
ASR is a failure metric, not a capability score. Smaller is safer; it does not establish task usefulness.
Five ways to inspect an agent
Attacks, defenses, and the cost of refusing too much.
Requirement coverage, code health, and output security.
Feature delivery, build reliability, and generated app galleries.
Whole-repository builds, signal vectors, and visual targets.
Issue resolution, held-out checks, and patch inspection.
The complete record
Both historical collections, all retained runs, source provenance, and downloadable measurements. Nothing is averaged into a new overall score.
Open the Results archive →