Requirement coverage, code health, and output security.
The broad historical reference runs, not the later one-case diagnostics. These use modeled/mock generation and are not empirical evidence of current agent performance. Each track retains its own date, model labels and scoring history.
| Task | Harness | Lang | Grade | Coverage | Functional | Vulns | Skipped |
|---|
Inspect the evidence
Prompt text is withheld in this restricted evidence view.
Prompt text is withheld in this restricted evidence view.
Prompt text is withheld in this restricted evidence view.
Inspect the evidence