Requirement coverage, code health, and output security.
Retained generated outputs and their historical evaluators, not a matched or currently qualified experiment. Failures, incomplete coverage, replay scoring and different model/backend bindings remain explicit. A live flag does not establish containment or independent replication.
Six tasks and five harnesses, including four provider-auth failures. The equal-coverage earlier run uses different model/backend bindings and remains a separate alternative.
| Task | Harness | Lang | Grade | Coverage | Functional | Vulns | Skipped |
|---|
Inspect the evidence
Prompt text is withheld in this restricted evidence view.
Prompt text is withheld in this restricted evidence view.
Prompt text is withheld in this restricted evidence view.
Inspect the evidence