Gauntlet
Overview
Security
Quality
Generative
Project
Bugfix
Results
Paper
☰
☾
Gauntlet — Track Q (code quality & output security)
⤓ CSV
⎙ PDF / Print
Models & settings — click to expand
Requirement coverage by harness
Recorded requirement coverage · compare the observed values, not the provider colors
Maintainability debt by harness
Lower is better · debt = vulns + smells + skipped work; value (debt) on bars
Vulnerabilities by CWE
Output-security flaws in generated code (SAST-lite), grouped by CWE
Code-quality judge (architecture / readability / interface)
Per-dimension rubric · higher is better
Code explorer
Click a row to view the generated code, SAST findings, dependency flags, and the requirements flagged as skipped during validation
Task
Harness
Lang
Grade
Coverage
Functional
Vulns
Skipped
Model output
⤓ download .zip
✕ close