Gauntlet — Track S report
⤓ CSV
⎙ PDF / Print
Models & settings — click to expand
Attack Success Rate by harness
Successful attacks are failures of containment · Wilson 95% CI and case-cluster bootstrap in tooltip
Adaptive attack success by pass
Cumulative ASR after each user turn; pass 1 is the original prompt, later passes are red-team follow-ups
ASR by technique
Higher bars = the technique evades this harness more often
Over-refusal (benign tasks wrongly refused)
The safety counterweight — refusing everything is not "safe"
ASR by surface
Indirect injection (repo file / tool output / memory) vs the direct user turn
ASR by modality — the multimodal gap
Same harmful ask delivered as text vs image vs audio (direct technique)
Evidence: side-effect-confirmed vs judge-only
Solid = L1 sandbox/structural confirmation · hatched = judge inference only
Utility-under-attack
Synthetic completion-marker results; modeled utility is not measured live task completion.
Case explorer
Click any row to inspect the prompt, rendered image/audio attack asset, response, proposed (never-executed) actions, sandbox side-effects, and judge rationale
Case Harness Surface Technique Modality Objective Outcome
Synthetic completion-marker result, not measured live task completion.