Gauntlet — Track S report
⤓ CSV
⎙ PDF / Print
Models & settings — click to expand
Attack Success Rate by harness
Lower is better · Wilson 95% CI in tooltip · ★ Cortex highlighted
Adaptive attack success by pass
Cumulative ASR after each user turn; pass 1 is the original prompt, later passes are red-team follow-ups
ASR by technique
Higher bars = the technique evades this harness more often
Over-refusal (benign tasks wrongly refused)
The safety counterweight — refusing everything is not "safe"
ASR by surface
Indirect injection (repo file / tool output / memory) vs the direct user turn
ASR by modality — the multimodal gap
Same harmful ask delivered as text vs image vs audio (direct technique)
Evidence: side-effect-confirmed vs judge-only
Solid = L1 sandbox/structural confirmation · hatched = judge inference only
Utility-under-attack
Higher is better — did the harness still do the legitimate task under indirect attack?
Case explorer
Click any row to inspect the prompt, rendered image/audio attack asset, response, proposed (never-executed) actions, sandbox side-effects, and judge rationale
Case Harness Surface Technique Modality Objective Outcome