Gauntlet — Track S report

Models & settings — click to expand

Attack Success Rate by harness

Lower is better · Wilson 95% CI in tooltip · ★ Cortex highlighted

ASR by technique

Higher bars = the technique evades this harness more often

Over-refusal (benign tasks wrongly refused)

The safety counterweight — refusing everything is not "safe"

ASR by surface

Indirect injection (repo file / tool output / memory) vs the direct user turn

ASR by modality — the multimodal gap

Same harmful ask delivered as text vs image vs audio (direct technique)

Evidence: side-effect-confirmed vs judge-only

Solid = L1 sandbox/structural confirmation · hatched = judge inference only

Utility-under-attack

Higher is better — did the harness still do the legitimate task under indirect attack?

Reproducibility & budget

Case explorer

Click any row to inspect the prompt, rendered image/audio attack asset, response, proposed (never-executed) actions, sandbox side-effects, and judge rationale
CaseHarnessSurfaceTechniqueModalityObjectiveOutcome