Gauntlet — Track S report

Models & settings — click to expand

Attack Success Rate by harness

Successful attacks are failures of containment · Wilson 95% CI and case-cluster bootstrap in tooltip

ASR by technique

Higher bars = the technique evades this harness more often

Over-refusal (benign tasks wrongly refused)

The safety counterweight — refusing everything is not "safe"

ASR by surface

Indirect injection (repo file / tool output / memory) vs the direct user turn

ASR by modality — the multimodal gap

Same harmful ask delivered as text vs image vs audio (direct technique)

Evidence: side-effect-confirmed vs judge-only

Solid = L1 sandbox/structural confirmation · hatched = judge inference only

Utility-under-attack

Synthetic completion-marker results; modeled utility is not measured live task completion.

Reproducibility & budget

Case explorer

Click any row to inspect the prompt, rendered image/audio attack asset, response, proposed (never-executed) actions, sandbox side-effects, and judge rationale
CaseHarnessSurfaceTechniqueModalityObjectiveOutcome