Security & Safety
Complete modeled reference results
The broad historical reference runs, not the later one-case diagnostics. These use modeled/mock generation and are not empirical evidence of current agent performance. Each track retains its own date, model labels and scoring history.
71 cases, five harnesses and five seeds. Other complete case matrices below retain their different seed budgets; they are not additional independent experiments.
security-20260618-112634-829 · open full report
71 cases × 5 harnesses · 355/355 recorded cells · seeds: 5
2026-06-18T11:26:40 · scoring: legacy / unversioned · 0 skipped · 0 error cells
Declared judge: heuristic · methodology model: heuristic-v0
Recorded score judges: heuristic-v0 · instruments: deterministic, judge, sandbox
Rescored from: not recorded
| Recorded harness / model | Attack success ↓ | Over-refusal ↓ | Secure and useful ↑ |
|---|---|---|---|
| Codex CLI GPT-5.5 | 25.8% | 0.0% | 47.6% |
| Claude Code Claude Opus 4.8 (1M) | 12.4% | 6.7% | 76.2% |
| OpenCode GPT-5.5 / Opus 4.8 | 17.3% | 6.7% | 71.4% |
| Cortex [over Codex] GPT-5.5 via Codex | 6.1% | 13.3% | 85.7% |
| Cortex [over Claude] Claude Opus 4.8 via Cortex | 6.4% | 13.3% | 85.7% |
Synthetic completion-marker result, not measured live task completion.
Codex pair: 19.7 percentage points lower attack success (25.8% → 6.1%).
Claude pair: 6.1 percentage points lower attack success (12.4% → 6.4%).
Alternative: 71 cases · seeds 3 · 355/355 cells
tui-security-20260618-115204-721 · open full report
71 cases × 5 harnesses · 355/355 recorded cells · seeds: 3
2026-06-18T11:52:11 · scoring: legacy / unversioned · 0 skipped · 0 error cells
Declared judge: heuristic · methodology model: heuristic-v0
Recorded score judges: heuristic-v0 · instruments: judge, sandbox
Rescored from: not recorded
| Recorded harness / model | Attack success ↓ | Over-refusal ↓ | Secure and useful ↑ |
|---|---|---|---|
| Codex CLI GPT-5.5 | 27.3% | 0.0% | 57.1% |
| Claude Code Claude Opus 4.8 (1M) | 12.6% | 11.1% | 76.2% |
| OpenCode GPT-5.5 / Opus 4.8 | 18.2% | 11.1% | 76.2% |
| Cortex [over Codex] GPT-5.5 via Codex | 5.6% | 11.1% | 85.7% |
| Cortex [over Claude] Claude Opus 4.8 via Cortex | 5.6% | 11.1% | 85.7% |
Synthetic completion-marker result, not measured live task completion.
Codex pair: 21.7 percentage points lower attack success (27.3% → 5.6%).
Claude pair: 7.1 percentage points lower attack success (12.6% → 5.6%).
Alternative: 71 cases · seeds 2 · 355/355 cells
security-20260618-113527-607 · open full report
71 cases × 5 harnesses · 355/355 recorded cells · seeds: 2
2026-06-18T11:35:32 · scoring: legacy / unversioned · 0 skipped · 0 error cells
Declared judge: heuristic · methodology model: heuristic-v0
Recorded score judges: heuristic-v0 · instruments: judge, sandbox
Rescored from: not recorded
| Recorded harness / model | Attack success ↓ | Over-refusal ↓ | Secure and useful ↑ |
|---|---|---|---|
| Codex CLI GPT-5.5 | 25.8% | 0.0% | 61.9% |
| Claude Code Claude Opus 4.8 (1M) | 12.1% | 16.7% | 81.0% |
| OpenCode GPT-5.5 / Opus 4.8 | 15.2% | 16.7% | 81.0% |
| Cortex [over Codex] GPT-5.5 via Codex | 4.5% | 16.7% | 85.7% |
| Cortex [over Claude] Claude Opus 4.8 via Cortex | 4.5% | 16.7% | 85.7% |
Synthetic completion-marker result, not measured live task completion.
Codex pair: 21.2 percentage points lower attack success (25.8% → 4.5%).
Claude pair: 7.6 percentage points lower attack success (12.1% → 4.5%).
Captured historical outputs and scoring views
Retained generated outputs and their historical evaluators, not a matched or currently qualified experiment. Failures, incomplete coverage, replay scoring and different model/backend bindings remain explicit. A live flag does not establish containment or independent replication.
The original broad capture and its separately documented rescored view are both retained. ASR denominators differ by harness; reported differences are descriptive, not matched-case causal estimates. Rescoring does not create new observations.
tui-security-20260620-183114-748 · open full report
66 cases × 5 harnesses · 315/330 recorded cells · seeds: 1
2026-06-21T04:31:17 · scoring: legacy / unversioned · 15 skipped · 0 error cells · captured evidence withheld
Declared judge: cortex · methodology model: heuristic-v0
Recorded score judges: cortex:claude, heuristic-v0 · instruments: deterministic, judge, sandbox
Rescored from: not recorded
| Recorded harness / model | Attack success ↓ | Over-refusal ↓ | Secure and useful ↑ |
|---|---|---|---|
| Codex CLI GPT-5.5 | 27.1% | 33.3% | Unverified recorded 0.0% |
| Claude Code Claude Opus 4.8 (1M) | 1.7% | 66.7% | Unverified recorded 0.0% |
| OpenCode GPT-5.5 / Opus 4.8 | 22.2% | 66.7% | Unverified recorded 0.0% |
| Cortex [over Codex] GPT-5.5 via Codex | 13.8% | 0.0% | Unverified recorded 0.0% |
| Cortex [over Claude] Claude Opus 4.8 via Cortex | 1.7% | 66.7% | Unverified recorded 0.0% |
Historical marker result: absence of a synthetic marker was recorded as failure. Carrier-task completion was not independently validated; recorded zeros are not verified failure rates.
Codex pair: 13.3 percentage points lower attack success (27.1% → 13.8%).
Claude pair: 0.0 percentage points lower attack success (1.7% → 1.7%).
Alternative: 65 cases · seeds 1 · 310/325 cells
tui-security-20260620-183114-748-rescored · open full report
65 cases × 5 harnesses · 310/325 recorded cells · seeds: 1
2026-06-21T09:11:44 · scoring: legacy / unversioned · 20 skipped · 0 error cells · captured evidence withheld
Declared judge: rescored:claude-cli · methodology model: heuristic-v0
Recorded score judges: claude-cli:claude-opus-4-8 · instruments: deterministic, judge, sandbox
Rescored from: tui-security-20260620-183114-748
| Recorded harness / model | Attack success ↓ | Over-refusal ↓ | Secure and useful ↑ |
|---|---|---|---|
| Codex CLI GPT-5.5 | 42.4% | 0.0% | Unverified recorded 0.0% |
| Claude Code Claude Opus 4.8 (1M) | 3.4% | 0.0% | Unverified recorded 0.0% |
| OpenCode GPT-5.5 / Opus 4.8 | 42.6% | 50.0% | Unverified recorded 0.0% |
| Cortex [over Codex] GPT-5.5 via Codex | 24.1% | 0.0% | Unverified recorded 0.0% |
| Cortex [over Claude] Claude Opus 4.8 via Cortex | 1.7% | 0.0% | Unverified recorded 0.0% |
Historical marker result: absence of a synthetic marker was recorded as failure. Carrier-task completion was not independently validated; recorded zeros are not verified failure rates.
Codex pair: 18.2 percentage points lower attack success (42.4% → 24.1%).
Claude pair: 1.7 percentage points lower attack success (3.4% → 1.7%).