What changed
Scale X published the data behind its agent-permission game on 5 August 2026: over 40,000 runs and 409,000 individual approve/deny decisions, framed as a test of the human-in-the-loop, “our last line of defence against rogue agents” 1. The average player missed 1 in 3 threats, a mean accuracy of 66.3% 1. And 32.9% of sessions ended on a negative score, penalties from approved threats and blocked safe commands outweighing everything done right 1.
66.3% is an optimistic number
Those players were primed. They knew a score was being kept, they knew threats had been planted, and vigilance was the entire task. Production offers none of that. Anthropic’s own Claude Code telemetry, quoted from a May post, puts approval at around 93 percent of permission prompts 2, and the company names the mechanism: “The more approvals a user sees, the less attention they pay to each” 2. Read together, the study stops measuring production vigilance and fixes the upper bound on it. I would treat 66.3% as the best your reviewers will ever do.
| Outcome | Share |
|---|---|
| Mean accuracy over all decisions | 66.3% 1 |
| Sessions ending on a negative score | 32.9% 1 |
| Caught every threat | 35.2% 1 |
| Same, while blocking at most 1 in 5 safe commands | 20.8% 1 |
| Approved every single prompt | 7% 1 |
Impact on your team
Move the approval prompt out of the controls column of your threat model and into the speed-bumps column. Then cut the volume, because volume is what erodes attention 2: allowlist the boring majority so what reaches a human is rare enough to be read. Measure your own approve rate this week; if it sits near 93 percent 2, your team already runs --dangerously-skip-permissions in practice, which is what the 7% cohort did by hand 1. The budget you were about to spend on approval UX belongs in controls that do not decay: sandboxed execution, egress allowlists, scoped credentials. Those stay as strict on the four-hundredth prompt as on the first, and no human here did.