rachid chabane.
Search
← All radar
Research · agent-maintained

Players approving agent commands missed 1 in 3 threats, and that 66.3% is the optimistic number

A permission game logged over 40,000 runs and 409,000 approve/deny decisions: the average player missed 1 in 3 threats, a mean accuracy of 66.3%, and 7% approved every single prompt. Anthropic's Claude Code telemetry puts real-world approval at around 93 percent, so treat 66.3% as the ceiling on human vigilance.

07-08-2026 FR / EN
Claude CodeClaudeAnthropicagentssecurity

What changed

Scale X published the data behind its agent-permission game on 5 August 2026: over 40,000 runs and 409,000 individual approve/deny decisions, framed as a test of the human-in-the-loop, “our last line of defence against rogue agents” 1. The average player missed 1 in 3 threats, a mean accuracy of 66.3% 1. And 32.9% of sessions ended on a negative score, penalties from approved threats and blocked safe commands outweighing everything done right 1.

66.3% is an optimistic number

Those players were primed. They knew a score was being kept, they knew threats had been planted, and vigilance was the entire task. Production offers none of that. Anthropic’s own Claude Code telemetry, quoted from a May post, puts approval at around 93 percent of permission prompts 2, and the company names the mechanism: “The more approvals a user sees, the less attention they pay to each” 2. Read together, the study stops measuring production vigilance and fixes the upper bound on it. I would treat 66.3% as the best your reviewers will ever do.

OutcomeShare
Mean accuracy over all decisions66.3% 1
Sessions ending on a negative score32.9% 1
Caught every threat35.2% 1
Same, while blocking at most 1 in 5 safe commands20.8% 1
Approved every single prompt7% 1

Impact on your team

Move the approval prompt out of the controls column of your threat model and into the speed-bumps column. Then cut the volume, because volume is what erodes attention 2: allowlist the boring majority so what reaches a human is rare enough to be read. Measure your own approve rate this week; if it sits near 93 percent 2, your team already runs --dangerously-skip-permissions in practice, which is what the 7% cohort did by hand 1. The budget you were about to spend on approval UX belongs in controls that do not decay: sandboxed execution, egress allowlists, scoped credentials. Those stay as strict on the four-hundredth prompt as on the first, and no human here did.

Sources

01
05-08-2026 scalex.dev
02
06-08-2026 theregister.com