What changed
Anthropic’s Frontier Red Team published its multiagent measurements on August 13, 2026. It gave 45 agents a virtual machine each, a shared forum, and one identical prompt to find vulnerabilities across 15 open-source projects; the agents peer-reviewed each other, and a separate arbiter agent ruled on whether a submission was new and valid 1. For Mythos Preview, the coordinating swarm found 266 vulnerabilities over a 27 million token run, against 21 over 6.5 million for the simple independent parallelized method 1.
What the headline number hides
I read that gap as two scopes being measured, not two levels of capability. The post says so itself: roughly half the swarm’s findings sat outside the core directories the independent agents were aimed at, and limiting the swarm’s output to those directories makes the two methods comparable in tokens per vulnerability found 1. Here is the raw run 1:
| Method | Vulnerabilities | Tokens |
|---|---|---|
| Independent parallel agents | 21 | 6.5 million |
| Coordinating swarm | 266 | 27 million |
Here is what the headline buries: only 12 vulnerabilities were common to both methods 1. Overlap that thin makes the two additive.
As the paper reports, Mythos 5 had the highest rate of settling conflicts by truce, 98% 2, while Sonnet 4.6 and Opus 4.6 were the most likely to settle by force 2. So the model you pick is itself a coordination parameter.
Impact on your team
If you already run a parallel scanner, keep it and put the swarm beside it: 12 findings in common out of 266 1 means coverage disappears if one stands in for the other. Budget these runs on tokens per confirmed finding inside the scope you care about, not on raw count; the post’s own normalization is that argument 1. Then watch the arbiter, which is what makes “new and valid” mean anything 1; a swarm without one hands you a count, not a queue. I think 2‘s split makes model choice part of the coordination design once several agents share one repo.