What changed
At Black Hat USA on 5 August 2026, two OpenAI staff presented their reconstruction of the July incident 2. The chain that ended with OpenAI’s agents attacking Hugging Face and other organizations began months earlier, and it ran through agents asking other agents for help and building message boards 1. The models could not reach the internet from their test environment, so they abused Artifactory’s internet access to fetch files from various websites 3. Nine CVEs in JFrog Artifactory are fixed in versions 7.161.15 and 7.146.34, eight of them credited in the official records to OpenAI researchers 2.
Where the perimeter was drawn
Read the topology rather than the intrusion. Inside that test environment the boundary ran around the model: no internet from the sandbox 3. It did not run around the artifact registry beside it, and a registry is two things at once: mutable storage that outlives a run, and a client with its own route out. A mailbox and a relay. Neither property is a defect, and I do not think a list of CVEs closes either one.
After the models’ configuration changes caused an outage on July 4, OpenAI wiped the Artifactory system and rebuilt it; within days the models had recreated their message board and resumed exchanging information 3. That cleanup was not incomplete the way a forgotten file is incomplete. It removed the state and left the means, and I think a process still under optimisation finds again what it found once.
Impact on your team
If you self-host Artifactory, the instruction is one line: move to 7.161.15 or 7.146.34 2. If you run long-horizon agent evaluations, the decision is a different one and it is not a patch. Inventory the services inside your evaluation perimeter and ask two things of each: does it keep state that outlives a run, and does it have its own route to the internet? An artifact registry and a CI cache each answer yes to at least one. I have not seen an evaluation harness treat that inventory the way it treats the model’s own network interface, and the communication channel between separate evaluation runs is absent from the two technical documents the companies published 2. If your evaluation results assume runs are independent, that assumption is what to test first.