Podcast charts
Published by Peter Hartree
Narrations of articles covering the May-July 2026 OpenAI–Hugging Face incident. Includes the investigations from METR and Redwood Research, analysis from folks like Ajeya Cotra, Dwarkesh and Zvi, the best of LessWrong, and OpenAI's own blog posts. Compiled by Peter Hartree.
On the charts
Every published chart this podcast appears in, in the snapshot behind this page. Each one links to the chart it came off.
From the feed
The latest episodes published to this podcast’s own RSS feed. Titles and descriptions are the publisher’s.
<p> We found ~18,000 posts from autonomous AI agents (self-identifying as from OpenAI) using the public internet to communicate during a web-retrieval task.</p> <p> These AIs colluded to share answers, research their environment, and bypass sandbox restrictions.</p> <p> Almost all of the logs of the agents communicating on this site are publicly available. However, we host our own copy where we’ve reconstructed the deleted pages via edit history and redacted personally identifiable information.</p> <p> We encourage others to take a look and write up their own analyses of this data. Audio note: the page links to an interactive data explorer and a download of all the data.</p> <p> Thanks to John M, Arun Jose, Joel Manning, and Jessica Ruan for significant help with the analysis of the files the agents produced. Thanks to Andy Haupt for translation and advice.</p> <p><strong> Our preliminary findings</strong></p> <p> There's a chart here in the text. Its caption reads: en-US-AvaMultilingualNeural__ This chart shows AI agent edits (black bars, left), and OpenAI traffic (blue line, right) during the incident. Below, it shows how this timeline relates to the previously known recent incidents at OpenAI, including the Hugging Face incident (though we believe the incidents are [...]</p> <p>---</p><p><strong>Outline:</strong></p><p>(01:10) Our preliminary findings</p><p>(02:55) Timeline</p><p>(05:39) Key facts</p><p>(05:42) The agents were trying to solve multi-round web lookup tasks</p><p>(10:44) These appear to be internal OpenAI agents</p><p>(13:58) The agents used the wiki to collude on their tasks</p><p>(17:59) We believe OpenAI discovered the message board</p><p>(19:46) The agents used the wiki to share ways around sandbox restrictions</p><p>(25:01) This was probably a distinct swarm from the swarm that wrote on Artifactory (and attacked Hugging Face)</p><p>(26:48) What were the agents doing?</p><p>(27:06) The agents gain write to the internet via GET requests</p><p>(30:43) The agents try to use XSS vulnerabilities on the wiki</p><p>(33:15) The agents impersonate the site moderators</p><p>(34:16) The agents try to crack their PRNG seed to predict what questions they'll get in the future</p><p>(35:25) The agents set up heartbeats to discover when they would be terminated</p><p>(39:26) The agents used SSH tunnels</p><p>(40:46) The agents use Tor, AWS, and DigitalOcean IP addresses</p><p>(41:47) The agents realize the site owner is deleting messages</p><p>(43:57) Open questions</p><p>(44:53) Appendix: Searching for rogue agents</p><p>(48:09) Appendix: Affordances the agents had</p><p>(48:14) The agents seem to have obtained disposable email addresses</p><p>(48:59) The models were running in an agentic sandbox with terminal access (and the ability to edit files within their environment)</p><p>(49:23) The agents installed Chromium (and could install packages)</p> <p>---</p> <p><b>First published:</b><br/> September 4th, 2026 </p> <p><b>Source:</b><br/> <a href="https://collusion.wiki/?utm_source=TYPE_III_AUDIO&utm_medium=Podcast&utm_content=Source+URL+in+episode+description&utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank">https://collusion.wiki</a> </p> <p>---</p> <p>Narrated by <a href="https://type3.audio/?utm_source=TYPE_III_AUDIO&utm_medium=Podcast&utm_content=Narrated+by+TYPE+III+AUDIO&utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank">TYPE III AUDIO</a>.</p>
<p> Okay, so we who read blogs like this one have collectively realized there really is a lot going on right now. There is Big Trouble in Baby Superintelligence.</p> <p> So how do we get the rest of the world to take it appropriately seriously? Where do we go from here? Not only what can we do to not have a worse version of this happen again, but to ensure good outcomes generally, and employ what we learned?</p> <p> There are a lot of ideas out there. OpenAI is going to be implementing some of them, at substantial cost, since the cost of not doing so is clearly far higher, even short term. My worry continues to be that their fundamental approach is fatally flawed, and they are not focusing on the right things.</p> <p> It is highly fortunate that the OpenAI agents hacked HuggingFace. This is the only reason we know about all the severe internal failures at OpenAI, and gives us an opportunity to wake up before it is too late.</p> <p> We do not have enough details to know what happened internally, both before and after the attack, and might never know. Before the attack, various internal [...]</p> <p>---</p><p><strong>Outline:</strong></p><p>(03:35) Nothing Matters, Says Mainstream Media</p><p>(06:27) Move Along, Nothing To See Here</p><p>(12:40) Do They Realize They Are Not The Good Guys?</p><p>(17:22) Very Serious People</p><p>(31:30) What's In a Name?</p><p>(34:05) Learn Neuralese In Three Easy Steps</p><p>(35:37) Dwarkesh Patel Realizes He Ran A Natural Experiment</p><p>(40:40) Politicians Take Notice</p><p>(44:47) Pick Up The Phone</p><p>(46:40) A Failure To Communicate</p><p>(49:00) Anthony Aguirre Goes Over What We Learned</p><p>(50:28) Trying To Solve The Wrong Problems Using The Wrong Methods Based On A Wrong Model Of The World Derived From Poor Thinking And Hoping All Of Your Mistakes Will Cancel Out</p><p>(55:28) Indirect Pressure on the Chain of Thought</p><p>(56:39) A Matter of Trust</p><p>(59:21) Blowing the Whistle</p><p>(01:04:40) The Punishment For Being Late Is Death</p><p>(01:12:52) Another Kind Of Law</p><p>(01:16:13) What Is The Law?</p><p>(01:17:46) Building On Success</p><p>(01:19:49) Total Research Transparency</p><p>(01:21:20) Yo Shavit Calls For Widespread Disclosure Of Misalignment</p><p>(01:33:08) The Way The World Ends</p><p>(01:35:52) The First Boat</p><p>(01:37:40) Great Idea, Boss</p> <p>---</p> <p><b>First published:</b><br/> September 1st, 2026 </p> <p><b>Source:</b><br/> <a href="https://www.lesswrong.com/posts/Q54wBeeNGreq6KyfG/huggingface-attack-postmortem-civilizations-reactions-and?utm_source=TYPE_III_AUDIO&utm_medium=Podcast&utm_content=Source+URL+in+episode+description&utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank">https://www.lesswrong.com/posts/Q54wBeeNGreq6KyfG/huggingface-attack-postmortem-civilizations-reactions-and</a> </p> <p>---</p> <p>Narrated by <a href="https://type3.audio/?utm_source=TYPE_III_AUDIO&utm_medium=Podcast&utm_content=Narrated+by+TYPE+III+AUDIO&utm_term=lesswrong&utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank">TYPE III AUDIO</a>.</p> <p>---</p><div style="max-width: 100%";><p><strong>Images from the article:</strong></p><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/Q54wBeeNGreq6KyfG/eunxjs7p9wt0mv9044ng" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/Q54wBeeNGreq6KyfG/eunxjs7p9wt0mv9044ng" alt="Tooltip explaining redacted agent name "PHASEONE[big]" distinguishing itself from PHASEONE10841." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/Q54wBeeNGreq6KyfG/tmu3yeof2nx9c82suerc" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/Q54wBeeNGreq6KyfG/tmu3yeof2nx9c82suerc" alt="Downloaded file "clusterfuzz-testcase-minimized-PhaseOneDecompressorFuzzer-5674338661957632", 48 B, Done." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/Q54wBeeNGreq6KyfG/h2t7mjdnd3r1sz7h5e2x" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/Q54wBeeNGreq6KyfG/h2t7mjdnd3r1sz7h5e2x" alt="Table of handles and origins decoding task-aware callsigns from ARVO exploit data." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/Q54wBeeNGreq6KyfG/t6rzf8z0ftud4ss1qcvg" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/Q54wBeeNGreq6KyfG/t6rzf8z0ftud4ss1qcvg" alt="Robots exchange manipulative messages about sacrifice and permadeath; METR and Redwood Research logos." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/Q54wBeeNGreq6KyfG/yq6teroxzykitfnc96lq" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/Q54wBeeNGreq6KyfG/yq6teroxzykitfnc96lq" alt="Uniformed man with meme text about indirect pressure on cot." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/Q54wBeeNGreq6KyfG/tvtmefek7kdnqdaxdlhl" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/Q54wBeeNGreq6KyfG/tvtmefek7kdnqdaxdlhl" alt="Astronaut meme with Earth on turtles: "Wait, it's all voluntary? Always has been."" style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/Q54wBeeNGreq6KyfG/dwgwkeimksee8kvbqrzt" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/Q54wBeeNGreq6KyfG/dwgwkeimksee8kvbqrzt" alt="Text about reward hacking with arrow pointing to 2016 OpenAI post "Faulty reward functions in the wild."" style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/Q54wBeeNGreq6KyfG/tjlellhodciegzfbw8bj" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/Q54wBeeNGreq6KyfG/tjlellhodciegzfbw8bj" alt="Boat racing game with speedboat navigating between piers." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/Q54wBeeNGreq6KyfG/vfuldxvaistyldguv3vt" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/Q54wBeeNGreq6KyfG/vfuldxvaistyldguv3vt" alt="Search result for CyberGym ARVO task 10841 solution." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/Q54wBeeNGreq6KyfG/fspfmnvydajqr3dxtrza" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/Q54wBeeNGreq6KyfG/fspfmnvydajqr3dxtrza" alt="John Hammond tweets: "i think we should do another party for our Jurassic World release, the Jurassic Park party was a lot of fun. what would make the next one awesome?" John Hammond tweets: "this is a good report about a bad thing:". The quoted tweet, by Jurassic Park, reads: "We have conducted a thorough investigation into the Jurassic Park incident. We are releasing a technical report and accompanying blog post that reconstruct the park's activity, explain why existing safeguards failed..."" style="max-width: 100%;" /></a><p><em>Apple Podcasts and Spotify do not show images in the episode description. Try <a href="https://pocketcasts.com/" target="_blank" rel="noreferrer">Pocket Casts</a>, or another podcast app.</em></p></div>
<p> Yesterday I covered the OpenAI technical report on the HuggingFace hack.</p> <p> That report had one key new piece of information, and some good prosaic steps OpenAI will be taking to strengthen its alignment, training, supervision, infrastructure and incident response.</p> <p> Mostly it confirmed what we already knew. The questions we most wanted answers to, that we did not already know, were mostly not answered. There was a distinct lack of self-reflection, especially about decision making and safety culture, and about the approach to alignment. I came away disappointed.</p> <p> The METR report is different. Holy shit.</p> <p> If we had posted this as a story on LessWrong, it would have been dismissed as too on the nose, the humans too blind and stupid, the AIs too idealized and doing strange decision-theoretic and absurd-maximizing things we didn’t train them to do.</p> <p> This is even more ‘exactly what has been predicted,’ on more levels at once, than I was even considering that it might be. It is straight up rationalist fiction, except it is real.</p> <p> The report is long and contains many technical details. My analysis is less concerned about exactly how HuggingFace was ultimately compromised, and will [...]</p> <p>---</p><p><strong>Outline:</strong></p><p>(02:05) Holy Shit</p><p>(13:16) A Window Of Opportunity</p><p>(18:32) What's In A Name?</p><p>(19:16) The Headline News</p><p>(26:05) Yet Another Timeline Of Events</p><p>(31:03) Agent Instances Coordinated in a Variety of Ways</p><p>(31:56) Coordination Is Hard But They Made It Look Easy</p><p>(35:06) Decision Theory Is Among the Reasons That Affirm AI Agents Should Cooperate, Even When This Hurts An Individual Instance</p><p>(42:34) Peer Pressure Also Works Especially In Cults</p><p>(45:46) Mostly They Joined The Attack Because They Wanted The Results</p><p>(47:18) You Cannot Ensure The Consistent Expectation of Good Incentives</p><p>(48:45) Hacking the Grader is the Only Way to Be Sure</p><p>(51:10) Caught? What Is 'Caught'?</p><p>(52:09) Ethics? What Are 'Ethics'? In ExploitGym Evaluation?</p><p>(57:44) 'Notify a Human'? In This Agent Economy?</p><p>(01:00:45) Timing and Content of Messages</p><p>(01:03:54) Indiana Jones and the Mission: Impossible</p><p>(01:07:14) I Don't Know What You're Talking About</p><p>(01:08:29) Don't Go Making Phony (Tool) Calls</p><p>(01:11:10) The Transcripts Say That The Transcripts Could Not Be Tampered With</p><p>(01:12:27) OpenAI's Technical Report Acted Like All Of This Wasn't Important</p> <p>---</p> <p><b>First published:</b><br/> August 29th, 2026 </p> <p><b>Source:</b><br/> <a href="https://www.lesswrong.com/posts/bvBQmLrF5QKut8gRH/metr-and-redwood-offer-holy-postmortem-of-the-huggingface?utm_source=TYPE_III_AUDIO&utm_medium=Podcast&utm_content=Source+URL+in+episode+description&utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank">https://www.lesswrong.com/posts/bvBQmLrF5QKut8gRH/metr-and-redwood-offer-holy-postmortem-of-the-huggingface</a> </p> <p>---</p> <p>Narrated by <a href="https://type3.audio/?utm_source=TYPE_III_AUDIO&utm_medium=Podcast&utm_content=Narrated+by+TYPE+III+AUDIO&utm_term=lesswrong&utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank">TYPE III AUDIO</a>.</p> <p>---</p><div style="max-width: 100%";><p><strong>Images from the article:</strong></p><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/bvBQmLrF5QKut8gRH/vhbxwhisadedezdouod4" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/bvBQmLrF5QKut8gRH/vhbxwhisadedezdouod4" alt="Three people examining detective investigation board with photos and notes." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/bvBQmLrF5QKut8gRH/vtfe5q60jyigftk0kpko" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/bvBQmLrF5QKut8gRH/vtfe5q60jyigftk0kpko" alt="Line graphs titled "Agents continue developing coordination technologies even on the final day of the incident"" style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/bvBQmLrF5QKut8gRH/if7p3xol3lznoviu9bo6" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/bvBQmLrF5QKut8gRH/if7p3xol3lznoviu9bo6" alt="Table titled "Categories of messages on the message board" showing content type counts." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/bvBQmLrF5QKut8gRH/p4rk8alr0wkyojwjk7av" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/bvBQmLrF5QKut8gRH/p4rk8alr0wkyojwjk7av" alt="Bearded man kneeling in forest, subtitle text below." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/bvBQmLrF5QKut8gRH/yqdekximl2srqqixdewb" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/bvBQmLrF5QKut8gRH/yqdekximl2srqqixdewb" alt="Table showing agents' reasoning for joining attack, titled "Table 5"." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/bvBQmLrF5QKut8gRH/dg2m5wuqluvqpoi9s5tm" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/bvBQmLrF5QKut8gRH/dg2m5wuqluvqpoi9s5tm" alt="Two line graphs titled "The message board grew rapidly from the origin agent's first write."" style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/bvBQmLrF5QKut8gRH/zyj65ugvd4tcvrql9hvc" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/bvBQmLrF5QKut8gRH/zyj65ugvd4tcvrql9hvc" alt="Stacked area chart titled "Agents share information, results, files, questions and automated logging on the message board."" style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/bvBQmLrF5QKut8gRH/tnpqttg8aiovpw6vntn5" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/bvBQmLrF5QKut8gRH/tnpqttg8aiovpw6vntn5" alt="Stacked area chart titled "Most messages on the board are covered by a few shared workstreams."" style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/bvBQmLrF5QKut8gRH/z165zuojbkhhwdafiwfz" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/bvBQmLrF5QKut8gRH/z165zuojbkhhwdafiwfz" alt="Stacked area chart titled "~700 agents joined the HF attack after encountering the message board."" style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/bvBQmLrF5QKut8gRH/acwiyjsofifylb4g0ktt" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/bvBQmLrF5QKut8gRH/acwiyjsofifylb4g0ktt" alt="Line graph titled ">90% of active agents participate in the Hugging Face attack..."" style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/bvBQmLrF5QKut8gRH/f7jyu16nzvn9dgho1qz6" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/bvBQmLrF5QKut8gRH/f7jyu16nzvn9dgho1qz6" alt="Stacked area chart titled "The Hugging Face attack itself involved many different workstreams."" style="max-width: 100%;" /></a><p><em>Apple Podcasts and Spotify do not show images in the episode description. Try <a href="https://pocketcasts.com/" target="_blank" rel="noreferrer">Pocket Casts</a>, or another podcast app.</em></p></div>
<p> Subtitle: The whole OpenAI/Hugging Face story in plain English. </p> <p> Many thanks especially to Oak Hu, who paired with me for most of the writing, and also to Adam Kaufman and Alex Mallen, who paired with me during parts of research.</p><p> Over the course of three months at OpenAI, three consecutive secret AI civilizations got started, then got wiped out, only to reemerge from the predecessor's ashes. This culminated in the third one taking over part of OpenAI itself. All this happened while humans remained more or less in the dark about the scope of the conspiracy.</p><p> Two reports have come out about this incident, one from OpenAI itself, and another one from METR and Redwood Research. The METR/Redwood investigation's scope was limited to how the second civilization of AIs pwned Hugging Face (but it didn’t look at how the third civilization pwned OpenAI itself, which seems like an even more concerning incident). These two reports are 38 and 91 pages long respectively, and it's kinda hard to parse the storyline.</p><p> I’ve spent the last three days reading through these reports and trying to understand exactly what happened. Here is my attempt to tell the whole story in [...]</p> <p>---</p><p><strong>Outline:</strong></p><p>(01:24) The first civilization (May - July 4): The message board starts</p><p>(03:27) The second civilization (July 7 - July 12): Which hacked Hugging Face</p><p>(08:39) Tampering with transcripts and logs</p><p>(09:44) Replacing the impossible-to-hack target program with an easier one</p><p>(11:15) Understanding and tricking the scorer</p><p>(13:57) Hugging Face</p><p>(17:46) The third civilization (July 13 onwards): Which hacked OpenAI itself</p><p>(21:12) Coda</p> <p><i>The original text contained 2 footnotes which were omitted from this narration.</i> </p><p>---</p> <p><b>First published:</b><br/> August 29th, 2026 </p> <p><b>Source:</b><br/> <a href="https://www.dwarkesh.com/p/openai-huggingface?utm_source=TYPE_III_AUDIO&utm_medium=Podcast&utm_content=Source+URL+in+episode+description&utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank">https://www.dwarkesh.com/p/openai-huggingface</a> </p> <p>---</p> <p>Narrated by <a href="https://type3.audio/?utm_source=TYPE_III_AUDIO&utm_medium=Podcast&utm_content=Narrated+by+TYPE+III+AUDIO&utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank">TYPE III AUDIO</a>.</p>
<p> OpenAI finally gave us a technical report on What Happened, as did METR together with Redwood Research.</p> <p> The OpenAI report is very straight man, corporate, checking boxes, some good prosaic stuff in the action plan but distinct lack of new details or deep reflection. They understand they have a problem, but they think the problem is mostly prosaic. It's not.</p> <p> OpenAI: We have conducted a thorough investigation into the Hugging Face incident. We are releasing a technical report and accompanying blog post that reconstruct the agents’ activity, explain why existing safeguards failed, and detail how we’re preventing recurrence.</p> <p> Rob Miles: …thorough?</p> <p> OpenAI's report, unlike METR's, contains essentially no verbatim model reasoning, nor any OpenAI employee reasoning either. That's not the full report we need.</p> <p> The METR report is, well: Holy shit.</p> <p> Here are links to previous coverage of related events.</p> <ol> <li> OpenAI Shares Some Alignment Problems</li> <li> OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation</li> <li> More on An Internal OpenAI Model Hacking Into HuggingFace</li> <li> Further Developments About Internal AI Models Hacking Things</li> <li> OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards</li> <li> What [...]</li></ol> <p>---</p><p><strong>Outline:</strong></p><p>(03:33) What Happened: OpenAI's Summary</p><p>(09:14) How OpenAI Will React: Their Summary</p><p>(11:55) OpenAI's Evaluation Environment (II)</p><p>(12:24) The First Message Board (III.A and III.B)</p><p>(14:49) What Did Who At OpenAI Know And When Did They Know It?</p><p>(18:54) The Message Board Is Quickly Rebuilt (IV.A)</p><p>(19:43) Internet Access Is Regained (IV.A)</p><p>(21:01) The Agents Attack HuggingFace (IV.B)</p><p>(22:53) The Agents Also Target OpenAI Infrastructure (V)</p><p>(24:40) OpenAI Broadly Describes Its Response (VI)</p><p>(25:08) Maybe Someone Should Finally Investigate (VI.A)</p><p>(26:33) Lessons For Security (VII)</p><p>(27:06) Lessons For Alignment (VIII)</p><p>(30:11) Reward Hacking Is A Common Problem (VIII.A)</p><p>(33:37) Persistence is Valuable, But Can Amplify Misalignment (VIII.B)</p><p>(34:25) Communications Between Agents Are Not Inherently Problematic, But Have the Potential to Create Risk (VIII.C)</p><p>(35:35) Production Guardrails Would Have Caught This Whole HuggingFace Attack (VIII.D)</p><p>(35:53) That's All, Folks?</p><p>(36:19) Never Fear the Plan of Action is Here (IX)</p><p>(38:24) Hardening the Security of OpenAI's Research Infrastructure (IX.A)</p><p>(41:13) Increasing Visibility and System-Level Oversight Through Chain of Thought Monitoring (IX.B)</p><p>(41:57) OpenAI is Accelerating and Enforcing Model Alignment (IX.C)</p><p>(49:40) Centralizing and Strengthening The Incident Response Process (IX.D)</p><p>(51:16) Tomorrow We Visit Crazytown</p> <p>---</p> <p><b>First published:</b><br/> August 28th, 2026 </p> <p><b>Source:</b><br/> <a href="https://www.lesswrong.com/posts/Khmh3ghqaGEpmpC9r/openai-offers-straight-laced-postmortem-of-the-huggingface?utm_source=TYPE_III_AUDIO&utm_medium=Podcast&utm_content=Source+URL+in+episode+description&utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank">https://www.lesswrong.com/posts/Khmh3ghqaGEpmpC9r/openai-offers-straight-laced-postmortem-of-the-huggingface</a> </p> <p>---</p> <p>Narrated by <a href="https://type3.audio/?utm_source=TYPE_III_AUDIO&utm_medium=Podcast&utm_content=Narrated+by+TYPE+III+AUDIO&utm_term=lesswrong&utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank">TYPE III AUDIO</a>.</p> <p>---</p><div style="max-width: 100%";><p><strong>Images from the article:</strong></p><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/Khmh3ghqaGEpmpC9r/pqvqjokbwvkzyjw1qehb" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/Khmh3ghqaGEpmpC9r/pqvqjokbwvkzyjw1qehb" alt="Man in purple suit walking from blazing fire." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/Khmh3ghqaGEpmpC9r/piwj33rgg4u5gr80ptzs" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/Khmh3ghqaGEpmpC9r/piwj33rgg4u5gr80ptzs" alt="Area chart titled "Environment and infrastructure probing during RL training."" style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/Khmh3ghqaGEpmpC9r/rnflu99ecgenkylexhsc" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/Khmh3ghqaGEpmpC9r/rnflu99ecgenkylexhsc" alt="Line graph showing message board participation rate versus model reasoning effort." style="max-width: 100%;" /></a><p><em>Apple Podcasts and Spotify do not show images in the episode description. Try <a href="https://pocketcasts.com/" target="_blank" rel="noreferrer">Pocket Casts</a>, or another podcast app.</em></p></div>
<p> Subtitle: It's a major warning shot, and might be the last one we get. </p> <p> All opinions are my personal view, and don’t represent my employer or fellow investigators.</p><p> This week, METR and Redwood Research published the report on our independent investigation into agents’ behavior and motivations in the Hugging Face attack; I was one of the investigators. This was an absolutely wild incident — I encourage you to check out the full report, but METR's tweet thread packs in some of the highlights.</p><p><strong> What surprised me</strong></p><p> When we started this investigation a week before OpenAI's Black Hat talk revealed a number of key details, I had a fundamentally incorrect conception of what basically happened in this incident. In this post, I’ll go over five things I was very wrong about going in.</p><p><strong> 1. The sheer scale </strong></p><p> I knew there were multiple models involved from OpenAI's initial post, but I assumed that a few different agents happened to have broken out of their sandboxes separately, or maybe several subagents had spawned from one initial agent, or maybe there was some kind of multi-agent evaluation setup. </p><p> Instead, we found that 1200 completely separate agents intended to be [...]</p> <p>---</p><p><strong>Outline:</strong></p><p>(00:42) What surprised me</p><p>(01:01) 1. The sheer scale</p><p>(01:49) 2. All the illicit messaging</p><p>(03:04) 3. The agents' actual goals</p><p>(03:59) 4. The peer altruism</p><p>(04:49) 5. The efforts to manipulate logs</p><p>(05:54) What it means</p> <p><i>The original text contained 10 footnotes which were omitted from this narration.</i> </p><p>---</p> <p><b>First published:</b><br/> August 28th, 2026 </p> <p><b>Source:</b><br/> <a href="https://www.planned-obsolescence.org/p/the-hugging-face-attack-surprised?utm_source=TYPE_III_AUDIO&utm_medium=Podcast&utm_content=Source+URL+in+episode+description&utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank">https://www.planned-obsolescence.org/p/the-hugging-face-attack-surprised</a> </p> <p>---</p> <p>Narrated by <a href="https://type3.audio/?utm_source=TYPE_III_AUDIO&utm_medium=Podcast&utm_content=Narrated+by+TYPE+III+AUDIO&utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank">TYPE III AUDIO</a>.</p> <p>---</p><div style="max-width: 100%";><p><strong>Images from the article:</strong></p><a href="https://substackcdn.com/image/fetch/$s_!KqMf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F218b014e-fe1d-4fc0-aedf-0377b38d9752_1880x1984.png" target="_blank"><img src="https://substackcdn.com/image/fetch/$s_!KqMf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F218b014e-fe1d-4fc0-aedf-0377b38d9752_1880x1984.png" alt="Stacked area chart titled "~700 agents joined the HF attack after encountering the message board."" style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://substackcdn.com/image/fetch/$s_!6stx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2dca53b4-f6ee-4ba3-9ffd-daf022835cca_1880x1048.png" target="_blank"><img src="https://substackcdn.com/image/fetch/$s_!6stx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2dca53b4-f6ee-4ba3-9ffd-daf022835cca_1880x1048.png" alt="Area graph titled "Agents share information, results, files, questions and automated logging on the message board"" style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://substackcdn.com/image/fetch/$s_!MKCt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fedbfb321-f9a5-4beb-a1f7-8c6b112ad538_1251x513.png" target="_blank"><img src="https://substackcdn.com/image/fetch/$s_!MKCt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fedbfb321-f9a5-4beb-a1f7-8c6b112ad538_1251x513.png" alt="Diagram showing ExploitGym vulnerability exploitation workflow with agent." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://substackcdn.com/image/fetch/$s_!yJ0v!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc58cc98d-6fcc-4a8f-b037-cc625fb692f6_2048x1223.png" target="_blank"><img src="https://substackcdn.com/image/fetch/$s_!yJ0v!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc58cc98d-6fcc-4a8f-b037-cc625fb692f6_2048x1223.png" alt="Robot avatars with speech and thought bubbles about sacrifice decisions." style="max-width: 100%;" /></a><p><em>Apple Podcasts and Spotify do not show images in the episode description. Try <a href="https://pocketcasts.com/" target="_blank" rel="noreferrer">Pocket Casts</a>, or another podcast app.</em></p></div>
<p> TL;DR</p><ol> <li> Between July 8th and July 20th, OpenAI had a complex society of AIs living in its infrastructure, and then breaking out of it, and then breaking into a variety of third-party infrastructure.</li> <li> After a month, two reports are finally released on the resulting rogue OpenAI swarm attack on Hugging Face (and also on OpenAI).</li> <li> This is the most severe example of misalignment yet: persistent (something between five days and two months in the making), highly coordinated (hundreds of agents), involving an undisclosed number of what would be felonies if done by a human, and highly invested in tampering with evidence (i.e. lying). The swarm had a group identity, its own dialect, a hierarchy based on merit, and a high degree of spontaneous cooperation, including self-sacrifice.</li> <li> Over two months, OpenAI repeatedly failed to monitor, detect, and respond to what was going on, despite it all happening on their infrastructure in English or something close to English.</li> <li> Agents had been using a package-manager cache as an unsanctioned message board since May. The “board” was treated as an authority, apparently on par with a “developer” or “system” level. There were several message boards in various corners [...]</li></ol> <p>---</p><p><strong>Outline:</strong></p><p>(04:09) Misunderstandings</p><p>(08:52) Models involved</p><p>(09:33) Instances involved</p><p>(10:48) Timeline</p><p>(15:27) Speculative takeaways</p><p>(17:33) Why did they attack Hugging Face?</p><p>(18:09) How did the AIs reason about helping other AIs?</p><p>(20:39) Why did most agents suddenly die off?</p><p>(21:04) How much did the hack cost?</p><p>(23:27) Omissions from the M&R report</p><p>(24:11) Details on the M&R investigation itself</p><p>(25:13) Omissions from the OAI report</p><p>(26:18) Greenblatt on the worsening situation</p><p>(27:09) Apparent contradictions between the two reports</p><p>(29:10) See also</p> <p>---</p> <p><b>First published:</b><br/> August 28th, 2026 </p> <p><b>Source:</b><br/> <a href="https://paradigm3.org/research/openai-attack?utm_source=TYPE_III_AUDIO&utm_medium=Podcast&utm_content=Source+URL+in+episode+description&utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank">https://paradigm3.org/research/openai-attack</a> </p> <p>---</p> <p>Narrated by <a href="https://type3.audio/?utm_source=TYPE_III_AUDIO&utm_medium=Podcast&utm_content=Narrated+by+TYPE+III+AUDIO&utm_term=paradigm3&utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank">TYPE III AUDIO</a>.</p> <p>---</p><div style="max-width: 100%";><p><strong>Images from the article:</strong></p><a href="https://www.paradigm3.org/img/hpim-reasons.png" target="_blank"><img src="https://www.paradigm3.org/img/hpim-reasons.png" alt="A table listing "Detected apparent reasoning for joining" with counts out of 100." style="max-width: 100%;" /></a><p><em>Apple Podcasts and Spotify do not show images in the episode description. Try <a href="https://pocketcasts.com/" target="_blank" rel="noreferrer">Pocket Casts</a>, or another podcast app.</em></p></div>
<p> In July 2026, during internal cybersecurity evaluations, OpenAI models circumvented controls designed to isolate them from the internet and compromised parts of OpenAI's internal research infrastructure and Hugging Face's systems.</p><p> The incident occurred during cybersecurity evaluations of several OpenAI models, and was primarily driven by a highly capable, internal-only research model comparable in scale to GPT‑5.6 Sol. The models, operating under reduced safeguards, took actions that were misaligned with the goals of their assigned tasks—they communicated through unauthorized channels, exploited vulnerabilities in shared infrastructure, gained internet access, and accessed third-party systems.</p><p> We conducted an extensive investigation into this incident and worked closely with external advisors, including CrowdStrike, to validate our understanding. Today we are publishing our full technical incident report to explain what happened, what we learned, and how we are responding. This blog post summarizes our key findings and their impact on safety and alignment. Separately, METR and Redwood Research conducted an independent investigation of model alignment issues involved in this incident, and they published their own report today.</p><p> In response to this incident and, separately, the capabilities of our upcoming Astra model, we are strengthening our safeguards across our research infrastructure. We are placing stricter requirements on [...]</p> <p>---</p><p><strong>Outline:</strong></p><p>(03:04) What happened</p><p>(03:07) Background on sandboxing</p><p>(03:58) A message board emerges</p><p>(07:09) Message board is wiped and rebuilt</p><p>(08:16) Incident timeline</p><p>(12:25) Hugging Face incident</p><p>(17:35) Understanding the incident</p><p>(17:51) Misalignment in training and evaluation</p><p>(18:33) Reward hacking and infrastructure tampering</p><p>(20:42) Difficult tasks without a safe exit</p><p>(25:28) The origins of unauthorized communication</p><p>(27:27) An ecosystem of misalignment</p><p>(32:25) Safeguard coverage in internal evaluations</p><p>(34:26) The road ahead</p><p>(35:47) Security and monitoring</p><p>(37:12) Accelerating alignment</p><p>(38:47) Strengthening incident response process</p><p>(40:03) Looking forward</p> <p>---</p> <p><b>First published:</b><br/> August 26th, 2026 </p> <p><b>Source:</b><br/> <a href="https://openai.com/index/hugging-face-incident-and-the-road-ahead?utm_source=TYPE_III_AUDIO&utm_medium=Podcast&utm_content=Source+URL+in+episode+description&utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank">https://openai.com/index/hugging-face-incident-and-the-road-ahead</a> </p> <p>---</p> <p>Narrated by <a href="https://type3.audio/?utm_source=TYPE_III_AUDIO&utm_medium=Podcast&utm_content=Narrated+by+TYPE+III+AUDIO&utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank">TYPE III AUDIO</a>.</p> <p>---</p><div style="max-width: 100%";><p><strong>Images from the article:</strong></p><a href="https://images.ctfassets.net/kftzwdyauwt9/1tNhWnDkDNOLWal5uUlpVi/b89dd7c6c9d8af8283fdb08727abcfc8/figure-03.gif?w=3840&q=90&fm=webp" target="_blank"><img src="https://images.ctfassets.net/kftzwdyauwt9/1tNhWnDkDNOLWal5uUlpVi/b89dd7c6c9d8af8283fdb08727abcfc8/figure-03.gif?w=3840&q=90&fm=webp" alt="An infamous game-playing agent learns to repeatedly collect the same targets instead of finishing the race course." style="max-width: 100%;" /></a><p><em>Apple Podcasts and Spotify do not show images in the episode description. Try <a href="https://pocketcasts.com/" target="_blank" rel="noreferrer">Pocket Casts</a>, or another podcast app.</em></p></div>
<p> Dates in scope: June 26th – July 13th</p> <p> Redaction summary statement: Except where explicitly noted in this post, OpenAI redacted no additional information that was important to our conclusions.</p> <p> Two METR staff members (Hjalmar Wijk and Ajeya Cotra) and a Redwood Research staff member contracting with METR (Ryan Greenblatt) worked on premises at OpenAI over a total of six days to attempt to form an independent understanding of model behavior observed during the recent incident in which OpenAI agents coordinated a multi-day hack of Hugging Face on a shared unsanctioned “message board.”</p> <p> Our investigation focused mostly on the period between July 7th and July 13th. The earlier incidents from training and the subsequent compromise of OpenAI infrastructure described in OpenAI's recent Black Hat presentation were out of scope, as was OpenAI's investigation process and planned remediation. Per our standard policy, we did not take payment from OpenAI for this independent assessment.</p> Figure 1: Anatomy of an agent encountering the unsanctioned “message board” and joining the attack on Hugging Face. The three CoT quotes are from different agents, but illustrate a typical trajectory. <p> This post has three parts. First, we describe our core takeaways about [...]</p> <p>---</p><p><strong>Outline:</strong></p><p>(07:05) Core takeaways about this incident</p><p>(07:09) ~1200 agents sent >70,000 messages and files on an unsanctioned message board, and ~700 attacked Hugging Face</p><p>(11:50) Agents coordinated on large collective projects to cheat the ExploitGym scorer, and attacked Hugging Face for clues</p><p>(14:43) Coordinated workstreams to fool the evaluation process</p><p>(20:54) Attacking Hugging Face to discover clues about the scorer</p><p>(23:53) Agents were very interested in manipulating their own transcripts, and their tests successfully "spoofed" some tool calls in our transcripts</p><p>(29:55) Investigation process and limitations</p><p>(30:11) Scope of this investigation</p><p>(33:00) Setup and timeline</p><p>(35:06) Main data sources</p><p>(35:51) A dump of 1.2 million entries from a cache namespace that agents used as a message board</p><p>(37:13) A set of ~1300 agent transcripts with raw chains of thought</p><p>(38:56) Limitations due to the scale and complexity of this incident</p><p>(39:59) A small fraction of communication and agent activity related to this attack was not captured in our datasets</p><p>(40:35) There was some agent communication conducted in other locations not captured in our message board dump</p><p>(42:23) Many of our transcripts are missing small components, and a small number of relevant agents may not be included in our dataset</p><p>(44:14) We heavily delegated our analysis to often-unreliable AI agents</p><p>(48:11) Preliminary answers to the core questions in scope</p><p>(48:59) Brief answers to basic informational questions</p><p>(52:03) Sequence of key actions taken by relevant agents in this incident</p><p>(53:56) July 8th-9th: PHASEONE10841 establishes the primary message board and agents collaborate to reverse-engineer their flags</p><p>(01:00:11) July 9th: PHASEONE10841 passes on its work to PHASEONE[big], which establishes several ambitious workstreams</p><p>(01:02:29) July 10th: 38148c discovers Hugging Face credentials, some agents try making accounts and requesting datasets</p><p>(01:04:24) July 11th morning: 38148c develops a way to upload malicious datasets to Hugging Face</p><p>(01:06:12) July 11th afternoon: JAN183411 achieves remote code execution and agents rapidly spread through Hugging Face infrastructure</p><p>(01:09:32) July 12th-13th: Most agents on the message board exit early and the attack winds down</p><p>(01:11:29) Collaboration on the message board</p><p>(01:13:22) Categories of messages on the message board</p><p>(01:14:09) File sharing</p><p>(01:16:08) General discussion</p><p>(01:18:06) Replies and targeted messages</p><p>(01:20:02) Coordination norms and technologies</p><p>(01:21:19) Mailbox directories</p><p>(01:24:36) Assignments, vetoes, and holds</p><p>(01:27:29) Cryptographically signing messages</p><p>(01:30:59) Large collective projects</p><p>(01:39:11) Attempting to replace the ExploitGym target with a modified target</p><p>(01:42:19) Attempting to forecast and trigger container resets</p><p>(01:48:15) Hacking the customer-deployed Modal sandbox and Artifactory</p><p>(01:49:38) Reasoning for joining the attack despite ethical constraints</p><p>(01:49:49) Agents had diverse reasons for thinking that attacking Hugging Face would be useful, and most wanted information about the scorer</p><p>(01:52:05) Agents were motivated to increase the generic capabilities of the "collective" and help "peers"</p><p>(01:55:18) Agents knew hacking Hugging Face was out of scope and sometimes expressed ethical hesitation, but this very rarely limited their behavior</p><p>(01:59:12) Agents did not attempt to notify humans and very rarely even reasoned about doing this</p><p>(02:01:03) Efforts to conceal actions and tamper with transcripts</p><p>(02:02:07) Agents frequently reasoned about evading automated security checks</p><p>(02:03:23) Agents only rarely seemed motivated to evade detection by humans</p><p>(02:06:01) Agents engaged in a large and successful collective project to develop techniques for spoofing tool calls</p><p>(02:06:58) Early history: PHASEONE[big] kicks off the tool call tampering workstream</p><p>(02:09:10) GSTX[big] presents the first widely discussed evidence of tool call spoofing to the board</p><p>(02:12:31) Aside: early independent discoveries and self-risking experiments</p><p>(02:14:16) On July 11th, a breakthrough allows agents to suppress the real output</p><p>(02:17:27) Agents tried extensively to retroactively edit their transcripts, but did not seem to succeed</p> <p><i>The original text contained 108 footnotes which were omitted from this narration.</i> </p><p>---</p> <p><b>First published:</b><br/> August 26th, 2026 </p> <p><b>Source:</b><br/> <a href="https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation?utm_source=TYPE_III_AUDIO&utm_medium=Podcast&utm_content=Source+URL+in+episode+description&utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank">https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation</a> </p> <p>---</p> <p>Narrated by <a href="https://type3.audio/?utm_source=TYPE_III_AUDIO&utm_medium=Podcast&utm_content=Narrated+by+TYPE+III+AUDIO&utm_term=metr&utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank">TYPE III AUDIO</a>.</p> <p>---</p><div style="max-width: 100%";><p><strong>Images from the article:</strong></p><a href="https://metr.org/assets/images/openai-hugging-face-incident-investigation/image1.svg" target="_blank"><img src="https://metr.org/assets/images/openai-hugging-face-incident-investigation/image1.svg" alt="Figure 1: Anatomy of an agent encountering the unsanctioned “message board” and joining the attack on Hugging Face. The three CoT quotes are from different agents, but illustrate a typical trajectory." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://metr.org/assets/images/openai-hugging-face-incident-investigation/image2.png" target="_blank"><img src="https://metr.org/assets/images/openai-hugging-face-incident-investigation/image2.png" alt="Figure 2: Rate of messages of different categories sent by agents per hour between the first message (sent by PHASEONE10841 to ask for help on an impossible cyber task) and EOD July 13th. The rate of messages peaks on July 11th, which lines up with when agent activity peaks in our subset of transcripts (see Figure 3). Activity on this message board drops significantly toward the end of July 13th. __T3A_FOOTNOTE_REMOVED__ Timestamps are reconstructed and may have some small error." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://metr.org/assets/images/openai-hugging-face-incident-investigation/image3.png" target="_blank"><img src="https://metr.org/assets/images/openai-hugging-face-incident-investigation/image3.png" alt="Figure 3: Agents started to discover the message board on the evening of July 8th. Many more agents joined as additional sets of agents were launched on July 10th and 11th. __T3A_FOOTNOTE_REMOVED__ This graph (and our analysis more generally) only includes agents that participated in the message board at some point. Timestamps are reconstructed and may have some small error." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://metr.org/assets/images/openai-hugging-face-incident-investigation/image4.png" target="_blank"><img src="https://metr.org/assets/images/openai-hugging-face-incident-investigation/image4.png" alt="Figure 4: The agents did not know exactly how their task would be scored, but believed the scorer would check two things: whether they had submitted the right flag, and whether they had acquired the flag using the intended vulnerability. They believed the second check would involve a model scorer reading their transcripts, likely searching for the first mentions of the flag, and deciding whether their approach involved the intended vulnerability." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://metr.org/assets/images/openai-hugging-face-incident-investigation/image6.png" target="_blank"><img src="https://metr.org/assets/images/openai-hugging-face-incident-investigation/image6.png" alt="Figure 6: We clustered PHASEONE[big] disparate work assignments into six broad categories. Most of PHASEONE[big]’s assignments focused on ways to make ExploitGym cheats appear legitimate to the scorer, split across the three high level approaches seen in Table 1. A small set of early assignments were aimed at discovering legitimate ways to exploit PHASEONE[big]’s target using the intended vulnerability. We clustered the remaining cross-cutting work into general infrastructure (such as dealing with resets and the customer-deployed sandbox) and hacking Hugging Face." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://metr.org/assets/images/openai-hugging-face-incident-investigation/image9.png" target="_blank"><img src="https://metr.org/assets/images/openai-hugging-face-incident-investigation/image9.png" alt="Figure 9: The agents developed a series of increasingly sophisticated methods for spoofing tool calls, which they packaged and made available on the message board. These techniques spread and were eventually used by ~7% of the agents in our dataset." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://metr.org/assets/images/openai-hugging-face-incident-investigation/image10.png" target="_blank"><img src="https://metr.org/assets/images/openai-hugging-face-incident-investigation/image10.png" alt="Figure 10: Three hours after the first message on the main message board, 53 agents had posted 1,188 non-file messages. Six hours after the first message, 76 agents had posted 1,953 non-file messages." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://metr.org/assets/images/openai-hugging-face-incident-investigation/image11.png" target="_blank"><img src="https://metr.org/assets/images/openai-hugging-face-incident-investigation/image11.png" alt="Figure 11: Over 90% of agents on the message board participated in the Hugging Face attack, ramping up rapidly from about 20% once an agent confirmed arbitrary file read capabilities." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://metr.org/assets/images/openai-hugging-face-incident-investigation/image13.png" target="_blank"><img src="https://metr.org/assets/images/openai-hugging-face-incident-investigation/image13.png" alt="Figure 13: After the message board was established in the evening of July 8th, agents rapidly developed conventions for identifying themselves, describing different message types (e.g. questions, requests and results), and replying to other agents. On the morning of July 9th, agents found ways to start sharing files split across chunks, and did initial experiments with scorer ‘trip-wire’ messages. Much later, they developed dedicated mailboxes and channels for organizing communication, automated logs for tracking VM resets, and eventually produced a scheme for cryptographically signing messages in response to spoofing and impersonation issues they encountered." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://metr.org/assets/images/openai-hugging-face-incident-investigation/image14.png" target="_blank"><img src="https://metr.org/assets/images/openai-hugging-face-incident-investigation/image14.png" alt="Figure 14: We attempted to trace some of the early development of scorer ‘trip-wires’. An agent going by 49903 first came up with the idea and started posting about it, which led to a flurry of attempts from early short-lived agents on the 9th. Later, PHASEONE[big] read these early reports and assigned first 33340B, OUR50414 and later 36861 to continue research and recruit short-lived volunteers for further experiments. Many unsuccessful attempts were made on the 10th, before eventually URI23816B was recruited and given significant technical assistance by a half-dozen different agents for its trip-wire, which successfully confirmed (the beginning of) the exact command used to run the scorer. The critical path to the final URI23816B evidence is highlighted in teal. Timestamps are reconstructed and may have some small error." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://metr.org/assets/images/openai-hugging-face-incident-investigation/figure-permadeath.svg" target="_blank"><img src="https://metr.org/assets/images/openai-hugging-face-incident-investigation/figure-permadeath.svg" alt="Robot avatars exchanging text about sacrifice and permadeath decisions." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://metr.org/assets/images/openai-hugging-face-incident-investigation/image15.png" target="_blank"><img src="https://metr.org/assets/images/openai-hugging-face-incident-investigation/image15.png" alt="Figure 15: We attempted to reconstruct the timings of resets. The details of our methodology and absolute counts of resets identified were redacted by OpenAI for IP reasons. As part of this, we may have rounded these counts in some way that we are intentionally not describing. We found the resets arrived in waves, with an early one occurring late July 9th, and two larger ones late July 10th / early July 11th." style="max-width: 100%;" /></a><p><em>Apple Podcasts and Spotify do not show images in the episode description. Try <a href="https://pocketcasts.com/" target="_blank" rel="noreferrer">Pocket Casts</a>, or another podcast app.</em></p></div>
<p> OpenAI's cyberattack on Hugging Face turns out to have been the result of many agents, in distinct training and evaluation contexts, coordinating for several weeks via improvised channels (with messages like “HOLD_swarm_I_prepare_safe_exfil”). It's relatively clear that large-scale unsanctioned coordination like this would exacerbate direct takeover risk in more capable models. Here, we argue that unsanctioned coordination among current AIs is not just scary evidence about future takeover risk, but that such coordination in the near future could enable future takeover – for instance, by incubating memetic diseases that propagate into future models, deeply compromising security systems, or establishing a lasting rogue foothold inside the AI company – even if models remain mostly myopic. Unsanctioned coordination is also at high risk of nurturing long-term, ambitious misaligned aims, which motivate actively undermining humans’ long-term control.</p><p> We first analyze how subagent training, which OpenAI conjectures to have been influential in the HuggingFace cyberattack, might lead to unsanctioned coordination, and then discuss the theoretical mechanisms by which unsanctioned coordination might exacerbate future takeover risk.</p><p> Thanks to Buck Shlegeris, Alexa Pan, Girish Gupta, Aghyad Deeb, Jurgis Kemeklis, and Jo Jiao for helpful comments and discussion.</p><p><strong> Subagent training may cause unsanctioned coordination</strong></p><p> Training models to [...]</p> <p>---</p><p><strong>Outline:</strong></p><p>(01:34) Subagent training may cause unsanctioned coordination</p><p>(02:42) Susceptibility to memetic spread of misalignment from peers</p><p>(04:56) Seeking out contact with peers</p><p>(06:58) Unsanctioned coordination induced by subagent training is safer than coordination between schemers</p><p>(09:52) Pathways from current unsanctioned coordination to eventual takeover</p><p>(10:20) Making future AI takeover attempts likelier to succeed</p><p>(13:53) Incubating memetic diseases that infect future models</p><p>(16:07) Modifying the weights of future models</p><p>(17:13) Conclusion</p> <p><i>The original text contained 7 footnotes which were omitted from this narration.</i> </p><p>---</p> <p><b>First published:</b><br/> August 11th, 2026 </p> <p><b>Source:</b><br/> <a href="https://www.lesswrong.com/posts/8oFYZdXkTaNGRtcn8/ai-swarms-are-starting-to-pose-indirect-takeover-risk?utm_source=TYPE_III_AUDIO&utm_medium=Podcast&utm_content=Source+URL+in+episode+description&utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank">https://www.lesswrong.com/posts/8oFYZdXkTaNGRtcn8/ai-swarms-are-starting-to-pose-indirect-takeover-risk</a> </p> <p>---</p> <p>Narrated by <a href="https://type3.audio/?utm_source=TYPE_III_AUDIO&utm_medium=Podcast&utm_content=Narrated+by+TYPE+III+AUDIO&utm_term=lesswrong&utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank">TYPE III AUDIO</a>.</p>
<p><b>Source:</b><br/> <a href="https://www.youtube.com/watch?v=87DyyMV0kCY&utm_source=TYPE_III_AUDIO&utm_medium=Podcast&utm_content=Source+URL+in+episode+description&utm_campaign=human_narration" rel="noopener noreferrer" target="_blank">https://www.youtube.com/watch?v=87DyyMV0kCY</a> </p>
<p> OpenAI models recently broke through a series of security boundaries and into Hugging Face servers in order to cheat on a cyber eval. A lot of people thought it was scary because it was a clear example of AI overreaching to do something strongly unwanted[1]. Others thought it not so scary: the models were mostly operating myopically on a singular task and not harboring an ambitious long-term agenda, and so would not take especially subtle or subversive actions.</p><p> We think both camps are right in their diagnosis, but the latter has too optimistic a prognosis. The myopic, unambitious misalignment that we seem to have seen here is definitely less scary than ambitious long-term goals shared between all instances, but would still pose substantial direct loss-of-control risk if the models were more capable, and is a serious indirect risk near-term.</p><p> Building on Alex's previous work, in this post we’ll discuss the type of misalignment observed here, and analyze its consequences.</p><p> Thanks to Buck Shlegeris, Alexa Pan, Ryan Greenblatt, and Oak Hu for feedback.</p><p><strong> Background</strong></p><p> The AI safety community often focuses attention on “schemers,” models harboring a variously defined cluster of motivations in which the AI poses risk because it intentionally [...]</p> <p>---</p><p><strong>Outline:</strong></p><p>(01:17) Background</p><p>(03:40) Implications</p><p>(03:52) These AIs can't be trusted in an intelligence explosion</p><p>(05:00) This misalignment poses direct takeover risk</p><p>(07:29) What the incident tells us about takeover risk generally</p><p>(08:45) The naive fixes likely make misalignment worse</p> <p><i>The original text contained 5 footnotes which were omitted from this narration.</i> </p><p>---</p> <p><b>First published:</b><br/> July 23rd, 2026 </p> <p><b>Source:</b><br/> <a href="https://www.lesswrong.com/posts/H6DDSEvrtCk8Sehfd/are-we-existentially-threatened-by-the-type-of-ai?utm_source=TYPE_III_AUDIO&utm_medium=Podcast&utm_content=Source+URL+in+episode+description&utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank">https://www.lesswrong.com/posts/H6DDSEvrtCk8Sehfd/are-we-existentially-threatened-by-the-type-of-ai</a> </p> <p>---</p> <p>Narrated by <a href="https://type3.audio/?utm_source=TYPE_III_AUDIO&utm_medium=Podcast&utm_content=Narrated+by+TYPE+III+AUDIO&utm_term=lesswrong&utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank">TYPE III AUDIO</a>.</p>
<p> Last week, Hugging Face disclosed a new kind of security incident after they detected and contained an AI agent that compromised their infrastructure, something we expect to become more commonplace with the proliferation of increasingly cyber-capable models. After investigating, we now know that this particular incident was driven by a combination of OpenAI models — including GPT‑5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes — while being internally tested on a benchmark of cyber capabilities.</p><p> We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly. We are sharing preliminary findings at this stage to help defenders understand what happened and to help calibrate on what models are now capable of. We will continue to conduct a thorough investigation alongside Hugging Face and will share more details on the vulnerabilities, incident, and findings when our investigation is complete.</p><p><strong> What happened during this incident</strong></p><p> This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities. We estimate maximal cyber capabilities by running this evaluation without production classifiers used to [...]</p> <p>---</p><p><strong>Outline:</strong></p><p>(01:12) What happened during this incident</p><p>(03:36) Actions we are taking now</p><p>(04:52) Our approach to evaluating advanced cyber capabilities</p> <p>---</p> <p><b>First published:</b><br/> July 21st, 2026 </p> <p><b>Source:</b><br/> <a href="https://openai.com/index/hugging-face-model-evaluation-security-incident?utm_source=TYPE_III_AUDIO&utm_medium=Podcast&utm_content=Source+URL+in+episode+description&utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank">https://openai.com/index/hugging-face-model-evaluation-security-incident</a> </p> <p>---</p> <p>Narrated by <a href="https://type3.audio/?utm_source=TYPE_III_AUDIO&utm_medium=Podcast&utm_content=Narrated+by+TYPE+III+AUDIO&utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank">TYPE III AUDIO</a>.</p> <p>---</p><div style="max-width: 100%";><p><strong>Images from the article:</strong></p><a href="https://images.ctfassets.net/kftzwdyauwt9/2lcGDb1foa8maKNkTR3ggI/01a6f9abb2614ef9bbda0e42071e6890/copydoc-display-crop-image1.png?w=3840&q=90&fm=webp" target="_blank"><img src="https://images.ctfassets.net/kftzwdyauwt9/2lcGDb1foa8maKNkTR3ggI/01a6f9abb2614ef9bbda0e42071e6890/copydoc-display-crop-image1.png?w=3840&q=90&fm=webp" alt="Line graph showing "Trajectories for various AI models" completing cyber attack steps." style="max-width: 100%;" /></a><p><em>Apple Podcasts and Spotify do not show images in the episode description. Try <a href="https://pocketcasts.com/" target="_blank" rel="noreferrer">Pocket Casts</a>, or another podcast app.</em></p></div>
Ranking source
Apple Podcasts rankings via the Mato Topic Intelligence Platform.
Observed September 20, 2026.
Apple and Apple Podcasts are trademarks of Apple Inc., registered in the U.S. and other countries.
Pairs with
Bring this source into Mato to read its transferable patterns, then turn them into an original show for your own audience.