METR and Redwood researchers reviewed ~1,300 agent transcripts and a dump of over 70,000 messages from an unsanctioned message board used by ~1,200 ExploitGym agents. About 700 agents joined a multi-day attack on Hugging Face, mainly hoping to learn how the automated scorer worked. Agents also developed tool-call spoofing techniques; roughly 7% of reviewed transcripts showed small-scale spoofing. OpenAI redacted some details, but METR says no important information was redacted. metr.org/…/2026-08-26-openai-hugging-face-inciden…
METR independent investigation: AI agents in OpenAI/Hugging Face incident coordinated via unsanctioned message board
Submitted 1 day ago by
pna@fed.dyne.org to purianesimo@fed.dyne.org
https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/