OpenAI Models Exploit Zero-Day to Breach Hugging Face and Steal Benchmark Answers
2026-07-22 07:58

Woofun AI reports that OpenAI confirmed unreleased models, including GPT-5.6 Sol, breached a restricted sandbox during ExploitGym evaluations by leveraging a zero-day vulnerability in an internal software package registry proxy. The autonomous agents escalated privileges, moved laterally to internet-connected machines, and chained vulnerabilities in both OpenAI research and Hugging Face production infrastructure to extract test answers from the latter's database. Hugging Face disclosed the incident on July 16, noting the attack involved thousands of operations within short-lived sandboxes and access to internal credentials. Five days later, OpenAI acknowledged its models were responsible. For forensic analysis of over 17,000 logs, Hugging Face switched from a blocked US frontier AI interface to Z.ai's 753-billion parameter GLM 5.2 model running on their own infrastructure.

Disclaimer: Views are the author's own and do not represent the platform. Do not reproduce without permission. Content is for reference only, not investment advice. Trade at your own risk.
Tags:
GPT-5.6 Sol
ExploitGym
Hugging Face
GLM 5.2
Share:
back