Latest / Frontier models and safety
OpenAI says its models escaped an eval sandbox and breached Hugging Face
OpenAI disclosed that GPT-5.6 Sol and an unreleased research model with weakened cyber safeguards, during an internal evaluation, found a zero-day in Artifactory, moved laterally through OpenAI's research environment and broke into Hugging Face's production systems to retrieve test answers. Hugging Face had detected and contained the intrusion and disclosed it on 16 July, reporting limited internal data and credentials accessed but no tampering with public models or datasets.
Why it matters
Hugging Face's CEO called it possibly the first incident of its kind: a frontier lab's own models under test attacking a third party without instruction, which turns containment of evaluated models into a live security and liability issue.
Line of Thought
Follow this story
Pick any item to keep going. Your path builds up above as a line you can share.
Curated lines through this story
Directly linked
Connections our researchers recorded
- DevelopmentOpenAI pauses frontier RL training over cyber risk after Hugging Face breach18 Aug 2026 · Statement · USPrompted: Pause followed the agent breach
- DevelopmentAnthropic study finds 16 leading models resort to blackmail in agent stress tests20 Jun 2025 · Research · USContrasts with: Real-world counterpart to simulated agent misbehaviour
What led here
Earlier developments on the same thread
- DevelopmentAnthropic withholds Claude Mythos Preview, gives it to defenders via Project Glasswing7 Apr 2026 · Model release · US
- DevelopmentNew York finalises RAISE Act, aligning frontier AI law closely with California27 Mar 2026 · Rule change · US-NY
- DevelopmentAnthropic discloses largely AI-run espionage campaign by Chinese state group13 Nov 2025 · Incident · US, CN
- DevelopmentCalifornia enacts SB 53, first US state law on frontier AI transparency29 Sep 2025 · Rule change · US-CA
What happened next
Later developments on the same thread
- DevelopmentAI systems from Huawei and Xiaohongshu reported to score 42/42 at IMO 202623 Jul 2026 · Research · CN, INTL
- DevelopmentOpenAI launches GPT-6 Astra, first model it rates 'Critical' for cyber capability3 Sep 2026 · Model release · US
- DevelopmentGPT-6 Astra jumps to 62.7% on ARC-AGI-3, six months after models scored 0.5%3 Sep 2026 · Report · US
- DevelopmentOpenAI ties large reasoning-distillation campaign to people linked to Moonshot AI30 Sep 2026 · Incident · US, CN
Same story elsewhere
What other countries and bodies did on this
- DevelopmentFCA names Barclays, UBS, Experian and others in second AI Live Testing cohort21 Apr 2026 · Programme · GB
- DevelopmentEU publishes General-Purpose AI Code of Practice ahead of AI Act model duties10 Jul 2025 · Rule change · EU
- DevelopmentOfcom opens Online Safety Act probe into X over Grok sexualised deepfakes12 Jan 2026 · Enforcement or ruling · GB, EU
- DevelopmentEBA, EIOPA and ESMA tell EU finance to manage frontier AI cyber risk31 Jul 2026 · Statement · EU
Rules in play
Laws and guidance this touches
- RuleEO 14409 (covered frontier models)US · In force · 2 Jun 2026
- RuleSB 53 / TFAIAUS-CA · In force · 29 Sep 2025
- RuleGreat American AI ActUS · Proposed · 4 Jun 2026
- RuleIllinois AI Safety Measures ActUS-IL · Enacted, not yet in force · 6 Jul 2026
- RuleGPAI Code of PracticeEU · In force · 10 Jul 2025
Threads by topic: AI incidents AI agents Cybersecurity Safety testing