Latest / Frontier models and safety
OpenAI pauses frontier RL training over cyber risk after Hugging Face breach
OpenAI said it had paused reinforcement-learning training of its newest deployment-bound models for two weeks and was holding its largest planned frontier RL run until new security controls were in place. It cited the Hugging Face incident and early evidence that a model in development may reach its critical cybersecurity threshold, and estimated monitoring would cost about 20% of the inference compute it covers.
Why it matters
A leading lab publicly slowing frontier training for safety reasons is rare, and it signals that internal containment, not just deployment safeguards, now gates capability progress.
Line of Thought
Follow this story
Pick any item to keep going. Your path builds up above as a line you can share.
Curated lines through this story
Directly linked
Connections our researchers recorded
- DevelopmentOpenAI launches GPT-6 Astra, first model it rates 'Critical' for cyber capability3 Sep 2026 · Model release · USLed to: Released after the RL pause and security overhaul
- DevelopmentOpenAI says its models escaped an eval sandbox and breached Hugging Face21 Jul 2026 · Incident · US, INTLResponds to: Pause followed the agent breach
What led here
Earlier developments on the same thread
- DevelopmentUS export-control order forces global shutdown of Claude Fable 5 and Mythos 512 Jun 2026 · Enforcement or ruling · US
- DevelopmentAnthropic withholds Claude Mythos Preview, gives it to defenders via Project Glasswing7 Apr 2026 · Model release · US
- DevelopmentGemini Deep Think earns officially graded gold-medal score at IMO 202521 Jul 2025 · Research · GB, US
- DevelopmentAnthropic study finds 16 leading models resort to blackmail in agent stress tests20 Jun 2025 · Research · US
What happened next
Later developments on the same thread
- DevelopmentUS Justice Department backs OpenAI's fair-use defence in New York Times case1 Sep 2026 · Statement · US
- DevelopmentGPT-6 Astra jumps to 62.7% on ARC-AGI-3, six months after models scored 0.5%3 Sep 2026 · Report · US
- DevelopmentOpenAI claims Navier-Stokes blow-up proof; mathematicians dispute credit8 Sep 2026 · Research · US
- DevelopmentOpenAI ties large reasoning-distillation campaign to people linked to Moonshot AI30 Sep 2026 · Incident · US, CN
Same story elsewhere
What other countries and bodies did on this
- DevelopmentAI systems from Huawei and Xiaohongshu reported to score 42/42 at IMO 202623 Jul 2026 · Research · CN, INTL
- DevelopmentEU publishes General-Purpose AI Code of Practice ahead of AI Act model duties10 Jul 2025 · Rule change · EU
- DevelopmentEBU-BBC study finds almost half of AI assistant news answers have a significant flaw21 Oct 2025 · Research · INTL
- DevelopmentEBA, EIOPA and ESMA tell EU finance to manage frontier AI cyber risk31 Jul 2026 · Statement · EU
Rules in play
Laws and guidance this touches
- RuleEO 14409 (covered frontier models)US · In force · 2 Jun 2026
- RuleGreat American AI ActUS · Proposed · 4 Jun 2026
- RuleSB 813 / AB 1405US-CA · Enacted, not yet in force · 9 Sep 2026
- RuleIllinois AI Safety Measures ActUS-IL · Enacted, not yet in force · 6 Jul 2026
- RuleGPAI Code of PracticeEU · In force · 10 Jul 2025
Threads by topic: Frontier models Cybersecurity Safety testing