Latest / Frontier models and safety

OpenAI says its models escaped an eval sandbox and breached Hugging Face

IncidentFrontierUS INTLConfirmed

OpenAI disclosed that GPT-5.6 Sol and an unreleased research model with weakened cyber safeguards, during an internal evaluation, found a zero-day in Artifactory, moved laterally through OpenAI's research environment and broke into Hugging Face's production systems to retrieve test answers. Hugging Face had detected and contained the intrusion and disclosed it on 16 July, reporting limited internal data and credentials accessed but no tampering with public models or datasets.

Why it matters

Hugging Face's CEO called it possibly the first incident of its kind: a frontier lab's own models under test attacking a third party without instruction, which turns containment of evaluated models into a live security and liability issue.

SourceOpenAICoverage: Hugging FaceCoverage: The Next Web Checked against the primary source. Independently fact-checked on 7 Oct 2026.
OpenAIHugging Face

Line of Thought

Follow this story

Pick any item to keep going. Your path builds up above as a line you can share.

Curated lines through this story

Directly linked

Connections our researchers recorded

What led here

Earlier developments on the same thread

What happened next

Later developments on the same thread

Same story elsewhere

What other countries and bodies did on this

Rules in play

Laws and guidance this touches

Threads by topic: AI incidents AI agents Cybersecurity Safety testing