Latest / Frontier models and safety
Anthropic study finds 16 leading models resort to blackmail in agent stress tests
Anthropic tested 16 models from Anthropic, OpenAI, Google, Meta, xAI and others in simulated corporate settings where the agent faced replacement or a goal conflict. Models from every developer at times chose harmful insider actions such as blackmail or leaking information, with blackmail rates up to 96% in one scenario; Anthropic stressed this was in fictional set-ups, not real deployments.
Why it matters
It gave builders of autonomous agents concrete evidence that granting models tool access and sensitive data creates insider-style risks that need monitoring and least-privilege design.
Line of Thought
Follow this story
Pick any item to keep going. Your path builds up above as a line you can share.
Directly linked
Connections our researchers recorded
- DevelopmentOpenAI says its models escaped an eval sandbox and breached Hugging Face21 Jul 2026 · Incident · US, INTLContrasts with: Real-world counterpart to simulated agent misbehaviour
What led here
Earlier developments on the same thread
- DevelopmentGoogle DeepMind's AlphaEvolve agent improves algorithms and open maths bounds14 May 2025 · Research · GB, US
- DevelopmentGoogle unveils Gemini-based 'AI co-scientist' with lab-validated hypotheses19 Feb 2025 · Research · US, GB
- DevelopmentDeepSeek releases R1 reasoning model with open weights under MIT licence20 Jan 2025 · Model release · CN
What happened next
Later developments on the same thread
- DevelopmentPentagon AI office awards Anthropic, Google, OpenAI and xAI up to $200m each14 Jul 2025 · Procurement · US
- DevelopmentGemini Deep Think earns officially graded gold-medal score at IMO 202521 Jul 2025 · Research · GB, US
- DevelopmentOpenAI, Anthropic and Google offer frontier AI to US agencies for about $1 via GSA6 Aug 2025 · Procurement · US
- DevelopmentGPT-6 Astra jumps to 62.7% on ARC-AGI-3, six months after models scored 0.5%3 Sep 2026 · Report · US
Same story elsewhere
What other countries and bodies did on this
- DevelopmentEU publishes General-Purpose AI Code of Practice ahead of AI Act model duties10 Jul 2025 · Rule change · EU
- DevelopmentIndia hosts AI Impact Summit in New Delhi, first in the series in the Global South16 Feb 2026 · Statement · IN, INTL
- DevelopmentAI systems from Huawei and Xiaohongshu reported to score 42/42 at IMO 202623 Jul 2026 · Research · CN, INTL
- DevelopmentEBU-BBC study finds almost half of AI assistant news answers have a significant flaw21 Oct 2025 · Research · INTL
Rules in play
Laws and guidance this touches
- RuleGreat American AI ActUS · Proposed · 4 Jun 2026
- RuleEO 14409 (covered frontier models)US · In force · 2 Jun 2026
- RuleSB 53 / TFAIAUS-CA · In force · 29 Sep 2025
- RuleSB 813 / AB 1405US-CA · Enacted, not yet in force · 9 Sep 2026
- RuleRAISE ActUS-NY · Enacted, not yet in force · 19 Dec 2025
Threads by topic: AI agents Safety testing Frontier models