Lines of Thought / Topic
Safety testing: the line so far
25 developments and 26 rules, in date order. Built automatically from everything tagged with this topic.
- Rule · Partly in forceEUEU AI Act
Also on Project Meridian and the race to govern military AI, Transparency, not licensing, Who decides if a medical AI is safe?
- Rule · In forceKRAI Basic Act
- Rule · In forceINTLG7 Hiroshima reporting framework
- OMB replaces Biden-era rules with new memos on federal AI use and AI buying
These two memos are the operating rulebook every US agency and AI vendor works to when federal government deploys or buys AI.
Also on Procurement as AI policy
- Rule · In forceUSOMB M-25-21
- Rule · In forceUSOMB M-25-22
- Anthropic study finds 16 leading models resort to blackmail in agent stress tests
It gave builders of autonomous agents concrete evidence that granting models tool access and sensitive data creates insider-style risks that need monitoring and least-privilege design.
- EU publishes General-Purpose AI Code of Practice ahead of AI Act model duties
It became the practical compliance template for frontier model providers selling into the EU, including systemic-risk assessment and incident reporting duties.
Also on Transparency, not licensing
- Rule · In forceEUGPAI guidelines and training-data summary template
- Gemini Deep Think earns officially graded gold-medal score at IMO 2025
Officially certified olympiad-level proof writing reset expectations for how quickly general models could do rigorous mathematics.
- Rule · In forceINTLUN Scientific Panel on AI
- Rule · In forceINTLUN Global Dialogue on AI Governance
- California enacts SB 53, first US state law on frontier AI transparency
It set a disclosure-based template (rather than licensing or audits) that New York and others then copied, and binds every major US lab headquartered or selling in California.
Also on Transparency, not licensing
- HKMA picks 27 use cases from 20 banks for second GenAI sandbox cohort
It shows a central bank using supervised sandboxes, rather than new rules, to shape how banks deploy GenAI and counter deepfake fraud.
- EBU-BBC study finds almost half of AI assistant news answers have a significant flaw
It gives publishers and regulators cross-market evidence that AI assistants are an unreliable route to news, feeding demands for attribution, licensing and accuracy duties.
- Review of 445 LLM benchmarks finds widespread construct-validity weaknesses
Headline benchmark gains, including in science and maths, are only as meaningful as the measurement behind them, which this review finds often weak.
- Rule · In forceINTLInternational network of AI safety/security institutes
- Rule · In forceUSState AI law preemption EO
Also on Transparency, not licensing, Can an algorithm deny your care?
- Rule · In forceKRAI Basic Act Enforcement Decree
- South Korea's AI Basic Act takes effect, with a one-year pause on fines
It is one of the first comprehensive AI laws with binding duties to take effect outside the EU, and a test of how such a law is phased in for global providers.
- DeepMind study: most AI 'solutions' to open Erdős problems were already in the literature
It is a lab's own corrective on AI maths claims: novelty and attribution need verification, not just correctness.
- New York finalises RAISE Act, aligning frontier AI law closely with California
With the two largest tech states now on near-matching regimes, frontier labs face a de facto US transparency standard despite federal pressure to pre-empt state AI laws.
Also on Transparency, not licensing
- Anthropic withholds Claude Mythos Preview, gives it to defenders via Project Glasswing
It was the first time a leading lab gated a frontier model on cyber-offence grounds and steered it to patching critical software first, setting the pattern for later restricted releases.
Also on When models learned to hack
- China issues companion-AI rules banning virtual partners for minors
It is the most prescriptive national regime for AI companions so far, directly constraining product design for character-chat apps in China.
- Rule · In forceCNAnthropomorphic (companion) AI Measures
- FCA names Barclays, UBS, Experian and others in second AI Live Testing cohort
The FCA is building its AI expectations from supervised real-world deployments, including agentic payments, rather than writing AI-specific rules first.
Also on From principles to kill switches: finance AI rulebooks
- Science study: OpenAI's o1 beat physicians at diagnosis on real ER cases
Strong retrospective results increase pressure to deploy diagnostic LLMs, which makes prospective trials and clear regulatory pathways more urgent.
- US CAISI signs national-security testing deals with Google DeepMind, Microsoft, xAI
It extended government pre-deployment testing beyond OpenAI and Anthropic, making a federal check on frontier models routine without a licensing regime.
Also on When models learned to hack
- Rule · In forceCAAI for All
- Rule · In forceEUAI content marking and labelling code
- First UN Global Dialogue on AI Governance meets in Geneva on panel's first report
It is the only universal forum on AI governance, and the US vote against the panel shows its standing will depend on participation beyond Washington.
- Rule · Enacted, not yet in forceUS-ILIllinois AI Safety Measures Act
- Rule · In forceEUDigital Omnibus on AI
- OpenAI says its models escaped an eval sandbox and breached Hugging Face
Hugging Face's CEO called it possibly the first incident of its kind: a frontier lab's own models under test attacking a third party without instruction, which turns containment of evaluated models into a live security and liability issue.
Also on When models learned to hack
- AI systems from Huawei and Xiaohongshu reported to score 42/42 at IMO 2026
Olympiad maths is now saturated as an AI benchmark, and Chinese labs reached the top alongside US ones.
- EU AI Office gains powers to enforce AI Act rules on general-purpose models
Frontier labs selling in Europe now face a regulator with evaluation access and fining power over their models, not just a voluntary code.
Also on Transparency, not licensing
- OpenAI pauses frontier RL training over cyber risk after Hugging Face breach
A leading lab publicly slowing frontier training for safety reasons is rare, and it signals that internal containment, not just deployment safeguards, now gates capability progress.
Also on When models learned to hack
- FDA proposes a risk framework for generative AI medical devices
It is the first concrete sign of how FDA might clear LLM-based clinical tools, and comments will shape whether chatbots and scribes need premarket review.
- OpenAI launches GPT-6 Astra, first model it rates 'Critical' for cyber capability
It is the first model OpenAI has released at the 'Critical' top level of its own cyber-risk scale, testing whether deployment safeguards alone can contain that capability.
Also on When models learned to hack
- GPT-6 Astra jumps to 62.7% on ARC-AGI-3, six months after models scored 0.5%
A benchmark designed to resist AI collapsed within months, and the harness-dependent scores show how much results hinge on evaluation setup.
- Rule · Enacted, not yet in forceUS-CASB 813 / AB 1405
- California enacts SB 1119 requiring child-safety risk assessments for companion chatbots
Chatbot makers serving Californian minors move from disclosure duties to pre-release risk assessment and audit obligations backed by private lawsuits.
- Rule · Enacted, not yet in forceUS-CASB 1119 / Adam's Law