Latest / Health and life sciences
Science study: OpenAI's o1 beat physicians at diagnosis on real ER cases
Researchers at Harvard Medical School and Beth Israel Deaconess reported in Science that OpenAI's o1 reasoning model outperformed physician baselines across clinical-reasoning benchmarks and, on 76 real emergency-department cases using raw health-record data, named the exact or a very close diagnosis more often than attending physicians. The authors said prospective clinical trials are needed before autonomous use.
Why it matters
Strong retrospective results increase pressure to deploy diagnostic LLMs, which makes prospective trials and clear regulatory pathways more urgent.
Line of Thought
Follow this story
Pick any item to keep going. Your path builds up above as a line you can share.
Curated lines through this story
What led here
Earlier developments on the same thread
- DevelopmentOpenAI launches ChatGPT Health, linking medical records and wellness apps7 Jan 2026 · Launch · US
- DevelopmentFDA loosens oversight of AI decision-support software and health wearables6 Jan 2026 · Rule change · US
- DevelopmentIllinois bars AI from delivering therapy without a licensed professional1 Aug 2025 · Rule change · US-IL, US
- DevelopmentGemini Deep Think earns officially graded gold-medal score at IMO 202521 Jul 2025 · Research · GB, US
What happened next
Later developments on the same thread
- DevelopmentFDA proposes a risk framework for generative AI medical devices18 Aug 2026 · Policy · US
- DevelopmentOpenAI pauses frontier RL training over cyber risk after Hugging Face breach18 Aug 2026 · Statement · US
- DevelopmentGPT-6 Astra jumps to 62.7% on ARC-AGI-3, six months after models scored 0.5%3 Sep 2026 · Report · US
- DevelopmentOpenAI launches GPT-6 Astra, first model it rates 'Critical' for cyber capability3 Sep 2026 · Model release · US
Same story elsewhere
What other countries and bodies did on this
- DevelopmentEU deal keeps AI medical devices under AI Act high-risk rules, delayed to 20287 May 2026 · Rule change · EU
- DevelopmentNHS England updates AI scribe guidance and launches supplier registry2 Apr 2026 · Policy · GB
- DevelopmentAI systems from Huawei and Xiaohongshu reported to score 42/42 at IMO 202623 Jul 2026 · Research · CN, INTL
- DevelopmentEBU-BBC study finds almost half of AI assistant news answers have a significant flaw21 Oct 2025 · Research · INTL
Rules in play
Laws and guidance this touches
- RuleFDA PCCP guidanceUS · In force · 4 Dec 2024
- RuleEU AI ActEU · Partly in force · 13 Jun 2024
Threads by topic: Clinical AI Safety testing