Latest / Health and life sciences

Science study: OpenAI's o1 beat physicians at diagnosis on real ER cases

ResearchHealthScienceUSConfirmed

Researchers at Harvard Medical School and Beth Israel Deaconess reported in Science that OpenAI's o1 reasoning model outperformed physician baselines across clinical-reasoning benchmarks and, on 76 real emergency-department cases using raw health-record data, named the exact or a very close diagnosis more often than attending physicians. The authors said prospective clinical trials are needed before autonomous use.

Why it matters

Strong retrospective results increase pressure to deploy diagnostic LLMs, which makes prospective trials and clear regulatory pathways more urgent.

SourceHarvard Medical School via NewswiseCoverage: STATCoverage: Becker's Hospital Review Checked against the primary source. Independently fact-checked on 7 Oct 2026.
Harvard Medical SchoolBeth Israel Deaconess Medical CenterOpenAI

Line of Thought

Follow this story

Pick any item to keep going. Your path builds up above as a line you can share.

Curated lines through this story

What led here

Earlier developments on the same thread

What happened next

Later developments on the same thread

Same story elsewhere

What other countries and bodies did on this

Rules in play

Laws and guidance this touches

Threads by topic: Clinical AI Safety testing