Lines of Thought

When AI started doing real mathematics

How far has AI got in research mathematics, and how far can its claims be trusted?

In about a year AI went from olympiad gold to refuting a famous conjecture and claiming a Millennium-problem case. Along the way, novelty checks, formal verification and credit disputes became as important as the results.

  1. ResearchScienceGB US
    Gemini Deep Think earns officially graded gold-medal score at IMO 2025

    Certified olympiad-level proofs in natural language set the baseline.

  2. ResearchScienceGB US
    Google DeepMind's AlphaEvolve agent improves algorithms and open maths bounds

    Search plus automatic verification yields new, checkable bounds.

  3. ResearchScienceGB INTL
    Review of 445 LLM benchmarks finds widespread construct-validity weaknesses

    A reminder that benchmark scores need valid measurement behind them.

  4. ResearchScienceGB US
    DeepMind study: most AI 'solutions' to open Erdős problems were already in the literature

    Many 'open' problems fell to literature search, not new ideas.

  5. ResearchScienceUS
    OpenAI model disproves Erdős's 1946 unit-distance conjecture

    First widely accepted historically significant AI result.

  6. ResearchScienceCN INTL
    AI systems from Huawei and Xiaohongshu reported to score 42/42 at IMO 2026

    Olympiad maths saturates as a benchmark, including for Chinese labs.

  7. ResearchScienceUS GB
    Claude completes first end-to-end Lean formal proof of Fermat's Last Theorem

    Machine-checked formalisation at the scale of Wiles's proof.

  8. ResearchScienceUS
    OpenAI claims Navier-Stokes blow-up proof; mathematicians dispute credit

    The largest claim yet, contested over credit and awaiting acceptance.

Where this line could go next

Connected developments not on this line