NetGoodIndexSubmit a correction

Benefit Ledger · verified · Mathematics

Gemini Deep Think and an OpenAI reasoning model both reach IMO 2025 gold-medal scores

In July 2025, Google DeepMind’s Gemini Deep Think was officially graded at 35/42 on the IMO 2025 paper, and OpenAI reported the same score from an experimental reasoning model graded by former medalists.

21 Jul 2025Tier 3 MajorMethodology 0.1

Current score

+1.53

10 base · Major (tier 3 of 5, 10 pts)
× 0.9000 attribution · Primary causal contribution
× 0.5000 evidence · External expert evaluation
× 0.5000 realization · Experimentally validated
× 0.7000 durability
Event-level product before credit split: 1.58

Independent gold-threshold performances are a major milestone (tier 3), not a solved open research problem. Attribution is high. Evidence is official grading for Gemini plus journalism and developer posts for OpenAI (0.50). Realization is demonstrated contest-style proofs. Durability is medium because models and protocols change quickly.

What happened

The 2025 International Mathematical Olympiad became the first year two independent general-purpose reasoning systems reported gold-medal-threshold scores (35 of 42, five of six problems) under contest-style 4.5-hour sessions and natural-language proofs. IMO coordinators graded DeepMind’s entry. OpenAI did not enter officially and used three former medalists. Both results are capability demonstrations with released writeups, not official contest medals. Human gold medalists still exist above and at this score.

Model attribution

OpenAI reasoning models

Experimental reasoning model scored 35/42 by former IMO medalists on the same paper.

Independent parallel result; not official IMO grading.

Attribution 0.8500 · Credit share 50% · OpenAI

Gemini

Officially graded IMO 2025 solutions scoring 35/42.

Independent system with IMO coordinator grading.

Attribution 0.9000 · Credit share 50% · Google DeepMind

Claims

  • Gemini Deep Think scored 35/42 on IMO 2025 as graded by IMO coordinators.

    outcome · supported

  • An OpenAI experimental model also scored 35/42 on the same problems under contest-like constraints.

    outcome · supported

  • Either company officially won an IMO gold medal as a contestant.

    significance · disputed

Sources

primary sources

independent sources

Secondary domains: Computer Science

Revision history

  • 13 Sep 2026 · 0.00 1.53

    Initial adjudicated seed score under methodology 0.1.