Net Good IndexSubmit a correction

Capability Ledger · verified · Mathematics

An OpenAI reasoning model reaches IMO 2025 gold-medal standard

In July 2025 OpenAI reported that an experimental reasoning model scored 35/42 on the IMO 2025 paper under contest-like constraints, graded by former medalists.

Tier 3 MajorMethodology 0.2
Submit a correction

Capability score

1.49

Does not count toward Net Good or the leaderboard.

10 base · Major (tier 3 of 5, 10 pts)
× 0.8500 attribution · Primary causal contribution
× 0.5000 evidence · External expert evaluation
× 0.5000 realization · Complete public manuscript
× 0.7000 durability · Useful for many years
Event-level product before credit split: 1.49

Independent gold-threshold IMO performance is a major capability milestone (tier 3), not Net Good. It is not shared with Gemini. Attribution is high. Evidence is journalism and developer posts rather than official IMO grading (0.50).

What happened

OpenAI did not enter IMO 2025 as a contestant. An experimental reasoning model produced natural-language solutions to five of six problems for 35 of 42 points, scored by three former medalists under contest-style time limits. Gemini Deep Think independently reached the same score with official IMO grading; that is a separate event. This is a capability demonstration on a contest paper whose answers already existed, not newly created mathematics, and does not count toward Net Good.

Model attribution

OpenAI reasoning models

Experimental reasoning model scored 35/42 by former IMO medalists on the 2025 paper.

Independent parallel result; not official IMO grading and not a collaboration with Gemini.

Attribution 0.8500 · Credit share 100% · OpenAI

Claims

  • An OpenAI experimental model scored 35/42 on IMO 2025 under contest-like constraints.

    outcome · supported

  • OpenAI officially won an IMO gold medal as a contestant.

    significance · disputed

Sources

independent sources

Secondary domains: Computer Science

Revision history

  • 14 Sep 2026 · 0.00 1.49

    Imported events under methodology 0.2.