AI Wins Gold in Math Olympiad
AI reaches new milestone by solving hardest school-level math problems
OpenAI recently announced that its generative AI model had achieved a Gold medal performance in the International Mathematical Olympiad (IMO). Shortly thereafter, Demis Hassabis—CEO of Google DeepMind and Nobel laureate—announced that their model has achieved the same performance. Last year, their specialized model called AlphaProof and AlphaGeometry could only manage a Silver.
This isn’t just another benchmark. The IMO fiercely tests pre-university students’ mathematical creativity, problem-solving skills, and logical depth. It does not test routine calculations; instead, it demands original, often elegant, ideas. The problems are notoriously hard even for some mathematicians, let alone machines.
The exam typically consists of six problems spread over two days, and examiners score each contestant out of 42 points. Achieving a Gold usually requires 30+ points, depending on that year’s difficulty. India has had an impressive history at the IMO, with students consistently ranking in the top 20 countries globally and winning Golds and Silvers several times..
So, what does it mean when an AI model secures a Gold in such a competition? It’s exciting news for more than one reason. It signals that generative AI—originally embraced for tasks like writing emails, generating code snippets, or searching resources online—can now solve some of the hardest intellectual problems, which everyone once believed required deep human intuition and creativity.
This doesn’t mean AI has become sentient, though that misunderstanding persists. What it does mean is that these models are now robust tools capable of complex reasoning, mathematical insight, and strategic problem-solving—far beyond their initial design goals.
What makes this even more remarkable is that these AI systems are general-purpose large language models (LLMs). Unlike AlphaGo or AlphaFold, which were trained with the explicit purpose of mastering Go or protein folding, these models were not specialised for maths. These models—trained on vast amounts of general text and code rather than specialized mathematical datasets—nevertheless solved the Olympiad problems without any internet access or human intervention.