How Nemotron AI Models Reached Gold-Medal Levels in Math and Coding
Researchers adapted Nemotron foundation models into high-level specialists for math and coding competitions. By combining specialized training with search and refinement loops, the systems reached gold-medal scores.
In short: Researchers adapted Nemotron foundation models into high-level specialists for math and coding competitions. By combining specialized training with search and refinement loops, the systems reached gold-medal scores.
Artificial intelligence systems have reached a new milestone by solving complex mathematical proofs and computer programming challenges at the level of top human competitors.
What happened, in plain words
Teams at Hugging Face and NVIDIA used Nemotron foundation models to create specialist systems for the International Olympiad in Informatics (IOI) and the International Mathematical Olympiad (IMO). Starting from Nemotron base models, the teams used supervised fine-tuning, reinforcement learning, and feedback-driven inference loops to generate, evaluate, and refine answers. The programming system was tested in a live, unsupervised run under human constraints, while the math system's proofs were graded by official IMO graders. Both systems achieved scores exceeding the official gold-medal thresholds.
Key points
- Four-part recipe for adaptation The teams started with a Nemotron base model, curated domain-specific problems with reasoning traces, applied training methods like supervised fine-tuning, and paired the specialist with an inference loop to evaluate and improve answers.
- Coding competition results For programming, researchers trained two specialists named Nemotron-3-Nano-CC and Nemotron-3-Ultra-CC. Using an iterative strategy called GenCorrect, the models improved their scores and crossed the gold threshold at IOI 2025.
- Math competition results For olympiad mathematics, the team trained specialists using datasets covering proof generation, refinement, and verification. The final system worked entirely in natural language without external tools and scored 30 out of 42 points.
- Shared resources for the community The project released training datasets, model checkpoints, code repositories, and benchmarks including Nemotron-IMO-Bench to help others reproduce and build upon the work.
Terms explained
- Foundation model — A large, general-purpose artificial intelligence model that can be adapted to many different specific tasks. Example: Think of it like a blank canvas that can be trained to paint portraits or landscapes.
- Supervised fine-tuning — A training process where an AI model learns by reviewing example problems and correct answers provided by humans. Example: Like a student studying solved homework problems to learn how to solve similar ones.
- Reinforcement learning — A method of training an AI through trial and error, rewarding it for correct steps and penalizing it for mistakes. Example: Like teaching a dog a trick by giving it treats when it performs correctly.
- Inference loop — A process where the AI generates an answer, checks it for errors, and uses that feedback to try again. Example: Like writing an essay, proofreading it, and rewriting sections to make it better before turning it in.
Why it matters
These experiments show a reproducible recipe for turning general AI models into reliable domain specialists. Releasing these models and code allows the broader community to study and build upon these advanced reasoning techniques.
What we still don't know
The IOI result was an unofficial, unsupervised benchmark and was not included in the official IOI ranking. The systems relied on specific competitions and datasets, and the source notes that adaptation methods and outcomes varied depending on the model scale.
Based on reporting from Hugging Face Blog. This is an independent explainer, written in our own words with AI assistance; Hugging Face Blog has not reviewed or endorsed it. Read the original for the full details.