Dev News Daily ENDE

NVIDIA fine-tunes Nemotron 3 to gold-level scores at both IOI 2026 and IMO 2026

NVIDIA published on 7 October, on the Hugging Face blog, how it specialised its Nemotron 3 models for two very different competitions: the International Olympiad in Informatics (IOI), which needs code that passes hidden tests under time and submission limits, and the International Mathematical Olympiad (IMO), which needs written proofs.

The two results.

  • IOI 2026: Nemotron-3-Ultra-CC, fine-tuned and run with an iterative generate-evaluate-refine loop the team calls GenCorrect, scored 535.4 of 600. The gold threshold was 361.12 and the top human score 498.27. NVIDIA stresses that this was a live run under the contestants' time, internet and submission constraints, but an unofficial, unsupervised benchmark that is not in the official ranking.
  • IMO 2026: a generate-verify-refine system built on Nemotron 3 Ultra checkpoints scored 30 of 42, above the official gold threshold of 29, with full credit on four of six problems. The proofs were graded by official IMO graders. The system worked in natural language, with no formal prover, tools or internet access.
NVIDIA fine-tunes Nemotron 3 to gold-level scores at both IOI 2026 and IMO 2026
NVIDIA fine-tunes Nemotron 3 to gold-level scores at both IOI 2026 and IMO 2026 — Dev News Daily

The recipe. The post describes four steps it says are reusable: start from a strong base model; curate domain problems and reasoning traces; apply standard post-training (supervised fine-tuning and, where useful, reinforcement learning); and pair the specialist with an inference loop that generates, checks and improves answers.

The numbers behind it. For programming the team curated 22,000 problems. The smaller Nemotron-3-Nano-CC (30 billion parameters, 3 billion active) went from 130 points on IOI 2025 before post-training to 280 after fine-tuning, 291 after reinforcement learning and 468 with GenCorrect. For mathematics the fine-tuning set held 414,890 examples over 15,818 proof problems, covering not only proofs but their verification and critique.

The post's own conclusion: specialisation and test-time compute work together — a better-trained model gives the search loop better candidates and better critics.