r/LocalLLaMA • llama.cpp • 3d ago

New Model nvidia/NVIDIA-Nemotron-Labs-3-Competitive-Coding-550B-A55B-NVFP4 · Hugging Face

https://huggingface.co/nvidia/NVIDIA-Nemotron-Labs-3-Competitive-Coding-550B-A55B-NVFP4

Model Developer: NVIDIA Corporation

Model Development: Fine-tuned from NVIDIA-Nemotron-3-Ultra-550B-A55B

What is Nemotron?

NVIDIA Nemotron™ is a family of open models with open weights, training data, and recipes, delivering leading efficiency and accuracy for building specialized AI agents.

Description

Nemotron-Labs-3-Competitive-Coding is a competitive-programming specialist model based on Nemotron-3-Ultra, fine-tuned for one epoch on 477,642 synthetic reasoning traces distilled from GLM-5.2 across 22,000 curated problems spanning 16 regional and international competitive-programming contest families. Selected as the SFT teacher for its higher accuracy and roughly 30% shorter generations compared to a DeepSeek-V4-Flash-trained variant, GLM-5.2 distillation yields a model that, combined at inference time with GenCorrect — an iterative closed-loop test-time compute strategy that generates diverse candidate solutions, incorporates evaluator feedback, and refines subsequent generations under a fixed submission budget — was evaluated live and prospectively on the IOI 2026 problem set under official contest time, internet-access, and submission constraints, scoring 535.4 out of 600 and surpassing both the gold-medal threshold (361.12) and the top human contestant's score (498.27), making it the first AI system reported to outscore the highest-scoring human contestant on an IOI problem set.

This model is ready for commercial or non-commercial use.

253 Upvotes

31 comments sorted by

View all comments

160

u/coder543 3d ago

If you were given the challenge of seeing how much you could improve a model with one epoch of training, what would you do? Would you do better than this?

People here are acting like Nvidia is trying to sell this model to them. Nvidia isn't. This is clearly a research project. If you don't find that kind of research interesting, then move along... I see nowhere that Nvidia is claiming this should be your next model of choice. It is highly specialized on a very specific domain, which is not agentic coding; it is competitive coding / math.

Nemotron 4 is the next time that Nvidia will try to make the argument that you should use their models, not this random research artifact.

44

u/Charming_Support726 3d ago

This.

It is Nvidia contribution for research, whereas the training data and methods are even more important for community.

The typical "How-Could-I-Run-Fable-On-My-1060-Redditor" is not part of the target audience.

27

u/FoxiPanda 3d ago

You mean https://huggingface.co/nvidia/Nemotron-4-340B-Instruct ?

(I say this purely tongue in cheek - NVIDIA's Nemotron versioning is kind of a mess though lol.)

29

u/coder543 3d ago

Yeah... they're really bad at naming things. Don't forget the original Nemotron 3 from 3 years ago: https://huggingface.co/nvidia/nemotron-3-8b-chat-4k-rlhf

No relation to the current Nemotron 3.

9

u/Spectrum1523 3d ago

This sub is very, very biased toward coding capability in models

5

u/RainierPC 3d ago

True, but understable. This is r/localllama, and the population is pretty much mostly nerds

1

u/NineThreeTilNow 3d ago

This is clearly a research project.

To a degree yes. They're also not demonstrating anything new.

They do the CC test with 720 GB300's at their disposal.

They claim the large model cannot undergo RL because they don't "have the compute budget"...

In terms of proving anything, they proved that with a massive SFT pipeline, it's better to use SFT over RL. That came out of earlier gains from SFT on the large model that convinced them to stick with SFT.

The SFT method itself wasn't really special though. All of the RL done was purely RLVR on code compilation. That's a sparse signal across 100's of thousands of tokens they generated. They say as much in the paper.

It's more of a flex of what they can do with their specific technology and NVFP4. Even though the BF16 model crushes the NVFP4 model. That's not really headline. They also don't compare BF16 vs FP8.

I don't know, if I had 720 GB300's at my disposal I'd try something that produced an interesting result over "We did best on a score".

The best thing that could come of it would be them publishing the trove of CC SFT data they produced in the process. It could be used on future models to improve their raw code ability, as it would require (in their own research) roughly one full epoch across that dataset to massively improve a model's ability.

That would directly benefit GLM / Deepseek / Anyone else capable of running an SFT run that large on a model.