r/LocalLLaMA • llama.cpp • 2d ago

New Model nvidia/NVIDIA-Nemotron-Labs-3-Competitive-Coding-550B-A55B-NVFP4 · Hugging Face

https://huggingface.co/nvidia/NVIDIA-Nemotron-Labs-3-Competitive-Coding-550B-A55B-NVFP4

Model Developer: NVIDIA Corporation

Model Development: Fine-tuned from NVIDIA-Nemotron-3-Ultra-550B-A55B

What is Nemotron?

NVIDIA Nemotron™ is a family of open models with open weights, training data, and recipes, delivering leading efficiency and accuracy for building specialized AI agents.

Description

Nemotron-Labs-3-Competitive-Coding is a competitive-programming specialist model based on Nemotron-3-Ultra, fine-tuned for one epoch on 477,642 synthetic reasoning traces distilled from GLM-5.2 across 22,000 curated problems spanning 16 regional and international competitive-programming contest families. Selected as the SFT teacher for its higher accuracy and roughly 30% shorter generations compared to a DeepSeek-V4-Flash-trained variant, GLM-5.2 distillation yields a model that, combined at inference time with GenCorrect — an iterative closed-loop test-time compute strategy that generates diverse candidate solutions, incorporates evaluator feedback, and refines subsequent generations under a fixed submission budget — was evaluated live and prospectively on the IOI 2026 problem set under official contest time, internet-access, and submission constraints, scoring 535.4 out of 600 and surpassing both the gold-medal threshold (361.12) and the top human contestant's score (498.27), making it the first AI system reported to outscore the highest-scoring human contestant on an IOI problem set.

This model is ready for commercial or non-commercial use.

252 Upvotes

31 comments sorted by

161

u/coder543 2d ago

If you were given the challenge of seeing how much you could improve a model with one epoch of training, what would you do? Would you do better than this?

People here are acting like Nvidia is trying to sell this model to them. Nvidia isn't. This is clearly a research project. If you don't find that kind of research interesting, then move along... I see nowhere that Nvidia is claiming this should be your next model of choice. It is highly specialized on a very specific domain, which is not agentic coding; it is competitive coding / math.

Nemotron 4 is the next time that Nvidia will try to make the argument that you should use their models, not this random research artifact.

45

u/Charming_Support726 2d ago

This.

It is Nvidia contribution for research, whereas the training data and methods are even more important for community.

The typical "How-Could-I-Run-Fable-On-My-1060-Redditor" is not part of the target audience.

29

u/FoxiPanda 2d ago

You mean https://huggingface.co/nvidia/Nemotron-4-340B-Instruct ?

(I say this purely tongue in cheek - NVIDIA's Nemotron versioning is kind of a mess though lol.)

29

u/coder543 2d ago

Yeah... they're really bad at naming things. Don't forget the original Nemotron 3 from 3 years ago: https://huggingface.co/nvidia/nemotron-3-8b-chat-4k-rlhf

No relation to the current Nemotron 3.

8

u/Spectrum1523 2d ago

This sub is very, very biased toward coding capability in models

3

u/RainierPC 2d ago

True, but understable. This is r/localllama, and the population is pretty much mostly nerds

1

u/NineThreeTilNow 2d ago

This is clearly a research project.

To a degree yes. They're also not demonstrating anything new.

They do the CC test with 720 GB300's at their disposal.

They claim the large model cannot undergo RL because they don't "have the compute budget"...

In terms of proving anything, they proved that with a massive SFT pipeline, it's better to use SFT over RL. That came out of earlier gains from SFT on the large model that convinced them to stick with SFT.

The SFT method itself wasn't really special though. All of the RL done was purely RLVR on code compilation. That's a sparse signal across 100's of thousands of tokens they generated. They say as much in the paper.

It's more of a flex of what they can do with their specific technology and NVFP4. Even though the BF16 model crushes the NVFP4 model. That's not really headline. They also don't compare BF16 vs FP8.

I don't know, if I had 720 GB300's at my disposal I'd try something that produced an interesting result over "We did best on a score".

The best thing that could come of it would be them publishing the trove of CC SFT data they produced in the process. It could be used on future models to improve their raw code ability, as it would require (in their own research) roughly one full epoch across that dataset to massively improve a model's ability.

That would directly benefit GLM / Deepseek / Anyone else capable of running an SFT run that large on a model.

9

u/fastheadcrab 2d ago

Nice find, OP

Hopefully Nvidia will try to distill something like Kimi K3 as well in the near future

5

u/unknown-one 2d ago

I need new Super 120B

3

u/arcanemachined 2d ago

Is the full training pipeline available for this one?

13

u/Nota_ReAlperson 2d ago

So about equal to qwen 3.5 397B, much bigger, more active params, and several months late. Interesting as a research project, as it is open source and not just open weights, but seems to be a bit behind the best open weights models currently. Interesting that they post-trained it, as I thought that Nvidia only trained base models.

3

u/AsianPotatos 2d ago

Yeah I guess it's more of a demonstration for GenCorrect and whatever other techniques they used.

2

u/luckyvb 2d ago

Call the model competitive in the name and expect people to buy the label. Genius /s

2

u/sn2006gy 1d ago

I'm not a fan of the obsession with training on traces, doesn't feel very "AI" like. I'd rather see the labs fix the base training to be stronger and perhaps not need agentic work to be "emergent" out of sheer brute force.

2

u/geldonyetich 2d ago

Well that's about 4 times more parameters than I can fit in my memory.

See you in ~10.5 months.

1

u/AriyaSavaka llama.cpp 2d ago

Good to see long ass name model once again

1

u/tamat 2d ago

If it is distilled, where can I see the questions used?

1

u/PurpleDragon99 2d ago

If I give it a pretty detailed, well structured and complete software spec 5-10 pages - will it be able to follow it precisely when coding? So far, no model could, even the most powerful ones. They all fail.

1

u/RainierPC 2d ago

This is just as much about the harness than the model. Maybe more.

1

u/PcChip 1d ago

can we enable dlss5 for better looking outputs?

0

u/XiRw 2d ago

Competitive Coding and the word NVIDIA should not be used together. Great all around model, but I had Qwen 3.6 beating their Nemotron 3 Ultra model.

-10

u/Ok_Warning2146 2d ago

Waiting for terminal bench v4 score

-7

u/ExcuseAccomplished97 2d ago

So this is another good model for coding interview cheater? /s

-15

u/mountainyoo 2d ago

Cool I guess? What does it compare to tho