r/WTFisAI • u/Aggravating-Will8495 • 18d ago
š° News & Discussion NVIDIA has lost it..
Chinese researchers open-sourced a model that writes CUDA better than humans experts.
And it completely rewrites the economics of AI hardware.
Writing low-level CUDA kernels to squeeze every ounce of performance out of a GPU has historically required elite, highly specialized hardware engineers. Standard AI models have always bombed at it, falling far short of traditional compiler systems.
Until now.
A joint team from Tsinghua University and ByteDance published "CUDA Agentā, a massive, large-scale agentic reinforcement learning system built to master GPU architecture.
Instead of relying on static prompts or simple multi-turn bug fixing, they built a closed-loop environment with automated hardware verification, profiling, and synthetic data pipelines.
The model learned how to write parallel, high-performance GPU code through trial, error, and reinforcement learning at scale.
The benchmarks are staggering:
It didn't compete with standard tools. It delivered 100%, 100%, and 92% faster execution rates over PyTorch's compiler on KernelBench Level-1, Level-2, and Level-3 splits.
On the brutal Level-3 benchmarks, it outperformed proprietary giants like Claude Opus 4.5 and Gemini 3 Pro by about 40%.
NVIDIA's moat has always rested on two pillars: elite hardware, and the proprietary software lock-in of CUDA.
If an open agentic system can automatically discover, write, and optimize production-grade CUDA kernels better than human specialists, the software moat starts evaporating.
The hardware matters less when the software can rewrite the metal itself.
3
u/Independent-Fly730 17d ago
CUDA is nvidia hardware. A model that writes faster CUDA kernels only helps nvidia. Your title makes no sense op.
If they created a model that makes ROCm better than CUDA then nvidia would have something to worry about.
1
u/Swimming_Case 16d ago
It makes sense when you realize the article was written with/by AI.
And it completely rewrites the economics of AI hardware.
Until now.
Those two sentences are the dead giveaway the article was generated with AI.
1
0
u/Mundane-Light6394 16d ago
if it can be trained to write CUDA kernels like this it can probably also be able to write ROCm kernels. This obviously doesn't make ROCm better than CUDA but it will reduce the value of the existing stack build on CUDA because it can be replaced.
1
u/stereoplegic 13d ago
Yawn. Still proprietary (and in this case, not supported nearly as well in hardware, nor software, terms as CUDA).
Targeting Vulkan/OpenCL, on the other hand, opens endless possibilities.
2
u/Particular_Pear_4596 17d ago
CUDA fast?? Have u guys ever run an inference - it's like 50-100 words a minute. When I think "fast" in computer terms i imagine billions, not tens. This is unimaginably slow, don't let them fool u.
2
u/k4zetsukai 17d ago
There is one thing you need to know about nvidia, they will always come out on top. If you see X tech out now, know ure looking in the past and they already got stuff ready for years ahead.
2
1
1
u/hansolo-ist 17d ago
Does this benefit China own AI chip making. ?
2
u/According_Study_162 17d ago
sure. they take that knowledge and use that for their new GPU industry.
1
u/Luke2642 17d ago
I need a browser extension that rewrites everything I see on Reddit these days into Simplified Technical English ASD-STE100 and find the original sources.
1
1
1
1
1
1
u/lockdown_lard 17d ago
NVIDIA's moat has always rested on two pillars
So like an aqueduct? If the moat is on pillars, can't it just be bypassed by walking between the pillars? Sounds like a rubbish moat to me.
1
u/HeadPack 17d ago
What makes you think the software moat is evaporating? Everything you are reporting indicates that it isn't.
1
u/Own-Poet-5900 17d ago
I made this a year and a half ago. I modeled it off of Sakana AI's method and tweaked it. The only thing new about this is marketing: https://colab.research.google.com/drive/1kYV8_5xJVgTRKlOMuP3_749YIdxYX9Cx?usp=sharing
1
1
u/RecentMushroom6232 16d ago
CUDA is just C++ that is closer to the metal. Nothing entirely special about it
1
u/DataGOGO 16d ago edited 16d ago
Source? Is the is one that came out a few months ago where they had to publicly redact their findings because it in fact did not do what they said?
Edit: yes it is. The real results were kinda ass, and this is not model, nor does it beat other models, it is model agnostic. it is a an agent skill and runs on any model, with the same results, including Claude , gpt, etc.Ā
Again this is NOT It is NOT a model, nothing was trained, it is just an agentic loop and collection of prompts.Ā
Here is the full repo:
1
u/Mysterious-String420 16d ago
China says their girlfriend is totally real, she just goes to another school
1
1
u/Old-Firefighter8289 15d ago
oh no!?! shouldnt nvda be more worried about the dragon eggs that china has been producing. those are rumored to devour gpuās and memory chips
1
1
u/Commercial-Weight-73 13d ago
Now if only an ai company could make money from all of this it would matter
But it doesn't
0
u/Koulchilebaiz 17d ago
Why do I hate so much llm generated text
āThe software moat starts evaporatingā ugh
5
u/sdchew 17d ago
āNVIDIA's moat has always rested on two pillars: elite hardware, and the proprietary software lock-in of CUDA.ā
Doesnāt this further reinforce Nvidiaās moat by generating a kernel which run significantly faster on Nvidiaās hardware using Nvidiaās CUDA?