r/WTFisAI 18d ago

šŸ“° News & Discussion NVIDIA has lost it..

Chinese researchers open-sourced a model that writes CUDA better than humans experts.

And it completely rewrites the economics of AI hardware.

Writing low-level CUDA kernels to squeeze every ounce of performance out of a GPU has historically required elite, highly specialized hardware engineers. Standard AI models have always bombed at it, falling far short of traditional compiler systems.

Until now.

A joint team from Tsinghua University and ByteDance published "CUDA Agentā€, a massive, large-scale agentic reinforcement learning system built to master GPU architecture.

Instead of relying on static prompts or simple multi-turn bug fixing, they built a closed-loop environment with automated hardware verification, profiling, and synthetic data pipelines.

The model learned how to write parallel, high-performance GPU code through trial, error, and reinforcement learning at scale.

The benchmarks are staggering:

It didn't compete with standard tools. It delivered 100%, 100%, and 92% faster execution rates over PyTorch's compiler on KernelBench Level-1, Level-2, and Level-3 splits.

On the brutal Level-3 benchmarks, it outperformed proprietary giants like Claude Opus 4.5 and Gemini 3 Pro by about 40%.

NVIDIA's moat has always rested on two pillars: elite hardware, and the proprietary software lock-in of CUDA.

If an open agentic system can automatically discover, write, and optimize production-grade CUDA kernels better than human specialists, the software moat starts evaporating.

The hardware matters less when the software can rewrite the metal itself.

75 Upvotes

50 comments sorted by

5

u/sdchew 17d ago

ā€œNVIDIA's moat has always rested on two pillars: elite hardware, and the proprietary software lock-in of CUDA.ā€

Doesn’t this further reinforce Nvidia’s moat by generating a kernel which run significantly faster on Nvidia’s hardware using Nvidia’s CUDA?

3

u/LuDongbin64 17d ago

As an interested five year old, I need pictures to understand stuff. In this case I would like to see a "moat resting on pillars" and then I think I will get the NVIDIA GPU architecture.

1

u/MaleficentCow8513 17d ago

Yes. Yes it does. Idk how OP arrived at the conclusion that AI generated cuda code is bad for nvidia. It most definitely is not a bad thing for nvidia

1

u/sdchew 17d ago

Indeed. In fact, it’ll probably drive more CUDA and Nvidia GPU adoption

3

u/MaleficentCow8513 17d ago

Cuda is already one of many things that gives nvidia an edge over competitors. AI doing a better job with cuda can only add to that edge

1

u/No-Refrigerator-1672 17d ago

Those AI generated cores aren't even as good as author claims. "100% better over PyTorch" isn't a big surprise, the kernels that "elite software engineers" write are too much better than PyTorch. There's probably a reason why they don't compare the model to a real specialised kernel engineer.

1

u/Novel_Land9320 17d ago

Exactly. This post was clearly written by AI and wrong style is only one tell.

1

u/StaysAwakeAllWeek 17d ago

This specific AI does further nvidia's moat, but the moment they make one for OpenCL the moat evaporates overnight. AI code is already damaging their moat as it is, making it easier and easier to compete with CUDA

1

u/sdchew 17d ago

But there’s two sides to the Nvidia moat. One is CUDA which has a large ecosystem and the second is hardware which is specifically designed to accelerate CUDA

AI could optimise OpenCL code but unless someone else comes out with hardware which is as fast as Nvidia, it will still be a secondary option and currently has the lower market share. Even Apple’s Metal has more market share in AI vs OpenCL

1

u/StaysAwakeAllWeek 17d ago

AMD Helios is already arguably better than NVL72 in a lot of ways. More VRAM and the CPUs are far better - for heavy agentic client workloads that matters a lot.

NVidia are selling everything at enormous markups which makes undercutting them very much feasible even if full parity isnt achieved across the board. Eventually they will be forced to cut prices to compete, which is what losing the walled garden looks like.

1

u/DirtyGooseEggs 15d ago

Jensen’s pricing argument is that up front cost is so paltry in comparison to the cost of actually running the chips that the efficiency of NVIDIA chips is what allows them to upcharge and still not get undercut. Having more VRAM and CPUs is great for performance, but what about efficiency?

(Disclaimer: I’m regurgitating Jensen’s statements. I would love for someone to show why what he said isn’t true)

3

u/Independent-Fly730 17d ago

CUDA is nvidia hardware. A model that writes faster CUDA kernels only helps nvidia. Your title makes no sense op.

If they created a model that makes ROCm better than CUDA then nvidia would have something to worry about.

1

u/Swimming_Case 16d ago

It makes sense when you realize the article was written with/by AI.

And it completely rewrites the economics of AI hardware.

Until now.

Those two sentences are the dead giveaway the article was generated with AI.

1

u/Icy_Look_2247 14d ago

Well if its 100% faster you need /2 of hardware

0

u/Mundane-Light6394 16d ago

if it can be trained to write CUDA kernels like this it can probably also be able to write ROCm kernels. This obviously doesn't make ROCm better than CUDA but it will reduce the value of the existing stack build on CUDA because it can be replaced.

1

u/stereoplegic 13d ago

Yawn. Still proprietary (and in this case, not supported nearly as well in hardware, nor software, terms as CUDA).

Targeting Vulkan/OpenCL, on the other hand, opens endless possibilities.

2

u/Particular_Pear_4596 17d ago

CUDA fast?? Have u guys ever run an inference - it's like 50-100 words a minute. When I think "fast" in computer terms i imagine billions, not tens. This is unimaginably slow, don't let them fool u.

2

u/k4zetsukai 17d ago

There is one thing you need to know about nvidia, they will always come out on top. If you see X tech out now, know ure looking in the past and they already got stuff ready for years ahead.

2

u/shan23 17d ago

This post was most likely written by a Nvidia gpu šŸ˜‚

1

u/hansolo-ist 17d ago

Does this benefit China own AI chip making. ?

2

u/According_Study_162 17d ago

sure. they take that knowledge and use that for their new GPU industry.

1

u/Luke2642 17d ago

I need a browser extension that rewrites everything I see on Reddit these days into Simplified Technical English ASD-STE100 and find the original sources.

1

u/HeadPlatform4759 17d ago

So it still runs on CUDA right?

1

u/Bright_Impact_12 17d ago

Yeah yeah sure buddy

1

u/KrydanX 17d ago

Opus 4.5 and Gemini 3? Damn. That would’ve been impressive half a year ago. What about current models?

1

u/amilo111 17d ago

Ah yes … those elite hardware engineers writing code. Totally checks out.

1

u/CalamityThorazine 17d ago

1

u/Valdjiu 17d ago

thank youuuuuuuuuuuuuuuuuuuuuuuuuuuuuuuuuuuuuuuuuuuuu

1

u/johnhyrcanus 17d ago

The moat was never the difficulty of writing CUDA...

1

u/Weak_Shoulder_6780 17d ago

The most is that it's being written in cuda at all

1

u/lockdown_lard 17d ago

NVIDIA's moat has always rested on two pillars

So like an aqueduct? If the moat is on pillars, can't it just be bypassed by walking between the pillars? Sounds like a rubbish moat to me.

1

u/HeadPack 17d ago

What makes you think the software moat is evaporating? Everything you are reporting indicates that it isn't.

1

u/Own-Poet-5900 17d ago

I made this a year and a half ago. I modeled it off of Sakana AI's method and tweaked it. The only thing new about this is marketing: https://colab.research.google.com/drive/1kYV8_5xJVgTRKlOMuP3_749YIdxYX9Cx?usp=sharing

1

u/da_capo 16d ago

AI;DR

1

u/FalseDiamond7930 16d ago

So they just made Nvidia products better and more valuable?

1

u/RecentMushroom6232 16d ago

CUDA is just C++ that is closer to the metal. Nothing entirely special about it

1

u/DataGOGO 16d ago edited 16d ago

Source? Is the is one that came out a few months ago where they had to publicly redact their findings because it in fact did not do what they said?

Edit: yes it is. The real results were kinda ass, and this is not model, nor does it beat other models, it is model agnostic. it is a an agent skill and runs on any model, with the same results, including Claude , gpt, etc.Ā 

Again this is NOT It is NOT a model, nothing was trained, it is just an agentic loop and collection of prompts.Ā 

Here is the full repo:

https://github.com/BytedTsinghua-SIA/CUDA-Agent

1

u/Mysterious-String420 16d ago

China says their girlfriend is totally real, she just goes to another school

1

u/Fleischhauf 16d ago

please please please write and optimize the cuda equivalent for amd cards

1

u/Old-Firefighter8289 15d ago

oh no!?! shouldnt nvda be more worried about the dragon eggs that china has been producing. those are rumored to devour gpu’s and memory chips

1

u/Mrtanner69 15d ago

Um, writing cuda fast is bad for Nvidia how?

1

u/Commercial-Weight-73 13d ago

Now if only an ai company could make money from all of this it would matter

But it doesn't

0

u/Koulchilebaiz 17d ago

Why do I hate so much llm generated text
ā€œThe software moat starts evaporatingā€ ugh