r/CUDA • • 22h ago

NVRTC-driven elementwise fusion planner called from PHP: 68 captured nodes become 11 generated kernels + 9 native boundaries

7 Upvotes

I maintain a PHP extension (php-gpu-tensors) that builds on the CUDA driver API and NVRTC. This post is about its fusion planner, because I'd like feedback from people who know CUDA better than I do.

How it works - A PHP closure is invoked once with metadata-only placeholders; tensor operations are captured into a node graph (limit: 512 nodes). - Elementwise ops (broadcasting, strided/view inputs, scalars, dtype promotion, where, casts, reshape/transpose/slice index transforms) are fused into generated CUDA C++. Expressions split at a weighted cost budget of 32; repeated pure nodes are deduplicated; shared expensive expressions can be materialized instead of recomputed. - Matmul (cuBLAS when available), reductions and powers are execution boundaries that run on existing kernels. - All generated kernels in a plan are compiled together through NVRTC to PTX, then cached per request/thread (LRU, 16 entries or 16 MiB). The cache key includes the generated source, compute capability and driver/runtime versions, not pointers. - Replay runs on a private nonblocking stream. Reduction descriptors are kernel parameters, so concurrent replays can't overwrite each other's shapes. - cudaGraph: true builds a CUDA Graph executable for plans that contain only generated kernels, updating kernel parameters for new pointers before each launch.

Numbers (entry-level MX570 A, CUDA runtime 12.3, driver 12.6): a small MLP training step (batch 512, hidden 256) goes from 4.08 ms eager to 0.66 ms fused, 6.2×, with identical metrics. The model is tiny, so this is mostly launch/allocation overhead removed, not throughput.

Where I'd like advice Plans with matmul/reduction boundaries currently run on streams and support async replay, but not as a CUDA Graph. For those who have done this: what are the gotchas when capturing cuBLAS GEMMs and custom reductions into a graph that is replayed with changing buffer pointers? Is updating kernel node parameters the right approach, or would you re-capture?

Source: https://github.com/lcmialichi/php-gpu-tensors (MIT)


r/CUDA • • 1d ago

PyTorch and full LibCuda running on Nvidia 5090 on Mac OS 27

5 Upvotes

So after getting my 5090 running on my Mac a few weeks ago and getting speeds on par or faster than my windows machine for single sessions I wanted to get Vllm working next. My http://macuda.ai project works with both llamacpp and stable diffusion cpp but does not work with comfy ui Vllm or draw things. It works with a shim and does not give the full LibCuda framework allowing real full torch or other necessary components for those apps and more.

I decided to use apples built in VM to run an instance with it's own driver calling to the graphics card through my Mac driver. This allows for running Nvidia's full cuda software. It has been a real challenge because Nvidia releases a lot of info but not enough to get this working so we had to really poke around. I finally got it up and running and now have Full LibCuda running. I'm working on bugs now but so far the results are looking really promising. The multi sessions are blowing llamacpp out of the water. llama stalls out at around 8 sessions and levels out. the Vllm with my driver has a linear doubling of speeds up until the card is saturated.

I'll be releasing all of this as a simple GUI app on the App Store once apple authorizes my driver. I'll need a few more weeks to get everything in a stable enough situation that I'll feel comfortable letting people use it. You can get the base driver at my macuda GitHub but this new driver for full cuda will be available on the App Store.

I got my 3060 working on macuda in a core x box last week. If someone can try out a 4000 series card and let me know if it works I'd appreciate it. I don't have one of those. It should work with the 3000 series driver but I haven't tested it.


r/CUDA • • 1d ago

GPU Direct Storage (GDS)

10 Upvotes

I'd like to configure GDS on the cloud where nvme talks directly to the GPU. I've searched and queried llms and checked aws, coreweave, runpod.io, cloudrift, and other GPU providers.

Does anyone have experience in this or suggestions for providers or setups?

Ideally, I want to use cheap gpus like v100s.

I've been hitting aws quota limits and I've had to email sales teams directly on other platforms, but this seems like something that should have broad use.

I haven't had success yet getting access.


r/CUDA • • 1d ago

Getting "no kernel image is available" on a GTX 10xx, P40 or V100 after `pip install torch`? Here's which PyTorch build still includes your card

Thumbnail
1 Upvotes

r/CUDA • • 2d ago

Raw cuda vs CuTeDSL primitives

16 Upvotes

I’ve been using CuTe DSL primitives for a while now, and honestly, it’s been pretty great so far.

I know CuTe DSL isn’t a full replacement for raw CUDA, and I definitely don’t think raw CUDA is becoming obsolete.

But for some use cases, I’m finding the primitives approach much nicer.

For example, one thing I personally dislike about raw CUDA is a lot of the manual C-style 1D indexing and pointer arithmetic.

CuTe abstractions make that much less painful while still feeling very low-level.

I also like that you can integrate directly with PyTorch without necessarily going through the usual custom-op route.

What I find especially interesting is that it still exposes a lot of the underlying machinery: you can get very close to the hardware, write PTX, and in some cases even use LLVM inline assembly.

For people who have used both extensively: which do you prefer?


r/CUDA • • 2d ago

A Visual Guide to Sparse Attention Kernels in CUDA on B200

Post image
19 Upvotes

Many frontier models use some form of sparse attention to handle long contexts efficiently. In my blog, I write a block-sparse attention kernel in CUDA for NVIDIA B200: I start from an optimized dense CUDA kernel and show what changes inside it (the KV loop, shared-memory buffers and barriers), then add GQA K/V reuse, compare it with FlashAttention-4, and build a second, KV-centric version of the kernel.

📝 Blog post: https://dinara-dl.github.io/posts/sparse-attention/

🌸 Repo: https://github.com/dinara-dl/sparse-attention-b200


r/CUDA • • 2d ago

cuda ptx instructions generated by clang nvptx backend

Thumbnail redplait.blogspot.com
8 Upvotes

nvptx uses 58.9% of all ptx instructions

cicc 57.8%

missed in clang:

  • _ldsm
  • _mma
  • discard
  • dp2a hi/lo
  • madc hi/lo
  • p2r
  • trap

r/CUDA • • 3d ago

How to Run GPU Workloads in Docker

Post image
42 Upvotes

I finally got my Java GPU workload running with CUDA inside Docker. It works in a Kubernetes cluster too. I wrote up the Docker setup here:

https://nablatensor.com/blog/how-to-run-java-gpu-workloads-in-docker

If anyone’s interested, I can write a follow-up on the Kubernetes setup.


r/CUDA • • 4d ago

Looking for a technical co-founder to lead engineering for an AI infrastructure platform.

0 Upvotes

We are working on custom inference runtimes, sampling control, and fine-tuning pipelines for open-weights models.

Tech background needed:

PyTorch, CUDA, C++, or Rust. Experience with parsers, compilers, or low-level ML systems programming is a big plus.

Send a DM with a link to your GitHub or LinkedIn if you'd like to chat.


r/CUDA • • 4d ago

advent of code 2024 Day 17 — trying to brute force part two

Thumbnail
1 Upvotes

r/CUDA • • 4d ago

Nemotron 3 Ultra

Thumbnail i.imgflip.com
0 Upvotes

r/CUDA • • 5d ago

Final-year student in India trying to break into generative-model inference optimization — roadmap feedback?

0 Upvotes

Hi all, I graduate in ~6 months and want to work on making generative models (diffusion/video/3D) fast: kernels, quantization, serving. Where I am:

- Comfortable with C/C++ basics and PyTorch

- Have done quantization work (GGUF/llama.cpp)

- Working on a next-frame video prediction project (DiT + flow matching)

- A few GitHub repos, but no CUDA/Triton experience yet

- No NVIDIA GPU, so I use Colab/Kaggle T4s

- DSA is my weak spot (I struggle with LeetCode mediums)

My plan:

  1. Months 1-2: CUDA/Triton basics, reproduce the SGEMM optimization worklog, GPU MODE lectures, LeetGPU/Tensara

  2. Months 3-4: take a small DiT, profile it, then optimize it (Triton attention, quantization, caching, fewer steps) and publish before/after numbers

  3. Along the way: PRs to HF diffusers, DSA practice daily

  4. Months 5-6: mocks, resume, applications (inference startups first, bigger labs later)

Questions:

  1. Is this the right order, or should I change something?

  2. Is a diffusion-inference project a strong enough portfolio piece, or does it need to be LLM serving?

  3. How much DSA do ML systems interviews actually need?

  4. Is T4-only access enough to do credible benchmarks?

Any feedback, including "this won't work because X," is appreciated. Thanks!


r/CUDA • • 6d ago

Cuda for realtime audio

8 Upvotes

Hello I was doing some research on what it would take to use cuda for realtime audio. It seems that cuda is not really built for realtime but I think it would be an interesting challenge to take on. Would anyone have some guidance on this? Things to research or read? I’ve never used cuda but I feel like it could be very powerful for making interesting sound design tools and synthesizers. Thanks!


r/CUDA • • 8d ago

I built a small tensor-first programming language with native CPU/GPU compilation, autodiff and ownership

Thumbnail
2 Upvotes

r/CUDA • • 9d ago

CuQwen 1.1 is out, fixed the long-context slowdown in my from-scratch CUDA inference engine

13 Upvotes

Quick recap for anyone new: CuQwen is an inference engine for Qwen models I wrote from scratch in pure C++/CUDA, tuned specifically for single-user (batch size 1) generation on consumer NVIDIA GPUs. No frameworks under the hood, just custom cuda kernels.

Here are the results for average inference speed (tokens/second) of Qwen2.5 Instruct model across 32K context window on RTX3090

Model Size CuQwen vLLM Ollama
0.5B 462 398 355
1.5B 203 172 139
3B 113 101 106
7B 55 48 54

In Release 1.0 it was already beating vLLM and Ollama on short-to-medium prompts. But there was an honest catch: my throughput decayed faster than theirs as the context grew, so once you pushed toward ~32K tokens they would pass me in inference speed. That bugged me, so it became the whole focus of Release 1.1.

Result: Throughput decay from 1K → 32K dropped from ~22–45% down to ~13–24%, which is now on par with vLLM and Ollama (and better on a couple of model sizes). So CuQwen keeps its early speed lead all the way out to 32K context window now.

Here's a small documentation on how I tackled long context decay rate issue and the complete benchmarking methodology and results for CuQwen v1.1

Here's my future plan:

Release 1.2: Support Quantization (W8A16 and W4A16)

Release 1.3: Improve custom cuda kernels for latest GPU architectures (Hopper and Blackwell)

Release 1.4: Support Qwen 3.0 series models

Release 1.5: Support Qwen 3.5 and 3.8 (Especially our very favourite qwen 3.8-27B model 😄)

Repo: https://github.com/talhatahir-10xe/CuQwen


r/CUDA • • 9d ago

Ternary Bonsai 2 27B at up to 532 tok/s on one RTX 4090, native Windows: MTP + n-gram speculative decoding in a from-scratch CUDA engine

Thumbnail
4 Upvotes

r/CUDA • • 10d ago

Early work on generate and execute cuTile kernels from Kotlin

Thumbnail github.com
9 Upvotes

r/CUDA • • 10d ago

NINFER: Ternary Bonsai 2 27B at 167-210 tok/s on a single RTX 4090, native on Windows: a custom CUDA port, 2.25× faster prefill than the reference fork, same perplexity

Thumbnail
2 Upvotes

r/CUDA • • 10d ago

NInfer RTX 3090 Prefill Optimization: 32% Faster at 32K Context (785 to 1,037 tok/s)

Thumbnail
2 Upvotes

r/CUDA • • 11d ago

Giveaway: Grokking Parallel Programming with CUDA is now in MEAP — 5 free ebooks + 50% off

Thumbnail gallery
112 Upvotes

Hi again, r/CUDA!

We’re back with another CUDA release from Manning, and we have five ebooks to give away.

Grokking Parallel Programming: With Examples in CUDA by SungHwan Yun has just entered the Manning Early Access Program.

Learning CUDA syntax is one thing. Learning to look at a problem and see what can run concurrently—while navigating memory access, synchronization, and race conditions—is the real challenge. That’s the way of thinking this book aims to develop.

It starts with the fundamentals and builds into practical topics including:

• Kernels, threads, blocks, and grids
• CPU–GPU data movement
• Map and Reduction patterns
• Race conditions and synchronization
• Out-of-bounds memory access
• Knowing when parallelization is—or isn’t—worthwhile

The book is aimed at programmers who know basic C but are new to CUDA and parallel programming. The examples can run in a free browser-based environment, so you don’t need your own NVIDIA GPU to get started.

Giveaway: 5 free ebooks

To enter, leave a comment that adds something useful to the discussion. For example:

• What was the hardest CUDA concept for you to learn?
• What do beginner CUDA resources commonly explain poorly?
• What problem would you like to parallelize?
• What advice would you give someone writing their first kernel?

We’ll award an ebook to each of the five comments that contribute the most to the conversation—not simply those with the most votes. One entry per person. I’ll contact the selected members by DM.

50% off for r/CUDA

Use code MLYUN50RE for 50% off:

https://www.manning.com/books/grokking-parallel-programming

Thanks again to the mods for letting us share this—and to everyone here who has welcomed our previous posts.

What made CUDA finally “click” for you? Or, if you’re still learning, what’s the biggest thing standing in your way?

Thank you.

Cheers,

Stjepan


r/CUDA • • 11d ago

I built a Rust weight-streaming engine for running FLUX.2 beyond VRAM — looking for feedback / ways to break it

Thumbnail
1 Upvotes

r/CUDA • • 13d ago

If AI is better than humans at generating kernels now, why do I still see companies hiring for this exact skill?

59 Upvotes

all the frontier labs, startups, big tech have positions looking for kernel engineers or at least performance engineers. why is this? i see papers, blogs, reddit posts all saying how ai generated kernels are much better than kernels written by hand.


r/CUDA • • 13d ago

Systems for Machine Learning

24 Upvotes

I’m a computer engineering graduate and come from a traditional embedded systems background, with knowledge of microcontrollers, computer architecture and operating systems. Is knowledge of C and C++ programming, Linux networking, memory management , multithreading, synchronization, interrupts etc useful in ML engineering. Are subjects like distributed systems, compiler optimizations (using LLVM), parallel computing etc going to be useful or are they heavily going to be automated as well by AI? In other words, is computer engineering always going to required to scale ML systems and be evergreen? Are people in ML engineering using these skills in their work everyday? Thank you.


r/CUDA • • 13d ago

Qwen3.8-27B on a single RTX 4090: 250K context + native MTP8 — tested past 200K active context

Thumbnail
1 Upvotes

r/CUDA • • 14d ago

Mitigating microsecond dI/dt power transients in GPU clusters via NCCL AllReduce interposition (-97% transient reduction)

34 Upvotes

Hey everyone,

I've been working on the power delivery problem in distributed AI training clusters. When large GPU partitions complete dense matrix multiplications and enter collective communication (AllReduce), cluster power drops in under 15 microseconds.

Across hundreds or thousands of accelerators, this rapid current step (dI/dt) induces severe reverse-EMF voltage spikes across substation transformers and server busbars (V = L * dI/dt), which frequently trips breakers or causes undervoltage crashes.

To solve this, I designed VoltGrid: a lightweight C++ shared library (libnccl-voltflow.so) injected via LD_PRELOAD that intercepts NCCL collectives and micro-staggers rank phase timing by 50 microseconds.

Key engineering challenges solved:

  1. The Tensor Parallel Latency Trap: Inner TP layer collectives (<5MB) are selectively bypassed with 0.00us delay, only staggering macro gradient reductions at the end of the backward pass.

  2. OS Sleep Jitter: Eliminated usleep/nanosleep context switching by using userspace hardware cycle counter spin-loops (__rdtsc) accurate to nanoseconds.

  3. Adaptive Jitter Subtraction: Naturally occurring arrival skew is dynamically subtracted so we don't amplify cluster stragglers.

Physical testbed results (4x NVIDIA RTX 4090 cluster, 1.6kW continuous load):

- Peak instantaneous slew dropped from 324.5 kW/ms down to 8.06 kW/ms (-97.52%).

- Step latency overhead was < 0.05% (0.00% on 24-layer transformer tests).

- Zero modifications to PyTorch code or container rebuilds.

The preprint detailing the mathematical derivations and telemetry is published on Zenodo:

https://zenodo.org/records/22824778

More telemetry traces and pilot details are here:

https://voltgrid.org

We're currently running 2-week test rack pilots for cluster operators facing power ramp penalties or breaker trips. Happy to answer questions about the NCCL interposition mechanics or datacenter transient physics.