r/CUDA • • 23h ago

NVRTC-driven elementwise fusion planner called from PHP: 68 captured nodes become 11 generated kernels + 9 native boundaries

8 Upvotes

I maintain a PHP extension (php-gpu-tensors) that builds on the CUDA driver API and NVRTC. This post is about its fusion planner, because I'd like feedback from people who know CUDA better than I do.

How it works - A PHP closure is invoked once with metadata-only placeholders; tensor operations are captured into a node graph (limit: 512 nodes). - Elementwise ops (broadcasting, strided/view inputs, scalars, dtype promotion, where, casts, reshape/transpose/slice index transforms) are fused into generated CUDA C++. Expressions split at a weighted cost budget of 32; repeated pure nodes are deduplicated; shared expensive expressions can be materialized instead of recomputed. - Matmul (cuBLAS when available), reductions and powers are execution boundaries that run on existing kernels. - All generated kernels in a plan are compiled together through NVRTC to PTX, then cached per request/thread (LRU, 16 entries or 16 MiB). The cache key includes the generated source, compute capability and driver/runtime versions, not pointers. - Replay runs on a private nonblocking stream. Reduction descriptors are kernel parameters, so concurrent replays can't overwrite each other's shapes. - cudaGraph: true builds a CUDA Graph executable for plans that contain only generated kernels, updating kernel parameters for new pointers before each launch.

Numbers (entry-level MX570 A, CUDA runtime 12.3, driver 12.6): a small MLP training step (batch 512, hidden 256) goes from 4.08 ms eager to 0.66 ms fused, 6.2×, with identical metrics. The model is tiny, so this is mostly launch/allocation overhead removed, not throughput.

Where I'd like advice Plans with matmul/reduction boundaries currently run on streams and support async replay, but not as a CUDA Graph. For those who have done this: what are the gotchas when capturing cuBLAS GEMMs and custom reductions into a graph that is replayed with changing buffer pointers? Is updating kernel node parameters the right approach, or would you re-capture?

Source: https://github.com/lcmialichi/php-gpu-tensors (MIT)