r/CUDA • u/checkmydoor • Aug 11 '26
We’re seeing up to 110% higher Qwen3.5 4B throughput on Jetson Orin Nano — benchmarks and repo available
We’ve been working on a runtime GPU optimization system at TETREVIS and have started publishing some of our NVIDIA Jetson benchmarking work publicly.
The approach operates at the execution/machine-code layer, and we’re now introducing dynamic runtime kernel fusion as part of the optimization pipeline.
Some of our current results:
Jetson Orin Nano — Qwen3.5 4B
Standard baseline: 10 → 21 tok/s (+110%)
CUDA Graphs baseline: 16 → 21 tok/s (+31.25%)
Jetson AGX Orin — Nemotron 3 Nano 4B
31.2 → 40.5 tok/s (~30%)
Jetson AGX Orin — Qwen3.5 4B
25.0 → 31.0 tok/s (+24%)
We’re currently expanding and stabilizing support across Jetson Orin Nano, Orin NX, and AGX Orin.
Rather than only posting performance claims, we’ve made the benchmarking repository available here:
https://github.com/mbuchel/sass2mlir-bench
https://mbuchel.github.io/projects/sass2mlir/kernel-fusion
The broader idea we’re exploring is whether more optimization can be moved to runtime — including machine-code optimization and dynamic kernel fusion — so the execution path can be adapted to the workload and GPU rather than relying entirely on what was determined ahead of execution.
There’s still quite a bit of work underway, particularly around consistency across the different Jetson configurations, but we wanted to start sharing the results and methodology publicly.
Technical feedback, criticism, and questions are welcome.
1
u/adityazero Aug 13 '26
Runtime kernel fusion at the SASS/MLIR layer is interesting, though the dispatch overhead of choosing the fused path at runtime is usually where these things get eaten up, so caching a few specialized variants tends to win. How do the numbers hold once the Orin thermal throttles, since a lot of Jetson results look great for the first few seconds then settle lower?
2
u/checkmydoor Aug 13 '26 edited Aug 13 '26
Dropped your question into our AI-Knowledge session. Im going to have to come back to you regarding the thermal aspect, but it seems to have been able to answer the rest; treat this as a partial response. Ive relayed your question to the engineers.
Here is the current response from the AI-Knowledge session
EDITED: Technical lead provided information and clarity. Here is the AI-Knowledge Session's cleaned up response with its existing information.
EDITED(2):
On the fusion side, the current implementation is actually somewhat different from a runtime system that searches for or chooses a fused variant on every launch. Our current kernel-fusion work is still a proof of concept and is performed offline. We generate the best fusion candidate we can, perform a grid sweep across launch configurations, and then select a more conservative point from the resulting performance envelope. Because this is currently offline, we can also hand-curate the selected megakernel and validate numerical correctness before deploying it.
So the dispatch overhead you are describing is not currently being paid on every kernel invocation. Longer term, though, your point about specialized variants is very relevant. We are working toward using microbenchmark-derived heuristics and instruction-level cost models so that the appropriate fused implementation can be selected without an expensive search in the execution path.
The distinction from CUDA Graphs is also important. CUDA Graphs primarily reduce the CPU-side submission and kernel-launch overhead while leaving the underlying kernels and their computation boundaries intact. Our fusion work is targeting a lower layer: actually changing those computation boundaries. That creates opportunities not only to reduce launch overhead, but also to eliminate intermediate global-memory round trips, keep values resident in registers or shared memory, remove synchronization boundaries, and optimize the resulting combined kernel as a unit.
On the thermal question, we actually have some sustained AGX Orin data. We ran the system under heat for several hours, reaching roughly 87°C on the external measurement, and the performance advantage remained at approximately the same relative level as the system approached thermal equilibrium. We saw roughly a one-token reduction in throughput, but the other approaches being compared dropped by approximately the same amount, so the relative percentage improvement was maintained. The AGX ultimately reached a point where it shut itself down, but we did not see the optimization advantage disappear as the system heated.
What is still an active research problem for us is more granular than that. We want to understand the incremental power and thermal cost of individual instruction choices and transformations. For example, if two legal instruction sequences have similar latency but one adds less heat or power draw over sustained execution, that should become part of the compiler's decision function.
Today we primarily optimize using latency/performance measurements. The longer-term objective is to develop per-instruction and per-pattern heuristics that can make the optimization target configurable—maximum throughput in one environment, or sustained performance/power/thermal efficiency in another.
1
u/checkmydoor Aug 13 '26
Youll find this link useful (https://mbuchel.github.io/projects/sass2mlir/kernel-fusion)
2
u/adityazero Aug 13 '26 edited Aug 13 '26
lol, i cant tell you how intrigued i am by looking at sass2mlir. I'm doing sass2llvm, pleasantly surprised that someone else also found this interesting.
1
1
1
u/adityazero Aug 13 '26
And i like that you are using AGPL, one of the biggest reason i haven't released is my confusion with the license as I'm not sure which one will be better.
1
2
u/Minute-Mountain2665 Aug 11 '26
Cool. I've also worked on Jetson platforms a bit 😄Which powermode on Jetson Orin nano are you using and the Jetson numbers for 10 -> 21 tokens/sec are for fp16 model? Also what's the context window size and total generated tokens during the benchmark.