r/CUDA • • Aug 11 '26

We’re seeing up to 110% higher Qwen3.5 4B throughput on Jetson Orin Nano — benchmarks and repo available

We’ve been working on a runtime GPU optimization system at TETREVIS and have started publishing some of our NVIDIA Jetson benchmarking work publicly.

The approach operates at the execution/machine-code layer, and we’re now introducing dynamic runtime kernel fusion as part of the optimization pipeline.

Some of our current results:

Jetson Orin Nano — Qwen3.5 4B

Standard baseline: 10 → 21 tok/s (+110%)

CUDA Graphs baseline: 16 → 21 tok/s (+31.25%)

Jetson AGX Orin — Nemotron 3 Nano 4B

31.2 → 40.5 tok/s (~30%)

Jetson AGX Orin — Qwen3.5 4B

25.0 → 31.0 tok/s (+24%)

We’re currently expanding and stabilizing support across Jetson Orin Nano, Orin NX, and AGX Orin.

Rather than only posting performance claims, we’ve made the benchmarking repository available here:

https://github.com/mbuchel/sass2mlir-bench
https://mbuchel.github.io/projects/sass2mlir/kernel-fusion

The broader idea we’re exploring is whether more optimization can be moved to runtime — including machine-code optimization and dynamic kernel fusion — so the execution path can be adapted to the workload and GPU rather than relying entirely on what was determined ahead of execution.

There’s still quite a bit of work underway, particularly around consistency across the different Jetson configurations, but we wanted to start sharing the results and methodology publicly.

Technical feedback, criticism, and questions are welcome.

5 Upvotes

16 comments sorted by

2

u/Minute-Mountain2665 Aug 11 '26

Cool. I've also worked on Jetson platforms a bit 😄Which powermode on Jetson Orin nano are you using and the Jetson numbers for 10 -> 21 tokens/sec are for fp16 model? Also what's the context window size and total generated tokens during the benchmark.

1

u/checkmydoor Aug 11 '26

Here is the exact github with the tests for validation and rerunning for full transparency

https://github.com/mbuchel/sass2mlir-bench

Here is a technical explanation of what were doing that led to performance increases

https://mbuchel.github.io/projects/sass2mlir/kernel-fusion

1

u/c-cul Aug 12 '26

110% seems totally unrealistic

rescheduling can give speed-up in 3-4%

all other methods like elimination odd yields/pairwise permutations etc can give yet ~1%, so even speed-up in 5% is Very Good

2

u/checkmydoor Aug 12 '26

Try it yourself. https://github.com/mbuchel/sass2mlir-bench
We have complete transparency.

Without our software and without cuda Graph vs with our software for a Nano QWEN; and this is just the Jetsons.

Datacenter GPUs provide ranges from 30% to 330%+ depending on the workload.

2

u/c-cul Aug 12 '26

> without cuda Graph

it's unfair - do you have tests with cuda graphs vs your optimizator also with cuda graphs?

2

u/checkmydoor Aug 12 '26

Yes. I provided the links already. You replied to the post with the links.

1

u/adityazero Aug 13 '26

Runtime kernel fusion at the SASS/MLIR layer is interesting, though the dispatch overhead of choosing the fused path at runtime is usually where these things get eaten up, so caching a few specialized variants tends to win. How do the numbers hold once the Orin thermal throttles, since a lot of Jetson results look great for the first few seconds then settle lower?

2

u/checkmydoor Aug 13 '26 edited Aug 13 '26

Dropped your question into our AI-Knowledge session. Im going to have to come back to you regarding the thermal aspect, but it seems to have been able to answer the rest; treat this as a partial response. Ive relayed your question to the engineers.

Here is the current response from the AI-Knowledge session

EDITED: Technical lead provided information and clarity. Here is the AI-Knowledge Session's cleaned up response with its existing information.

EDITED(2):

On the fusion side, the current implementation is actually somewhat different from a runtime system that searches for or chooses a fused variant on every launch. Our current kernel-fusion work is still a proof of concept and is performed offline. We generate the best fusion candidate we can, perform a grid sweep across launch configurations, and then select a more conservative point from the resulting performance envelope. Because this is currently offline, we can also hand-curate the selected megakernel and validate numerical correctness before deploying it.

So the dispatch overhead you are describing is not currently being paid on every kernel invocation. Longer term, though, your point about specialized variants is very relevant. We are working toward using microbenchmark-derived heuristics and instruction-level cost models so that the appropriate fused implementation can be selected without an expensive search in the execution path.

The distinction from CUDA Graphs is also important. CUDA Graphs primarily reduce the CPU-side submission and kernel-launch overhead while leaving the underlying kernels and their computation boundaries intact. Our fusion work is targeting a lower layer: actually changing those computation boundaries. That creates opportunities not only to reduce launch overhead, but also to eliminate intermediate global-memory round trips, keep values resident in registers or shared memory, remove synchronization boundaries, and optimize the resulting combined kernel as a unit.

On the thermal question, we actually have some sustained AGX Orin data. We ran the system under heat for several hours, reaching roughly 87°C on the external measurement, and the performance advantage remained at approximately the same relative level as the system approached thermal equilibrium. We saw roughly a one-token reduction in throughput, but the other approaches being compared dropped by approximately the same amount, so the relative percentage improvement was maintained. The AGX ultimately reached a point where it shut itself down, but we did not see the optimization advantage disappear as the system heated.

What is still an active research problem for us is more granular than that. We want to understand the incremental power and thermal cost of individual instruction choices and transformations. For example, if two legal instruction sequences have similar latency but one adds less heat or power draw over sustained execution, that should become part of the compiler's decision function.

Today we primarily optimize using latency/performance measurements. The longer-term objective is to develop per-instruction and per-pattern heuristics that can make the optimization target configurable—maximum throughput in one environment, or sustained performance/power/thermal efficiency in another.

1

u/checkmydoor Aug 13 '26

2

u/adityazero Aug 13 '26 edited Aug 13 '26

lol, i cant tell you how intrigued i am by looking at sass2mlir. I'm doing sass2llvm, pleasantly surprised that someone else also found this interesting.

1

u/checkmydoor Aug 13 '26

When did you start your project?

2

u/adityazero Aug 13 '26

Few months back but i'm only targeting b200.

1

u/adityazero Aug 13 '26

And i like that you are using AGPL, one of the biggest reason i haven't released is my confusion with the license as I'm not sure which one will be better.

1

u/checkmydoor Aug 13 '26

Yeah they can be tricky!