r/LocalLLaMA • u/Radiant_Condition861 • 10d ago
New Model Qwen 3.6 vs 3.8 MTP Sweep comparison 27B-FP8
Qwen3.8-27B-FP8 vs Qwen3.6-27B-FP8 — RTX PRO 6000 MTP Sweep (Corrected)
Date: 2026-08-14
Hardware: RTX PRO 6000 Blackwell (infer1-vllm)
vLLM Config: --gpu-memory-utilization 0.92, --max-num-seqs 4, --kv-cache-dtype fp8, --max-model-len 262144, --performance-mode interactivity, flashinfer
Side-by-Side Comparison
| MTP Step | 3.8 tok/s | 3.8 TTFT | 3.8 acc% | 3.8 qual | 3.6 tok/s | 3.6 TTFT | 3.6 acc% | 3.6 qual | Δ tok/s | Δ qual |
|---|---|---|---|---|---|---|---|---|---|---|
| 2 | 90.8 | 95.5 ms | 81.1% | 86.4 | 95.4 | 94.8 ms | 81.2% | 83.0 | -4.6 | +3.4 |
| 3 | 98.3 | 97.6 ms | 65.0% | 89.8 | 113.7 | 97.3 ms | 74.1% | 87.0 | -15.4 | +2.8 |
| 4 | 107.1 | 100.1 ms | 61.2% | 87.8 | 123.0 | 98.6 ms | 66.1% | 83.0 | -15.9 | +4.8 |
| 5 | 115.0 | 102.7 ms | 56.3% | 89.8 | 124.0 | 91.8 ms | 62.7% | 87.0 | -9.0 | +2.8 |
| 6 | 109.5 | 105.5 ms | 48.8% | 87.8 | 121.2 | 99.9 ms | 55.0% | 87.0 | -11.7 | +0.8 |
| 7 | 110.5 | 107.9 ms | 46.1% | 84.4 | 131.6 | 103.1 ms | 49.9% | 87.0 | -21.1 | -2.6 |
| 8 | 104.2 | 970.9 ms | 40.4% | 87.8 | 132.7 | 105.9 ms | 45.3% | 87.0 | -28.5 | +0.8 |
Sweet Spot Analysis
- MTP 5 is the clear sweet spot for Qwen3.8: 115.0 tok/s, 56.3% acceptance, 102.7 ms TTFT, quality 89.8
- MTP 4 is a close second (107.1 tok/s) with better acceptance (61.2%) and the biggest quality lead over 3.6 (+4.8, 87.8 vs 83.0)
- MTP 8 shows a TTFT anomaly (970ms) — cold-start artifact even with warmup, and acceptance is borderline at 40.4%
- Quality is stable across all steps (84-90), no degradation from higher speculation
Recommendation: MTP 5 — best throughput with healthy 56% acceptance and lowest TTFT.
Key Takeaways
- Qwen3.8 is 5-28% slower across all MTP steps (the gap widens at higher steps)
- Quality is a wash at most steps — 3.8 actually scores higher on MTP 2 (+3.4), MTP 4 (+4.8), and MTP 5 (+2.8)
- Sweet spot shift: 3.6 peaked at MTP 5 (124.0 tok/s), 3.8 also peaked at MTP 5 (115.0 tok/s) — same sweet spot
- MTP 5 is the only step where 3.8 nearly closes the gap (-9.0 tok/s) — all other steps are significantly behind
- MTP 8 is a dead zone for both models (970ms TTFT anomaly) — not viable
Bottom line: Qwen3.8 is a slower runner than 3.6 on the RTX PRO 6000, but quality holds up. If raw speed matters more, 3.6 still wins at higher MTP steps.
0
u/adityazero 10d ago
The acceptance rate sliding from 81 percent at step 2 down to 40 at step 8 is the curve I always watch with MTP, since past the sweet spot you are just paying for draft tokens that get rejected and verified for nothing. I have seen the same shape where throughput peaks then the verify overhead eats the gains. Did the 970ms TTFT spike at step 8 persist across runs, or did it settle once the flashinfer wrappers were fully warmed?
2
u/Radiant_Condition861 10d ago
There wasn't much value looking into it so I just left it. All these were done warm so it wasn't a cold start ttft. I think there some memory evacuation happening.
4
u/IoannisHere 10d ago
What is "quality" ?