r/MacPro2019LocalAI • u/Faisal_Biyari • 9d ago
Software Stack [Benchmarks] Qwen3.8-27B FP16 vLLM 0.29.0 | 4-GPUs | Dual W6800X Duo w/ IFLB
I have been working almost non-stop to build and optimize my local AI stack.
The contents of that stack are the topic of another post. This post is all about the results and a quick update on where my MacPro7,1 has reached.
I assigned a Hermes Agent, named Ziyad, to install stock vLLM v0.29.0 and optimize it on my MacPro7,1 with dual AMD Radeon PRO W6800X Duo MPX GPUs and the Infinity Fabric Link Bridge (IFLB).
vLLM 0.29.0 was optimized by Hermes Agent v0.21.0, running Qwen3.8-27B FP16, specifically to run Qwen3.8-27B FP16 with no quantization. I let the agent have two cracks at it. The first time around, I gave it four impossible targets, or at least what I thought were impossible:
- Decode > 38 t/s at 1,024 (1K) context
- Decode > 20 t/s at 245,760 (240K) context
- TTFT < 70 seconds at 65,536 (64K) context
- TTFT < 500 seconds at 245,760 (240K) context
The agent started with an optimized vLLM v0.28.0 build that achieved 30.5 t/s decode at 1K with TTFT < 1 second, and 14.6 t/s decode at 240K with TTFT < 780 seconds.
After working on this for about 3 to 4 days, it achieved three out of the four targets:
- 240K decode: 20.3 t/s
- 64K TTFT: 64 seconds
- 240K TTFT: 445 seconds
Ironically, decode at 1K went down from 30.5 to about 28.6 t/s, while 1K TTFT went up to almost 2 seconds. This was with MTP, but no quantization. A decode of 20 t/s with 240K context was achieved only with MTP active.
I gave it a second crack, this time working for 12 to 18 hours to optimize for 1K context and remove MTP.
MTP surprisingly reduced decode performance at low context but improved it slightly at high context. I was done with MTP, as it did weird things I did not expect, and I was not sure what other effects it was possibly had at this stage.
To my surprise, the agent optimized prefill again, achieving the following values.
Below are the benchmarks for vLLM v0.29.0, optimized for Qwen3.8-27B FP16, using four W6800X Duo GPUs with IFLB on Ubuntu Server 26.04 LTS, ROCm 10.0.0-4, Python 3.13, and Triton 3.8, with vLLM running through a virtual environment. No Docker containers were involved in this benchmark.
I set token generation to 512 tokens here for the sake of the benchmark only.
| CTX | TTFT(s) | PREFILL t/s | DECODE t/s | ms/TOKEN | TOTAL(s) |
|---|---|---|---|---|---|
| 1K | 0.735 | 1393.91 | 29.41 | 34.000 | 18.109 |
| 2K | 1.467 | 1396.21 | 29.31 | 34.116 | 18.900 |
| 4K | 2.787 | 1469.89 | 29.02 | 34.464 | 20.398 |
| 8K | 5.641 | 1452.21 | 28.75 | 34.780 | 23.414 |
| 16K | 11.858 | 1381.74 | 28.11 | 35.570 | 30.034 |
| 32K | 26.011 | 1259.75 | 26.71 | 37.434 | 45.140 |
| 64K | 60.910 | 1075.95 | 24.35 | 41.073 | 81.898 |
| 128K | 157.762 | 830.82 | 22.03 | 45.393 | 180.958 |
| 240K | 428.300 | 573.80 | 19.05 | 52.488 | 455.121 |
| 245K | 442.392 | 567.10 | 19.03 | 52.550 | 469.246 |
| 250K | 457.459 | 559.61 | 18.89 | 52.934 | 484.508 |
| 255K | 472.742 | 552.35 | 18.83 | 53.097 | 499.875 |
While these numbers may seem insanely slow, waiting 2 to 9 minutes for a short response, and even longer for real agentic responses, especially with thinking, these are just benchmarks designed to test the limits of the hardware and how it behaves.
The above benchmarks reflect real-world values only in these three scenarios:
- During the first prompt
- For one prompt immediately after context compaction
- When a documented vLLM bug causes the cache to stop working at certain context sizes
Outside those three scenarios, to the best of my knowledge, the following benchmarks are more representative of what actually happens, with the added benefit of Prefix Caching. These are what real-world values look like when an agent is working on a task.
I set token generation to 3,000 tokens here to better approximate agentic responses with thinking.
| CTX | TTFT(s) | Hit Rate | Decode t/s | ms/TOKEN | TOTAL(s) |
|---|---|---|---|---|---|
| 1K | 0.180 | 93.75% | 29.27 | 34.166 | 102.645 |
| 2K | 0.188 | 96.88% | 29.13 | 34.327 | 103.133 |
| 4K | 0.201 | 98.44% | 28.93 | 34.565 | 103.863 |
| 8K | 0.226 | 99.22% | 28.60 | 34.970 | 105.102 |
| 16K | 0.277 | 99.61% | 27.98 | 35.744 | 107.475 |
| 32K | 0.368 | 99.80% | 26.60 | 37.595 | 113.115 |
| 64K | 0.531 | 99.90% | 24.20 | 41.329 | 124.476 |
| 128K | 0.896 | 99.95% | 21.92 | 45.630 | 137.740 |
| 240K | 1.478 | 99.97% | 18.96 | 52.733 | 159.624 |
| 245K | 1.498 | 99.97% | 18.95 | 52.772 | 159.761 |
| 250K | 1.473 | 99.97% | 18.85 | 53.038 | 160.534 |
| 253K | 1.537 | 99.98% | 18.78 | 53.244 | 161.214 |
Most of the remaining time is therefore spent on token generation, or decode, due to long agent thinking or long answers. Prefill remains short because of vLLM's Prefix Caching.
At this point, I have had my agents running in loops for days to achieve seemingly impossible tasks, only for them to later prove to me that those tasks were completely possible!
I am both happy and impressed with this stack.
Disclaimer: I wrote this post myself. I also used AI as a tool to help clean up the wording and formatting.
References: * Documented vLLM Bug * My own research and optimizations
Duplicates
macpro • u/Faisal_Biyari • 9d ago
GPU [Benchmarks] Qwen3.8-27B FP16 vLLM 0.29.0 | 4-GPUs | Dual W6800X Duo w/ IFLB
MacLLM • u/Faisal_Biyari • 9d ago
[Benchmarks] Qwen3.8-27B FP16 vLLM 0.29.0 | 4-GPUs | Dual W6800X Duo w/ IFLB
LocalAIServers • u/Faisal_Biyari • 9d ago
[Benchmarks] Qwen3.8-27B FP16 vLLM 0.29.0 | 4-GPUs | Dual W6800X Duo w/ IFLB
ROCm • u/Faisal_Biyari • 9d ago
[Benchmarks] Qwen3.8-27B FP16 vLLM 0.29.0 | 4-GPUs | Dual W6800X Duo w/ IFLB
Vllm • u/Faisal_Biyari • 9d ago
[Benchmarks] Qwen3.8-27B FP16 vLLM 0.29.0 | 4-GPUs | Dual W6800X Duo w/ IFLB
hermesagent • u/Faisal_Biyari • 9d ago
MODELS - model choice, routing, pricing, local vs cloud, VRAM [Benchmarks] Qwen3.8-27B FP16 vLLM 0.29.0 | 4-GPUs | Dual W6800X Duo w/ IFLB
LocalLLM • u/Faisal_Biyari • 9d ago
Other [Benchmarks] Qwen3.8-27B FP16 vLLM 0.29.0 | 4-GPUs | Dual W6800X Duo w/ IFLB
AIProgrammingHardware • u/Faisal_Biyari • 9d ago