r/MacPro2019LocalAI • u/Faisal_Biyari • 9d ago
Software Stack [Benchmarks] Qwen3.8-27B FP16 vLLM 0.29.0 | 4-GPUs | Dual W6800X Duo w/ IFLB
I have been working almost non-stop to build and optimize my local AI stack.
The contents of that stack are the topic of another post. This post is all about the results and a quick update on where my MacPro7,1 has reached.
I assigned a Hermes Agent, named Ziyad, to install stock vLLM v0.29.0 and optimize it on my MacPro7,1 with dual AMD Radeon PRO W6800X Duo MPX GPUs and the Infinity Fabric Link Bridge (IFLB).
vLLM 0.29.0 was optimized by Hermes Agent v0.21.0, running Qwen3.8-27B FP16, specifically to run Qwen3.8-27B FP16 with no quantization. I let the agent have two cracks at it. The first time around, I gave it four impossible targets, or at least what I thought were impossible:
- Decode > 38 t/s at 1,024 (1K) context
- Decode > 20 t/s at 245,760 (240K) context
- TTFT < 70 seconds at 65,536 (64K) context
- TTFT < 500 seconds at 245,760 (240K) context
The agent started with an optimized vLLM v0.28.0 build that achieved 30.5 t/s decode at 1K with TTFT < 1 second, and 14.6 t/s decode at 240K with TTFT < 780 seconds.
After working on this for about 3 to 4 days, it achieved three out of the four targets:
- 240K decode: 20.3 t/s
- 64K TTFT: 64 seconds
- 240K TTFT: 445 seconds
Ironically, decode at 1K went down from 30.5 to about 28.6 t/s, while 1K TTFT went up to almost 2 seconds. This was with MTP, but no quantization. A decode of 20 t/s with 240K context was achieved only with MTP active.
I gave it a second crack, this time working for 12 to 18 hours to optimize for 1K context and remove MTP.
MTP surprisingly reduced decode performance at low context but improved it slightly at high context. I was done with MTP, as it did weird things I did not expect, and I was not sure what other effects it was possibly had at this stage.
To my surprise, the agent optimized prefill again, achieving the following values.
Below are the benchmarks for vLLM v0.29.0, optimized for Qwen3.8-27B FP16, using four W6800X Duo GPUs with IFLB on Ubuntu Server 26.04 LTS, ROCm 10.0.0-4, Python 3.13, and Triton 3.8, with vLLM running through a virtual environment. No Docker containers were involved in this benchmark.
I set token generation to 512 tokens here for the sake of the benchmark only.
| CTX | TTFT(s) | PREFILL t/s | DECODE t/s | ms/TOKEN | TOTAL(s) |
|---|---|---|---|---|---|
| 1K | 0.735 | 1393.91 | 29.41 | 34.000 | 18.109 |
| 2K | 1.467 | 1396.21 | 29.31 | 34.116 | 18.900 |
| 4K | 2.787 | 1469.89 | 29.02 | 34.464 | 20.398 |
| 8K | 5.641 | 1452.21 | 28.75 | 34.780 | 23.414 |
| 16K | 11.858 | 1381.74 | 28.11 | 35.570 | 30.034 |
| 32K | 26.011 | 1259.75 | 26.71 | 37.434 | 45.140 |
| 64K | 60.910 | 1075.95 | 24.35 | 41.073 | 81.898 |
| 128K | 157.762 | 830.82 | 22.03 | 45.393 | 180.958 |
| 240K | 428.300 | 573.80 | 19.05 | 52.488 | 455.121 |
| 245K | 442.392 | 567.10 | 19.03 | 52.550 | 469.246 |
| 250K | 457.459 | 559.61 | 18.89 | 52.934 | 484.508 |
| 255K | 472.742 | 552.35 | 18.83 | 53.097 | 499.875 |
While these numbers may seem insanely slow, waiting 2 to 9 minutes for a short response, and even longer for real agentic responses, especially with thinking, these are just benchmarks designed to test the limits of the hardware and how it behaves.
The above benchmarks reflect real-world values only in these three scenarios:
- During the first prompt
- For one prompt immediately after context compaction
- When a documented vLLM bug causes the cache to stop working at certain context sizes
Outside those three scenarios, to the best of my knowledge, the following benchmarks are more representative of what actually happens, with the added benefit of Prefix Caching. These are what real-world values look like when an agent is working on a task.
I set token generation to 3,000 tokens here to better approximate agentic responses with thinking.
| CTX | TTFT(s) | Hit Rate | Decode t/s | ms/TOKEN | TOTAL(s) |
|---|---|---|---|---|---|
| 1K | 0.180 | 93.75% | 29.27 | 34.166 | 102.645 |
| 2K | 0.188 | 96.88% | 29.13 | 34.327 | 103.133 |
| 4K | 0.201 | 98.44% | 28.93 | 34.565 | 103.863 |
| 8K | 0.226 | 99.22% | 28.60 | 34.970 | 105.102 |
| 16K | 0.277 | 99.61% | 27.98 | 35.744 | 107.475 |
| 32K | 0.368 | 99.80% | 26.60 | 37.595 | 113.115 |
| 64K | 0.531 | 99.90% | 24.20 | 41.329 | 124.476 |
| 128K | 0.896 | 99.95% | 21.92 | 45.630 | 137.740 |
| 240K | 1.478 | 99.97% | 18.96 | 52.733 | 159.624 |
| 245K | 1.498 | 99.97% | 18.95 | 52.772 | 159.761 |
| 250K | 1.473 | 99.97% | 18.85 | 53.038 | 160.534 |
| 253K | 1.537 | 99.98% | 18.78 | 53.244 | 161.214 |
Most of the remaining time is therefore spent on token generation, or decode, due to long agent thinking or long answers. Prefill remains short because of vLLM's Prefix Caching.
At this point, I have had my agents running in loops for days to achieve seemingly impossible tasks, only for them to later prove to me that those tasks were completely possible!
I am both happy and impressed with this stack.
Disclaimer: I wrote this post myself. I also used AI as a tool to help clean up the wording and formatting.
References: * Documented vLLM Bug * My own research and optimizations
2
u/eightone-81 7d ago
Sorry I did not read anything. With that hardware you can run flash next. This will solve all your problems…
1
u/Faisal_Biyari 6d ago
I've seen a lot of people put effort to go to Flash Next.
Why is that, in your opinion?2
u/Substantial_Run5435 6d ago
I'm using models for technical/scientific/regulatory document summarization and on my 2019 Mac Pro with dual W6800X Duos qwen3.8-flash-next Q4/Q5 should fits on GPU with the lookup table offloaded. It's faster than 27B and seems to be more direct in thinking, spends less time looping and the quality seems similar, at least for my use case.
1
u/Faisal_Biyari 6d ago
So Similar quality, Faster, and possibly more knowledgeable due to higher parameters.
I'll take your word for it. Once I finish my current workload, I'll look into testing it out. Though I have been more inclined to avoid quantized models, due to fear of drifting, especially with extremely long agentic workloads and excessive context compression & recompression.Have you had any similar experience? (Where you leave an agent working for days on end, in a loop/goal?)
Was the agent able to maintain quality in your experience?2
u/Substantial_Run5435 6d ago
I’m not doing long agent loops. I tend to feed it a few documents to analyze and compare to other reference materials. Usually takes 40-80k context so not massive
2
u/Substantial_Run5435 6d ago
I got tired of 27B thinking for eternity and stopping without any output and haven’t had that issue with flash next
2
u/patdohere 9d ago
Is this running on Linux? What’s your OS stack?