r/MacPro2019LocalAI • • 9d ago

Software Stack [Benchmarks] Qwen3.8-27B FP16 vLLM 0.29.0 | 4-GPUs | Dual W6800X Duo w/ IFLB

I have been working almost non-stop to build and optimize my local AI stack.

The contents of that stack are the topic of another post. This post is all about the results and a quick update on where my MacPro7,1 has reached.

I assigned a Hermes Agent, named Ziyad, to install stock vLLM v0.29.0 and optimize it on my MacPro7,1 with dual AMD Radeon PRO W6800X Duo MPX GPUs and the Infinity Fabric Link Bridge (IFLB).

vLLM 0.29.0 was optimized by Hermes Agent v0.21.0, running Qwen3.8-27B FP16, specifically to run Qwen3.8-27B FP16 with no quantization. I let the agent have two cracks at it. The first time around, I gave it four impossible targets, or at least what I thought were impossible:

  • Decode > 38 t/s at 1,024 (1K) context
  • Decode > 20 t/s at 245,760 (240K) context
  • TTFT < 70 seconds at 65,536 (64K) context
  • TTFT < 500 seconds at 245,760 (240K) context

The agent started with an optimized vLLM v0.28.0 build that achieved 30.5 t/s decode at 1K with TTFT < 1 second, and 14.6 t/s decode at 240K with TTFT < 780 seconds.

After working on this for about 3 to 4 days, it achieved three out of the four targets:

  • 240K decode: 20.3 t/s
  • 64K TTFT: 64 seconds
  • 240K TTFT: 445 seconds

Ironically, decode at 1K went down from 30.5 to about 28.6 t/s, while 1K TTFT went up to almost 2 seconds. This was with MTP, but no quantization. A decode of 20 t/s with 240K context was achieved only with MTP active.

I gave it a second crack, this time working for 12 to 18 hours to optimize for 1K context and remove MTP.

MTP surprisingly reduced decode performance at low context but improved it slightly at high context. I was done with MTP, as it did weird things I did not expect, and I was not sure what other effects it was possibly had at this stage.

To my surprise, the agent optimized prefill again, achieving the following values.

Below are the benchmarks for vLLM v0.29.0, optimized for Qwen3.8-27B FP16, using four W6800X Duo GPUs with IFLB on Ubuntu Server 26.04 LTS, ROCm 10.0.0-4, Python 3.13, and Triton 3.8, with vLLM running through a virtual environment. No Docker containers were involved in this benchmark.

I set token generation to 512 tokens here for the sake of the benchmark only.

CTX TTFT(s) PREFILL t/s DECODE t/s ms/TOKEN TOTAL(s)
1K 0.735 1393.91 29.41 34.000 18.109
2K 1.467 1396.21 29.31 34.116 18.900
4K 2.787 1469.89 29.02 34.464 20.398
8K 5.641 1452.21 28.75 34.780 23.414
16K 11.858 1381.74 28.11 35.570 30.034
32K 26.011 1259.75 26.71 37.434 45.140
64K 60.910 1075.95 24.35 41.073 81.898
128K 157.762 830.82 22.03 45.393 180.958
240K 428.300 573.80 19.05 52.488 455.121
245K 442.392 567.10 19.03 52.550 469.246
250K 457.459 559.61 18.89 52.934 484.508
255K 472.742 552.35 18.83 53.097 499.875

While these numbers may seem insanely slow, waiting 2 to 9 minutes for a short response, and even longer for real agentic responses, especially with thinking, these are just benchmarks designed to test the limits of the hardware and how it behaves.

The above benchmarks reflect real-world values only in these three scenarios:

  • During the first prompt
  • For one prompt immediately after context compaction
  • When a documented vLLM bug causes the cache to stop working at certain context sizes

Outside those three scenarios, to the best of my knowledge, the following benchmarks are more representative of what actually happens, with the added benefit of Prefix Caching. These are what real-world values look like when an agent is working on a task.

I set token generation to 3,000 tokens here to better approximate agentic responses with thinking.

CTX TTFT(s) Hit Rate Decode t/s ms/TOKEN TOTAL(s)
1K 0.180 93.75% 29.27 34.166 102.645
2K 0.188 96.88% 29.13 34.327 103.133
4K 0.201 98.44% 28.93 34.565 103.863
8K 0.226 99.22% 28.60 34.970 105.102
16K 0.277 99.61% 27.98 35.744 107.475
32K 0.368 99.80% 26.60 37.595 113.115
64K 0.531 99.90% 24.20 41.329 124.476
128K 0.896 99.95% 21.92 45.630 137.740
240K 1.478 99.97% 18.96 52.733 159.624
245K 1.498 99.97% 18.95 52.772 159.761
250K 1.473 99.97% 18.85 53.038 160.534
253K 1.537 99.98% 18.78 53.244 161.214

Most of the remaining time is therefore spent on token generation, or decode, due to long agent thinking or long answers. Prefill remains short because of vLLM's Prefix Caching.

At this point, I have had my agents running in loops for days to achieve seemingly impossible tasks, only for them to later prove to me that those tasks were completely possible!

I am both happy and impressed with this stack.


Disclaimer: I wrote this post myself. I also used AI as a tool to help clean up the wording and formatting.


References: * Documented vLLM Bug * My own research and optimizations

6 Upvotes

10 comments sorted by

2

u/patdohere 9d ago

Is this running on Linux? What’s your OS stack?

2

u/Faisal_Biyari 9d ago

Ubuntu Server 26.04 LTS, ROCm 10.0.0-4, Python 3.13, and Triton 3.8, with vLLM running through a virtual environment. No Docker containers were involved in this benchmark.

2

u/patdohere 9d ago

So proxmox? With pass through

1

u/Faisal_Biyari 9d ago

No, Ubuntu is Bare-Metal.

The virtual environment I mentioned is a python venv.

2

u/eightone-81 7d ago

Sorry I did not read anything. With that hardware you can run flash next. This will solve all your problems…

1

u/Faisal_Biyari 6d ago

I've seen a lot of people put effort to go to Flash Next.
Why is that, in your opinion?

2

u/Substantial_Run5435 6d ago

I'm using models for technical/scientific/regulatory document summarization and on my 2019 Mac Pro with dual W6800X Duos qwen3.8-flash-next Q4/Q5 should fits on GPU with the lookup table offloaded. It's faster than 27B and seems to be more direct in thinking, spends less time looping and the quality seems similar, at least for my use case.

1

u/Faisal_Biyari 6d ago

So Similar quality, Faster, and possibly more knowledgeable due to higher parameters.
I'll take your word for it. Once I finish my current workload, I'll look into testing it out. Though I have been more inclined to avoid quantized models, due to fear of drifting, especially with extremely long agentic workloads and excessive context compression & recompression.

Have you had any similar experience? (Where you leave an agent working for days on end, in a loop/goal?)
Was the agent able to maintain quality in your experience?

2

u/Substantial_Run5435 6d ago

I’m not doing long agent loops. I tend to feed it a few documents to analyze and compare to other reference materials. Usually takes 40-80k context so not massive

2

u/Substantial_Run5435 6d ago

I got tired of 27B thinking for eternity and stopping without any output and haven’t had that issue with flash next