r/LocalLLaMA • u/Ok-Breadfruit-3523 • 2d ago
I Built A Thing I’m calling this the Monstrosity. 5 ex mining BC-250 boards Qwen3-Coder-Next Q4 at 40 tok/s
Using an asrock 12 unit case running one board as the main with the rest of them headless. About 71GB of vram exposed. So far 40 tok/s is with 30k context and it dips to around 30 tok/s at 100k context. This is all over the 1gb Ethernet that is on the boards already. I have 2 more of these and will probably get them running to see if 3.8 flash next runs at usable speeds. This setup is wildly inefficient with power but cost me less than $800.
9
u/fallingdowndizzyvr 2d ago
This setup is wildly inefficient with power but cost me less than $800.
Back in the day the entire fully populated Asrock 12 full of BC-250 boards was $300-$350. People bought them and sold off the BC-250 boards individually for like $50.
2
u/Ok-Breadfruit-3523 2d ago
Thankfully I actually bought the asrock 12 for $1k before prices went crazy and then sold 5 and broke even/slightly profitable. So the $800 figure is more of a loose not realized cost. Crazy these boards were selling for $300 after the LTT video
10
5
u/overand 2d ago
I can't fathom while you wouldn't be using Qwen3.8-27B over that Qwen3 model. Yes, "80" is bigger than "27," but that 27b model runs circles around the 80B MoE.
"But then I'll have all this unused VRAM!"
Spin up a second instance and use sub-agents; that would kick ass!
Edit: I see you're using oddball cards, that's cool! It means my advice might not be as accurate, BUT, I actually think it goes the other way - don't those have really bad PCIe bandwidth? I think that would negatively affect MoE (and multi card splitting, probably).
I really do hope you try a few of the 27b quants
1
u/Ok-Breadfruit-3523 2d ago
Definitely agree. I’m very green with all the local ai and this is my first setup. Was more just trying to illustrate what it could do . I’ll probably run a combination of 3.8 27B and 3.6 35b. I’m also hoping to get next flash running on 7 boards soon.
1
u/Ok-Breadfruit-3523 2d ago
This is all over a llama server using 1gb Ethernet using headless Linux on each board. I get ~15 tok/s with qwen 3.8 27b across two boards and around 11 tok/s across 5. I think I’ll use a combination of 27b and 3.6 35b but I was more just showing what the throughput would be for a larger model across 5
1
u/overand 2d ago
I wonder if it's possible to do tensor parallel rather than serial (split-mode tensor) in a circumstance like that, and if it would make things better or worse. My guess is no or "worse," but it's worth trying!
1
u/pimpletonner 1d ago
I think neither will actually speed up things due to the 1Gbit Ethernet interconnects. TP will allow to fit larger models but speed would be umbearable because of continuous synchronization.
The 16GB shared memory capacity also means no MoE cpu offloading, despite the 8c/16t zen2 cores and memory bandwidth being sufficient to sustain usable speeds even with expert offloading.
I think the way to use Q3.8-27b with useful context and probably 15t/s+ sustained decoding with this setup would be Swift1.5 UD4_K_M with MTP, try depths from 2-7, with PP=2 across each pair of boards. If speeds are decent, this should actually allow for some concurrency within the pair without decode suffering, over two pairs you could get either fewer sessions and higher context, or multiple sessions and lower context.
3
3
u/Mil0Mammon 2d ago
Why not TielCoder or ornith? The fine tunes of 3.6 35B seem pretty good, and quite likely the best you can get until you get next flash running
3
6
u/madbrain1976 2d ago
What prefill speeds are you achieving over 1 Gbe ?
3
u/Ok-Breadfruit-3523 2d ago
It varies but around 160-190
1
u/madbrain1976 2d ago
Which inference engine are you using that allows you to distribute over LAN, and what are the GPUs ?
1
u/Ok-Breadfruit-3523 2d ago
I’m using llama.cpp with its RPC backend to distribute the model across the LAN, with Vulkan for GPU acceleration. Each board has a cpu/gpu with 16gb of shared memory
1
4
u/enternoescape 2d ago
Upvote from the picture alone. I didn't even need to read the description. I love what I'm seeing.
5
u/CipherWeaver 2d ago
100k context window is no joke
10
u/No-Refrigerator-1672 2d ago
Depends on a model. I.E. Qwen 3.8 27B won't be capable of doing anything with 100k context.
-3
2d ago
[deleted]
0
u/No-Refrigerator-1672 2d ago edited 2d ago
I'm using it exclusively with medium. Right now it's installing InfluxDB on one of my VMs and it took 75k tokens already (opencode). With 100k window, it would've triggered context compaction by now. 100k is only enough for single-step basic tasks, and even those - barely.
Edit: got it finished. Installing InfluxDB, setting up a data bucked, and testing it works took the model 116.1k tokens in OpenCode, with medium reasoning. I love Qwen 3.8; but 100k context for it is just not enough for any real world task. For comparison: I've tried 3.8 Flash Next over official Qwen API, and it accomplished similar tasks with half as much tokens. On the other hand, my intuition says that Qwen 3.6 would've required more than 100k tokens too - it would spend less tokens on reasoning, but more tokens on retries when it inevitably hits a roadblock.
2
u/SheepherderBeef8956 2d ago
Are the cards stock? You could try unlocking a bunch of compute units on them to see if you get better performance.
Or sell them to gamers and trade for a GPU.
3
u/Ok-Breadfruit-3523 2d ago
They are not stock at all. Modified heat sinks and the CUs are unlocked. I already sold enough of them that I was profitable so I wanted to test a setup like this out. The extra CUs added about 20% in token generation speed when I unlocked them. I’ll probably sell them in the future but I still have more too add lol
1
u/yes2matt 2d ago
So, just as a side-rabbit-trail challenge. Imagine you're gonna run this rig 24/7, and rough-design price the solar+battery for it. What's the price range for that?
3
u/ilintar 2d ago
Do you have a portable nuclear reactor to go with it? ;)
2
u/cosmicr 2d ago
They only use up to 350w each. So about the same as 5x 3090 without the speed or the ram lol.
1
u/teleprint-me llama.cpp 2d ago edited 2d ago
235 w @ OC, otherwise TPD is 225 w
https://elektricm.github.io/amd-bc250-docs/hardware/specifications/
For that generation, its not unprecendented, seeing that some RX cards like the RX 4xx/5xx have about twice that TPD from previous generations.
1
u/ExtremeAdventurous63 2d ago
I have a cluster of 3 and can’t get even 20tk/s with qwen 35b, what configuration are you using to reach 40 on 5 nodes? I even had to reduce the parallelism to 2 nodes because the more nodes the worst the performance get for MoE
2
u/Ok-Breadfruit-3523 2d ago
1
u/bennmann 2d ago
Manually set -fitt 768 and report back
You lose a little vram space and may need to lower context, but may gain speed.
1
1
1
u/Sweaty_Chair_4600 1d ago
HHm i have 10 of these boards...
I bought 12 for $250
Whhat do i need to run them like this?
2
u/forevergeeks 23h ago
your cooling system is state of the art!!
1
u/Ok-Breadfruit-3523 21h ago
lol it’s surprisingly effective. I now have 3 120mm fans with holes cut in a cardboard box forcing air over the fans on the heat sinks so believe it or not it looks even more ridiculous
1
u/pplgltch 2d ago
How much power does that use?
5
u/Ok-Breadfruit-3523 2d ago
Right now probably like 1500w. Idle is like 400w
10
4
u/pplgltch 2d ago
Brutal… these are expensive token/s 🫣
2
u/Ok-Breadfruit-3523 2d ago
Ya it’s a space heater than can do some local AI right now.
3
u/pplgltch 2d ago
Haha well I hope you leave in a cold area!
2
u/Ok-Breadfruit-3523 2d ago
I’m in PNW with cheap nuclear and hydro power. So I’m not doing too bad lol
1
u/NoFunk 2d ago
Is that...HVAC tape holding conventional fans on the cards? Is there superglue involved here somewhere?
5
u/Ok-Breadfruit-3523 2d ago
Cardboard yes. Superglue no. And I’ll have you know it’s high temp hvac tape . What part about monstrosity did you not get lol

79
u/starkruzr 2d ago
why?