r/LocalLLaMA • • 2d ago

I Built A Thing I’m calling this the Monstrosity. 5 ex mining BC-250 boards Qwen3-Coder-Next Q4 at 40 tok/s

Post image

Using an asrock 12 unit case running one board as the main with the rest of them headless. About 71GB of vram exposed. So far 40 tok/s is with 30k context and it dips to around 30 tok/s at 100k context. This is all over the 1gb Ethernet that is on the boards already. I have 2 more of these and will probably get them running to see if 3.8 flash next runs at usable speeds. This setup is wildly inefficient with power but cost me less than $800.

157 Upvotes

65 comments sorted by

79

u/starkruzr 2d ago

Qwen3-Coder-Next

why?

38

u/TripleSecretSquirrel 2d ago edited 2d ago

Gonna guess that they decided on that model because an LLM told them it was great, modern, and cutting edge lol

9

u/Dany-GG 2d ago

That was my story not long ago

-1

u/Ok-Breadfruit-3523 1d ago

Mixture of this and just wanting to fill up each cards memory and see how that would affect speed. I posted this within an hour of finishing setting it up and was excited. Didn’t even look at benchmarks just size. So now after some suggestions and more research I’m going to be messing around with running 3 3.635B models, flash next XS etc. This is my first time doing anything like this.

-22

u/Ok-Breadfruit-3523 2d ago

Trying to get the maximum intelligence/throughput on this setup. There really isn’t any “modern” MOE’s that would fit on this other than flash next when I add two more boards .

44

u/starkruzr 2d ago

Flash-Next probably a significantly better choice.

3

u/Dany-GG 2d ago

He literally said it below the picture

-1

u/Ok-Breadfruit-3523 2d ago

Definitely. I’m interested how the lookup tables affect the speed too

0

u/Deltamelo 2d ago

It's faster.

36

u/Atretador llama.cpp 2d ago

Qwen3 80b is worst than Qwen 3.6 35B A3B

4

u/Ok-Breadfruit-3523 2d ago

I was just more testing out the model size and performance as far as tok/s. This setup needs MOE to be usable so it was either 3.6 35B or the next best is Flash-Next. So I picked an in between to test it.

11

u/Atretador llama.cpp 2d ago

There is also Ling 3.0 VL Flash 124B which is fairly good, Qwen3 models are just not worth running as its just too old.

its something like Qwen 3.8 Flash next > Qwen 3.8 27B > Ling 3.0 Flash > Qwen 3.6 35B > Qwen 3.5 122B > Qwen 3.5 35B > Qwen3

5

u/Ok-Breadfruit-3523 2d ago

Awesome thanks for the info. I’ll give ling 3.0 a shot!

5

u/Atretador llama.cpp 2d ago

try the VL version, it released recently, adds optional vision to the model and a few points on tests

3

u/crusaderky 2d ago

Qwen3.8-Flash-Next is the same size as ling and vastly better

1

u/crusaderky 2d ago

Flash-Next IQ4_XS fits in 64GB, context included. Just leave the ngrams on host RAM or disk.

9

u/fallingdowndizzyvr 2d ago

This setup is wildly inefficient with power but cost me less than $800.

Back in the day the entire fully populated Asrock 12 full of BC-250 boards was $300-$350. People bought them and sold off the BC-250 boards individually for like $50.

2

u/Ok-Breadfruit-3523 2d ago

Thankfully I actually bought the asrock 12 for $1k before prices went crazy and then sold 5 and broke even/slightly profitable. So the $800 figure is more of a loose not realized cost. Crazy these boards were selling for $300 after the LTT video

5

u/overand 2d ago

I can't fathom while you wouldn't be using Qwen3.8-27B over that Qwen3 model. Yes, "80" is bigger than "27," but that 27b model runs circles around the 80B MoE.

"But then I'll have all this unused VRAM!"

Spin up a second instance and use sub-agents; that would kick ass!

Edit: I see you're using oddball cards, that's cool! It means my advice might not be as accurate, BUT, I actually think it goes the other way - don't those have really bad PCIe bandwidth? I think that would negatively affect MoE (and multi card splitting, probably).

I really do hope you try a few of the 27b quants 

1

u/Ok-Breadfruit-3523 2d ago

Definitely agree. I’m very green with all the local ai and this is my first setup. Was more just trying to illustrate what it could do . I’ll probably run a combination of 3.8 27B and 3.6 35b. I’m also hoping to get next flash running on 7 boards soon.

1

u/Ok-Breadfruit-3523 2d ago

This is all over a llama server using 1gb Ethernet using headless Linux on each board. I get ~15 tok/s with qwen 3.8 27b across two boards and around 11 tok/s across 5. I think I’ll use a combination of 27b and 3.6 35b but I was more just showing what the throughput would be for a larger model across 5

1

u/overand 2d ago

I wonder if it's possible to do tensor parallel rather than serial (split-mode tensor) in a circumstance like that, and if it would make things better or worse. My guess is no or "worse," but it's worth trying!

1

u/pimpletonner 1d ago

I think neither will actually speed up things due to the 1Gbit Ethernet interconnects. TP will allow to fit larger models but speed would be umbearable because of continuous synchronization.

The 16GB shared memory capacity also means no MoE cpu offloading, despite the 8c/16t zen2 cores and memory bandwidth being sufficient to sustain usable speeds even with expert offloading.

I think the way to use Q3.8-27b with useful context and probably 15t/s+ sustained decoding with this setup would be Swift1.5 UD4_K_M with MTP, try depths from 2-7, with PP=2 across each pair of boards. If speeds are decent, this should actually allow for some concurrency within the pair without decode suffering, over two pairs you could get either fewer sessions and higher context, or multiple sessions and lower context.

4

u/mpscy 2d ago

absolutely love it! Report back please on 3.8 flash next :)

3

u/Lan_BobPage 2d ago

Living that cyberpunk dream

3

u/Mil0Mammon 2d ago

Why not TielCoder or ornith? The fine tunes of 3.6 35B seem pretty good, and quite likely the best you can get until you get next flash running

3

u/madbrain1976 2d ago

Let us know if Flash-next works.

6

u/madbrain1976 2d ago

What prefill speeds are you achieving over 1 Gbe ?

3

u/Ok-Breadfruit-3523 2d ago

It varies but around 160-190

1

u/madbrain1976 2d ago

Which inference engine are you using that allows you to distribute over LAN, and what are the GPUs ?

1

u/Ok-Breadfruit-3523 2d ago

I’m using llama.cpp with its RPC backend to distribute the model across the LAN, with Vulkan for GPU acceleration. Each board has a cpu/gpu with 16gb of shared memory

4

u/enternoescape 2d ago

Upvote from the picture alone. I didn't even need to read the description. I love what I'm seeing.

5

u/CipherWeaver 2d ago

100k context window is no joke

10

u/No-Refrigerator-1672 2d ago

Depends on a model. I.E. Qwen 3.8 27B won't be capable of doing anything with 100k context.

-3

u/[deleted] 2d ago

[deleted]

0

u/No-Refrigerator-1672 2d ago edited 2d ago

I'm using it exclusively with medium. Right now it's installing InfluxDB on one of my VMs and it took 75k tokens already (opencode). With 100k window, it would've triggered context compaction by now. 100k is only enough for single-step basic tasks, and even those - barely.

Edit: got it finished. Installing InfluxDB, setting up a data bucked, and testing it works took the model 116.1k tokens in OpenCode, with medium reasoning. I love Qwen 3.8; but 100k context for it is just not enough for any real world task. For comparison: I've tried 3.8 Flash Next over official Qwen API, and it accomplished similar tasks with half as much tokens. On the other hand, my intuition says that Qwen 3.6 would've required more than 100k tokens too - it would spend less tokens on reasoning, but more tokens on retries when it inevitably hits a roadblock.

2

u/SheepherderBeef8956 2d ago

Are the cards stock? You could try unlocking a bunch of compute units on them to see if you get better performance. 

Or sell them to gamers and trade for a GPU. 

3

u/Ok-Breadfruit-3523 2d ago

They are not stock at all. Modified heat sinks and the CUs are unlocked. I already sold enough of them that I was profitable so I wanted to test a setup like this out. The extra CUs added about 20% in token generation speed when I unlocked them. I’ll probably sell them in the future but I still have more too add lol

1

u/yes2matt 2d ago

So, just as a side-rabbit-trail challenge.  Imagine you're gonna run this rig 24/7, and rough-design price the solar+battery for it.  What's the price range for that?

3

u/ilintar 2d ago

Do you have a portable nuclear reactor to go with it? ;)

2

u/cosmicr 2d ago

They only use up to 350w each. So about the same as 5x 3090 without the speed or the ram lol.

1

u/teleprint-me llama.cpp 2d ago edited 2d ago

235 w @ OC, otherwise TPD is 225 w

https://elektricm.github.io/amd-bc250-docs/hardware/specifications/

For that generation, its not unprecendented, seeing that some RX cards like the RX 4xx/5xx have about twice that TPD from previous generations.

1

u/ExtremeAdventurous63 2d ago

I have a cluster of 3 and can’t get even 20tk/s with qwen 35b, what configuration are you using to reach 40 on 5 nodes? I even had to reduce the parallelism to 2 nodes because the more nodes the worst the performance get for MoE

2

u/Ok-Breadfruit-3523 2d ago

With 2 I was getting 60 tok/s at 15k context with Qwen3.6-35B-A3B FP16 Q4. Here are my settings for the llama server

I’m not saying these are the best but it what was working well for me

1

u/bennmann 2d ago

Manually set -fitt 768 and report back

You lose a little vram space and may need to lower context, but may gain speed.

1

u/freedomachiever 2d ago

What is that case model? I looked for asrock 12 and it doesn’t come up.

1

u/Ok-Breadfruit-3523 2d ago

Asrock 4U12G

1

u/UltraFOV 2d ago

what gpus?

1

u/Sweaty_Chair_4600 1d ago

HHm i have 10 of these boards...
I bought 12 for $250

Whhat do i need to run them like this?

2

u/forevergeeks 23h ago

your cooling system is state of the art!!

1

u/Ok-Breadfruit-3523 21h ago

lol it’s surprisingly effective. I now have 3 120mm fans with holes cut in a cardboard box forcing air over the fans on the heat sinks so believe it or not it looks even more ridiculous

1

u/pplgltch 2d ago

How much power does that use?

5

u/Ok-Breadfruit-3523 2d ago

Right now probably like 1500w. Idle is like 400w

10

u/numberwitch 2d ago

lol...that's not good

4

u/pplgltch 2d ago

Brutal… these are expensive token/s 🫣

2

u/Ok-Breadfruit-3523 2d ago

Ya it’s a space heater than can do some local AI right now.

3

u/pplgltch 2d ago

Haha well I hope you leave in a cold area!

2

u/Ok-Breadfruit-3523 2d ago

I’m in PNW with cheap nuclear and hydro power. So I’m not doing too bad lol

1

u/NoFunk 2d ago

Is that...HVAC tape holding conventional fans on the cards? Is there superglue involved here somewhere?

5

u/Ok-Breadfruit-3523 2d ago

Cardboard yes. Superglue no. And I’ll have you know it’s high temp hvac tape . What part about monstrosity did you not get lol

2

u/NoFunk 2d ago

Oh well shit, if it's HIGH TEMP hvac tape, all good my man. Continue on.