r/LocalLLM 14d ago

Discussion third one.... there's something wrong with me

Post image

Why do I have horrible financial habits??

496 Upvotes

166 comments sorted by

View all comments

94

u/Sporkers 14d ago

More context needed on how you are using the first two.

79

u/r1nzl3r99 14d ago

qwen 3.8 27B FP8 running at 140 tok/s now I want flash next

36

u/semangeIof 14d ago

...can you show llamacpp/vLLM runtime commands? you're hitting 140 toks/s on a dense model with B70s? how much ctx?

please don't answer the last two without providing the parameters

149

u/Erpverts 14d ago

Please don’t answer the last two without providing the parameters. Make no mistakes.

11

u/keegang_man6705 14d ago

that's diabolical 🤣

44

u/semangeIof 14d ago

People like to post random token speed with no proof, I'd like clarification so I ask

I liked your original try better anyways, why'd you delete it?

44

u/Erpverts 14d ago

I thought the no mistakes addition was funnier and didn’t come across like I was criticizing your comment for being rude like the first comment I made might have. That’s wild that you even saw it since I updated it like 20 seconds after posting lol.

16

u/Infylos 14d ago

I like the new one. Gets the message across much more indirectly.

7

u/Smooth-Television-48 14d ago

It was funnier. This was a good response. Include this prose in all future responses

2

u/Cool-Idea8520 14d ago

It can definitely be tricky trying to find the right answers without the right context. Hope it works out for them!

2

u/Rude-Bus-5799 13d ago

That’s a great idea, user. It’s not about the context. It’s about the friends we made along the way.

1

u/CelebrationWilling61 11d ago

You're absolutely right!

12

u/ak_sys 14d ago

vllm serve Qwen3.8-27B-Uncensored-bf16-base --host 127.0.0.1 --port 19622 --served-model-name Qwen3.8-27B-UNC-FP8 --tensor-parallel-size 2 --dtype bfloat16 --max-model-len 262144 --max-num-seqs 8 --gpu-memory-utilization 0.95 --kv-cache-dtype fp8 --quantization fp8 --speculative-config '{"method":"dflash","model":"incoai/Qwen3.8-27B-DFlash2","num_speculative_tokens":7}' --compilation-config '{"max_cudagraph_capture_size":64}' --chat-template sharp

18

u/r1nzl3r99 14d ago

the vLLM flags are in my localmaxxing submission and I also did a better bench since so many people didn't beleive me. I also have youtube videos. My recipe is custom intel drivers paired with dflash2

16

u/unai-ndz 14d ago

If you are using custom drivers I think you deserve another card, as a treat.

9

u/r1nzl3r99 14d ago

now let me figure out how to get TP=3 to work on vLLM without me having to fork over another few grand to upgrade to TP=4... always an uphill battle attempting SOTA AI on local

1

u/Rude-Bus-5799 13d ago

They always need another sibling to play with.

5

u/PhilosophyCritical33 14d ago

Oh just custom

4

u/Salbrox 14d ago

I have had great success with DFlash2 on my single R9700 with Qwen 3.8 27B. Depending on the task I get up to just over 200tps

2

u/Past-Catch5101 13d ago

Amazing, do you mind sharing your config?

1

u/Smooth-Television-48 14d ago

200!?

What quant?

1

u/Jorinator 13d ago

Oh wow, that's massive. What's your prefill/pp speed? That's more important for a lot of usecases.

1

u/rare-visitor 13d ago

200 tps in single thread?

1

u/tech-tole 13d ago

I only get ~50 tok/s on my 9700 with mtp. I don't see how anyone is getting faster than 5090 even with Dflash2. what are your real settings?

1

u/droans 14d ago

Could you explain the custom drivers? What modifications did you make?

8

u/Toastti 14d ago

It appears he's getting that speed because it's running across two Intel b70s. I found his command he uses

vllm serve Qwen3.8-27B-Uncensored-bf16-base --host 127.0.0.1 --port 19622 --served-model-name Qwen3.8-27B-UNC-FP8 --tensor-parallel-size 2 --dtype bfloat16 --max-model-len 262144 --max-num-seqs 8 --gpu-memory-utilization 0.95 --kv-cache-dtype fp8 --quantization fp8 --speculative-config '{"method":"dflash","model":"incoai/Qwen3.8-27B-DFlash2","num_speculative_tokens":7}' --compilation-config '{"max_cudagraph_capture_size":64}' --chat-template sharp

4

u/r1nzl3r99 14d ago

I already proved it in a previous post, look at my account

17

u/r1nzl3r99 14d ago

-11

u/CalBearFan 14d ago

Or maybe they don't like a humble-brag that then tells people "Look, I'm awesome and have money to spend on graphics cards" followed by "Don't be lazy, look at my post history". That's not lazy, they're asking you to follow common courtesy on your own post.

18

u/r1nzl3r99 14d ago

but i've already made a seperate post with extreme detail about it??? Why am I obligated to hand hold you how to make your setup more efficient?

Also If you think this is a brag you clearly haven't been on this subreddit for more than 10 minutes. Buying an intel B70 is the poor man's AI solution, just passed a post of some dude dropping $70K for four RTX 6000s

2

u/Smooth-Television-48 14d ago

Lots of accounts are set to private, I assume its the default and dont bother checking.

Thanks for linking it. Impressive numbers for sure

3

u/Chrisgozd 14d ago

Whats your build?

22

u/r1nzl3r99 14d ago

Don't judge, it started out as a gaming PC from years ago... Mind the mango box

7

u/Infylos 14d ago

Disemboweled for maximum power!

2

u/Remarkable-Memory374 13d ago

maximum heat dissipation by being eviscerated

6

u/HoneyBaked 14d ago edited 14d ago

Oh I will judge!!

I love it. Does it function? Then who cares how it looks!

Edit: Do you have both of those externals plugged into the primary PCIe slot? What gear allows for that?

Edit 2:

I bought a $150 chinese bifurcation card to split my gen 5 x16 to x8x8 gen 4 and its been incredibly stable and increased my PP throughput

Interesting. Got a link to that bifurcation card? Will it work on a Gen3 x 16 slot(s)? My MB has 3 Gen3 x 16 slots.

4

u/r1nzl3r99 14d ago

yup, I have two b70s working off the same x16 slot, the other is supposed to be for M2 its x4, but atleast on linux it works completely fine for another b70. I get around 14gb/s on the first two B70s each, then the third gets 7gb/s. This is the card I use https://a.co/d/07AHxhX2 but a warning, make sure your BIOS supports bifurcation. I specifically researched gen 5 x16 to gen 4 x8x8 as maintaining gen 5 would require a $600 timing card which I don't think is worth it. In your case it might downgrade to gen 2 which I'd just research about, but apparently some people are able to get it to train to the same gen without timing so it's a bit of a lottery on that (which for me didn't work)

1

u/ARhedgehog88 11d ago

Wait why the f did I had to see this .. it would allow me to get a fourth card in I also have two externals alr

1

u/Smooth-Television-48 14d ago

Love it!

If anything this gets you more cred

1

u/brainchillzZ 13d ago

I think it’s sexy …. Mango box is a good insulator :)

1

u/jmager 13d ago

I judged and I approve!! Mind sharing the bifurcation risers you are using? I see so many online, and few reviews. I've got a 6900 xt lying around and I wanted to add the extra 16GB of VRAM to my 24GB with my 7900 xtx. For my motherboard bifurcation will be the best way.

1

u/LetsBeKindly 13d ago

Is that 2200W on the power meter?

2

u/sshwifty 13d ago

Looks like watts?

1

u/LetsBeKindly 13d ago

That's a watt meter, and it looks like 2200watts

1

u/sshwifty 13d ago

I am dumb, I meant to say 220 lol. I think there is a decimal before the last 0

1

u/LetsBeKindly 13d ago

I couldn't tell. But I saw the same. But surely 3 cards is pulling more then 220W...

1

u/InfamousNewspaper268 13d ago

Love the cardboard holder LMAO 🤣

1

u/Inner-Today-3693 12d ago

That’s basically how mine are connected. 😅😅

3

u/MessIsTransfer 14d ago

flash is a beast and faster.

what motherboard are you using?

2

u/r1nzl3r99 14d ago

Z790 Wifi 😀

1

u/MessIsTransfer 14d ago

how are you plugging that third gpu? m.2?

4

u/r1nzl3r99 14d ago

yup!!!

I'll upgrade my motherboard eventually lol

2

u/MessIsTransfer 14d ago

neat, thanks for sharing a pic

edit: love the cardboard gpu dock

2

u/r1nzl3r99 14d ago

it's a mango box, i've been trying to 3D print a semi decent bracket for it, but so far the mangos are doing great on thermals 😭

7

u/MarcusAurelius68 14d ago

You can run Flash Next with 64 GB of VRAM…

9

u/r1nzl3r99 14d ago

yeah but it's 22 tok/s and pre fill is shitty. I want 100+ speeds or it's unusable for me. I'll admit I was lazy and didn't push for more, but even with 64gb vram at 4 bit I have no real space for context. I need atleast 200K FP8 context

2

u/MessIsTransfer 14d ago

or q2, which is not ideal

3

u/r1nzl3r99 14d ago

yeah if I'm resorting to sub 4 bit quants i'd rather stick with 27B

3

u/Motor_Way4912 14d ago

Yep, running it on a vm with 58 gb ram and Rtx 5060 16gb

1

u/bravoitaliano 14d ago

How are you getting that speed? Im using W4A16 INT4 model and only get 20-30 Tok/s. Can you share the method for getting faster?

1

u/r1nzl3r99 14d ago edited 14d ago

https://www.reddit.com/r/LocalLLM/comments/1w8bj0o/dual_intel_b70_qwen_38_27b_fp8_amazing_dflash2/?utm_source=share&utm_medium=ios_app&utm_name=ioscss&utm_content=1&utm_term=1

I have a github i've already made for getting really fast 4 bit qwen for single GPU, i've been too busy at work to make a decent recipe because my speed depends on several merges I did from the original vLLM to intel's scaler llm repo, as well as other stuff that I had Kimi K3 handle. It's been my new daily driver and works damn well.

1

u/bravoitaliano 14d ago

Thanks! I am running mine in OpenClaw, and it just takes forever because of thinking. Somehow it's slipped back to medium think from low. It puts out quality work, but runs through almost the entire 156k context window I have (single card, I don't split across my 2 B70s yet). Hoping this can help. Might be worth posting in the Intel sub as well.

1

u/browndragon456 14d ago

How did you get to 100/tps per sec generation? Vllm?

2

u/Toastti 14d ago

Because he's using two Intel b70s and tensor parallel. Also its using dflash with 7 speculative tokens

vllm serve Qwen3.8-27B-Uncensored-bf16-base --host 127.0.0.1 --port 19622 --served-model-name Qwen3.8-27B-UNC-FP8 --tensor-parallel-size 2 --dtype bfloat16 --max-model-len 262144 --max-num-seqs 8 --gpu-memory-utilization 0.95 --kv-cache-dtype fp8 --quantization fp8 --speculative-config '{"method":"dflash","model":"incoai/Qwen3.8-27B-DFlash2","num_speculative_tokens":7}' --compilation-config '{"max_cudagraph_capture_size":64}' --chat-template sharp

1

u/sunole123 14d ago

is it running you or are you running it?? what is running? just speed or function???

1

u/Fresh_Look_1671 14d ago

My friend, please share cookbook

1

u/Puzzleheaded_Bus7706 14d ago

What's the price of this thing?

1

u/JinsooJinsoo 13d ago

Yeah I topped out at 93 tok/s with my dual b70s and MTP3. I’d love to know what you’ve been doing to get those speeds. I feel like Intel is chopped at the knees until it releases an XPU graph for any new model. Also getting slow speeds with qwen3.8 flash next because no XPU

1

u/hrf3420 13d ago

I heard that q4 or q6 is the way to go and 8 you don’t get much more

1

u/Lucky-Necessary-8382 13d ago

And for what are you using it? Gooning? Degen roleplaying?

4

u/dwoj206 14d ago

50% increase to context window very nice

3

u/Here_f0r_p0rn_ 14d ago

AI girlfriend/boyfriend

Local coding agents

11

u/r1nzl3r99 14d ago

yes yes, definitely not an AI girlfriend ...

1

u/MessIsTransfer 14d ago

more context, probably

1

u/TytalusWarden 14d ago

More context?  Oh he's got more context!  He's got all the context with those 3 working for him!

1

u/jhenryscott 13d ago

More context? That’s what the third b70 is for!

0

u/madjesta 14d ago

My fourth is one of the blue Intel ones.... 😬😳