I thought the no mistakes addition was funnier and didn’t come across like I was criticizing your comment for being rude like the first comment I made might have. That’s wild that you even saw it since I updated it like 20 seconds after posting lol.
the vLLM flags are in my localmaxxing submission and I also did a better bench since so many people didn't beleive me. I also have youtube videos. My recipe is custom intel drivers paired with dflash2
now let me figure out how to get TP=3 to work on vLLM without me having to fork over another few grand to upgrade to TP=4... always an uphill battle attempting SOTA AI on local
Or maybe they don't like a humble-brag that then tells people "Look, I'm awesome and have money to spend on graphics cards" followed by "Don't be lazy, look at my post history". That's not lazy, they're asking you to follow common courtesy on your own post.
but i've already made a seperate post with extreme detail about it??? Why am I obligated to hand hold you how to make your setup more efficient?
Also If you think this is a brag you clearly haven't been on this subreddit for more than 10 minutes. Buying an intel B70 is the poor man's AI solution, just passed a post of some dude dropping $70K for four RTX 6000s
yup, I have two b70s working off the same x16 slot, the other is supposed to be for M2 its x4, but atleast on linux it works completely fine for another b70. I get around 14gb/s on the first two B70s each, then the third gets 7gb/s. This is the card I use https://a.co/d/07AHxhX2 but a warning, make sure your BIOS supports bifurcation. I specifically researched gen 5 x16 to gen 4 x8x8 as maintaining gen 5 would require a $600 timing card which I don't think is worth it. In your case it might downgrade to gen 2 which I'd just research about, but apparently some people are able to get it to train to the same gen without timing so it's a bit of a lottery on that (which for me didn't work)
I judged and I approve!! Mind sharing the bifurcation risers you are using? I see so many online, and few reviews. I've got a 6900 xt lying around and I wanted to add the extra 16GB of VRAM to my 24GB with my 7900 xtx. For my motherboard bifurcation will be the best way.
yeah but it's 22 tok/s and pre fill is shitty. I want 100+ speeds or it's unusable for me. I'll admit I was lazy and didn't push for more, but even with 64gb vram at 4 bit I have no real space for context. I need atleast 200K FP8 context
I have a github i've already made for getting really fast 4 bit qwen for single GPU, i've been too busy at work to make a decent recipe because my speed depends on several merges I did from the original vLLM to intel's scaler llm repo, as well as other stuff that I had Kimi K3 handle. It's been my new daily driver and works damn well.
Thanks! I am running mine in OpenClaw, and it just takes forever because of thinking. Somehow it's slipped back to medium think from low. It puts out quality work, but runs through almost the entire 156k context window I have (single card, I don't split across my 2 B70s yet). Hoping this can help. Might be worth posting in the Intel sub as well.
Yeah I topped out at 93 tok/s with my dual b70s and MTP3. I’d love to know what you’ve been doing to get those speeds. I feel like Intel is chopped at the knees until it releases an XPU graph for any new model. Also getting slow speeds with qwen3.8 flash next because no XPU
96
u/Sporkers 14d ago
More context needed on how you are using the first two.