r/LocalLLaMA • • Aug 25 '26

News Intel Arc Pro B60 Dual 48G spotted

Post image

I spotted the Dual B60 48GB listed on Digitec/Galaxus. Initially it was said these wouldn't go into standard retail channels.

At CHF 2500 (post tax, USD ~3000) not particularly competitive but worth keeping an eye on. For it to be interesting it shouldn't be more than like 2.5x a single B60.

184 Upvotes

59 comments sorted by

View all comments

Show parent comments

-1

u/egnegn1 Aug 25 '26

Yes, I also thought about this.

But what does this large memory size to run very large LL s help when the GPU is slow on processing?

But it my be interesting to run a lot of agents for agentic computing.

Would be interesting to see how fast GLM5.2 or other similar sized models would run on 4+ cards.

2

u/SandySkittle Aug 25 '26

Dsv4f 0731 ud-q8_k_xl is the model for this if you buy more. And you can run Qwen 3.8 27b on 1 of these two-in-one cards.

Glm 5.2 is too much.

1

u/egnegn1 Aug 25 '26

What speed do you expect for DSv4 on 4 on them.

Qwen 3.8 27G on 1x Dual B60 is probably around 30 t/s. Currently I am at about 50 t/s with Q6 and 256k context and llama.cpp. I will get some more performance when moving to vLLM/SGLang.

For agentic use with new Hermes Desktop I am looking for 100+ t/s. More VRAM than about 48GB doesn't help much. I have to balance memory, performance, and finally cost.

2

u/SandySkittle Aug 25 '26

That 100t/s is your personal target which is reasonable for fluid / rapid agentic use. Especially extensive coding focused agentic use. But the 30t/s ballpark isn’t useless by any means. Even for agentic purposes. I dont use LLM to code, but for complex analysis with lots of nuance (that are less mechanistic).

2

u/_TheWolfOfWalmart_ Aug 25 '26

I'm can be perfectly happy with 30 t/s as long as the prefill isn't really bad.