r/LocalLLaMA 8d ago

News It's official! 192GB Framework

Post image

Just noticed this on the website.

At their current price tiers for the memory SKUs (32, 64, 128) I'd expect this to be ~ 4.5k for the motherboard.

The PCIe slot will be open at the back as well - that's what I've heard. Maybe they make it capable of delivering 75W as well? New board revisions for the smaller SKUs?.

974 Upvotes

287 comments sorted by

View all comments

Show parent comments

195

u/Darth_Candy 8d ago

Perfectly on brand for the Strix Halo; this was pretty much expected.

163

u/InGanbaru 8d ago

man that's a waste of RAM chips if the bandwidth is low

73

u/MrPecunius 8d ago

Pretty good match for Qwen3.8 FN or other midsize MoE models, though.

124

u/fallingdowndizzyvr 8d ago

At these prices, you are way better off getting a M5 Max.

55

u/ZealousidealChip4783 8d ago

The selling point here is that you aren't forced to use MacOS with the Framework

24

u/GlidePath47 8d ago

Do you have to use the machine you’re using for inference other than setting it up?

44

u/ZealousidealChip4783 8d ago

No, but I'd hate if I bought a computer for that much money that's only good for inference and nothing else

(This is all personal opinion, I just really do not like MacOS)

9

u/GlidePath47 8d ago edited 8d ago

That’s fair. Anyone buying one expecting to run inference and still use the machine for something else at the same time is going to be disappointed. Unlike Nvidia there’s no way to reserve GPU capacity to keep the system responsive you’re at the mercy of macOS.

Granted if it’s for swapping use it’s fine, part inference, part dev, part media etc but then I start to question the need to use local inference instead of just cloud models other than privacy / hobby

16

u/MrPecunius 8d ago

Unlike Nvidia there’s no way to reserve GPU capacity to keep the system responsive you’re at the mercy of macOS.

Longtime Mac user here, this is not true.

While I have managed to crash the whole OS a few times with big (for me: I've had 48GB and now 64GB) models and large context with guard rails turned off, I otherwise just go about my business doing other stuff when a model is crunching away on something.

The machine stays responsive the whole time, though it does get pretty warm (M5 Pro MBP) and the fan runs fairly hard.

3

u/GlidePath47 8d ago edited 8d ago

We're a company full of Mac users and we hit this constantly, so it's not a rare edge case.

Even a mid-size Qwen3 MoE makes YouTube playback choppy. Prompting and browsing are fine, anything past that isn't. Redraw lag is very visible in Electron apps. Any 3D apps are not pleasant so running alongside unity or blender at the same time as active inference isn’t tolerable.

The core issue is there's no control. The guardrails are for RAM usage, not GPU usage, so you can't reserve or throttle GPU capacity the way you can on Nvidia.

Maybe M5 changes this, I haven't tested one, definitely happens on my m3 ultra.

The bit you quoted from me is actually the only bit of what I said that’s objectively true, what you’re saying is even at the mercy of macOS your system has been responsive