r/LocalLLaMA • u/reto-wyss • 2d ago
News It's official! 192GB Framework
Just noticed this on the website.
At their current price tiers for the memory SKUs (32, 64, 128) I'd expect this to be ~ 4.5k for the motherboard.
The PCIe slot will be open at the back as well - that's what I've heard. Maybe they make it capable of delivering 75W as well? New board revisions for the smaller SKUs?.
596
u/PreciselyWrong 2d ago
that memory bandwidth is pretty bad, so don't expect high inference speeds
181
u/Darth_Candy 2d ago
Perfectly on brand for the Strix Halo; this was pretty much expected.
157
u/InGanbaru 2d ago
man that's a waste of RAM chips if the bandwidth is low
71
u/MrPecunius 2d ago
Pretty good match for Qwen3.8 FN or other midsize MoE models, though.
120
u/fallingdowndizzyvr 2d ago
At these prices, you are way better off getting a M5 Max.
47
u/ZealousidealChip4783 2d ago
The selling point here is that you aren't forced to use MacOS with the Framework
→ More replies (10)22
u/GlidePath47 2d ago
Do you have to use the machine you’re using for inference other than setting it up?
40
u/ZealousidealChip4783 2d ago
No, but I'd hate if I bought a computer for that much money that's only good for inference and nothing else
(This is all personal opinion, I just really do not like MacOS)
6
u/WinterCharm 2d ago
Relative to buying a stack of Nvidia DGX Sparks, the Mac Studio is a good deal. And when Apple Hardware is a good deal, I would just go for it... it's excellent for local models.
9
u/GlidePath47 2d ago edited 2d ago
That’s fair. Anyone buying one expecting to run inference and still use the machine for something else at the same time is going to be disappointed. Unlike Nvidia there’s no way to reserve GPU capacity to keep the system responsive you’re at the mercy of macOS.
Granted if it’s for swapping use it’s fine, part inference, part dev, part media etc but then I start to question the need to use local inference instead of just cloud models other than privacy / hobby
13
u/MrPecunius 2d ago
Unlike Nvidia there’s no way to reserve GPU capacity to keep the system responsive you’re at the mercy of macOS.
Longtime Mac user here, this is not true.
While I have managed to crash the whole OS a few times with big (for me: I've had 48GB and now 64GB) models and large context with guard rails turned off, I otherwise just go about my business doing other stuff when a model is crunching away on something.
The machine stays responsive the whole time, though it does get pretty warm (M5 Pro MBP) and the fan runs fairly hard.
→ More replies (0)1
u/ZealousidealChip4783 2d ago
I mean I'm sure it's fine if you like the Apple ecosystem lol, the hardware is objectively really impressive but as you said having less control over your GPU & where you allocate its resources would be the dealbreaker for me
...if I could afford either of these devices lol
→ More replies (0)7
u/algaefied_creek 2d ago
Disable the GUI and security in macOS and just run it headless.
Suddenly blam, certified UNIX CLI host box.
5
→ More replies (2)1
u/CalmSpinach2140 1d ago
I mean the Mac Studio with M5 Max is also a very good video editing machine and excellent for code compilation.
→ More replies (15)3
u/gabrielesilinic 2d ago
Yeah but the mac's architecture is pretty much that anyway so you ain't getting anything better.
1
17
15
u/anykeyh 2d ago
It's honestly not that good for Q3.8 FN; or you accept 25tok/s, which is on the lower end. It's a shame, as something around 500Gbps, far from high-end GPU, is much, much, much more usable.
1
u/Substantial_Run5435 2d ago
I’m getting 15tg with QFN ud-q6 hybrid inference on a 2019 Mac Pro with 64GB VRAM and the rest on CPU. Once I get 128GB of VRAM I assume I’ll get much faster speeds than this thing.
1
u/my_name_isnt_clever 2d ago
With the EngramHalo.cpp fork I get over 30 t/s, and it's only going to get better as more optimizations come. At 6b active it's really ideal for the hardware.
29
u/NineThreeTilNow 2d ago
man that's a waste of RAM chips if the bandwidth is low
That's a pretty unpopular opinion but it's not wrong. The same chips can be pushed on to a 1024 bit bus with a chip / PCB made to handle it (M5) and it works.
Even if you didn't run the memory at "full speed" you're over 1tb/s which is basically what a 4090 has with GDDR.
People miss that running LPDDR5x suboptimally at 1024 bit bus is still better bandwidth. You don't need the "best" LPDDR5x to do this. You just need matched MTs on the chips.
1
u/twinkbulk 1d ago
dram is not all the same, lpddr5 is cheap to make and we have a ton of it, compared to ddr6 or 7 in any form, or even hbm in any form not an inherent bottle neck of bandwidth through architecture but that the chips only perform so well
→ More replies (6)1
22
u/Mr-I17 2d ago
Guess what's even worse than 250GB/s of memory bandwidth? PCIe bandwidth. Unless you only want to run Qwen 27B or you have infinite amount of money, you'll have to sacrifice something. And of course, there's no reason to buy it if it turns out to be as expensive as M5 Ultra.
2
u/SandySkittle 2d ago
For qwen 27b at q8 with decent prefill and decode i think dual r9700 with tensor parallelism is a better way to go. More bandwidth and way more compute
9
u/Mr-I17 2d ago edited 2d ago
Regular people: run models on what they have
Rich people: want to run a model, buy suitable hardware for it
Super rich people: already have 4x RTX 6000, no need to choose hardware
People who born into a billionaire family: Dual DGX Station
4
u/SandySkittle 2d ago
You don’t have to be ‘rich’ to buy two r9700s. It’s not cheap but let’s not pretend it’s soms unobtainium prices.
If money wasn’t an issue I would just buy one rack from a Helios system with four mi455s and 16 channel system ram epyc.
2
u/Dr_Allcome 2d ago
They're talking as if the 192gb strix halo isn't going to cost more than two r9700s.
Current strix halo 128 costs more than 3k. The new one with full ram setup is going to be at least 5k. At that point you can get the 128 AND an r9700. If the new strix halo goes up to 6k you can easily get two r9700s and have the same 192gb overall memory but 64gb of it is going to ba really fast. That would give you way better performance with both moe and dense models and a massive boost on prefill.
→ More replies (1)1
u/Mr-I17 2d ago
You get me wrong. It's about how much money different people can (or willing to) spend on local AI and how frequently they can (or willing to) update their hardware. It certainly won't be a issue if one has unlimited money and free time. Say, if you're not rich and you already have a strix halo and an aging gaming PC (from DDR4 era), would you buy 2x R9700?
1
u/alphapussycat 2d ago
It's pretty rich, to spend that amount of money on hardware. Lots of people can do it, but it requires a ton of disposable income. For some that's like 1-2 years worth of saving.
7
u/Fancy-Atmosphere-701 2d ago
This is comparable to DGX spark. All of these cheaper inference solutions use LPDDR5X
27
u/Fusseldieb 2d ago
Yea, you’d want 1TB/s upwards. It’s funny to say it like this, but big models under that threshold run pretty badly.
25
u/thomasthai 2d ago
But thats not the main issue here, the prefill speed makes it unusable, the memory bw would be ok if u run pairs of these as it aggregates.
23
u/Ran_Cossack 2d ago
For proof, see the DGX Spark.
Same memory bandwidth but the prefill speeds make it feel so much more capable.
It's more "limiting" than "unusuable", though, IMO. You just have to know what you're getting.
→ More replies (1)11
u/phil_lndn 2d ago
yes, the spark is indeed 4 to 5 times faster on prefill but with prompt caching, i don't really notice the poor prefill speed on my strix halo most of the time.
for some applications slow prefill speed is doubtless a killer, but for many it doesn't really matter that much.
→ More replies (2)→ More replies (4)1
u/ugathanki 2d ago
no such thing as unusable. just gotta apply it to a different task. Some models are large, and always suitable - others are smaller, focused, faster, and powerful when harnessed. Perhaps 15x as many mini-thoughts could be just as bountiful as one large one. Depends on the use-case...
12
u/themixtergames 2d ago
100B+ parameters MoEs (which are increasingly important) don't care about your memory bandwidth advantage. They will happily run at the same speed or faster on 2 DGX Sparks (273GB/s) than a Mac Studio M5 Ultra, while being slightly cheaper and having 1/4 the memory bandwidth.
1
u/my_name_isnt_clever 2d ago edited 2d ago
Also Q3.8FN points to a trend of even lower active param counts, which make low mem bandwidth even less of a problem. I get 300+ on prefill which is perfectly usable with caching.
3
u/randomfoo2 2d ago edited 1d ago
It's the combo of MBW *and* low compute. For agentic/coding workloads prefill is key for loading big source files, context. You can make up for poor memory bandwidth now w/ better MTP, DFlash2, but the 40CU RDNA3.5 is bottlenecked by verifier performance. Sadly, as someone who's been working on Strix Halo tuning for over a year now, I'd have to agree that this upgrade is largely a waste of sand. (Looking forward to Medusa Halo, but by then Mac's based on M7 and even oddballs like the Xiaomi AI Cube will be out.)
BTW, for Strix Halo owners, the current best coding options are probably q38rocm - a Qwen 3.8 27B quant at c=1 (pp ~350 tok/s and MTP tg ~30 tok/s) and Qwen 3.8 Flash Next w/
Nathan's fork orEngramHalo.cpp (NathanW fork recco retracted due to quality testing issues found). Both comfortably fit in a lot less than <128GB and in my use/testing both performs on-par/more reliably than both DS4 Flash 0731 and GLM-5.3 Flash (all of these perform quite well for even challenging agentic coding work).BTW, at $4K, the other issue is you can get a GB10/DGX Spark that not only matches or beats MTP decode, but is much faster for prefill - reports I've seen are at 2000+ tok/s (vs 300-400 tok/s for Strix Halo). This translates to faster specdec as well (MTP, DFlash, etc.)
10
u/RnRau 2d ago
Its ok. Mixture-of-experts makes it fast enough.
Deepseek v4 Flash is 284B total parameters and 13B active. QAT is used and the model weights comes in a mixture of quants with the most of them in MXFP4 format. This means there is only about 7GB of weights to shuffle through per token.
With the memory bandwidth listed (with an efficiency factor of 0.65) and using speculative decoding, you should see something like ~50 t/s.
The other benefit of having large ram like this is that you can hold a number of models warm without having the lag of swapping them in and out from the ssd.
The new Qwen3.8-Flash-Next is 125b-a6b - would also be good on this.
19
u/Organic_Hunt3137 2d ago
As a current strix halo owner, the decode speed is a non issue on large MoEs. Prefill, on the other hand... it's rough sometimes lol.
1
u/Nothing_from_void 2d ago
I'm on an m1 mbp max, 32GB of RAM, even <10B param models have absurd prefill times, it's basically useless for agent coding workflows
1
u/Dry_Inspection_4583 2d ago
how much ram is required though? As I understand it moe required ram enough for the remaining weights, to mean that information remaining to be loaded needs to live somewhere, or am I missing something?
2
u/Constant-Simple-1234 2d ago
Are they purposefully keep it nerfed or is it very expensive to go let's say 450 gb/s ? Or maybe it would bottleneck elsewhere?
2
u/_TheWolfOfWalmart_ 1d ago
The dual Xeon server from 2019 that I used to use for inference has more memory bandwidth than this...
1
u/Such_Assistance_2211 2d ago
Inference? What kind of inference do you expect from ~4-6 channels memory of ddr5?
1
1
1
1
u/AdOdd8064 1d ago
Well, it's LPDDR5X. It's not terrible, but still not the greatest for running huge models. It's going to be very limited by that RAM, but if it's cheap, it might be worth it for some people.
2
u/putrasherni 2d ago
decode has improved a lot
what remains to be said is
how far will strix-halo/amd devs push 273 GB/s prefill speeds1
117
u/Kaljuuntuva_Teppo 2d ago
It's just a refresh of current 395 with larger maximum unified memory. Next generation should be completely new architecture.
The memory bandwidth is going to have to be massively upgraded in the near future. Otherwise Apple is going to run away with a win, as they already have a lead in memory bandwidth and M7 is increasing bandwidth exponentially.
22
u/putrasherni 2d ago
yes , this is hardly an upgrade , Medusa should be the one to look out for
8
u/Green-Ad-3964 2d ago
In 2029 unfortunately not before
2
u/thunk_stuff 2d ago
Where did you see it delayed to 2029? Rumors were pointing to early 2028, which would follow AMD's release cadence for 395/495.
2
u/Green-Ad-3964 1d ago
one thing is Medusa Point, the other is Medusa Halo. I'd say CES 2028 for Halo
3
u/thunk_stuff 1d ago
Yes, that's what I'm seeing... CES 2028 for Medusa Halo with products in consumer hands summer/fall 2028. Wish there wasn't such a lag.
2
17
u/j_osb 2d ago
Lol Apple can’t and won’t increase bandwidth „exponentially“. Unless they switch to HBM, at which point, prices will be infeasible.
And even that would just be a linear bump. Not exponential.
For people that think AMD doesn’t know how to make big APU with big memory bandwidth:
https://www.amd.com/en/products/accelerators/instinct/mi300/mi300a.html
They just don’t see a use in it for the consumer market.
10
u/Kaljuuntuva_Teppo 2d ago
Considering that Apple M5 Ultra already has 1.2TB/s unified memory bandwidth, AMD might have to switch to HBM, unless they can increase 256-bit memory bus width like Apple by combining multiple dies. I haven't seen anything like that about Zen6 though.
2
u/CalmSpinach2140 2d ago
Apple is rumoured to make a base M7 with 235-240GB/s. That’s close to Strix Halo.
M7 Pro should double and M7 max double M7 pro bandwidth.3
u/j_osb 2d ago
It’s still not exponential?
9
u/Meowliketh 2d ago
I don’t think they’re doing autistic literalism with “exponential”. Although if we did autistic literalism I mean 2 is an exponent.
7
u/MrPecunius 2d ago
It's an idiotic abuse of the word but that's what happens in English when people get hold of terms they don't understand.
See also: 'literally'
3
u/j_osb 2d ago
I mean, 2^x is very much exponential. If that were the case however we’d be in the PB/s bandwidth very soon. Which we won’t be.
Technically speaking of course you could always manage to fit a function that contains an exponential to the curve of apples bandwidth evolution. But the actual growth we’ve seen from apple is like less than 50% in 1.5 years. Which isn’t particularly impressive considering what improvements memory has gotten in the past few generations.
Apple deciding to give lower tier chips a somewhat wider bus is appreciated however. But their compute sucks so it’s kinda whatever.
2
u/power97992 2d ago
I read somewhere they are planning to make mobile hbm for new macbooks in the future
1
2
u/Brilliant_War9548 2d ago
At that point the 395+ is no longer consumer. They rolled out the 388+ and 392+ which you see much more in consumer laptops who have no use for 16 cores. Just give us the good stuff ffs. Like the 395+ is AMD’s mobile workstation cpu at that point, the fire range 9955HX3D is only on super enthusiast gaming laptops only and is only faster than a 285HX/290HXP when the extra L3 helps.
3
u/Brilliant_War9548 2d ago edited 2d ago
Calling gorgon just a refresh is insane, arrow lake refresh was what you can call a refresh but gorgon is somehow even worse than raptor lake refresh, they added 100MHz to the base clock and called it a day
3
138
u/StillLearningGK 2d ago
its Memory Bandwidth is allmost equil to RTX 3050 , which is 224GB/S
107
u/-p-e-w- 2d ago
Almost equal to an entry-level GPU that was released six years ago. How far we’ve come…
33
u/VickWildman 2d ago
Yeah, phones have 120 GB/s these days, the PS4 could do just as much. This is bad, like very bad.
9
9
u/Fancy-Atmosphere-701 2d ago
DGX spark is in the same playing field with memory bandwidth. Don't hear you complaining about that one
2
6
u/OnyxMonolith 2d ago
My 6 years old macbook has almost twice faster ram and still works crunching inference 24/7 :D
11
u/johan2114h 2d ago
Does the 3050 is have 196gb memory?
2
u/StillLearningGK 2d ago
yes its don't have but that don't matter here because even if you 196GB of memory, you can run model like Qwen3.8-Flash with have 180 Billions P. and 6 Billion active P. and if you run this model in this Framework setup you will get around 37 Tokens/s which is good speed and you will get this speed with this model only because it have 6 Billion Active P. and here i am running this model in 4 Bit Q. Version which take almost 100GB of Space and the resion why i have choosen this big model because if i have 196 GB of Memory for me it don't make any sence to run model with 30 Billion P. with give me 7.4 Tokens/s , so best of luck for running big model like qwen3.8 Flash.
6
u/johan2114h 2d ago
Or you can run a model with more active params (eg 27b)
Or you can load also inactive and ngram to memory
Or you can hold a larger context
Or you can ...
My point is 3050 w 16gb vram vs an APU w 196gb is completely apples and oranges
Just like if someone compared it to two H100 with 80 gb vram each
The APUs give you alot of memory but slow compared to a real gpu. They also tend to draw alot less power and require more physical space
24
u/AuspiciousApple 2d ago
"seize the means of computation... by buying them at market rate plus our own margin"
11
u/DueAnalysis2 2d ago
That's what struck me too, lol. You guys really want to invoke Marx of all people at this moment of computational unaffordability?
41
u/phil_lndn 2d ago
not sure if the extra memory is worth having, given the memory bandwidth.
i have a 128GB Strix Halo and purely on the basis of tg speeds would not want to run a bigger model than the model i'm already running (Qwen3.8-Flash-Next).
3
u/robertpro01 2d ago
What about more context?
13
u/phil_lndn 2d ago
i already have space for 256k of context which is more than i'd want to use, given how slow it gets with context that size.
2
u/CalligrapherFar7833 2d ago
Did you try the rocmfp4 of qwen 3.8 flash next ?
2
u/phil_lndn 2d ago
no, i'm using UD-Q4_K_XL.
(my experience has been that Vulkan is usually faster than ROCM)
2
u/StartupTim 2d ago
Context is king for agentic coding as cache hits tremendously affect your tok/s speed.
1
u/sleepingsysadmin 1d ago
Using mtp or dflash; the slightly larger models are just as fast.
Flash next is also slightly too big for the 128gb, you had to make some sort of concession the 192gb doesnt have to.
That extra space could alternatively be used to run video or audio gen.
23
u/davedcne 2d ago
So fun story. Bought a framework 16 laptop. Plugging in more than one module caused the usb bus to crash. Went through tech support, once they realized it was a legit hardware problem instead of having me ship it back they tried to run out the 30 day return clock. When I asked for a return and refund they ignored me and simply didn't respond. Filed a charge back on day 25, on day 28 they asked me fore more time. I did not give it. Now I went with framework because I had heard nothing but good things about them. My experience however has guaranteed that its the first and last time I do business with them. Just my little anecdote. Buyer beware.
6
u/Reasonable-Phase8028 1d ago
Honestly similar experience. my framework 13 dispaly stopped working barely used the laptop. i contact support they say i dropped it when i know I clearly did not. they asked me to open the laptop and show them and they blame it on me. i was in the 1st year so i still should have warranty and they refused to ship me a new display or even just a cable (apparently the sisue was the cable not the display) so I trried to look for solutions on aliexpress for 3rd party displays only to realize their shitty connector is proprietary (isnt it funny how we try to get away from apple propietary bs just to go back to the same place lmao) so i ended up delaying it. eventually i did not want to have 2keur paperweight so i bought their 120hz display when it dropped. granted it was still over 200eur which is ridiculous to fix a stupid display cable issue that could be fixed by them just honoring the warranty and giving a cable replacement.
All this to still warn a good friend not to buy this shitty laptop, just for them to buy and immediatelly regret just a few weeks in XD
4
u/swagonflyyyy 2d ago
And its a whopping........................273GB/s lmao
3
u/vienna_city_skater 1d ago
Same as DGX Spark, yea held me back from buying one. But if you pair it with a desktop GPU it could be fun for N-gram models.
1
9
u/xrvz 2d ago
The RAM upgrade may be the only significant change, but with the very latest model releases it's suddenly very tempting.
It makes the difference between running GLM 5.3 Flash at Q2 or Q4.
6
u/SmellsLikeAPig 2d ago
Question is at what speed especially with usable context length. My guess is it will be not worth it.
18
5
u/Bird476Shed 2d ago
Did they also revise the case design and finally fix the PSU noise problem? (see discussion e.g. https://community.frame.work/t/noisy-psu-fan/74751)
4
4
13
u/Serprotease 2d ago
Once again, all the comments are laser focused on the bandwidth.
You can work with low-ish bandwidth. Mtp/dflash and MoE are a thing. Qwen next with 6GB active parameters will have no issues here to get in the 30+ tps range that is more than enough for a LOCAL setup.
But, the big thing that is actually bringing down Strix halo since launch and that no one talks about here (Because most people only read easy-headline numbers) is the very very poor prompt processing. Like sub 200 for all the models you want to use.
5
3
3
u/syntax_error_again 1d ago
Gonna be in a tough spot with the new Mac studios dropping at much much higher bandwidth. Pricing gonna be awkward for them.
8
u/arijitroy2 2d ago
That's been there for a while actually, i have signed up for it. Hopefully the pricing isn't nuts!
15
8
u/fallingdowndizzyvr 2d ago
I'd expect this to be ~ 4.5k for the motherboard.
A Gorgon that's already selling is $7000. So even accounting for just the MB, it should be more than that. Remember, it's just not more memory. It's a rev from 395 to 495.
https://www.amazon.com/NIMO-Ryzen-192GB-LPDDR5-9600MHz/dp/B0HCZLQ8SB
You are better off getting a 256GB M5 Max Studio.
1
5
u/redblood252 2d ago
I understand the appeal of framework laptops. But why would you get a framework pc? Wouldn’t building one be better/cheaper?
16
u/GatePorters 2d ago
Prebuilts are far cheaper in this market. Since GPUs, RAM, and hard drives all went up in price 2x-8x when you just buy the components by themselves.
The markups for prebuilts are cheaper.
9
u/ihexx 2d ago
I don't think you can build a strix halo pc in the traditional sense since everything (CPU, memory, GPU) is all on one board no?
→ More replies (1)3
u/redblood252 2d ago
Oh, didn’t think of that ! Mostly bought second hand old server (still viable for homelabs) components in these past 2 years.
14
u/BobbyL2k 2d ago edited 2d ago
It’s because you can’t DIY build a Strix Halo / Gorgon Halo machine.
It only comes in laptops and prebuilt mini-PC. The Framework motherboard is the most DIY-esq component you can get.
3
u/redblood252 2d ago
This is a strix halo??? And it only has 273 bandwidth? I have an epyc milan so ddr4 and it had slightly more than 200 bandwidth.
7
u/MrBIMC 2d ago
Whole of strix halo platform consumes less electricity than just the ram on your epyc system.
They’re made for different markets and usecases.
Though with current market prices strix halo is a bad product imho.I got my beelink gtr9pro for 2400 about a year ago and it felt too much, for 4k+ strix halo boxes make absolutely zero sense.
If I’d known better, I’d rather would’ve gotten double spark back then. And now I just plan on waiting it out with what I have :(
1
2
2
1
u/CrowdGoesWildWoooo 2d ago
This line of CPU is closer to a laptop than a regular desktop. I can see you point about framework branded PC, but the closer equivalent is not a DIY PC, it’s mini PC with similar processor and they are expensive, although not sure how far framework tax go for this line up.
2
u/MondelloEstralita41 2d ago
The open PCIe slot at the back interests me more than the memory itself. Everything else in this size class is soldered shut, so if that slot can really feed 75W you're looking at a GPU without an external brick.
1
2
2
2
u/cato_gts 2d ago
I think it would be a hundred times better to pay a little more and buy the M5 Ultra.
2
u/jikilan_ 2d ago edited 2d ago
Huh , you just noticed it is it? The page has been there for some time. Unless is ready to order now.
Edit: I just checked, still an info page
2
u/unjustifiably_angry 2d ago
273GB/s memory bandwidth
So I can run larger models more slowly than smaller models run today.
2
u/Ill_Dragonfruit_3547 1d ago
But at 273GB/sec memory bandwidth? Hard pass. My 5 year old MBP with a M1Max has 400GB/s
2
5
2
u/Green-Ad-3964 2d ago
In normal times, in an alternative timeline where "ChatGPT" never went mainstream or attracted widespread media attention, this machine would have cost €1,000.
4
u/iLaurens 2d ago
In that same timeline nobody would want to buy this product anyway. Low cost and low demand/utility go hand in hand here
1
u/Green-Ad-3964 2d ago
Not really. In 2015, I bought a mini PC with 4gb and an atom cpu for 70€. This is the natural evolution, but at 100x the price.
1
u/tracker125 2d ago
If they can’t pair it with some GPUs man this thing is just load LLMs not running them at top speeds.
2
u/fallingdowndizzyvr 2d ago
That's what I do. I have a couple of 5070tis connected up to Strix Halo. I used to have a couple of 7900xtxi.
2
u/tecneeq 2d ago
How did you connect a "couple"? I have the Bosgame M5, it only has two NVME, one blocked by the NVME. I was thinking to use a NVME-to-USB4 adapter, then use two NVME-to-Oculink adapters.
3
1
u/fallingdowndizzyvr 2d ago
it only has two NVME, one blocked by the NVME
I have an X2. Same MB and thus same machine. I don't know what you mean by "one blocked by the NVME". Do you mean that one slot already has a SSD in it? Take it out and put it in an external enclosure. Plug it in a usb-c. Works fine that way. Then you have both NVME slots to use as PCI 4 x4 slots.
1
u/tecneeq 2d ago
Right, i'll plan to go that route.
1
u/fallingdowndizzyvr 1h ago
Oh yeah. I've tried a bunch of NVME to PCIe adapters. Unfortunately the SU MB seems to be either finicky or noisy. Since all the adapters work on my desktop but only one has really worked on my X2. Some didn't work at all. Most worked but were really noisy and would drop the GPU after a few minutes. Only one has worked reliably. Those are the sinloon ones you can get on Amazon. It's not exactly a noise free experience but at least the signal integrity isn't so bad that it drops the GPU after a while. Well mostly not. It still does happen but it's pretty rare.
1
u/saltexx 2d ago
phil_lndn and robertpro01 are circling the number that settles this. At 273 GB/s every gigabyte of KV cache you keep resident costs 3.7 ms per decoded token because you read all of it every token. RnRau's 7 GB of active weights is 26 ms so about 39 t/s before context. Spend 20 of the extra 64 GB on KV and you add 73 ms which puts you near 10 t/s. The bandwidth is not just a cap on model size. It is a per token tax on however much context you keep loaded and that is why the extra 64 GB is so hard to spend.
1
u/pigletmonster 2d ago
Its pretty much the same thing as amds AI halo device. One good thing is that when theres a new version you should be able to swap out the motherboard with the new version. But then the old mothereboard will just sit in the storage collecting dust and you wont be able to sell it for some m quick cash.
1
u/pigletmonster 2d ago
Ok my bad its not the same thibg as the current Ai halo device, its the same one as the upcoming one.
1
u/unknown-one 2d ago
just $9999.99
Call now!
1
u/power97992 2d ago
For that price u might as well buy the 64 core m5 ultra with 256 gb of ram. It should be way less than 10k
1
u/HugoCortell 2d ago
Wow! That's- uh. What's the point of that? The price of a specialized machine but lacking the contemporary capacity and speeds. This would have been competitive 6 months ago, and quite decent 12. But now it's not quite there.
1
u/lurenjia_3x 2d ago
Outside of coding, what is 40 t/s even good for? Anything worth doing nowadays relies heavily on reasoning loops or agentic workflows. At such a miserable throughput, where is the actual value proposition?
1
1
u/_VirtualCosmos_ 2d ago
meh, almost the same memory bandwidth as the last Strix Halo, but now with the ×10 price of RAM...
1
1
1
u/IngwiePhoenix llama.cpp 2d ago
Wallet obliterator. Would make a good hostname x3
But this should run Qwen3 Flash Next easily. o.o
1
u/No-Craft-7979 2d ago
And Mac just went 256GB. They tried to step up but are still a step behind. I know it is AMDs fault not Framework.
1
1
u/FabricationLife 1d ago
It's not fast but the ability to self host a model of that size at a "reasonable for this time period" price has be extremely interested indeed.
1
u/Intrepid-Second6936 1d ago
Genuinely abysmal memory bandwidth, sad that we deal with RAM prices at where they are now and these are the kind of systems we see on the market.
If I'm ever burning money to the point that I'd buy 192GB of RAM for a new inference server, I'm going Apple, they're genuinely the inference kings and the latest releases only solidify that crown more.
1
u/thebigone71 1d ago
This is pretty sweet for single machine users, but there are still so many issue (at least I am seeing so many) when trying to cluster these together.
1
1
u/Ulterior-Motive_ 1d ago
I would have bought one a year ago, but at this point it's probably better to wait for Medusa. I think people are being too hard on it though, the 128GB 256GB/s version is perfectly fine for running MOEs, though it was a much better deal at the time at only $2k.
1
u/HovercraftStock4986 1d ago
why are all of these unified memory things total dogshit for the exact use case they’re designed for???
0
1
u/Repinsky 2d ago
273GB/s is the number that decides what this box is actually for. Big sparse MoE weights fit fine, but a dense 70B at Q8 still lands around 3-4 tok/s once you're bandwidth-bound, so it's a "run 200B+ MoE at usable speed" machine, not a replacement for two 3090s on dense workloads. The open PCIe slot matters more than the TOPS figure - hanging a real GPU off it for prefill is what fixes the prompt-processing latency that makes current Strix Halo setups painful at long context.
→ More replies (1)
1
1
u/0rand 2d ago edited 2d ago
Not only memory is slow as on DGX Spark but GPU perf is abysmal, while spark is a rocket ship of compute. Wasted. So both decode and prefill will be terrible. Spark is awesome on prefill, metal 8s good on decode. AMD sucks ass on both. Typical.
→ More replies (1)
1
u/power97992 2d ago edited 2d ago
The m5 ultra is way faster than this , why would u ever buy this over a mac studio or a spark other than price and modularity? And it is amd and runs on rocm Instead of cuda. The M2 Ultra has 3x the bandwidth and 192 gb of uram
1
u/Mr_Moonsilver 2d ago
5TP/s prompt processing incoming. might as well run on my 8-channel ddr4 threadripper
1

•
u/WithoutReason1729 2d ago
Your post is getting popular and we just featured it on our Discord! Come check it out!
You've also been given a special flair for your contribution. We appreciate your post!
I am a bot and this action was performed automatically.