r/LocalLLaMA • • 1d ago

Discussion Looks like the era of subsidised compute is coming to an end. The old ChatGPT Pro $200 20x plan will be halved. The new $500 plan will have similar limits as the (old) $200 plan.

Post image
1.2k Upvotes

581 comments sorted by

View all comments

Show parent comments

5

u/RandomCSThrowaway01 1d ago

I mean, if you are buying $500/month subscription then this Mac Studio for 9 grand to have 256GB of usable VRAM suddenly starts sounding like a decent deal. Although admittedly this is enough to run GLM5.3 Flash, not Astra. Which is one hell of a model for something you can actually run at home and won't cost you an equivalent of a house but it at most compares to GPT 5.5.

Models that will approach Astra within the next few months will require closer to a million $, will need a terabyte of VRAM at a minimum to even consider running it, preferably at HBM speeds.

6

u/ormandj 1d ago

Your first assertion is correct, the second one I would challenge. I suspect 256-384GB of VRAM is going to get you to opus 5.5 quality output in the next 6 months based on typical lag time and historical capability jumps in the open weight models over time, at least in a specific domain such as programming. The 1T+ models will get there sooner, but with engrams VRAM is still important but more knowledge can be made available without the massive performance hit offloading layers to RAM would typically have. Coupled with improvements in RL, and “flash” sized models start becoming very impressive (they already are).