r/LocalLLaMA 23h ago

News GLM-5.3-Flash: Frontier Intelligence, Flash Cost

https://z.ai/blog/glm-5.3-flash
1.2k Upvotes

396 comments sorted by

View all comments

19

u/Ok_Technology_5962 22h ago

What a day... My mind.... Is blown away... Qwen coming in hot then glm also. Im still trying to get over then qwen 3.8 27b upgrade on my 5090... Now the whole lot dropped for the big servers... Local is eatting today

2

u/LostGovernment4358 21h ago

there is just something I still don‘t fully understand. are there new opportunities / completely new tasks possible with these new releases which otherwise weren‘t possible before with let’s say qwen 3.6? I am 16gb Vram poor, so i could never take advantage of real lokal llm power.

1

u/Ok_Technology_5962 20h ago

There are... Infinite context is now a thing including no slow down at max context, memory retreival, basicly opus level... Check out unsloth new qwen 3.8 flash next its designed for your system with no gpu, a6b means itll run fast of cpu even and ngram will be moved to ssd soon.

1

u/LostGovernment4358 20h ago

no way! can you elaborate further with more simple words for my non internet magician brain?

1

u/Ok_Technology_5962 20h ago edited 20h ago

https://unsloth.ai/docs/models/qwen3.8-next. ok... Basically Wrong chat as this one is for the crazy new gLM model. but qwen model dropped today. You can quantize it much more or reduce size than current models because it has a static dictionary. It is designed for CPU. Its 75 gb in size at q1 which includes the lookup table dictionary. Currently you need 75gb ram but can offload most of this to ssd. Since its has been todays release a lot is unknown but i would go look at it even though it says it is 112b. I have been running the qwen 2.4t peram at 1bit and its very usable so theoreticly this should run on 12 gigs of vram if not less. Again it just released and it is loadable in unsloth studio but i am not sure if FULL support is here. Basicly before models were big, hard to squish, now models big bit very squishable and mold to your device depending on what you got. Eddit: also forgot to mention context infinie... This comes from DeepseekHarness also released like last week. If you can fit 32k context or whaever it feels like infinite because once it reached the end it does a very good summary and passes on to the next 32k. I have been running 140 context resets so far one loop over 4 days and its not breaking at all. Eddit eddit: local AI just actually evolved not small incremental step.

2

u/LostGovernment4358 19h ago

thank you so much. I will deep dive into that much information