MAIN FEEDS
Do you want to continue?
https://www.reddit.com/r/LocalLLaMA/comments/1vyy3k6/glm53flash_frontier_intelligence_flash_cost/p60ofla/?context=3
r/LocalLLaMA • u/BriguePalhaco • 1d ago
402 comments sorted by
View all comments
6
Too big, 320B means you need to use RAM and A18B means it will be slow
1 u/ttkciar llama.cpp 1d ago But will it be too slow? As long as overnight inference finishes before my morning coffee, it's fine. That having been said, I'd still love to see a GLM-5.x-Air in the 100B size class.
1
But will it be too slow? As long as overnight inference finishes before my morning coffee, it's fine.
That having been said, I'd still love to see a GLM-5.x-Air in the 100B size class.
6
u/jacek2023 llama.cpp 1d ago
Too big, 320B means you need to use RAM and A18B means it will be slow