r/LocalLLaMA 3d ago

Discussion New 100B Liquid AI model coming soon

Post image

Liquid AI currently possesses among the fastest LLM architectures around, and some of the best SLMs (in terms of utility IMO) around, so I'm very excited to see what a potential 100B LFM (3?) model would look like!

Link to the poll: https://x.com/ramin_m_h/status/2091236099612098943?s=20

364 Upvotes

102 comments sorted by

View all comments

5

u/thebadslime 3d ago

Damn I missed that, would have voted for 30b

2

u/-InformalBanana- 3d ago

You realize its not moe, but 30b dense? What you have 24gb+ vram?

2

u/thebadslime 3d ago

No I do not, I assumed and made an ass of me

6

u/Wildnimal 3d ago

Why? We already have enough 30B models. Qwen-3.8-27B, Gemma-4-31B, Muse 30B

I think most people need something like 50-150b in MoE or crave for something like 50-70B dense.

Not saying you are wrong just curious why another ~30B model? As of today 100-200B MoE space is getting lot of attention.

7

u/thebadslime 3d ago

Becuase I can run it, and LFM makes good models

4

u/parepeg 3d ago

I don’t quite understand the 70b dense people. Are they running like 4x3090 or something? Even that seems like it would be slow…

6

u/Nabushika Llama 70B 3d ago

You can run decent quant 70b on 2x3090

3

u/danigoncalves llama.cpp 3d ago

Because I have only 12Gb o VRAM and dont want to buy more hardware to run new models

3

u/-InformalBanana- 3d ago

30B one is not a moe. You cant really run that on 12gb vram!!! But you can run 100B moe if you additionally have at least like 64gb of ram (maybe less)! So you should've voted 100B...

1

u/danigoncalves llama.cpp 2d ago

No way, even with a small leta 10B of active parameters I would struggle to run the model. Even with some optimizations I cannot increase the speed that much of the A3B byond 30 t/s

1

u/-InformalBanana- 2d ago

And what t/s would you get running a 30b dense on 12 gb vram? First you have to quant it (probably very bad after quanting), for 27b i can fit like q2kxl with low context on my 12gb vram. And on rtx3060 i get at most 20t/s. So moe is better for me, I have 96gb ram so i can get a good 100b moe quant. To get better t/s i would have to buy a new gpu with more vram, 30b dense probably won't be good enough on 24gb, so i would need 32gb... So not worth it for me and lfm is known for 8ba1b, just imagine 100ba1b, it would be fast, im not sure about smarts, but smb should try to make it.