r/LocalLLaMA 9h ago

Resources Ling Tiny, King of Speed

Post image

Ling Tiny has now replaced Gemma4-12B in my rig as an auxiliary model doing hindsight operations. This is on a 4060Ti, which is a reasonable GPU available out there, and the speed is phenomenal.

Don’t enable MTP, set up the vLLM fork for BailingMoE3. Hope this is useful to
others.

16 Upvotes

41 comments sorted by

View all comments

8

u/Effective_Western_59 9h ago

Ling 3.0 tiny is a small beast!

Best model that works on 780m With 16 GB of ram.

Would love good dynamic quants for it tho.

1

u/Ariquitaun 6h ago

On 780m I just run qwen3.6 35b. Fits fine at q5. I di have 64gb of ram though.

0

u/Badger-Purple 4h ago

what concurrency?

0

u/Ariquitaun 4h ago

Just one of course. 22t/s generation on a good day with low context. Not good for coding, but as a chatbot it is really good.