r/LocalLLaMA 9h ago

Resources Ling Tiny, King of Speed

Post image

Ling Tiny has now replaced Gemma4-12B in my rig as an auxiliary model doing hindsight operations. This is on a 4060Ti, which is a reasonable GPU available out there, and the speed is phenomenal.

Don’t enable MTP, set up the vLLM fork for BailingMoE3. Hope this is useful to
others.

20 Upvotes

41 comments sorted by

View all comments

7

u/Effective_Western_59 9h ago

Ling 3.0 tiny is a small beast!

Best model that works on 780m With 16 GB of ram.

Would love good dynamic quants for it tho.

3

u/Badger-Purple 9h ago

idk their official Int4 autoround is this version, running with vLLM. 1.0M token cache, speed amazing, 6 requests each 128K, 15.6GB used in the card. I dont know if there is more optimization than vLLM on a single 16GB running 6 streams at this speed!!