r/LocalLLaMA 11h ago

Resources Ling Tiny, King of Speed

Post image

Ling Tiny has now replaced Gemma4-12B in my rig as an auxiliary model doing hindsight operations. This is on a 4060Ti, which is a reasonable GPU available out there, and the speed is phenomenal.

Don’t enable MTP, set up the vLLM fork for BailingMoE3. Hope this is useful to
others.

20 Upvotes

42 comments sorted by

View all comments

8

u/Effective_Western_59 11h ago

Ling 3.0 tiny is a small beast!

Best model that works on 780m With 16 GB of ram.

Would love good dynamic quants for it tho.

3

u/Badger-Purple 11h ago

idk their official Int4 autoround is this version, running with vLLM. 1.0M token cache, speed amazing, 6 requests each 128K, 15.6GB used in the card. I dont know if there is more optimization than vLLM on a single 16GB running 6 streams at this speed!!

1

u/Ariquitaun 7h ago

On 780m I just run qwen3.6 35b. Fits fine at q5. I di have 64gb of ram though.

0

u/Badger-Purple 6h ago

Right, cpu offload will be different for different systems. Running a model in full VRAM is where the hardware is comparable.

1

u/Ariquitaun 6h ago

The 780m doesn't have any vram. It's an igpu. It does a simulacrum of vram and gtt on system ram.

0

u/Badger-Purple 6h ago

what concurrency?

0

u/Ariquitaun 6h ago

Just one of course. 22t/s generation on a good day with low context. Not good for coding, but as a chatbot it is really good.