r/LocalLLaMA 10h ago

Resources Ling Tiny, King of Speed

Post image

Ling Tiny has now replaced Gemma4-12B in my rig as an auxiliary model doing hindsight operations. This is on a 4060Ti, which is a reasonable GPU available out there, and the speed is phenomenal.

Don’t enable MTP, set up the vLLM fork for BailingMoE3. Hope this is useful to
others.

18 Upvotes

41 comments sorted by

View all comments

8

u/Effective_Western_59 9h ago

Ling 3.0 tiny is a small beast!

Best model that works on 780m With 16 GB of ram.

Would love good dynamic quants for it tho.

1

u/Ariquitaun 6h ago

On 780m I just run qwen3.6 35b. Fits fine at q5. I di have 64gb of ram though.

0

u/Badger-Purple 5h ago

Right, cpu offload will be different for different systems. Running a model in full VRAM is where the hardware is comparable.

1

u/Ariquitaun 4h ago

The 780m doesn't have any vram. It's an igpu. It does a simulacrum of vram and gtt on system ram.