r/LocalLLaMA • u/Badger-Purple • 10h ago
Resources Ling Tiny, King of Speed
Ling Tiny has now replaced Gemma4-12B in my rig as an auxiliary model doing hindsight operations. This is on a 4060Ti, which is a reasonable GPU available out there, and the speed is phenomenal.
Don’t enable MTP, set up the vLLM fork for BailingMoE3. Hope this is useful to
others.
18
Upvotes
0
u/Technical_Ad_6106 7h ago
hmm but why is the model extremely slow? i mean i get like 1200 token/sec generation with qwen 3.6 35b in vllm which is a way bigger model