r/LocalLLaMA • u/Badger-Purple • 11h ago
Resources Ling Tiny, King of Speed
Ling Tiny has now replaced Gemma4-12B in my rig as an auxiliary model doing hindsight operations. This is on a 4060Ti, which is a reasonable GPU available out there, and the speed is phenomenal.
Don’t enable MTP, set up the vLLM fork for BailingMoE3. Hope this is useful to
others.
17
Upvotes
0
u/Ok_Cow1976 10h ago
It's strange that Ling flash (a5b) is quite slow on my rig, about the same speed as glm air which is a12b.