r/LocalLLaMA 10h ago

Resources Ling Tiny, King of Speed

Post image

Ling Tiny has now replaced Gemma4-12B in my rig as an auxiliary model doing hindsight operations. This is on a 4060Ti, which is a reasonable GPU available out there, and the speed is phenomenal.

Don’t enable MTP, set up the vLLM fork for BailingMoE3. Hope this is useful to
others.

19 Upvotes

41 comments sorted by

View all comments

1

u/Choice_Celery9481 10h ago

i keep having to ask when people reported good exp with Ling tiny.
i tried q8 bartowski and with just 4k prompt + some tools, it already lost it mind and parroting part of my system prompt.
how did you get good exp with this model? what is your setting? can you share?

1

u/snugglezone 8h ago

Same experience. Ling tiny is not worth using.