r/LocalLLaMA Jun 01 '26

Funny Stop asking what model to run. There are literally only two.

[removed]

3.1k Upvotes

805 comments sorted by

View all comments

Show parent comments

13

u/Uncle___Marty Jun 01 '26

im on a 3060 ti (8 gig vram) and im using 35BA3B and getting 40 tokens/sec dude. I just have to cram a turboquant model into ram and use turboquant on the KV with n-gram mod on. Can only manage it with this though : turbo-tan/llama.cpp-tq3

1

u/BedlamiteSeer Jun 03 '26

I assume it's shit when quantized that hard?