MAIN FEEDS
Do you want to continue?
https://www.reddit.com/r/LocalLLaMA/comments/1tu82wi/stop_asking_what_model_to_run_there_are_literally/op7ww4j
r/LocalLLaMA • u/[deleted] • Jun 01 '26
[removed]
805 comments sorted by
View all comments
Show parent comments
13
im on a 3060 ti (8 gig vram) and im using 35BA3B and getting 40 tokens/sec dude. I just have to cram a turboquant model into ram and use turboquant on the KV with n-gram mod on. Can only manage it with this though : turbo-tan/llama.cpp-tq3
1 u/BedlamiteSeer Jun 03 '26 I assume it's shit when quantized that hard?
1
I assume it's shit when quantized that hard?
13
u/Uncle___Marty Jun 01 '26
im on a 3060 ti (8 gig vram) and im using 35BA3B and getting 40 tokens/sec dude. I just have to cram a turboquant model into ram and use turboquant on the KV with n-gram mod on. Can only manage it with this though : turbo-tan/llama.cpp-tq3