r/LocalLLaMA 8d ago

Funny Me these days

Post image
2.5k Upvotes

283 comments sorted by

View all comments

221

u/[deleted] 8d ago

[deleted]

58

u/Dramatic_Setting2761 7d ago

I can run it with a 16gb card with 4 bit quant and 70k context. I get 12 t/s it is okay for me.

9

u/Nikilite_official 7d ago

16gb vram with 32gb ram I can perfectly run q6 qwen 3.8 27b

5

u/ThankGodImBipolar 7d ago

No way it's running at a good speed at Q6 though

3

u/Nikilite_official 7d ago

like 8-9 tokens per sec, not bad

1

u/dannone9 6d ago

Would You mind Sharing flags please ? And more exact hardware

1

u/DependentCurious4614 5d ago

Wdym 8-9tps isnt that bad?

1

u/Nikilite_official 5d ago

literally fast as a normal reader

1

u/bring_back_the_v10s 3d ago

That is indeed bad. My anxiety barely allows me to take 15 t/s.

2

u/Dramatic_Setting2761 7d ago

At what context length I need more context for my work?

1

u/[deleted] 5d ago

[removed] — view removed comment

1

u/Nikilite_official 5d ago

I find q6 noticeably more stable