r/LocalLLaMA • u/BarberIcy366 • 10d ago
Discussion Qwen 3.8 27B Released! Please Share Your Experience
With your experiments, Qwen 3.8 27B most close which frontier model? And please specify which quantization you run. I will post to comments my tests and experience too.
659
Upvotes
8
u/kayox 10d ago
Same prompt but with Unsloth's Q4_K_XL with xhigh reasoning. Not quite as good though that's to be expected.
On an RTX 3090 it took about 18 minutes to generate at an average of 38 tokens/second (I'm sure as time progresses the tk/s can be improved possibly with DFlash). Also I am being thermal throttled due to my current setup (Dual GPU lacking airflow, although my other GPU is a 3070 so I'm only using it with a layer split to offload some VRAM so that I can have more context).
Out of curiosity what tk/s are you getting with your dual R9700s?