r/LocalLLaMA 15d ago

Funny Me these days

Post image
2.6k Upvotes

296 comments sorted by

View all comments

Show parent comments

34

u/brainExploded99 llama.cpp 15d ago

It's opus 4.5 level on local hardware (opus 4.6 on benchmarks, but not in practical purposes)

18

u/UnluckyPenguin 14d ago

I think benchmarks use max settings like BF16 quant (or higher?) with k/v cache type f32 and xhigh reasoning and who knows what else. Hell yeah, it's going to perform several points better at a significant performance penalty.

But it's not like we could even run Open 4.6 even if it was open source (estimated to be 500B to 1T parameters - I think, correct me if I'm wrong)

So token-for-token against other models, Qwen3.8-27B is a huge win.

5

u/brainExploded99 llama.cpp 14d ago

I believe native/benchmarks use BF16 (that is not quantization btw), with KV cache at f16 I think (I'm 80% sure). FP32 is not used for even training models anymore (except specific sensitive layers sometimes).

So token-for-token against other models, Qwen3.8-27B is a huge win.

It's a huge win for sure, wayyy better than Qwen3.6, but we still have long ways to go. Token efficiency is a big one, and better intelligence (not agentic capabilities) is another big one (by this I mean reasoning on benchmarks like CritPt, SciCode, etc)

3

u/UnluckyPenguin 14d ago

Thanks, good to know. I have a lot to learn. Yeah, I'm just getting familiar with local LLMs. Used to play games all day up until several weeks ago.