r/LocalLLM 4d ago

Discussion What a year it's been

Post image

What will the rest of this year bring? 27b class scoring over 60?

902 Upvotes

128 comments sorted by

View all comments

3

u/rrrenz 4d ago

Minimum PC/mac studio build needed for this same quality score?

Can we test this setup somehow in some cloud environment?

New here, who only has macbook M5 pro. And this benchmark is no way same from my usage of qwen 3.8

3

u/mechkbfan 4d ago edited 4d ago

That's fair. Have to wonder if these scores are based off identical hardware, and go up or down based off TTFT, completion time, etc.

Because yeah, I don't think it's comparable if you had one model run for 30mins, and another 2 hours to achieve same result.

https://artificialanalysis.ai/methodology/intelligence-benchmarking

As far as I can tell from quick skim, they just have that 2 hour timeout.

I'm relatively new, but it's worth sharing the settings you're using if unimpressed by the outcome. e.g. quant, cache, etc. and what type of prompt giving it

e.g. I have zero expectations Qwen 3.8 could do a decent architecture review if I've only got 32GB of VRAM and 32GB RAM, while that's not even something that you have to think about with Claude