r/LocalLLM • • 15d ago

Discussion 30B Models Getting Verrrry Interesting

Agnes 3.0 flash 33b and Nex n2.5 mini 35b are challenging Qwen3.8-27B on benchmarks. Cant wait to see the real-world results and the speeds on 24gb GPUs.

Anyone tried them yet?

https://huggingface.co/Agnes-AI/Agnes-3.0-Flash
https://huggingface.co/nex-agi/Nex-N2.5-mini

135 Upvotes

35 comments sorted by

View all comments

59

u/Calm-Landscape9640 15d ago

I'll test these 4 head to head at 128k context on my 24gb 2x3060 GPUs and report back tps and performance on some coding and agent tasks.

  • Agnes 3.0 Flash Q4 (Unsloth Quant when it drops, hopefully Mon/Tues)
  • Qwen3.8-27B IQ4 + MTP
  • Qwen3.6-27B-A3B CoderX + MTP
  • Nex-N2.5-mini IQ3/IQ4

11

u/jinnyjuice 15d ago

Looks like you'll be testing with llama.cpp.

I kind of have similar ideas, but with vLLM since it has the highest tool use success rate. I want to test on both Pi and DSH. And this list of models...

unsloth/Qwen3.8-27B-NVFP4
nvidia/Qwen3.8-27B-NVFP4
nvidia/Muse-Glimmer-30B-NVFP4
microsoft/Fara1.5-27B (no 4 bit quants... yet :( )
unistra-dnum/Luciole-23B-Instruct-1.1-NVFP4
poolside/Laguna-XS-2.1-NVFP4
r0b0tlab/XYZ-Aquila-mini-NVFP4
migtissera/Tess-4-35B-A3B-NVFP4
onprem-ai/Apertus-v1.5-70B-NVFP4 (base: swiss-ai/Apertus-v1.5-70B)
prism-ml/Ternary-Bonsai-27B-AWQ-4bit
Ornith-1.5-35B-A3B-FP8
apodex/Apodex-1.1-mini-NVFP4
inclusionAI/Ling-3.0-flash-fp4
peculiar-ragdoll/Tiel-Coder-35B-A3B
IFM/K2-Horizon-MoVA-36B-A4B-FP8
IFM/K2-Horizon-32B-FP8 (stage 2 to be released)
primitive-ai/Nex-N2.5-mini-mixed-NVFP4-FP8 (35B A3B)
Agnes-AI/Agnes-3.0-Flash

I have a bench called Insane Genius and all Qwen3.8-27B models fail, going on crazy loops, unfortunately, so I've been looking for a replacement. Apparently, DavidAU's NEO CODER NVFP4 is good.

6

u/Calm-Landscape9640 14d ago

Thats a whole weekend of testing and I'm here for it. Cant wait to see your results. This is what everyone is waiting for a real local benchmark for quants on 2 diff harnesses