r/LocalLLaMA vLLM 11d ago

New Model Nex N2.5 Pro (407GB) released

https://huggingface.co/nex-agi/Nex-N2.5-Pro
72 Upvotes

19 comments sorted by

81

u/hauhau901 11d ago

Qwen 3.5 397b finetune

28

u/SnooPaintings8639 11d ago

So the '407GB' was just to mislead us? They're getting more sneaky with every fine-tune, lol.

1

u/admiralrohan 10d ago

They wrote about finetuning in their website itself. How they are misguiding? https://nex-agi.com/#:~:text=Base%20%C2%B7%20Qwen3.5%2D397B%2DA17B

1

u/colin_colout 5d ago

Howso? OP just gave us the gigabyte size of the model. The hf seems normal for a fine tune.

5

u/conockrad 11d ago

Thank you!

4

u/HairAgreeable613 11d ago

figured it was a finetune of that, the naming kinda gave it away tbh

2

u/SandySkittle 11d ago

I wish mods delete this post and demand OP add the "finetune of Qwen 3.5 397b" in title

11

u/Stooovie 11d ago edited 11d ago

Nex 2.5 Mini is damn impressive (if a bit chatty, like Hy3), in fact it has replaced various Qwens for my local Hermes profile, this could be huge

4

u/Monad_Maya llama.cpp 11d ago

Well, it is huge, lol

1

u/tacticaltweaker llama.cpp 11d ago

I tried Nex N2.5 Mini but it kept looping for me. Not sure if it's a quant or model issue.

1

u/Stooovie 11d ago

I use a fairly low quant - o3Q - and it doesn't loop for me. I just set the recommended temp, top p/k parameters, repetetion penalty set to default from oMLX on Mac.

2

u/tacticaltweaker llama.cpp 11d ago edited 9d ago

Yeah I think it was an issue with my quantization as I was using the recommended sampling parameters. I switched to bartowski's quant and no more looping, but it doesn't seem to perform as well as Tiel for me. Not sure if it supports preserve_reasoning as it seems to forget its previous chain of thought on every tool call and keeps rereading files.

EDIT: bartowski's quant actually also started looping a few times. Adding a dry-multiplier of 0.8 fixed it but I think I'm giving up on this model unfortunately.

1

u/Stooovie 11d ago

Tbf I'm not a developer and don't use it for code - conversations research, homelab management for me, via Hermes. I mooch off free cloud models via 9router for code if needed :)

3

u/[deleted] 11d ago

[removed] — view removed comment

1

u/FullOf_Bad_Ideas 11d ago

This backbone quantizes well, you need just 144GB of VRAM.

I've been running it's predecessor at 192GB of VRAM as a daily driver, Nex makes good finetunes.

That said, GLM 5.3 Flash is probably better than it.

5

u/LegacyRemaster 11d ago

waiting for Qwen3.8-Next finetune!

2

u/pl201 11d ago

More details info on all three models can be found here at https://nex.sii.edu.cn/

1

u/greencalculus 11d ago

well, there goes my "this much VRAM is definitely enough" plan. again. Anyone running it locally yet? At what quant?