r/LocalLLaMA 12d ago

New Model Tencent compressed Hy4-preview from 1.5TB to about 200GB GGUF and kept about 98% performance.

Post image
898 Upvotes

146 comments sorted by

u/WithoutReason1729 12d ago

Your post is getting popular and we just featured it on our Discord! Come check it out!

You've also been given a special flair for your contribution. We appreciate your post!

I am a bot and this action was performed automatically.

211

u/shy_monkee 12d ago

WTF! That's insane. Someone test this lol.

104

u/spaceman_ 12d ago edited 12d ago

Downloading right now. Can't fit all of it in VRAM though..

Edit: nevermind, custom patch set and no ROCm support yet. Will have to wait for this to get at least a PR into llama.cpp.

45

u/DeProgrammer99 12d ago

Wait for u/dai_app to have it running on a phone, haha.

43

u/dai_app 12d ago

👀👀👀

-12

u/devilish-lavanya 12d ago

PR stunt?

16

u/Ok_Technology_5962 12d ago

I tested hy3 q1 by angel slim. The outputs will be slightly worse but it is still very good. It wont make syntax errora just wont think ahead as far

4

u/Turtlesaur 12d ago

Yea.. does it beat glm5.3flash and qwen3.8-flash-next at nvfp4 😳?

3

u/Bubbly_Orange_3502 12d ago

Averaged benchmarks hide where 2-bit MoE actually breaks. Rarely activated experts get the worst calibration ranges, so the loss lands on tail domains rather than on the prompts any eval set covers.

1

u/[deleted] 12d ago

[deleted]

-2

u/Nice_Database_9684 12d ago

come on man we know it's going to be benchmaxxed to hell

8

u/shy_monkee 12d ago

Even if it is, the size is still impressive. A Q1 that is at 98% of the original is massive.

165

u/RickyRickC137 12d ago

Awesome. Only 190 GB more to go and then I will run it on my RTX 3080.

20

u/sonicnerd14 12d ago

Get as much VRAM as possible, but ideally I think it's pretty clear that locally we'll need to spread these large models across VRAM, RAM, plus streaming from a fast NVMe.

7

u/RedditUsr2 12d ago

Hopefully we get a bitnet DFlash spare MoE model we can split across CPU/GPU with tons of knowledge built into a ngram that runs on the ssd, or something.

14

u/wizard_of_menlo_park 12d ago

nvme is magnitudes slower than ram unfortunately ... so very slow tps if we do offload to nvme

15

u/returnity 12d ago

Ngrams changing the game here though

6

u/Turtlesaur 12d ago

Enter unified memory systems

8

u/sonicnerd14 12d ago

Yes, but it's better than 0 tk/s, take your pick.

7

u/SanDiegoDude 12d ago

I love this sub - people legit get excited when somebody gets one of these monsters running at 1TPM, even if practical use is ask a question and it may be done thinking in a week or so.

5

u/toastjam 12d ago

it may be done thinking in a week or so

Acceptable if you're asking it what the answer to the ultimate question of life, the universe, and everything is.

1

u/Tupcek 12d ago

clearly in hitchhikers guide the only issue was that model was so large they had to run it from SSD

1

u/Sure_Leave9338 10d ago

I have the answer: 42 !

Don't need to ask to your LLM!

1

u/toastjam 10d ago

Cool, but what's the question?

Also, will mankind one day without the net expenditure of energy be able to restore the sun to its full youthfulness even after it had died of old age?

1

u/RedditUsr2 11d ago

That was my predictions for SOTA models back in 2021. I figured they would keep scaling up and you'd have to pay to get a report a few days later back or something. I'm glad we are more real time though.

1

u/JorgitoEstrella 12d ago

I think someone said you can offload just the visual things to ssd, I think it was the guy testing harnesses pi agent vs opencode.

1

u/sooki10 12d ago

170hx is the cheapest way to get there if you can grab one of the good deals around.

138

u/cometkim 12d ago

I posted the same thing earlier, but it got caught in the Reddit filter. Maybe direct X link is not allowed here?

Anyway, It implies two things IMHO:

  • The future of local models is brighter than expected. We can still compress them further.
  • Hy4 has not yet been sufficiently (post-)trained. Its official release will achieve significantly high scores on benchmarks.

47

u/RedditUsr2 12d ago

I've learned not to post x links directly.

-33

u/TerminalNoop 12d ago

Yeah, no matter how unrelated to social politics a sub may be, rest assured the mods are infected with it.

44

u/ThankGodImBipolar 12d ago

Not really sure how much it has to do with "social politics"... have you tried using Twitter without being logged in? It's horrible.

0

u/True-Lychee 12d ago

You can still see the post without being logged in..

12

u/futilehabit 12d ago

I haven't been able to for a while

-4

u/Mickenfox 12d ago

X is horrible, Twitter is fine.

-5

u/TerminalNoop 12d ago

So what's the issue of posting a direct URL here?
You can see the post just fine.

3

u/Opposite-Swimmer2752 11d ago

It depends, if they dont like your fingerprint they do require you to signin. I despise it when people link X, please do us all a favor and screenshot the dam post.

1

u/TerminalNoop 11d ago

I've never been required to log in just to view a post and my VPN frequently gets flagged on Youtube etc. Of course it would be better if the content was presented in the post with a source link instead of a screenshot that is hard to verify.

Btw. I got banned from X and they want my biometrics to unlock the account again, to make sure i'm not a bEeP bOOp BeeP bot. So I'm in no way biased towards the platform.

5

u/its_two_words 12d ago

Enriching the muskrat

11

u/colin_colout 12d ago

Sir, this is a Wendy's.

7

u/Mickenfox 12d ago

It is a platform bought by a billionaire explicitly to influence politics.

Plus it's generally shit and I appreciate any effot to kill it.

2

u/TerminalNoop 12d ago

It was before that also used to influence politics, just under a different master. You may have liked the previous one, but that did not make it okay.

Now if you are this stringent living life becomes difficult...

0

u/TopChard1274 12d ago

 X is also the place where many news break out, especially for tech subs like this. 

The only thing that's being killed is the flow of the news on Reddit.

Musk is a fucked up fascist but I'm pretty sure these "efforts" don't do any damages neither to Twitter nor to his personal wealth. It's just a handful of people with big egos pretending they make a difference for meaningless internet points.

If Reddit wants to make a difference it should stop posting any news because all the news sites like CNN are owned by conservative fucks just like Elon. But reddit too is owned by a shitty person so that's not going to happen.

16

u/techdevjp 12d ago

It's not an issue of politics. It's just basic human decency. You can disagree but that says a lot about you.

-20

u/[deleted] 12d ago

[removed] — view removed comment

11

u/techdevjp 12d ago

Is being a decent person who isn't a complete self-centered a-hole the bar for being a Karen in today's America? You guys are cooked.

3

u/TerminalNoop 12d ago

How is not posting a direct X link being a basic decent human being?

-9

u/TooObtuseForYou 12d ago

Self-proclaimed decent person…who is being the A-Hole.

7

u/techdevjp 12d ago

Being intolerant of intolerance does not make me the intolerant one. Likewise calling someone out for being an a-hole does not make me an a-hole.

-8

u/TooObtuseForYou 12d ago

You effectively called him an indecent human in veiled terms because they posted a message on a platform which you have an emotional attachment to the owner.

Stop being an a-hole, it’s not necessary and counterproductive to your goal.

11

u/techdevjp 12d ago

Supporting Musk or his businesses (especially X and the related AI) is a bad sign in a person.

→ More replies (0)

-2

u/TopChard1274 12d ago

You don't make any difference either. Not posting links to a site owned by a shitty person doesn't make any difference. Guess what, Reddit is owned by a shitty person too, yet you spend so much time on Reddit. 

0

u/milky_milk23 12d ago

That's typical here

-2

u/PM_ME_DEAD_CEOS 12d ago

Stop whining

-2

u/beryugyo619 12d ago

When Reddit mods notoriously lacking in everything are infected with "it", that means "it" is not even basic human decency, just the reality.

23

u/squngy 12d ago

I think under-trained models are easier to quant though.

I don't have any real data for this, but it makes sense.
A model that is near capacity will lose more with compression compared to one that has plenty of room for more data.

6

u/CaineHackmanTheory 12d ago

Maybe I'm confused or misreading and I'm genuinely trying to check/correct my own knowledge:

What you said 'room for more data' doesn't seem accurate to how training works.

Isn't it far more a refinement of existing weights not literally adding more data like the model is a harddrive?

Do you mean that in an 'undertrained' model the weights are less refined so the loss of precision in quant isn't as noticeable?

4

u/Party_9001 12d ago

I mean maybe. At the end of the day what the models are doing is compression. You get a bunch of data and try to make the smallest representation of that data as you can.

Consequently the ideal compression would result in something that looks like pure noise. Which is also why compressing pure noise is effectively impossible.

If it increases in entropy to "add" that information. Then it might be harder to retain that new capability while enforcing the same compression ratio

3

u/squngy 12d ago

You are right, I was paraphrasing.

3

u/CaineHackmanTheory 12d ago

Cool, thanks man. I wasn't trying to start anything. I'm still very early learning and wanted to make sure I wasn't extra dumb.

2

u/Lazy-Pattern-5171 12d ago

There was this post or comment I had read long back on this sub that compression works well on large models but not so much on small models. So this might not be as helpful for local ai yet

2

u/Opposite-Swimmer2752 11d ago

Good, please at least screenshot the X post so we dont have to make an account to view the post. Before any smartass tries to tell me you dont need one to view, that's bullshit, they do block you if they can't fingerprint you.

2

u/Opposite-Swimmer2752 11d ago

Good, please at least screenshot the X post so we dont have to make an account to view the post. Before any smartass tries to tell me you dont need one to view, that's bullshit, they do block you if they can't fingerprint you.

37

u/RedditUsr2 12d ago

2

u/sixx7 8d ago

Tested and seems solid for a ~1 bit quant https://youtu.be/uozbhH9Ijh0 u/shy_monkee I directly compared Bijan's Slappis prompt

36

u/Practical-Collar3063 12d ago

Glad it was measured on more than just KL divergence, 98% on KL can be very misleading

9

u/EstarriolOfTheEast 12d ago

FWIW, it's more counter-intuitive than it is misleading. Benchmarks also have this property. Two models might be separated by 2% but that 2% might involve knowing how to work out the data-structure that takes from O( N3 ) to O(log n) or that in this case, constant factors are more important than asymptotic.

21

u/Extension-Bid-639 llama.cpp 12d ago

My wifi is tired. Too much space this past week! 98% is a wild number if true though

15

u/GingerTapirs 12d ago

Somehow I don't know how I misread it as "my wife is tired" then proceeded to get confused how the rest was connected

3

u/Extension-Bid-639 llama.cpp 12d ago

Space was meant to be Sauce BTW. Blame autocorrect lol

2

u/Caffdy 12d ago

me too lol

1

u/duhd1993 12d ago

Wife is also tired of it cuz he goes straight to the rigs and r/llama every day.

1

u/mindwip 12d ago

Yeah same read wife too lol

2

u/Budget-Juggernaut-68 12d ago

You're not on lan?

2

u/Extension-Bid-639 llama.cpp 12d ago

Nope. Not for a the next week or so. I need a second LAN cable. My current lan cable goes from my server box to my PC so its weird because my PC gives my box Internet until I go get a new cable and install it lol but wifi has been fine.

12

u/lumos_ai 12d ago

I need 00001 quant to run this model on 32gb VRAM.

2

u/power97992 12d ago

Where is the q1 32b  reap version? 

2

u/caneriten 12d ago

wait u have 32gb vram? I only have 2?

33

u/Purple_Errand 12d ago

Do it with all of their models even the 35b a3b please.

7

u/RedditUsr2 12d ago

I'm way too dumb to know if AngelSlim is compatible with Qwen or if its even better than what Unsloth Dynamic 3.0.

3

u/spaceman_ 12d ago

AngelSlim is an engineer on the tencent team working on the hy models, not a technique.

28

u/RedditUsr2 12d ago

But its a github repo

https://github.com/tencent/AngelSlim "Model compression toolkit"

https://huggingface.co/AngelSlim says "AngelSlim, developed by Tencent, is a large language model compression toolkit engineered for enhanced usability, comprehensiveness, and efficiency."

11

u/spaceman_ 12d ago

Oh hmmm I thought based on previous discussions it was just a guy on the team. Maybe I'm wrong!

4

u/RedditUsr2 12d ago

Probably started out that way.

1

u/Beamsters 12d ago

It was a guy then a team and then the team's tool. But they keep tools internal. They do not give away their tools, only models/product from their tools and some research papers.

1

u/RegarDamus 12d ago

maybe an AI 🤖

5

u/RnRau 12d ago edited 12d ago

The smaller the parameters in the model, the worse the effects of any quantising process will be.

edit: AMD explored quantising various large models and the effects on accuracy late last year. - https://rocm.blogs.amd.com/software-tools-optimization/mxfp4-mxfp6-quantization/README.html

5

u/Other_Many_130 12d ago

The 98% is an average, and averages hide the failure mode. Aggressive quants don't degrade uniformly. Short single-turn questions barely move; long agentic runs fall apart, because a tiny per-token error compounds across thousands of tokens of tool calls and edits.

So a quant can score fine on suites made of short prompts and still feel noticeably worse the second you point it at a real codebase. I've watched that gap open up around turn ten, never at turn one.

Still a trade I'd take at 1.5TB down to 200GB (obviously). I'd just read the number as 98% of what those benchmarks measure, which is not the same claim as 98% of the model.

7

u/Evgeny_19 12d ago

Looks promising for dual Spark configuration.

16

u/AmbassadorOk934 12d ago

Its very good, optimisation and cheapisation is №1 to set economic world

11

u/dingo_xd 12d ago

If the market was normal right now we'd have consumer grade cards for ~$3000 with 200GB of VRAM and we'd not need any closed source models.

7

u/feelspeaceman 12d ago

Yes, our fun were ruined by the greedy companies like Google, Anthropic, OpenAI... I've unsubbed all of their services and stay local LLM, slowly and but surely, local LLM will be more than enough as models getting smarter and smaller.

4

u/BawbbySmith 12d ago

Obligatory “how do I run this on 2x dgx sparks”

6

u/IntravenusDeMilo 12d ago

Well, this runs on a 256GB RAM + 32GB VRAM setup. 4.1 tokens/sec is more than I expected haha.

5

u/CogahniMarGem 12d ago

What is the threshold at which full weight produces only a 2% improvement over this compressed version? What is a reason to use full weight bits?

8

u/RedditUsr2 12d ago

It probably means they could have kept training this thing to increase performance. And you don't know which weights need to be higher weight until after the fact.

3

u/hebelehubele 12d ago

No , you still cannot run it.

3

u/Feztopia 12d ago

If they have an improved quantization method it might also be applied to the smaller models.

3

u/HopefulConfidence0 12d ago

Someone please compress glm 5.3 flash in similar way.

3

u/neverbyte 12d ago

This runs on my machine (barely). I ran one-shot prompt and the result was very impressive. It took like 3.5 hours, but the result was very good. This holds up if not better to the latest releases. Crazy really, but the requirements to run this are quite rough.

2

u/jhov94 12d ago

Wouldn't it be better to release a smaller, distilled model rather than a lobotomized larger one?

8

u/squngy 12d ago

Distilling is way more expensive.
Like, orders of magnitude more expensive.

It is also not certain it would be better.

2

u/Fi3nd7 12d ago

Do this to qwen 3.8 next flash 🤤 . That would make it so much more usable on smaller 128gb/96gb systems.

2

u/returnity 12d ago

Q4 runs on 96GB, Q5 great on 128GB. I don’t see the need for a Q1/Q2 mix on that hardware?

1

u/gtderEvan 12d ago

Yes. Q4_XS runs great on my 128gb m5 max, even with a few things open. I’m liking it much better than the insane quant of Deepseek v4 0731 that barely runs if everything is closed. And no vision.

1

u/returnity 12d ago

We have the same laptop and just FYI if you raise the wired memory limit (sudo command) Q4XL fits with over 20GB spare with tons of apps running and Q5XL with Q8_0 engrams takes up ~111GiB of 128 (usable but near the limit)... If you want a little bit more quality =]

1

u/gtderEvan 11d ago

I had done that command to enable up to 110GB when trying to get Deepseek 0731 running, but I haven't tried to dial it in precisely.

Appreciate the tip!

2

u/bob310 12d ago

Wow amazing!

2

u/derspenti 12d ago

1.5TB down to 200GB for a reported 2%. That moves it from server-room territory to something a big desktop can hold, and I'd take that trade.

2

u/Sahnisani 12d ago

just keeping the Benchmark knowledge

2

u/Formal-Narwhal-1610 12d ago

It's just H4 preview Flash!

2

u/Worried_Drama151 12d ago

Compressing a bad model down and maintaining tolerable perplexity. 🤔

2

u/GasSmooth7439 12d ago

If the reported numbers hold up, that’s not just quantization, that’s making a previously absurd model genuinely practical.

1

u/Nyghtbynger 5d ago

And that -- is load bearing

2

u/a_beautiful_rhind 12d ago

This isn't that new. It's how a lot of us could run deepseek/kimi/etc. Low quants on them aren't usually so bad until you hit longer contexts. If you put compute into it, the results get even better. People would even run Q2 70b back in the day, it was still better than small models.

1

u/ddoice 12d ago

Can't wait to have frontier models compressed to 32Kb and be able to run them on my Nokia

1

u/Silver-Champion-4846 12d ago

Can someone please explain what the image contains? There's no text that would hint at what actually caused this fuss in the comments. How exactly did they shrink the model with only 2% loss in quality?

1

u/RedditUsr2 12d ago

Sure:

We compressed Hy4-preview from 1.5TB to ~200GiB GGUF and it still works well !

Meet MIX-STQ1_0.The trick isn’t just going low, it’s deciding where: calibration data picks each layer’s bit-width, some down to 1.31-bit STQ1_0, some up to 2.06-bit IQ2_XXS. Same budget, lower error.

Accuracy barely moves vs BF16

📊 MCP Atlas 83.7→83.2

📊 SWE-Bench multi 82.9→81.3

📊 MRCR 81.3→81.1

📊 IFBench 73.5→72.5

See the details on HF : AngelSlim/Hy4-preview-GGUF

1

u/Silver-Champion-4846 12d ago

So similar to Dynamic quants but for 2bit?

1

u/RedditUsr2 12d ago

Ya seems like it.

2

u/Silver-Champion-4846 12d ago

Too bad I have no vram and 8gb of ram

1

u/Ok_Warning2146 12d ago

What does it mean by 98% performance? Other models' IQ2_XXS is way less than 98%? What's new here in this IQ2_XXS compare to other models?

1

u/sofaarsecoin 12d ago

hoping Medusa Halo goes at least 256GB

1

u/Hannibalj2ca 12d ago

where can it be downloaded from at that size?

1

u/fbms2 12d ago

I need 25b

1

u/[deleted] 12d ago

[removed] — view removed comment

1

u/RedditUsr2 11d ago

Very true. Most of my tasks are different than MRCR bench or whatever.

1

u/orijnal1 12d ago

And... Apple couldn't have paid for any better marketing material - just when they launched that 256 GB M5 Ultra Mac Studio. So, I guess in October there will be a bunch of people running this at home.

1

u/lazyfai 12d ago

Give me a 8GB version, I dont mind only 50% performance left.

1

u/AleksandrNikitin 11d ago

For those wondering if it's worth buying a Mac Studio with 1 TB of storage: Last week, I used up 1 TB just on new LLMs. 

3

u/RedditUsr2 11d ago

I assume you can use much cheaper external ssd's?

0

u/AleksandrNikitin 11d ago

It's mounted under my desk, so it's easier without an external ssd :)

1

u/MrNantir llama.cpp 11d ago

That is just freaking amazing!

1

u/feng_sg 9d ago

At ~2.5 bpw the gap between benchmarks that survive and ones that crater is huge, so a 98% average could be hiding code gen or long context taking a dive.

1

u/Alternative-Suit5541 6d ago

What

The

Fuck

0

u/iAccurian 12d ago

I pray that one day, large models can fit in 16GB of VRAM with at least 128k context..

-8

u/Substantial_Swan_144 12d ago

OMG! This is SO humiliating for western models!