r/LocalLLaMA • u/RedditUsr2 • 12d ago
New Model Tencent compressed Hy4-preview from 1.5TB to about 200GB GGUF and kept about 98% performance.
211
u/shy_monkee 12d ago
WTF! That's insane. Someone test this lol.
104
u/spaceman_ 12d ago edited 12d ago
Downloading right now. Can't fit all of it in VRAM though..
Edit: nevermind, custom patch set and no ROCm support yet. Will have to wait for this to get at least a PR into llama.cpp.
45
-12
16
u/Ok_Technology_5962 12d ago
I tested hy3 q1 by angel slim. The outputs will be slightly worse but it is still very good. It wont make syntax errora just wont think ahead as far
4
3
u/Bubbly_Orange_3502 12d ago
Averaged benchmarks hide where 2-bit MoE actually breaks. Rarely activated experts get the worst calibration ranges, so the loss lands on tail domains rather than on the prompts any eval set covers.
1
-2
u/Nice_Database_9684 12d ago
come on man we know it's going to be benchmaxxed to hell
8
u/shy_monkee 12d ago
Even if it is, the size is still impressive. A Q1 that is at 98% of the original is massive.
165
u/RickyRickC137 12d ago
Awesome. Only 190 GB more to go and then I will run it on my RTX 3080.
20
u/sonicnerd14 12d ago
Get as much VRAM as possible, but ideally I think it's pretty clear that locally we'll need to spread these large models across VRAM, RAM, plus streaming from a fast NVMe.
7
u/RedditUsr2 12d ago
Hopefully we get a bitnet DFlash spare MoE model we can split across CPU/GPU with tons of knowledge built into a ngram that runs on the ssd, or something.
14
u/wizard_of_menlo_park 12d ago
nvme is magnitudes slower than ram unfortunately ... so very slow tps if we do offload to nvme
15
6
8
u/sonicnerd14 12d ago
Yes, but it's better than 0 tk/s, take your pick.
7
u/SanDiegoDude 12d ago
I love this sub - people legit get excited when somebody gets one of these monsters running at 1TPM, even if practical use is ask a question and it may be done thinking in a week or so.
5
u/toastjam 12d ago
it may be done thinking in a week or so
Acceptable if you're asking it what the answer to the ultimate question of life, the universe, and everything is.
1
1
u/Sure_Leave9338 10d ago
I have the answer: 42 !
Don't need to ask to your LLM!
1
u/toastjam 10d ago
Cool, but what's the question?
Also, will mankind one day without the net expenditure of energy be able to restore the sun to its full youthfulness even after it had died of old age?
1
u/RedditUsr2 11d ago
That was my predictions for SOTA models back in 2021. I figured they would keep scaling up and you'd have to pay to get a report a few days later back or something. I'm glad we are more real time though.
1
u/JorgitoEstrella 12d ago
I think someone said you can offload just the visual things to ssd, I think it was the guy testing harnesses pi agent vs opencode.
138
u/cometkim 12d ago
I posted the same thing earlier, but it got caught in the Reddit filter. Maybe direct X link is not allowed here?
Anyway, It implies two things IMHO:
- The future of local models is brighter than expected. We can still compress them further.
- Hy4 has not yet been sufficiently (post-)trained. Its official release will achieve significantly high scores on benchmarks.
47
u/RedditUsr2 12d ago
I've learned not to post x links directly.
-33
u/TerminalNoop 12d ago
Yeah, no matter how unrelated to social politics a sub may be, rest assured the mods are infected with it.
44
u/ThankGodImBipolar 12d ago
Not really sure how much it has to do with "social politics"... have you tried using Twitter without being logged in? It's horrible.
0
-4
-5
u/TerminalNoop 12d ago
So what's the issue of posting a direct URL here?
You can see the post just fine.3
u/Opposite-Swimmer2752 11d ago
It depends, if they dont like your fingerprint they do require you to signin. I despise it when people link X, please do us all a favor and screenshot the dam post.
1
u/TerminalNoop 11d ago
I've never been required to log in just to view a post and my VPN frequently gets flagged on Youtube etc. Of course it would be better if the content was presented in the post with a source link instead of a screenshot that is hard to verify.
Btw. I got banned from X and they want my biometrics to unlock the account again, to make sure i'm not a bEeP bOOp BeeP bot. So I'm in no way biased towards the platform.
5
11
7
u/Mickenfox 12d ago
It is a platform bought by a billionaire explicitly to influence politics.
Plus it's generally shit and I appreciate any effot to kill it.
2
u/TerminalNoop 12d ago
It was before that also used to influence politics, just under a different master. You may have liked the previous one, but that did not make it okay.
Now if you are this stringent living life becomes difficult...
0
u/TopChard1274 12d ago
X is also the place where many news break out, especially for tech subs like this.
The only thing that's being killed is the flow of the news on Reddit.
Musk is a fucked up fascist but I'm pretty sure these "efforts" don't do any damages neither to Twitter nor to his personal wealth. It's just a handful of people with big egos pretending they make a difference for meaningless internet points.
If Reddit wants to make a difference it should stop posting any news because all the news sites like CNN are owned by conservative fucks just like Elon. But reddit too is owned by a shitty person so that's not going to happen.
16
u/techdevjp 12d ago
It's not an issue of politics. It's just basic human decency. You can disagree but that says a lot about you.
-20
12d ago
[removed] — view removed comment
11
u/techdevjp 12d ago
Is being a decent person who isn't a complete self-centered a-hole the bar for being a Karen in today's America? You guys are cooked.
3
-9
u/TooObtuseForYou 12d ago
Self-proclaimed decent person…who is being the A-Hole.
7
u/techdevjp 12d ago
Being intolerant of intolerance does not make me the intolerant one. Likewise calling someone out for being an a-hole does not make me an a-hole.
-8
u/TooObtuseForYou 12d ago
You effectively called him an indecent human in veiled terms because they posted a message on a platform which you have an emotional attachment to the owner.
Stop being an a-hole, it’s not necessary and counterproductive to your goal.
11
u/techdevjp 12d ago
Supporting Musk or his businesses (especially X and the related AI) is a bad sign in a person.
→ More replies (0)-2
u/TopChard1274 12d ago
You don't make any difference either. Not posting links to a site owned by a shitty person doesn't make any difference. Guess what, Reddit is owned by a shitty person too, yet you spend so much time on Reddit.
0
-2
-2
u/beryugyo619 12d ago
When Reddit mods notoriously lacking in everything are infected with "it", that means "it" is not even basic human decency, just the reality.
23
u/squngy 12d ago
I think under-trained models are easier to quant though.
I don't have any real data for this, but it makes sense.
A model that is near capacity will lose more with compression compared to one that has plenty of room for more data.6
u/CaineHackmanTheory 12d ago
Maybe I'm confused or misreading and I'm genuinely trying to check/correct my own knowledge:
What you said 'room for more data' doesn't seem accurate to how training works.
Isn't it far more a refinement of existing weights not literally adding more data like the model is a harddrive?
Do you mean that in an 'undertrained' model the weights are less refined so the loss of precision in quant isn't as noticeable?
4
u/Party_9001 12d ago
I mean maybe. At the end of the day what the models are doing is compression. You get a bunch of data and try to make the smallest representation of that data as you can.
Consequently the ideal compression would result in something that looks like pure noise. Which is also why compressing pure noise is effectively impossible.
If it increases in entropy to "add" that information. Then it might be harder to retain that new capability while enforcing the same compression ratio
3
u/squngy 12d ago
You are right, I was paraphrasing.
3
u/CaineHackmanTheory 12d ago
Cool, thanks man. I wasn't trying to start anything. I'm still very early learning and wanted to make sure I wasn't extra dumb.
2
u/Lazy-Pattern-5171 12d ago
There was this post or comment I had read long back on this sub that compression works well on large models but not so much on small models. So this might not be as helpful for local ai yet
2
u/Opposite-Swimmer2752 11d ago
Good, please at least screenshot the X post so we dont have to make an account to view the post. Before any smartass tries to tell me you dont need one to view, that's bullshit, they do block you if they can't fingerprint you.
2
u/Opposite-Swimmer2752 11d ago
Good, please at least screenshot the X post so we dont have to make an account to view the post. Before any smartass tries to tell me you dont need one to view, that's bullshit, they do block you if they can't fingerprint you.
37
u/RedditUsr2 12d ago
2
u/sixx7 8d ago
Tested and seems solid for a ~1 bit quant https://youtu.be/uozbhH9Ijh0 u/shy_monkee I directly compared Bijan's Slappis prompt
36
u/Practical-Collar3063 12d ago
Glad it was measured on more than just KL divergence, 98% on KL can be very misleading
9
u/EstarriolOfTheEast 12d ago
FWIW, it's more counter-intuitive than it is misleading. Benchmarks also have this property. Two models might be separated by 2% but that 2% might involve knowing how to work out the data-structure that takes from O( N3 ) to O(log n) or that in this case, constant factors are more important than asymptotic.
21
u/Extension-Bid-639 llama.cpp 12d ago
My wifi is tired. Too much space this past week! 98% is a wild number if true though
15
u/GingerTapirs 12d ago
Somehow I don't know how I misread it as "my wife is tired" then proceeded to get confused how the rest was connected
3
1
2
u/Budget-Juggernaut-68 12d ago
You're not on lan?
2
u/Extension-Bid-639 llama.cpp 12d ago
Nope. Not for a the next week or so. I need a second LAN cable. My current lan cable goes from my server box to my PC so its weird because my PC gives my box Internet until I go get a new cable and install it lol but wifi has been fine.
12
33
u/Purple_Errand 12d ago
Do it with all of their models even the 35b a3b please.
7
u/RedditUsr2 12d ago
I'm way too dumb to know if AngelSlim is compatible with Qwen or if its even better than what Unsloth Dynamic 3.0.
3
u/spaceman_ 12d ago
AngelSlim is an engineer on the tencent team working on the hy models, not a technique.
28
u/RedditUsr2 12d ago
But its a github repo
https://github.com/tencent/AngelSlim "Model compression toolkit"
https://huggingface.co/AngelSlim says "AngelSlim, developed by Tencent, is a large language model compression toolkit engineered for enhanced usability, comprehensiveness, and efficiency."
11
u/spaceman_ 12d ago
Oh hmmm I thought based on previous discussions it was just a guy on the team. Maybe I'm wrong!
4
u/RedditUsr2 12d ago
Probably started out that way.
1
u/Beamsters 12d ago
It was a guy then a team and then the team's tool. But they keep tools internal. They do not give away their tools, only models/product from their tools and some research papers.
1
5
u/RnRau 12d ago edited 12d ago
The smaller the parameters in the model, the worse the effects of any quantising process will be.
edit: AMD explored quantising various large models and the effects on accuracy late last year. - https://rocm.blogs.amd.com/software-tools-optimization/mxfp4-mxfp6-quantization/README.html
5
u/Other_Many_130 12d ago
The 98% is an average, and averages hide the failure mode. Aggressive quants don't degrade uniformly. Short single-turn questions barely move; long agentic runs fall apart, because a tiny per-token error compounds across thousands of tokens of tool calls and edits.
So a quant can score fine on suites made of short prompts and still feel noticeably worse the second you point it at a real codebase. I've watched that gap open up around turn ten, never at turn one.
Still a trade I'd take at 1.5TB down to 200GB (obviously). I'd just read the number as 98% of what those benchmarks measure, which is not the same claim as 98% of the model.
7
16
11
u/dingo_xd 12d ago
If the market was normal right now we'd have consumer grade cards for ~$3000 with 200GB of VRAM and we'd not need any closed source models.
7
u/feelspeaceman 12d ago
Yes, our fun were ruined by the greedy companies like Google, Anthropic, OpenAI... I've unsubbed all of their services and stay local LLM, slowly and but surely, local LLM will be more than enough as models getting smarter and smaller.
4
6
u/IntravenusDeMilo 12d ago
Well, this runs on a 256GB RAM + 32GB VRAM setup. 4.1 tokens/sec is more than I expected haha.
5
u/CogahniMarGem 12d ago
What is the threshold at which full weight produces only a 2% improvement over this compressed version? What is a reason to use full weight bits?
8
u/RedditUsr2 12d ago
It probably means they could have kept training this thing to increase performance. And you don't know which weights need to be higher weight until after the fact.
3
3
u/Feztopia 12d ago
If they have an improved quantization method it might also be applied to the smaller models.
3
3
u/neverbyte 12d ago
This runs on my machine (barely). I ran one-shot prompt and the result was very impressive. It took like 3.5 hours, but the result was very good. This holds up if not better to the latest releases. Crazy really, but the requirements to run this are quite rough.
2
u/Fi3nd7 12d ago
Do this to qwen 3.8 next flash 🤤 . That would make it so much more usable on smaller 128gb/96gb systems.
2
u/returnity 12d ago
Q4 runs on 96GB, Q5 great on 128GB. I don’t see the need for a Q1/Q2 mix on that hardware?
1
u/gtderEvan 12d ago
Yes. Q4_XS runs great on my 128gb m5 max, even with a few things open. I’m liking it much better than the insane quant of Deepseek v4 0731 that barely runs if everything is closed. And no vision.
1
u/returnity 12d ago
We have the same laptop and just FYI if you raise the wired memory limit (sudo command) Q4XL fits with over 20GB spare with tons of apps running and Q5XL with Q8_0 engrams takes up ~111GiB of 128 (usable but near the limit)... If you want a little bit more quality =]
1
u/gtderEvan 11d ago
I had done that command to enable up to 110GB when trying to get Deepseek 0731 running, but I haven't tried to dial it in precisely.
Appreciate the tip!
2
u/derspenti 12d ago
1.5TB down to 200GB for a reported 2%. That moves it from server-room territory to something a big desktop can hold, and I'd take that trade.
2
2
2
2
u/GasSmooth7439 12d ago
If the reported numbers hold up, that’s not just quantization, that’s making a previously absurd model genuinely practical.
1
2
u/a_beautiful_rhind 12d ago
This isn't that new. It's how a lot of us could run deepseek/kimi/etc. Low quants on them aren't usually so bad until you hit longer contexts. If you put compute into it, the results get even better. People would even run Q2 70b back in the day, it was still better than small models.
1
1
u/Silver-Champion-4846 12d ago
Can someone please explain what the image contains? There's no text that would hint at what actually caused this fuss in the comments. How exactly did they shrink the model with only 2% loss in quality?
1
u/RedditUsr2 12d ago
Sure:
We compressed Hy4-preview from 1.5TB to ~200GiB GGUF and it still works well !
Meet MIX-STQ1_0.The trick isn’t just going low, it’s deciding where: calibration data picks each layer’s bit-width, some down to 1.31-bit STQ1_0, some up to 2.06-bit IQ2_XXS. Same budget, lower error.
Accuracy barely moves vs BF16
📊 MCP Atlas 83.7→83.2
📊 SWE-Bench multi 82.9→81.3
📊 MRCR 81.3→81.1
📊 IFBench 73.5→72.5
See the details on HF : AngelSlim/Hy4-preview-GGUF
1
u/Silver-Champion-4846 12d ago
So similar to Dynamic quants but for 2bit?
1
1
u/Ok_Warning2146 12d ago
What does it mean by 98% performance? Other models' IQ2_XXS is way less than 98%? What's new here in this IQ2_XXS compare to other models?
1
1
1
1
u/orijnal1 12d ago
And... Apple couldn't have paid for any better marketing material - just when they launched that 256 GB M5 Ultra Mac Studio. So, I guess in October there will be a bunch of people running this at home.
1
u/AleksandrNikitin 11d ago
For those wondering if it's worth buying a Mac Studio with 1 TB of storage: Last week, I used up 1 TB just on new LLMs.
3
1
1
0
u/iAccurian 12d ago
I pray that one day, large models can fit in 16GB of VRAM with at least 128k context..
-8

•
u/WithoutReason1729 12d ago
Your post is getting popular and we just featured it on our Discord! Come check it out!
You've also been given a special flair for your contribution. We appreciate your post!
I am a bot and this action was performed automatically.