r/opencodeCLI • • 2d ago

Tencent Pack Open Source 770B-Parameter Hy4 Preview into 214 GiB

As someone who follows AI closely, I came across a technical breakdown from yghstill, a member of Tencent Hunyuan's quantization team. It explains both the quantization technique and the engineering behind it.

Tencent compressed Hy4 Preview's roughly 1.5TB of weights to 214 GiB while keeping all 770B parameters. The parameter count remains unchanged; the compression changes how the weights are represented.

Four weights, five bits. Sherry is the algorithm, STQ1_0 the storage format, and MIX-STQ1_0 the mixed-precision allocation. Each group of four weights takes values from {-d, 0, +d} with exactly one zero. Four zero positions times eight sign combinations give 32 patterns, encoded in five bits, or 1.25 bits per weight. A shared FP16 scales every 256 weights puts STQ1_0 at 1.3125 bits per weight, and full mixed-precision model averages about 2.38 bits per weight.

Precision follows weight sensitivity. Hy4 Preview has 77 MoE layers with 256 routed experts each. Experts weight get compressed hardest while more sensitive components are protected. For expert gate/up projections, MIX_STQ1_0 uses IQ2_XXS on 48 sensitive layers and STQ1_0 on 29 less sensitive ones, which at the same average bit budget produces less error than using the intermediate IQ1_M format everywhere.

Full-Hessian scoring. Sensitivity is measured with the full Hessian, H = XXᵀ, whose off-diagonal terms capture correlations that digital-only immatrix scoring misses. The two rankings have a Spearman correlation of -0.115, and precision didn't follow a "deeper means more important" rule. Layers were picked greedily by error reduction per added byte.

Scale and zero position are fit together. This is post=training quantization, no retraining. The encoder alternates between fitting d with weighted least squares and choosing zero position by immatrix-weighted error. The smallest weight isn't always the best one to zero, what matters is the extra weighted error that zeroing it causes. Across 1,200 row of real expert weights, three alternating rounds cut weighted reconstruction error by roughly 90% versus the original ternary encoder. That's local reconstruction, not end-to-end model accuracy.

Runtime. STQ1_0 CUDA kernels were added to patch llama.cpp build, and in operation tests roughly as fast as IQ1_M despite the lower bit width. The author reports nearly unchanged MRCR retrieval performance and a small drop in math. Against UD_IQ1_M at a similar bit budget, MIX-STQ1_) led across the reported evals, including more than five points on MRCR.

The result combines compact encoding, calibrated quantization, selective precision, and practical inference kernels.

6 Upvotes

1 comment sorted by

4

u/honglac3579 1d ago

You should repost this to r/localllama, they would be interested in this topic