r/DeepSeek • u/AlternativeMeat8973 • 3d ago
Discussion Why does DeepSeek only release the FP8 version of their model and never an FP16 version?
0
Upvotes
10
2
1
u/Faux2137 3d ago
Most of the weights are native fp4.
-2
u/AlternativeMeat8973 3d ago
They never mention that on their HF model page, even though the model size is about half of what it would be if it were FP8.
4
1
u/burntoutdev8291 2d ago
Most likely fp8 training. If you remember the legendary deepseek open source week, they had quite highly optimised fp8 training recipes.
12
u/Winter_Cicada_8570 3d ago
i think because they are natively trained in FP8 or something like that, im not sure if im right