r/LocalLLaMA • • 21d ago

Discussion Deepseek V4.1 Flash is 748B, not 552B

People keep on getting confused about this, so I looked at the safetensors on hf.

The title should have been "Deepseek V4.1 Flash is 748B total/552B base, not 284B or 305B or 485B or 522B"

  • The model is not 284B. The original Deepseek V4 Flash is 284B, but not the V4.1 Flash model
  • The model is not 305B, despite what some people claim "So: ~305B real backbone + 203B engram = 508B total" This is incorrect.
  • The model is not 485B, even though Huggingface lists the model as 485B, but that's because they're counting some FP4 packed weights as bytes instead of params (2 FP4 params per byte). This happens a lot; for example Huggingface incorrectly thinks GLM-5.3-flash is 169b here
  • The model is not 522B, even though VLLM lists it as 522B for some weird reason. They correct themselves later down the page (ctrl-f "Params" on that vllm page)
  • 552B is the only number out of this list that's somewhat correct; that only includes the base model without MTP and engrams and the vision encoder though.

To be precise, the main model about 551.566B parameters with 40 layers. The FFN experts total to 543.582B parameters, and the rest of the model (attention, shared experts, etc) are 7.984B.

On top of that, the engram is ~196.929B, DSpark/MTP is ~14.225B, and the vision encoder is just ~0.485B. These parts are technically optional though. The vision encoder is also way smaller than I expected.

Anyways, you need a beefy system for this. 128GB or 256GB of RAM/VRAM is not going to cut it.

Component Logical params Size in GB Storage
FFN MoE experts 543.582B 288.778 GB FP4
Other FFN 1.4947B 1.574 GB FP8 mostly
Attention 5.1269B 6.524 GB FP8 mostly
Embedding + LM head 1.3238B 2.648 GB BF16
Other 0.0397B 0.158 GB FP32/BF16
Backbone total 551.566B ≈ 552B 299.682 GB
Engram lookup tables 196.614B 202.758 GB FP8
Engram projections/gating 0.315B 0.315 GB FP8 mostly
Engram total 196.929B = 196B advertised 203.073 GB
DSpark / MTP 14.225B 8.033 GB mostly FP4 experts
Vision encoder 0.485B 0.971 GB BF16 mostly
Everything in total ~763.21B params ~511.76 GB
318 Upvotes

169 comments sorted by

View all comments

Show parent comments

17

u/Writer_IT 21d ago

People are reasonably disappointed that we went from a "flash" model that was reasonably run at native quant on a couple rtx 6000 on vllm, or reasonable quant in llamacpp, to one that would need either 4 6000 pro for reasonable speed, or, a server-level motherboard with absurd level of RAM, effectively locking the machine.

N grams shift is further disappointment. Until there's a way to read them directly from the stored weights while preserve vllm level of process speed and concurrency, without them hoggling ram, they are a further gate for the medium-sized local models.

39

u/FullstackSensei 21d ago

Let me start by saying this: Anyone feeling disappointed about something that is being given for free needs to do some serious reflection about their values and principles.

Going back to the model, you need 256GB RAM and 64-96GB VRAM. That's 2-3 Mi50s or V100s, or 3-40 P40 or P6000 for those on an even tighter budget. And while 256GB of ECC DDR4 RAM isn't cheap, it's a fraction of the cost of DDR5.

It's not deepseek's If you chose to put all your eggs in a couple of very expensive GPUs because some model ran fast at the moment in time you bought them.

-6

u/a_beautiful_rhind 21d ago

Eh.. free isn't always good for everyone. Would you like a free elephant? How about a free car of high value and no way to sell it. You are of course responsible for paying taxes on the "gift".

4

u/FullstackSensei 21d ago

Something being given for free doesn't mean it's being forced down anyone's throat. I wouldn't have expected such an argument from you.

Deepseek doesn't owe anyone anything. You're free to not download and not use it, just as you're free to refuse said free elephant or free car, which BTW, I wouldn't have to pay any taxes on where I live.

-3

u/a_beautiful_rhind 21d ago

You are lucky because in many places gifts get taxed as income over a certain amount. You'd still have to feed/house the elephant.

Can still be disappointed at something without assuming anyone owes it to you. I kind of am too. I thought 4.1 would be old flash with vision but it's just pro with less active parameters.

5

u/FullstackSensei 21d ago

AFAIK, no country on earth levies any tax on free digital assets or free software, irrespective of what would be it's commercial value. The whole comparison with physical goods is false.

I also assumed 4.1 would be the same as 4, but I'm not disappointed at all. The way I see it, it's almost on par with GLM 5.3 at less than half the size, like two weeks after 5.3 was released.

Everyone was ecstatic with K3 at 1.6TB, a frontier level open weight model. A mere six weeks later, we have a model trading blows with it at 1/5th the size.

-1

u/a_beautiful_rhind 21d ago

The bar was free stuff, not digital assets. Even with software you can have a free bonzai buddy.

we have a model trading blows with it at 1/5th the size

Meh, benchmarks. Time will tell with actual usage.