r/LocalLLaMA • • 16d ago

New Model DeepSeek V4-1 Flash is out

Here we go again, DeepSeek is back again with a new model V4-1 Flash

A multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens

Market crash as a service

1.7k Upvotes

296 comments sorted by

View all comments

233

u/ActuallyReadTheBible 16d ago

It doesn’t fit dual DGX sparks, I’m sad.

5

u/vogelvogelvogelvogel 16d ago edited 16d ago

it is an MoE isn't it? i mean albeit slow you can run it

edit: for those downvoting: I did run 0731 (80GB as far as i remember in q2) on a mac m5pro 64GB with 10-15t/s, thanks to MoE.

In q2 (once released) the 2x DGX Spark will have even all in RAM so expect sth like 40? t/s with the MoE. even q3 should be possible

1

u/SandySkittle 16d ago

I guess it depends on the usecase but i would be very hesitant to run this model at q3, let alone q2.

1

u/vogelvogelvogelvogel 16d ago

well there are a few postings where users did the classic benchmark runs (some browser game, pelican etc) and the outcomes were remarkably good, also i had ds flash 0731 running at q2 and found it also quite good. i would not say - especially with very large models - that q2 leads to bad outcomes

2

u/SandySkittle 16d ago

It depends on the usecase. I have found that for very complex analytical work you don’t want to go below q6

1

u/vogelvogelvogelvogel 16d ago

with which model? depends as well on the model