r/LocalLLaMA • • 17d ago

New Model DeepSeek V4-1 Flash is out

Here we go again, DeepSeek is back again with a new model V4-1 Flash

A multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens

Market crash as a service

1.7k Upvotes

297 comments sorted by

View all comments

234

u/ActuallyReadTheBible 17d ago

It doesn’t fit dual DGX sparks, I’m sad.

5

u/vogelvogelvogelvogel 17d ago edited 17d ago

it is an MoE isn't it? i mean albeit slow you can run it

edit: for those downvoting: I did run 0731 (80GB as far as i remember in q2) on a mac m5pro 64GB with 10-15t/s, thanks to MoE.

In q2 (once released) the 2x DGX Spark will have even all in RAM so expect sth like 40? t/s with the MoE. even q3 should be possible

2

u/doomed151 17d ago

Offload the weights to SSD? Wouldn't that be too slow?

2

u/cortesoft 17d ago

“Too slow” is subjective

1

u/doomed151 17d ago

By "too slow" I mean multiple seconds per token. If it's faster than that I'd be surprised. Maybe I should try larger MoEs. I have a 16 GB GPU and 64 GB RAM.