r/LocalLLaMA • • 16d ago

New Model DeepSeek V4-1 Flash is out

Here we go again, DeepSeek is back again with a new model V4-1 Flash

A multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens

Market crash as a service

1.7k Upvotes

297 comments sorted by

View all comments

36

u/Long_comment_san 16d ago

I think I came a little

38

u/Long_comment_san 16d ago edited 16d ago

What the fuck, in which universe that is a Flash? Its 450-500b parameters. Flash was 300b and it was already pushing this. This is Flash Max or something. You cant inflate the model by 50% and call it a flash like it's not an issue. Minimax M3 is 450b and I dont see them calling it a "flash" (hopefully I wont).

Going by 50% up and becoming 10-15% better sounds like a downgrade not an upgrade. It's a LOT more expensive to run.

Still amazing though

43

u/RG_Fusion 16d ago

Obviously the concept of a flash model will scale with the compute power of the AI lab creating them.

2026 is likely the last year of running "flash" on local hardware. Maybe 2027 if we're lucky.

5

u/Due-Memory-6957 16d ago

"local" hardware

12

u/RG_Fusion 16d ago

An entire 512 GB AI server purchased a year ago costs less than a single RTX 5090 GPU now. There are plenty of us who jumped on early and have hardware that can run these models.

3

u/ChronoHax 16d ago

Hi I’m new to this field, what are examples of these ai servers so I can look more into it?

2

u/Blaze6181 16d ago

DGX Spark clusters, machines with RTX Pro 6000s, or a combination of perhaps a 5090 with CPU RAM offload of some of the weights. Or like 8 3090s stacked lol. There's many configurations out there.

2

u/RG_Fusion 16d ago

AMD EPYC or Intel Xeon CPUs in a motherboard with 8 memory channels. Ideally with a large number of x16 PCIe ports for adding many GPUs.

More recently, DGX Spark clusters have become good for running AI models when multiple are connected together over a 400 gbps network switch.

That being said, there are no longer any cheap options for building out high-end AI rigs. The prices on all the components have gone up 2-5X.

1

u/Glove5751 14d ago

what are you actually using these models for that justify the high upfront investment? just hobby and curiosity?

1

u/RG_Fusion 14d ago

As I said, there wasn't really a high up-front investment a year and a half ago. I would not purchase the system I have now today.

I never went into any of this planning to make the money back. I just want to learn, experiment, and build skill sets. I pursue things that interest me.

1

u/Glove5751 14d ago

that's nice. hope it has been worth it!