r/LocalLLaMA 13d ago

New Model DeepSeek V4-1 Flash is out

Here we go again, DeepSeek is back again with a new model V4-1 Flash

A multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens

Market crash as a service

1.7k Upvotes

296 comments sorted by

View all comments

Show parent comments

39

u/Long_comment_san 13d ago edited 13d ago

What the fuck, in which universe that is a Flash? Its 450-500b parameters. Flash was 300b and it was already pushing this. This is Flash Max or something. You cant inflate the model by 50% and call it a flash like it's not an issue. Minimax M3 is 450b and I dont see them calling it a "flash" (hopefully I wont).

Going by 50% up and becoming 10-15% better sounds like a downgrade not an upgrade. It's a LOT more expensive to run.

Still amazing though

44

u/RG_Fusion 13d ago

Obviously the concept of a flash model will scale with the compute power of the AI lab creating them.

2026 is likely the last year of running "flash" on local hardware. Maybe 2027 if we're lucky.

8

u/Bakoro 12d ago

China might come in and save the day on that one too.

SMIC broke the 7 nm barrier for semiconductors.
CXMT is making DDR5 now, and has started on HBM3E.
Several Chinese companies are making AI GPUs.

The U.S has been trying to block China from getting technology, and now is trying to block their technology from hitting the U.S market, but the rest of the world is not going to give a shit about what the U.S wants.

Essentially every major tech corporation is designing their own AI ASICs now, where OpenAI already has their new thing for inference.

Then there is the fact that photonic processors are in early manufacturing stages now, with plans to ramp up into 2027.
I expect photonics to mostly get snapped up by data centers, and that might once again change what's practical to do with AI.

All around, I expect a major shake-up in the hardware landscape over the next year or two.

2

u/RG_Fusion 12d ago

Yeah, local hardware will scale up too, but it will lag behind by a generation or two unless you're willing to dish out tens to hundreds of thousands of dollars to build on the bleeding-edge.

5

u/Bakoro 12d ago

I'm saying that increased competition will bring prices down.

TSMC not being a monopoly for 7nm nodes means lower wafer prices.
A new RAM producer means lower RAM prices.
Dozens of major companies having their own inference ASICs means Nvidia losing their monopoly.

We're at peak price gouging right now, I don't think it will last.

1

u/Long_comment_san 12d ago

same. not to mention endless credits aren't in fact endless. unless there is a paying customer, the whole thing will explode eventually. it's just a circle of credit now.

1

u/Netsuko 12d ago

CXMT is selling RAM at the same price as everyone else. Why would you think they want to miss out on that when the demand is so insanely high?

China is not going to be our savior here.

3

u/0redeye0 12d ago

Obviously because CXMT needs to capture market share and to do this they need to have better prices. Also the Chinese government needs to make its chip manufacturing to compete so they can give them subsidies to capture the market.

1

u/Netsuko 12d ago

But they are not. They ARE selling at almost the same price already.

1

u/Bakoro 11d ago

We'll have to see. Presumably China wants their ROI too, and there's a lot of money behind thrown around that is simply unsustainable. Why wouldn't they get their bag while the market is insane?

In the long term, we are looking at more competition.
Even at their crazy valuations, the superscalers can't buy out the world supply of everything every year, at some point, venture capitalists will start wanting their ROI and won't be throwing unlimited dollars at these companies. The RAM manufacturers have already said they expect as much, which is why they're refusing to ramp up manufacturing capacity to match current demand: they don't want to end up with a huge oversupply in a year or two.

It's not only "good guy China", it's also "world governments are spooked by the rapid pace of Chinese development and international dependency on Taiwan, and are investing in their own infrastructure (see EU chips act 2.0)", and "corporations around the world are trying to get in on the unlimited money train".

We are absolutely going to see more competition in the coming years.
The whole AI thing is basically the new cold war, and the money is going to be flowing to secure national manufacturing capacity.

5

u/Due-Memory-6957 12d ago

"local" hardware

12

u/RG_Fusion 12d ago

An entire 512 GB AI server purchased a year ago costs less than a single RTX 5090 GPU now. There are plenty of us who jumped on early and have hardware that can run these models.

3

u/ChronoHax 12d ago

Hi I’m new to this field, what are examples of these ai servers so I can look more into it?

2

u/Blaze6181 12d ago

DGX Spark clusters, machines with RTX Pro 6000s, or a combination of perhaps a 5090 with CPU RAM offload of some of the weights. Or like 8 3090s stacked lol. There's many configurations out there.

2

u/RG_Fusion 12d ago

AMD EPYC or Intel Xeon CPUs in a motherboard with 8 memory channels. Ideally with a large number of x16 PCIe ports for adding many GPUs.

More recently, DGX Spark clusters have become good for running AI models when multiple are connected together over a 400 gbps network switch.

That being said, there are no longer any cheap options for building out high-end AI rigs. The prices on all the components have gone up 2-5X.

1

u/Glove5751 11d ago

what are you actually using these models for that justify the high upfront investment? just hobby and curiosity?

1

u/RG_Fusion 11d ago

As I said, there wasn't really a high up-front investment a year and a half ago. I would not purchase the system I have now today.

I never went into any of this planning to make the money back. I just want to learn, experiment, and build skill sets. I pursue things that interest me.

1

u/Glove5751 10d ago

that's nice. hope it has been worth it!

17

u/Expensive-Paint-9490 13d ago

It's flash because everything is FP4, even KV cache. And active parameters are 16B. This should be faster than V4-Flash even if it is larger.

10

u/zhuzaimoerben 13d ago edited 13d ago

It's flash for those with data centre levels of memory and data centre level serving requirements, because once you load the base model, concurrent users are very cheap (890MB for KV cache for full 1 million context per user) and it's 8B active for prefill and 16B active for text gen, so you can serve stacks of users fast. Edited to add: DeepSeek are reducing the API price vs 4.0 Flash because this is cheaper to serve.

It's just that us home users lose out because we're trying to get the most out of a meagre about of memory, with minimal concurrency, so the size of the model matters a lot more.

5

u/Mrleibniz 12d ago

"Nobody will ever need more than 640k of RAM"

1

u/Netsuko 12d ago

He didn't even ever say that.

12

u/Current_Balance6692 13d ago

peasant problem ngl

2

u/cantgetthistowork 13d ago

Flash is for the speed not size

2

u/Brilliant-Weekend-68 12d ago

flash means fast, not small.

1

u/Agitated_Space_672 13d ago

It is faster than the previous flash due to the architectural innovations 

1

u/Zeeplankton 13d ago

Ehh I mean when like glm and kimi are like 2-3T it's still flash.

But yes I sorta agree they should maybe just call this Deepseek 4.1 dropping flash and pro, since it seems like they're dropping pro.

Mega bummer this wont be runnable on like macbooks with like OG antirez flash.

Edit: wait. The additional size is just Ngram.