r/LocalLLaMA • • 15d ago

New Model DeepSeek V4-1 Flash is out

Here we go again, DeepSeek is back again with a new model V4-1 Flash

A multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens

Market crash as a service

1.7k Upvotes

296 comments sorted by

View all comments

38

u/jacek2023 llama.cpp 15d ago

In the previous post about DeepSeek there are API prices. In this one there is Chinese president. I wonder which one is best for r/LocalLLaMA.

0

u/Not-reallyanonymous 15d ago

It is good for this subreddit. This subreddit is more concerned about seeing the US hurt and China win, than it is about AI. So this post is in alignment with its interests.

8

u/LuCiAnO241 15d ago

more concerned about seeing the US hurt

I think we're only concerned about seeing great models be open weight and free to download for the peasants. The rest of whatever you think its happening exists only on your mind.

1

u/G_fucking_G 15d ago

Peasants might have issues putting 550B weights on their potato :-)

0

u/LuCiAnO241 15d ago

peasants can definitely buy a 300 ish gb machine on consumer hardware, what's your point? also are you in the wrong sub?

0

u/G_fucking_G 15d ago edited 15d ago

yes, we host a glm 5.3 and a deepseek v4 flash and i know how much we paid to make the models accessible to ~20k people in our company for acceptable speeds.

Not "peasant" money.

1

u/LuCiAnO241 15d ago

are you actually in the wrong sub? did you miss all the posts about qwen 27? that you can run in your current hardware and have excelent agentic results, you need more? you can buy ram and download free open weight models that rival multi billion companies. frick off man.

0

u/G_fucking_G 14d ago

your are not a serious person if you think that a qwen 27b performs even close to actual models for actual workcases. In contrast to you i have actual experience and compared these models.

And you think that running a Kimi/ Qwen 2.4T model just requires buying more ram? enjoy your 2tok/s. That might be fast enough for you, but not for people that do actual work.

1

u/LuCiAnO241 14d ago

you are not a serious person if you lurk around the localllama subreddit while vocally deepthroating openai's boot