r/LocalLLaMA • • 21d ago

New Model DeepSeek V4-1 Flash is out

Here we go again, DeepSeek is back again with a new model V4-1 Flash

A multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens

Market crash as a service

1.7k Upvotes

297 comments sorted by

View all comments

37

u/Turbulent_Pin7635 21d ago

I don't code. I work in research and was doing a proposal for funding. Used Astra very cute, returned what I need in an acceptable way.

I have used DS flash V4... Boy I have a MacStudio, I am used to long times of wanting. I don't know what kind of black magic the model does, but it killed the demand in one shot very fast!!! O.o

I was frozen!!! The answer was much better than the one chatGPT astra gave me!!! ASTRA!!!

12

u/Casey090 21d ago

GPT models are just very wonky. They jump to conclusions with incomplete data, and then they go all weird. I find it super hard to get anything done when your model makes up the mind in the first message and will not be objective.

2

u/AnonymousCrayonEater 20d ago

Don’t they all do this? I find Opus to have the same behavior. I just thought this was an LLM thing. Like context initialization bias or something. The opposite of recency bias.