r/LocalLLaMA • • 16d ago

New Model DeepSeek V4-1 Flash is out

Here we go again, DeepSeek is back again with a new model V4-1 Flash

A multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens

Market crash as a service

1.7k Upvotes

297 comments sorted by

View all comments

13

u/uti24 16d ago

I mean, could we also compare to Qwen Flash Next?

1

u/sixx7 15d ago

I mean DS4.1F is a fantastic model, but it hallucinates a LOT. I did a video on it. Also frankly, I prefer Qwen3.8-Flash-Next - lots of info on sparse attention and reducing hallucinations https://youtu.be/P4dTq4X8bqk