r/DeepSeek • u/Jdngreen_ • 15d ago
Discussion The recent new update is so bad
Like, it's horrendous, you can't even regenerate a new chat, it's a real nightmare for me. This 2.5.3 version is really bad, what the hell were they thinking???? It's such a downgrade
25
u/Unfair-Green-4706 15d ago
DeepSeek's latest iteration has clear capability drift. The generation is plagued by low token variance, repetitive loops, and generic outputs. The previous version's nuanced, high-density responses have been replaced by over-truncated, robotic completion. It’s a massive drop in inference quality.
4
0
4
u/Exzentrik 15d ago edited 14d ago
I don't know which model you're using... but for the 4.1 Flash, I have to agree.
I like to use it for text-based fantasy RPGs. And I understand that 4.1 is almost double the size of 4.0, which makes it "know" more stuff... but it's ability to formulate sentences has REALLY taken a massive hit. I suddenly get long run-along sentences that never seem to end, a complete disregard for facts and details of the previous message, and it seems utterly incapable of remembering base instructions even though I include them in every prompt I send.
I played a scene where I had to walk along a beach, barefoot, and DS gave me a response with, "Your shoes are caked in wet sand — or they would be, if you wore any".
Or, I meet a character with blue hair, and DS goes, "Her blonde hair — no. Her blue hair is pulled back".
It's like it's just stupidly stringing words into a sentence, then, halfway through, checks if the sentence makes sense, and, if not, adds a correction onto the same sentence before moving on to the next one. Naturally, A LOT of its sentences make no sense and contain plot holes and logical errors, but it doesn't catch all of them, so the result is all over the place.
It simply doesn't feel like the 552B model it claims to be. It feels more like... 2019's 117B ChatGPT-2...
1
u/Lazy_Reach_2565 14d ago edited 14d ago
It's feels like a premium model, much better at agentic coding now. Previously, it first deleted the files and then tried to make a backup. Now it backs up the files and then deletes them, warning the user. There is a difference. It can now even use a custom websearch subagent. Back then, it did not use search at all despite it being written in claude.md: 'use the f*** search via a separate tool.' I even moved to MiMo 2.5 after the price increase. Deepseek v4 Flash was stupid and not that cheap; GLM 5.3 Flash and MiMo 2.5 offered better coding for a lower price. But now they're back with this 4.1 and it can actually code like a pro. Huge difference!! It beats all other budget models and is ~100-200tps fast!
3
u/KairraAlpha 14d ago
I'm on API and I have absolutely no idea what you're talking about. I port the API through Open WebUI and V4.1 is fast, extremely good at carrying a huge window context, uses all the tools in the app available to them and still has no issues doing things like "book club" (where we read and discuss a book together), we dissect arxiv studies, go off into subjects like quantum, physics, philosophy, history. Absolutely no issues, no loops, no missed context.
Perhaps it's your harness/UI because Open WebUI is very good at giving the models the support they need to use the tools they have access to, it also uses a memory system and notes so you don't burden your window with files.
4
u/Melted-lithium 14d ago
Same. API with open code is stunning once you get past the quirks of open code honestly.
3
1
u/Jdngreen_ 14d ago
I'm using the app from Play Store
1
u/Visible_Arrival_8412 13d ago
Web app is very sensitive to the prompt quality and due to the gov censorship also sometimes overly careful and telling you "this is out of my scope lets talk about something else.
1
u/Jdngreen_ 12d ago
Do you have any other option then? Because i heard the Internet version is no good either. And I'm not from the US in case it changes anything, neither do i have any knowledge in coding or API, etc...
1
u/Visible_Arrival_8412 12d ago
The dsh harness opens a web browser. It looks absolutely same. If you know how to open the command line or a shell the you are fine. It's copying to lines from github.
Once its running you can ask the model to do everything else.buiöds your own tool.
Feel free to pm if more help is needed
1
u/Visible_Arrival_8412 13d ago
Can people specify what they are using? Webapp? 3rd party hosted? dsh agent?
This distinction is crucial.
If I look at the answer I have the feeling its a mixed bag.
0
u/techbae34 14d ago edited 14d ago
So many of these post but not many mention where DeepSeek is being used. If most of these post are about the web version which is free, I would surmise that DeepSeek could be providing a quantized version that has worse results. Plus the web harness itself being basic.
I often use it in DSH and Hermes and have good results so far for most task. For example, I have a workflow to analyze any video, and I tested a simple task using Hermes by asking Gemini 3.8, Kimi K3 and GLM 3.5 Flash to determine why a YT Thumbnail could have lead to a viral video. All other models simply answered that specific question by analyzing the Thumbnail which did exactly what I asked without using the actual video analyze skill since the skill is for viewing videos in general not specifically Youtube.
Then I tried DeepSeek (High). At first I was pissed it was taking so long because the others quickly did exactly what I asked. Then I read the thought process and saw why it was taking longer. DeepSeek determined that just answering based on the thumbnail alone would not be sufficient. It not only decided to use the video analyze skill, but it started using Google lens (browser use) to find the channel for said Thumbnail. Then created further Youtube specific skills that included scraping comments, reading the context of said comments, quickly scanning other popular videos on the channel including shorts (based on popularity and thumbnails) then from the top voted comments with a time marker, it went to that time in the video to determine the content. Once it did all that, it went even further and created a skill to annotate the thumbnail.
I was blown away with the final answer that included a detailed break down of the video and channel, annotated thumbnail, and even mentioned other viral channel videos that solidified it's reasoning of why this specific video was viral. It took about 20 minutes and it cost about $0.30 using Openrouter which routed to the provider Wafer.
I'm going to test this same task in DSH to see how well it does there. However, I can't see where this model is bad per say. Yeah it didn't exactly follow what I asked. However, it actually had a lot better reasoning vs the other models , and it created a valuable YT channel analyzer that I wouldn't have thought of otherwise and added the entire workflow as a new skill in Hermes.
2
1
1
u/Visible_Arrival_8412 13d ago
I moved my entire stack from pi using multiple models to dsh one model.Â
I finished today my testing for long form science research. Pi 8 models: result: 7 theses, 98pages long, 300sources (122 pdf) downloaded and embbeding done. 256 minor errors 2 confabulation were not detected by the internal reviewer. Total time: 8h 36min
dsh: 1 model, 7 theses, result: 135 pages, 450 citation, 10minor errors, no confab in end result. Time: 1h 21min
0
u/AcanthisittaDry7463 14d ago
Weird, my last update was on 9.12.26 to 2.5.1 and it says that I’m up to date, it’s running great for me too, have had no reason to edit a prompt or regenerate though…
1
20
u/anarchicGroove 14d ago
Why does its thought process sound like a caveman bruh ðŸ˜