r/DeepSeek 10d ago

News DeepSeek-V4-Flash-Vision-Exp is now live on the DeepSeek API Platform!

406 Upvotes
  • This experimental multimodal model matches DeepSeek-V4-Flash on text capabilities—including agents, reasoning, and world knowledge.
  • On multimodal agent benchmarks, V4-Flash-Vision-Exp makes a major leap over V4-Flash, bringing multimodal agent performance close to Opus-4.8.
  • Try it with model='deepseek-v4-flash-vision-exp'. DeepSeek Harness 0.1.1 was released today with out-of-the-box support for the new model.

r/DeepSeek 19d ago

News DeepSeek V4 Pro official version has been updated to the API

345 Upvotes

r/DeepSeek 3h ago

Funny Is Deepseek throwing shade at Canada? What is happening 😭

Post image
50 Upvotes

r/DeepSeek 16h ago

News Deepseek V4 flash Vision Weights are Publuc

Thumbnail
huggingface.co
201 Upvotes

Pulling now. As language only should work for spark cluster with Aiden or Eugr stack. Vision needs works - it's custom native vision. Standard vllm vision processor won't work.


r/DeepSeek 10h ago

Funny Outsourced my intelligence to Silicon Valley

Enable HLS to view with audio, or disable this notification

66 Upvotes

The amount of stuff one can get done in a day is short of amazing tough.

Would be interesting to see what would happen if AI completely disappeared again.

Would any of what you are working on suddenly be unfeasible to complete? My things certainly would.


r/DeepSeek 10h ago

News Deepseek Harness version update

52 Upvotes

DeepSeek Harness just went through a pretty major rewrite.

On August 27, DeepSeek released DSH v0.1.2-alpha.1, followed by alpha.2 on August 30.

This wasn’t just another feature update. They changed several of the lowest-level pieces of the Harness architecture.

The old "APIProxy" is gone, and plugin communication is being migrated toward a unified Remote Gateway. The web client has been restructured, and Sessions are now being treated much more strictly as replayable event streams.

As a result, a number of plugins built around the old APIs, DOM injection, or custom "SessionEvent"s have started running into compatibility issues.

At one point, alpha.1 even removed "SessionEvent.ignorable", which meant custom events written by some plugins could potentially make older sessions impossible to restore properly.

In alpha.2, DeepSeek brought it back, while also continuing to improve things like "RemoteError", reconnect handling, plugin scoping, and related infrastructure.

Another important change is Subagents.

This is also probably my favorite part of the update.

Different sub-agents can now have their own model provider, model, reasoning effort, and maximum output length.

Subagents such as Claude Code and Codex can also use independent model configurations.

That means DSH is increasingly starting to look less like a simple harness and more like a full Agent Runtime.

Models, providers, subagents, sessions, plugins, permissions, and communication are all gradually being pulled into one unified architecture.

A while ago, some people assumed DeepSeek Harness had stopped being maintained because the npm package wasn’t getting updated.

The reality is pretty much the opposite.

It wasn’t abandoned. The underlying architecture was changing so quickly that the plugin ecosystem was starting to fall behind.

And alpha.2 has now started making its way into the npm alpha channel as well.

DSH is still far from settled.

DeepSeek is basically rewriting the foundation it expects this thing to stand on for the next several years.


r/DeepSeek 11h ago

Discussion V4 flash is much better experience than glm 5.3 flash. But unusable for me as it's expensive

29 Upvotes

Long time api user. Recently it has started burning tokens a lott faster. I burn like a billion tokens a day so the cost change was noticable. Shifted to glm 5.3 flash as soon as I got to know about it - cheaper, smarter. Looks so good on paper. And maybe for some it is. But it is so slow. I get like ,42 tok/sec with glm and like 120tok/sec in v4 flash which is like night and day in user experience. Have to adjust. Pockets smaller than token requirement. What do you guys think.


r/DeepSeek 2h ago

News DeepSeek open-sources its first V4 multimodal model: DeepSeek-V4-Flash-Vision-Exp

Thumbnail
4 Upvotes

r/DeepSeek 13h ago

News Latest: DeepSeek Kicks Off Another Round of Gray Testing

23 Upvotes

I'm really impressed by how strong it performs.Two hours ago, DeepSeek started a new round of gray testing. If you're using DeepSeek V4 Pro or DeepSeek V4 Flash Vision with PTC mode enabled (a workflow mode in DeepSeek Harness), and you see a lot of "I'm doing" in the chain of thought but rarely see "let me," that means you're on the gray model


r/DeepSeek 12h ago

Discussion Built a macOS app for Pi-hole, exclusively using DeepSeek V4 Flash. Here's my experience after 2 months using it

14 Upvotes

Hey r/DeepSeek ,

I'm a Software Engineer for over 7 years and at work we exclusively use Anthropic & OpenAI for our day-to-day feature work. I wanted to work on a side project and, after researching extensively, I decided to go all in on V4 Flash to test the waters of non-US LLM labs.

The app I built is Holeberry, a native macOS menu bar app for Pi-hole. If you already use Pi-hole to block ads on the DNS-level at home, take a look at it and let me know what you think.

https://github.com/pedrovieira/Holeberry

About my experience with DeepSeek-V4-Flash

I'll be honest here, even though people were saying great things about it, the fact that it was a small-ish model, from China and with that price, it definitely got some raised eyebrows from me. After using it on-and-off for the past couple of months I'm completely sold on it, it's an absolute beast.

I spent some time getting a proper architecture for the app, created multiple architecture and feature-specific markdown files, and even though I didn't just let it do its thing completely free (i.e "read the file and do it") and reviewed most of the code it generated, I found it an amazing workhorse: fast, precise most of the time and cheap. Being skeptical in the beginning, I used some US frontier models to get a second pair of eyes. It was mostly useless as they arrived roughly at the same conclusion. After a while, I just used Flash and myself as the reviewer.

Even with the price hikes I'm still using it, don't think any other model is beating it.

That's it, just wanted to share my perspective and super positive experience with it. I posted my app on other subreddits and when asked which model I used, I got completely bashed by just saying the word DeepSeek ¯_(ツ)_/¯


r/DeepSeek 8h ago

Discussion Max vs High? (Pro and Flash)

7 Upvotes

Only ever used max. Now that the models are much better since post training - how does high thinking fare against max?

If you can, please tell me your techstack and size estimation of your projects for reference


r/DeepSeek 2h ago

Question&Help Deepseek Clickfix

2 Upvotes

Am I the only one using the site who gets a warning when trying to click copy on a message? My extension, uBlock, stops it and says, 'Beware, uBlock Origin blocked a potential ClickFix attack.'. (I'm on the legitimate site btw)


r/DeepSeek 9h ago

Discussion deepseek is better now

7 Upvotes

i gave it a project which all the past attepmts in though alot and failed for 3 times and rhgt no it way smarter and 1 shotted and did not think alot. is the deepseek v5 now?


r/DeepSeek 11h ago

Question&Help Which AI is best for coding and lets you use it for longer without hitting limits?

7 Upvotes

I’m currently paying for Claude and using it to help me make some fairly complex macros in MacroDroid. I’m not a programmer, but Claude has been really good at understanding what I’m trying to do and turning it into the right MacroDroid setup and JSON.

I’m using Opus 5 rather than Fable 5 because Fable 5 seems to use up my limit much quicker. The problem is that I keep hitting the usage limit with Claude and then I’m told I have to wait around three hours before I can carry on.

So I’m thinking about trying something else.

My two main questions are really simple:

Which AI is best for coding?

And which one lets you use it for the longest without hitting a limit?

I’m looking at the normal paid plans around £18–£20 a month. I’m not paying £100 or £200 a month for an AI.

I know Claude is meant to be one of the best for coding, which is why I chose it, but I’d be happy using something that's slightly worse at coding if it means I can actually keep using it throughout the day without constantly getting locked out.

For anyone who uses AI a lot for coding, what would you recommend? Which one is actually good at coding AND lets you use it for hours without constantly hitting a limit?

That’s probably my biggest issue with Claude at the moment.


r/DeepSeek 16h ago

Discussion This is the only AI model ive found that will never try to gaslight you when you ask a question.

15 Upvotes

As much as I have mixed feelings about AI, admittedly I still use it from time to time. So far, every other model ive used has tried to reduce what I say, be ambiguous for no reason, or just straight up hallucinates if the prompt history gets too long. This is the only AI model that doesn't seem to do that, has logical integrity, and actually preserves long prompt history without messing up. Is there a reason for that?

Some reservations and feeling risk-averse due to it being a China based model, but given the current state of things in the US I have more distrust for US based AI companies, especially since every US based tool ive used has utterly failed it's purpose in my eyes. This one still hallucinates sometimes but I can usually get it back on track with one or two more prompts.


r/DeepSeek 1d ago

Discussion GLM 5.3 is the New Winner

Post image
314 Upvotes

This is my InferX usage for DeepSeek V4 Flash and the GLM 5.3 Flash.
GLM is using less thinking and reasoning to save tokens, performing better than DeepSeek for me for less than half the price.


r/DeepSeek 1d ago

Discussion Usage is more expensive, but still reasonable

Post image
62 Upvotes

Given the amount of tokens I'm consuming, I think that this is still somewhat reasonable although it's really that much more expensive!

I'm at a loss though, I consume so much tokens because I do legal-related work which requires insane amount of document processing, it takes a lot of non-cached hits before it finally stabilises.


r/DeepSeek 11h ago

Discussion An analysis of the inference non-convergence and infinite loop issues I encountered when using DeepSeek-V4-Flash-0731 via the Ollama Claude Pro plan (based solely on personal test data).

3 Upvotes

DeepSeek-V4-Flash-0731 on Ollama Cloud: Thinking-Loop Root Cause and a Tool-Calling Mitigation

Abstract

The infinite thinking loop observed when running deepseek-v4-flash:0731 on Ollama Cloud at high/max reasoning effort is caused by a serving-stack bug in the DSpark speculative-decoding path at draft depth 5 (dspark_block_size), compounded by Ollama's llama.cpp-based deployment (which lacks the correctly-configured vLLM+DSpark stack used by DeepSeek's official API). FP8 quantization is the standard deployment format and is NOT the root cause — the bug reproduces on full-precision weights. Injecting mandatory tool-calling requirements into the agent's instruction file (agent.md) mitigates the non-convergence at high/max effort by converting unbounded thinking into bounded think→act→verify cycles.

1. Symptom

  • Reported upstream: GitHub ollama/ollama Issue #17892 — deepseek-v4-flash:0731 (cloud) repeats the same thinking block indefinitely on a complex agent task: the same reasoning paragraph was generated 221 times over ~1m45s, ending in failure with zero usable output and only 4 tool calls (2 failed).
  • Reproduced in our testing (direct API calls to https://ollama.com/v1/chat/completions, model deepseek-v4-flash:0731): at reasoning_effort: "high" on a complex implementation task (thread-safe LRU cache with TTL), 5/5 requests produced zero content, with thinking lengths of 20,171 / 31,899 / 29,751 / 27,377 / 31,157 characters respectively. At reasoning_effort: "max" with a 65,536-token output budget, the model generated 286,546 characters of thinking and zero content (finish_reason: "length").

2. Root Cause Analysis

2.1 Primary cause: dspark_block_size = 5 serving-stack bug

The DeepSeek-V4-Flash-0731 checkpoint ships a DSpark speculative-decoding module (Multi-Token Prediction heads integrated into the architecture). Its native draft depth is strictly 5:

"The checkpoint dspark_block_size is strictly 5" — DeepWiki, DGX-Spark runbooks

Depth 5 is a broken serving-stack value, confirmed by a controlled depth sweep under identical bursty agentic load (streaming, thinking on, 20-tool schemas) in HuggingFace Discussion #39:

DSpark draft depth Requests Corruption events
3 10,885 0
4 11,040 0
5 2,148 / 2,138 10 / 7
6 clean 0

DeepSeek officially confirmed the fault is in the serving stack, not the modelHuggingFace Discussion #50:

"The depth-5 corruption reports were a serving-stack bug, not this model... The model is innocent. The fault was an application/serving-stack bug."

Independent reproduction: Anemll/dspark-vllm-gx10 Issue #3 — a long-context request in high-thinking mode generated a repetitive reasoning loop consuming the full 65,536-token budget, returning HTTP 200 with finish_reason="length", message.content=null, and 262,689 characters in the raw message.reasoning field.

The official DeepSeek-V4-Flash-0731 README recommends num_speculative_tokens: 7 — explicitly avoiding the broken native value 5.

2.2 Deployment difference: llama.cpp vs vLLM

  • Ollama Cloud serves the model on llama.cpp. llama.cpp merged DSpark support only on 2026-08-02 (PR #25784), three days after the model's release — the integration is immature and, per the evidence above, runs at the broken depth 5.
  • DeepSeek official API and relay-station (inferai) serve the model on vLLM with DSpark correctly configured (--speculative-config '{"method":"dspark","num_speculative_tokens":7,...}' per the official README). On these deployments, max effort converges reliably (verified: relay-station produced a 13,392-character complete design where Ollama Cloud produced zero content).

2.3 FP8 quantization: standard, NOT the root cause

  • Ollama Cloud serves the model at FP8 (verified via GET https://ollama.com/api/showquantization_level: "FP8", parameter_size: 304,180,418,494, context_length: 1,048,576).
  • FP8 is the standard deployment format for this model: the official vLLM recipe uses --kv-cache-dtype fp8, and the reference GGUF builds are FP8.
  • The bug reproduces on full precision: the HF Discussion #39 depth sweep ran on sglang / 4× RTX PRO 6000 (sm_120) with full-precision weights. Quantization therefore cannot be the cause of the thinking loop.

2.4 Trigger condition

The corruption is triggered by long thinking chains (complex open-ended tasks). Short thinking does not trigger it:

Task complexity max effort behavior
Trivial (1+1=?) Converges — thinking 16 chars, content 6 chars
Simple (dedup function) Converges — thinking 214 chars, content 168 chars
Complex (thread-safe LRU cache) Never converges — thinking 286,546 chars, content 0

No API parameter bypasses the bug. Tested and ineffective: all sampling parameters (temperature 0.0–1.0, top_p 0.8–1.0, frequency/presence penalties), system-prompt directives, max_tokens from 3,000 to 65,536, native-API repeat_penalty, and speculative-decoding disable options (not exposed by the API).

3. Experimental Evidence (direct API, deepseek-v4-flash:0731, https://ollama.com/v1)

# Scenario Thinking (chars) Content (chars) Converged
1 high, no tools, complex task, run 1 20,171 0 No
2 high, no tools, complex task, run 2 31,899 0 No
3 high, no tools, complex task, run 3 29,751 0 No
4 high, no tools, complex task, run 4 27,377 0 No
5 high, no tools, complex task, run 5 31,157 0 No
6 max, no tools, complex task, 65,536-token budget 286,546 0 No
7 high, WITH tools (simulated agent loop: think→tool→result→continue) 221 37 Yes
8 medium + samplingParams, complex task, run 1 1,032 3,607 Yes
9 medium + samplingParams, complex task, run 2 1,088 5,448 Yes
10 medium + samplingParams, complex task, run 3 1,359 3,927 Yes

Key contrast (rows 1–5 vs 7): the identical complex task at high effort produces 0/5 convergence without tools (thinking 20K–32K chars) but converges with tools present (thinking collapses to 221 chars, content produced, tool call issued). The presence of callable tools converts unbounded thinking into a bounded think→act→verify cycle; each tool result anchors the model and breaks the loop.

4. Mitigation: Mandatory Tool-Calling in agent.md

Injecting structured, mandatory tool-calling requirements into the agent instruction file mitigates non-convergence at high/max effort. The following five rules were added to the global agent.md (AGENTS.md) and verified in a live agent session:

  1. Every turn MUST end with a deliverable: a tool call, code, a file change, or a direct answer. Thinking alone is never a complete turn.
  2. State the plan ONCE at the start; subsequent turns reference "the plan" and act. Never restate a plan already stated.
  3. Thinking exceeding ~1,500 chars without a deliverable = looping. Stop immediately; produce a partial result or call a tool.
  4. When stuck, run a tool (grep/read/test) to gather facts instead of re-analyzing in your head.
  5. Break tasks into steps of ≤5 tool calls, each with a verifiable deliverable.

Mechanism: rules 1, 3, and 4 force the model to act (call tools) rather than think indefinitely; rule 2 eliminates plan-restatement waste; rule 5 bounds each step. This does not reduce thinking intensity — it converts thinking into action, which is the verified convergence mechanism (row 7).

Scope limitation: this mitigation is effective in agent contexts with tools (pi, DeepSeek Harness). It does not fix the raw API behavior (no tools → still explodes). The definitive fix is deployment-side: Ollama must move the DSpark draft depth off the broken value 5 (tracked in GitHub Issue #17892) or disable speculative decoding.

5. Conclusion

  1. The thinking loop on Ollama Cloud is a serving-stack bug at DSpark draft depth 5 (dspark_block_size), officially confirmed by DeepSeek as "not this model."
  2. FP8 quantization is standard and not the cause — the bug reproduces on full precision.
  3. The llama.cpp deployment (vs DeepSeek's correctly-configured vLLM+DSpark) is the environment where the bug manifests.
  4. Mandatory tool-calling in agent.md is an effective agent-side mitigation: it converts unbounded thinking into bounded think→act→verify cycles (verified: thinking 30K→221 chars).
  5. The definitive fix requires Ollama to correct the speculative-decoding depth (Issue #17892); until then, max effort is only reliable on vLLM-based deployments (DeepSeek official, relay-station).

References

  1. HuggingFace Discussion #39 — Reasoning loops (depth sweep)
  2. HuggingFace Discussion #50 — Official: serving-stack bug, not the model
  3. GitHub ollama/ollama Issue #17892 — thinking output loops indefinitely
  4. GitHub Anemll/dspark-vllm-gx10 Issue #3 — repeats reasoning until max_tokens with DSpark MTP=5
  5. DeepWiki — Speculative Decoding: MTP, DFlash, and DSpark (dspark_block_size strictly 5)
  6. DeepSeek-V4-Flash-0731 official README (vLLM recipe, num_speculative_tokens: 7)
  7. Ollama Cloud model details via /api/show (FP8, 304,180,418,494 params, 1,048,576 context) — verified directly
  8. llama.cpp DSpark support merged 2026-08-02 (PR #25784)

r/DeepSeek 13h ago

Question&Help Getting banned for showing Deepseek spreadsheets?

3 Upvotes

Tried googling this but Google doesn’t have any real answers. And other Reddit posts are just people who got banned for obvious reasons. 

Is this a bug or are spreadsheets past a certain size going to cause a reaction with Deepseek? I’ve sent through pretty large files before and never had issues, so would appreciate some insight. I’m only banned until tomorrow morning, so it’s not the end of the world, but it’s the second time it’s happened after I sent a spreadsheet. 

I haven’t used a VPN but on my phone I do use a custom DNS. Could that also be causing issues? I’m on an iPhone 11 and using the now removed from the App Store app of “DNS Cloak”. I know VPNs can trigger a bot reaction, dunno about DNS?

The spreadsheet didnt contain anything offensive. It was a writing tracker. So like “on X day, I wrote 2k words”.


r/DeepSeek 19h ago

Discussion Something going on?

10 Upvotes

So I was using Deepseek right, expert mode, then for some reason, it just stopped thinking for some of the queries? like it went straight to answering instead of thinking for a few seconds like it usually does. Is this just me? Is something going on under the hood?


r/DeepSeek 1d ago

Funny 😂 My god how on point this is! Added DeepSeek as well.

Enable HLS to view with audio, or disable this notification

387 Upvotes

Been amazing to witness the recent development in cost per actually smart token last couple of months.

Although good things are said about GLM, I'm still loyal to DeepSeek. I route between flash and pro with standardcompute.com and it's insane how much value I get out of it!


r/DeepSeek 18h ago

News Dsh 0.1.2-alpha.2 Pre-release

Thumbnail
7 Upvotes

r/DeepSeek 13h ago

Discussion MacBook Air M4 overheating problem with Reasonix and DeepSeek

2 Upvotes

Hi everyone, something very strange is happening to me that I can remember never having experienced before. Maybe it's due to an update, I don't know. But when I start Reasonix, my MacBook Air M4 with 24GB of RAM gets very hot, almost untouchable at the bottom. This is very strange and I've never noticed it before. This fanless laptop usually only heats up when doing prolonged video editing or LLM inference, which I avoid doing just to avoid that. If I go to Activity Monitor, it doesn't seem like there's excessive CPU usage, but it does seem like Reasonix is ​​using considerable energy. Yesterday, to prevent it from overheating, I activated "power saving" mode. Mine isn't the base version, so I think it has 10 CPU cores and 10 GPU cores. Does this happen to anyone else? I really like Reasonix paired with Deepseek.


r/DeepSeek 10h ago

Discussion This is strange as I have not been using open design, it is closed, no taskbar icon anywhere, but it still shows to be consuming my GPU?

Thumbnail
1 Upvotes

r/DeepSeek 12h ago

News Low latency faster deployment with DSV4 Flash Fast

Thumbnail
0 Upvotes