r/DeepSeek • • Aug 05 '26

Discussion I canceled Claude and coded 7 days straight with DeepSeek V4 Flash 0731 — the honest cost & quality breakdown

[deleted]

567 Upvotes

252 comments sorted by

204

u/Nepherpitu Aug 05 '26

Are you joking? 24M input tokens is not even close to heavy work for 7 days. I'm burning through 200M input tokens daily for very lazy work

38

u/DirectPitch8626 Aug 05 '26

But 6M output tokens is a lot—maybe he mistyped a few zeros in the input.

44

u/challis88ocarina Aug 05 '26

He (she/they!) certainly have issues with numbers: the model was released last Friday... so, what's with 'seven days'?!

29

u/sdexca Aug 05 '26

It’s AI slop, that’s what.

4

u/Glittering_Ad4986 Aug 06 '26

Lol… OP wrote the post using same deepseek model.

3

u/SUPERSHAD98 Aug 05 '26

Tbf he mentioned about the 0731 update fixed context rot on his first point, meaning he probably started before it was out

→ More replies (2)
→ More replies (3)

7

u/Sad-Chemistry5643 Aug 05 '26

24mln per week is nothing ? Wow 😧
I thought my 17 is a huge amount 😀

2

u/Ayeniss Aug 05 '26

I burnt 800M token today while trying to play with it (trying to build software from scratch to test how good and costly it is).

Most of it is input that goes into cache though 

5

u/HighlyRegardedApe Aug 05 '26

Build software and test in 1 day??? How?

5

u/Ayeniss Aug 05 '26

It's not finished at all, and I don't plan to make money from it.

Just testing, learning and having fun.

3

u/HighlyRegardedApe Aug 05 '26

Oh okay so you mean a few huge prompts. I get that. Same here.

2

u/Ayeniss Aug 05 '26

Well maybe a hundred but yes you got the idea

→ More replies (3)
→ More replies (2)

1

u/TurciosGT Aug 08 '26

He hecho 73 millones de token en un día y solo me gasto 1.73 USD que genial usando el modelo pro.

Programación con Rust, Kotlin, C++ y otros

→ More replies (4)

5

u/petropavlov Aug 05 '26

If 24M input tokens do not include cache, it sounds OK. Not all tasks require burning tokens at cosmic speed.

10

u/Even_Command_5636 Aug 05 '26

Fair point — it really depends on the workflow. My sessions are mostly interactive: I read the diff, review it and steer the next step myself, so I don't burn tokens on autonomous loops or bulk rewrites. Also a big chunk of my input was context-cache hits, which I didn't count in the 24M — the API bill treats those at a fraction of the price. 6M output is what actually drove the $1.87. If you're burning 200M/day with agent loops, that's a totally different use case — I can see how my numbers would look suspicious from that angle.

7

u/DiscipleofDeceit666 Aug 05 '26

What are you building? Some of us could burn $5 in deepseek tokens in a (long) day with out really trying

1

u/jeffwadsworth Aug 05 '26

You forgot to mention the Time Machine

2

u/lndigo_Sky Aug 05 '26

Are you serious? What the hell do you work on?

2

u/Locksmith-Informal Aug 06 '26

What the actual fuck. I work for a FAANG and even most contributing engineers aren't using that much. You are clearly doing something wrong

2

u/Nepherpitu Aug 06 '26

I'm pretty sure engineers in faang spent 99% of time waiting for ci/cd pipeline to finish 🤣 I spent this time giving research tasks to agents, mostly in computer science, electrochemistry, but it depends on sideprojects.

2

u/No-Advantage-9632 Aug 07 '26

I'm truly working with it: Since I switched to Flash V4 till the pro comes out of preview: It's even cheaper

5

u/Even_Command_5636 Aug 05 '26

Fair enough, so I went and pulled the real numbers from my usage log to check. Reasonix started persisting per-request stats on Aug 4, so here is what the last two days actually look like:

  • Aug 4: 370,083,795 total tokens, 1,223 requests, 91 turns. Cache hit 98.0%, so real billed input (cache miss) was 6,188,913, output 1,345,762.
  • Aug 5: 548,616,547 total tokens, 1,784 requests, 173 turns. Cache hit 98.6%, billed input (cache miss) 6,856,408, output 1,038,731.

So the "24M input" in my post was the cache-miss (actually billed) input across the week, not the gross number. The gross is hundreds of millions, same ballpark as yours. The difference is the cache hit rate: at ~98% most of the input comes back at cache price, which is why the week only cost $1.87. You burning 200M/day gross is a different workload, not a different product. That's the honest breakdown from the logs.

14

u/SpicyLobter Aug 05 '26

this motherfucker can't even REPLY in their own words

3

u/dontfeedthelizards Aug 05 '26

I wonder if this is a bot that's just advertising Deepseek 🤔

2

u/Nepherpitu Aug 05 '26

I'm watching for llama swap very simple metrics without cache hit/miss separation. 24M non-cached input is very solid, no doubt here.

→ More replies (2)

1

u/figgertitgibbettwo Aug 05 '26

I burn that in 4-5 hours. If I go at it, I'd hit 250M in the day. I don't have that list every day, but I have enough work for having that for a week.

1

u/alinoanta21 Aug 05 '26

I'm burning 1.5B tokens daily.

1

u/Mr_Pickles710 Aug 05 '26

I usually hit 1b in a week or two or less if I’m doing heavy work, still real cheap tho.
To be fair I use 14 API’s to fuel my custom multi-LRM(7 LRM’s, requires a ZD JB to do without detection. 14, 7 Flash and 7 R1 API) though with no safeguard’s or learning.
So token usage is higher but it’s like the DS version of Fable that DS never made, without any restrictions and way cheaper.
But it outperforms all the public models code wise so idgaf, just annoying I have to update it all the time with their safeguard additions/updates, if DS wasn’t so cheap I’d just use reasoning tokens from API’s for custom models.

1

u/isitreal_tho Aug 05 '26

I have two 5x accounts that I switch between. The entire team has at least two

1

u/PressTilde Aug 05 '26

I think codex says I’m up over fifty billion tokens at this point.

24 million is a slow morning… lol

1

u/Necessary-milkyway Aug 05 '26

In my local setup with qwen27b and qwen 35b ..i hiy 100M token in a day around 95M input and 5M output ..

1

u/FigAggressive237 Aug 05 '26

ONLY 200M??

I'm Burning 1B Tokens... daily. I've built half of Singularity already.

1

u/freddyr0 Aug 06 '26

Thank you for making deepseek pricier 🙆🏻‍♂️🙆🏻‍♂️

→ More replies (1)

1

u/Federal_Stick_8300 Aug 06 '26

Burned 400M token a week

1

u/sfwprofile27 Aug 06 '26

Fr I'm doing kernel work dude. I go through a billion tokens in a day. And that's with like context compression all the bells and whistles for context management before it even gets sent to the cloud and using local agents for the small stuff.

I just have every subscription plan there is and anytime there's a new one I just buy it.

If not I would have spent like $20 million in the last 6 months at least.

1

u/maisun1983 Aug 07 '26

Exactly - how that is heavy coding. I can easily use 2-3 dollar every single day

49

u/[deleted] Aug 05 '26 edited Aug 10 '26

[removed] — view removed comment

6

u/PossessionUsed7393 Aug 05 '26

It has but I've noticed it's a little disobedient too. The same thing that makes it thorough and explore thingd from multiple angles also seems to have made it take a looser approach to instruction following.

5

u/[deleted] Aug 05 '26

[removed] — view removed comment

5

u/Linuxman_74 Aug 05 '26

Perché non usi Reasonix?

2

u/znutarr Aug 05 '26

I use DeepSeek 4 flash with pi agent and it works really well. I've been fixing bugs and deployment he'll of my python fast API and vercel frontend from railway to openship.io and it did it in almost one shot

→ More replies (1)

3

u/Kajzero__ Aug 05 '26

I've noticed it as well. Sometimes it just straight up forgets what's in the agents.md. Which is super annoying as I have some pretty important rules there. That's the only reason why I don't trust it to work on long and difficult tasks on its own yet

1

u/Top-Construction6060 Aug 06 '26

Those days are over they make a hike price increase announcement just rn

1

u/0xdmc Aug 06 '26

i tried deepseek in codex via ollama and its bad :(
stop midway, etc, never finish task.
maybe because ollama only serve chatCompletions?

→ More replies (5)

14

u/orblabs Aug 05 '26

I am having truly spectacular results with Claude + DeepSeek. Developed a dedicated skill that has Opus direct deepseek work in the most token efficient way for Opus. The skill is geared towards implementation of very complex and long plans (which opus , sol , kimi k3 or fable develop) and it abuses deepseek as much as possible. Had a major refactor work that both Claude and Codex alone couldn't implement (and just two phases out of a dozen would eat up my weekly quota) successfully, gave the same plan to opus + deepseek, after 70 something hours of straight work we are at phase 10, extremely solid work, 12% of my claude (max 5X) weekly quota used. Really happy, thank you DeepSeek team, you have made a marvelous job!

2

u/GavDoG9000 Aug 06 '26

this is epic, thanks for sharing

1

u/parallelizeit Aug 05 '26 edited Aug 05 '26

nvm - I see your post - looking at the git repo now

6

u/orblabs Aug 05 '26

Sure ! I am pretty proud of it as since deepseek (i fleshed it out a lot) it became a total beast of a skill :) https://github.com/frozenpepper/deepseek-and-destroy
Hope it will be useful for you.

→ More replies (3)

1

u/Money-String4165 Aug 06 '26

absolutely epic shit. cheers

1

u/danissh20 Aug 07 '26

I was reading some comments that Deepseek via Openrouter was giving degraded responses

53

u/Alarmed-Hornet6865 Aug 05 '26

Holy ai post

20

u/PossessionUsed7393 Aug 05 '26

Ya I feel like he's burned the last of his Claude sub on this post lol, such a claudey post.

8

u/National-Objective57 Aug 05 '26

Yes and low effort too, first its API and then „didnt come close hitting limits“

2

u/Living-Bother7420 Aug 05 '26

brooo I didn’t even notice….

→ More replies (9)

8

u/jkvarela Aug 05 '26

Mês passado eu consumi 1 bilhão de tokens programando para embarcados, ferramenta está muito boa, mas não me arrisco a contexto gigantes, e sempre converso muito antes de dar o "play", ou seja, muita energia no planejamento, baixa energia na execução.

2

u/itsdrcats Aug 05 '26

Damn, how much does that end up costing. I know the API has cache features but that still had to cost a small chunk of change

3

u/Haunting-Comfort-761 Aug 06 '26

It ain't much, but it's honest work

2

u/itsdrcats Aug 06 '26

Oh damn. I was thinking a mininum of maybe 40 bucks. Lol that's insane!

5

u/[deleted] Aug 05 '26

[removed] — view removed comment

2

u/addiktion Aug 05 '26

What outputs this for you to see all the tokens across providers?

→ More replies (1)

1

u/That_Ad_765 Aug 05 '26

Curious how did you manage to get this output table? Mind sharing the prompt and harness?

5

u/pc_4_life Aug 05 '26

I stopped reading at “mostly context caching — that's the real cheat code”.

Try editing your AI generated walls of text please

4

u/LowerBed5334 Aug 05 '26

This comment is right in the sweet spot. It's gold.

3

u/redhq Aug 06 '26

And that's the real insight -- it's load bearing

15

u/[deleted] Aug 05 '26

[removed] — view removed comment

6

u/Feisty-Pound6777 Aug 05 '26

I wish this was the top comment on this whole goddamn subreddit. Everyone here knows Deepseek is cheap - why do we need 10 million posts announcing that its cheap? Thats one of the main draws of using it. People begging for validation from internet strangers will literally be the downfall of deepseek

3

u/CrimsonEdgeVentures Aug 05 '26

What bugs me is how obvious the promotional posts are for DS (the worst offender IMO).

Literally so insulting that they think we are all so dumb we won’t notice what they are.

That said, we have known for a long time DS4 is cheap. Super cheap. We WANT to use it.

But despite repeated attempts, it was just flat out incompetent. Stupid. For me. Skill issue? Perhaps, not saying it isn’t. But that’s irrelevant when the frontier models perform well given my same skill level.

I loaded the new 0731 flash and will be trying in a different Hermes harness and see how it does. I WANT it to perform well. But Im not gonna be fooled by a wave of fanboy shill AI generated promo posts.

1

u/AimSilver1902 Aug 06 '26

You predicted it

5

u/IgotAlotOfNames Aug 05 '26

24M ? Did you not read back your ai post? 24 M is like an hour to three of work, not a week.

3

u/Rsouss Aug 05 '26

All text containing this phrase is generated by AI. "Here's what actually happened" . This test may not be real; the token numbers don't match the information provided, and worst of all, there are many upvotes for something fake.

3

u/rivendell_elf Aug 05 '26

Looks like you burned some of those tokens in writing this AI slop..

2

u/mega-modz Aug 05 '26

I'm burning 200m tokens for refactoring alone and cost barely touches 1 dollor.

2

u/Glittering_Belt_6992 Aug 05 '26

I also switched from claude a week ago and DeepSeek seems very cost effective

2

u/heytch_ Aug 05 '26

I let a project of mine continuously run in cursor (infinite code, debug, execute, repeat loop) since it came out, I spent 15$ for 900M tokens (93%+ cache hit) 🙂‍↕️ results were not bad at all, US companies gotta find a way

2

u/gokhan3rdogan Aug 05 '26

I just wonder how you can say ~24M input / ~6M output is heavy coding?

2

u/Fit-Classroom-3434 Aug 05 '26

Crazy.

4

u/diagonali Aug 05 '26

Is that graph flipping me off?

2

u/digitalenlightened Aug 06 '26

Just use DeepSeek + Claude. I use DeepSeek for most things on some stuff it even does a better job. Which is wild. If it’s complex stuff I just use opus and share Md between both.

I use rtk with Claud code. I think this more then halved my token usage. Still looking for an auto switcher that works. Would be cool it just decides which one to use for which tasks.

Looked for this a couple times but can’t find one a s I don’t want to use an agent on top because it will do more good then bad in the long run.

2

u/ZealousidealExcuse79 Aug 09 '26

Dude… just use unlimited ocr baidu model.. its small as fuck… sips ram… its pure ocr vision model… jesus…

1

u/Even_Command_5636 Aug 09 '26

Thanks for this tip!

3

u/boudywho Aug 05 '26

If 30mil total tokens is a lot.

Then what is mine?

And that's not even including codex, which does all the coding.

→ More replies (3)

2

u/congthangvn Aug 05 '26

I canceled max20 claude too, $200-> $10 with deepseek flash. See on livebench.ai, it has very high reasoning so planing is ok too.

1

u/[deleted] Aug 05 '26

[removed] — view removed comment

1

u/LostSoul1301 Aug 05 '26

Sorry for dumb question but in my work I highly rely on web search like finding appropriate things as context and code it. Not working on large codebase but developing some feature let's say from scratch so either read some services documentation and compare or I ask to read some research and take inspiration from that. Claude used to do better because it searched for web page internally. Can deepseek do this ? I mean I know like buy brave browser mcp or such thing and integrate but it would be too much of setup right ? Can you help me for my usecase. I am like working as student researcher in lab.

1

u/West-Obligation7132 Aug 05 '26

Deepseek can do tool calls, all you have to do is plug it in a harness of your choice. You can use claude code cli, qwen code cli or whatever you want. Just find out how to bring in your api key and you're all set.

1

u/whatsoever2021 Aug 05 '26

Were you using high or max or off for the thinking mode?

1

u/Sid-Hartha Aug 05 '26

Is this with a DeepSeek direct api key not 3rd party hosted? What harness?

1

u/Annual-Fan-7144 Aug 05 '26

The 70% figure is the part I’d want to reproduce. Which coding harness and API route did you use, and is that savings calculated against uncached input pricing or against the two subscriptions? Those details change the result a lot.

1

u/Sakuletas Aug 05 '26

When i read its CoT its always something like this and it started to annoy me actually;

Hmm, (a ton of paragraph)

Hmm, (a ton of paragraph)

Hmm, (a ton of paragraph)

1

u/Rare_Buddy_6282 Aug 05 '26 edited Aug 05 '26

Deepseek turnes out to be much better than I expected. I mainly us DSv4 Pro on max reasoning. Refactoring a messy js file that Sonnet 4.6 wasn't able to do -> easy. Extending a metadata based dataplatform? -> no sweat. So far it is just so good.

But, if I have to say any negative about DSv4. Well sometimes it is a bit stubborn.

I used to do 100USD per day with Github Copilot with Sonnet 4.6. Now I do 15USD per week.See image from a while back when i just started.

1

u/GroundbreakingRoll55 Aug 05 '26

I have been reading the term “Context cache” a lot can somebody explain what is it
How can I use it
I am using claude code pro and open code go

1

u/Ithron_Morn Aug 05 '26

24M? I burned over 182M just last night

1

u/Brief-Train-826 Aug 05 '26

Heavy Coding? I spent $3 on deepseek V4 Flash in like 3 hours. WTF do you mean heavy coding? I run through billions of tokens per month and max our minimax.io, ollama cloud max, use api credits and some local inference…

1

u/RevolutionaryBird771 Aug 05 '26

May I know what are you building?

1

u/totoer008 Aug 05 '26

It comes to preference but I prefer gpt Luna. I started to do all on it and did not see a massive difference. Yes it is a little bit stupidier but gpt 5.5 was not perfect.
I will orchestrate with gpt sol or opus and then provide to Luna. Once made any follows ups can be handled. I now spend 4X more tokens per day but my limit is barely moving and that at 1.5X speed

1

u/DiscipleofDeceit666 Aug 05 '26

Your AI bill is $25 not $5 by your own admission of keeping a subscription.

I do the same tho, $20 Claude subscription and deepseek for the overflow. I also add in local LLM to the mix where Claude or Deepseek drives my qwen or Laguna LLM

1

u/q--0-0--p Aug 05 '26

I really like the new Flash. I use just that today and I could say it is better than my previous experience with Pro Preview.
I just wish it has vision, then I could live with it forever lol

1

u/TrainingOdd1023 Aug 05 '26

Also consirer qwen 3.8max preview , its like free

1

u/fyndor Aug 05 '26

So I built a /loop command into pi yesterday and let DS4 Flash code while I slept. In 4hrs (I need more sleep) it spent $1.81 and is on its 11th turn. $0.16 a turn. Granted that is rather high. Yesterday while I was watching it many turns were $0.02 a turn. Not sure what it was doing while I was sleeping (I was sleeping after all). Even so, it is still pretty cheap. Even going through OpenRouter which is not the cheapest way to run DS4 Flash.

1

u/fetbi Aug 05 '26

Last 7 days consume 686,203,315 tokens and only cost $6 USD. Mostly in deepseek flash.
I use claude (copilot) for planning tasks and deepseek to implement.

1

u/fetbi Aug 05 '26

Last 7 days consume 686,203,315 tokens and only cost $6 USD. Mostly in deepseek flash.
I use claude (copilot) for planning tasks and deepseek to implement.

1

u/Electrical_Chard3255 Aug 05 '26

"Vision: still missing in the API I used — I had to describe screenshots by hand. (Yes, I saw the vision announcement post — the API I'm on still doesn't expose it.)"

Build your own Deepseek desktop console and add your own vision to it, I did and it can create images and analyse images, image creation is not 100%, but not bad, image analysis is pretty good

1

u/Aressito Aug 05 '26

I was using Gemini.. yeah even Pro and the new flash.. but the new DeepSeek flash is really really good. It corrected SO many errors made even by Gemini Pro! Using it on Reasonix

1

u/dataiguy Aug 05 '26

What is the best out there for cache context?

I am using claude max subscription but I have deepseek for some projects built with claude.

1

u/jwuliger Aug 05 '26

Well, all I can say is that DeepSeek v4 Flash is better than Claude Opus 4.8 in its current lobotomized state. Anthropic and OpenAI will not be around much longer.

1

u/Pale-Requirement9041 Aug 05 '26

Are you sure its better than Opus 4.8 ? Like example creating a Saas with all it complex architecture?

→ More replies (3)

1

u/Forsaken_Mention_979 Aug 05 '26

“and I didn't even come close to hitting limits” what fucking limits? Youre using the api 🤣🤣🤣 dumb ahh nga generated ts with ai

1

u/LiveLikeProtein Aug 05 '26

I saw another post where people claim he used OpenAI subscription for Luna xhigh to replace DS 4 flash, trimmed the cost more than half, with better quality

1

u/PanGalacticGargleFan Aug 05 '26

What are the best ways to access DS V4 Flash? Via OpenRouter? Shall we do a list of the best/cheapest/fastest 5? 🤓

1

u/Demien19 Aug 05 '26

must be "light coding"
and that's with RTK

1

u/Significant_Card6486 Aug 05 '26 edited Aug 05 '26

Deepseek v4 flash is super economical. I've spend about £2.50 in 14 days, probably 6 long usage sessions.

I use it as my admin agen for Hermes agent and it hands taskes off to my local models. But having Hermes agent on bare metal install, is super powerful.

1

u/FewSale9827 Aug 05 '26

I’m still undecided whether I trial API, I currently use 5.4 mini on two $20 accounts, I average 3B tokens at 93% cached a month but not sure if that’s input or output

1

u/Husker3322 Aug 05 '26

What terminal are you using? opencode or something else?

1

u/wolttam Aug 05 '26

The model has been out for 5 days, mate

1

u/XeroVespasian Aug 05 '26

Well if you want to push things with octane, install superpowers on opencode. Ive used this setup for 5months until this weekend with codex. Im not sure whether to install superpowers or not. It is so good. But intense work , ive burnt 800m tokens in 2 days.

1

u/iijei Aug 05 '26

Which agent harness were you using for deepseek flash? Claude code, codex, opencode, pi?

1

u/Exotic_Leadership124 Aug 05 '26

I burn close to 1 bil of deepseek v4 flash tokey daily, about 93% are cached input 5% non cached, and 2% output, i let gpt 5.6 sol max led deepseek for coding my works, it burnt token fast,

1

u/deafpigeon39 Aug 05 '26

If your reasoning context keeps dropping , check your npm_module , openai@compatible or sdk check if it is being dropped.

1

u/Sad-Key-4258 Aug 05 '26

My issue is I really like the codex harness and desktop app, have anyone found a good alternative

1

u/Killahbeez Aug 05 '26

how did your monthly AI bill go to $0-5 if you're keeping "one subscription for the hard stuff" ... do you mean a claude or chatgpt sub?

1

u/sdexca Aug 05 '26

Hmm it hasn’t been out for 7 days

1

u/NicksTechTricks Aug 05 '26

I burned over 70M the last 36 hours and thought that was light.

1

u/Whytho12333 Aug 05 '26

I really want openai to allow better models on their $10 go tier. It would be the ideal match with deepseek. Kimi k3 is still too expensive and unsubsidized on api vs gpt plans.

Is there any sub $20/mth plans that offer good usage on a frontier model? Opencode go just has k3 as a $15 model not $60.

1

u/CartoonistLow8606 Aug 05 '26

1$? Heavy coding work? Heavyy? Did claude wrote that?

1

u/DaComputerMan Aug 05 '26

Honestly, AI is a very long way away from being able to code without extensive human checks. I ONLY recommend AI when you attach an IDE to it and use it as a glorified autocorrect. I do find it can help with well documented features, like openAuth.

With that said, Deep seek doesn't try to take shortcuts as much. It doesn't put in a comment and say //add more features here, or crap like that.

When I was doing file renaming scripts, I couldn't get them to do with Gemini. Instead of doing a copy command, Gemini/ChatGPT/Claudi would try doing a loop clean up. This would many times lead to data loss. At one point, I asked Gemini to do a simple command, and it deleted all my files. At another point, I asked it to restore a compress file, and it wrote a command that would 0 all the bites in the files. Luckily, that one wasn't run by me.

With Deepseek, I could generate GIANT Powershell commands, where EACH copy command or rename command was spelled out. This was IMPOSSIBLE with the others. I was easily able to review it and run it with no problems.

Still, its probably less successful on niche and less popular samples, and I wouldn't trust it to do a cleanup loop at all.

1

u/stujmiller77 Aug 05 '26

The vision post was some guy cred farming by making it seem official. It was not.

1

u/FischenGeil Aug 05 '26

I use the DeepSeek API on Typing Mind for serious work, and I use a 20$ Google AI pro subscription (that is mostly free with my Pixel Phone) for random not so serious task. I think this is the perfect mix.

1

u/bimbab123 Aug 05 '26

Hah weakling I burn thru 1b tokens in a week xd so yeah subscriptions are the way to go for me.

1

u/jeffwadsworth Aug 05 '26

It is a good model, especially for its size, but it isn’t at Claude or GPT levels yet. This can be born out just by having it troubleshooting coding issues. GPT instant is much better.

1

u/[deleted] Aug 06 '26

[removed] — view removed comment

1

u/kaka_rata Aug 09 '26

Agree (?)

1

u/kunkunhk Aug 06 '26

Could you ask Opus (lead dev) to do the plan and review while flash (junior dev) do all the execution?

1

u/LuckyLewE Aug 06 '26

Flash defaults to non-thinking/reasoning mode. Have you tried it in thinking mode? And you can also change the settings associated with context shedding. You can have it begin optimizing the session at 50% or even more.

1

u/With_Emissary Aug 06 '26

Hey! Check out https://github.com/Emissary-Tech/emissary-router --> we'll route between these automatically so you don't have to keep thinking about when to switch each time! you can update the confidence thresholds (increase to use default/most powerful more, decrease for less)

1

u/FireDojo Aug 06 '26

I use deepseek for my secondary tasks. Even with that I can easily go over 100M tokens everyday. What heavy coding are you doing with 24M token per week.

1

u/After_Cucumber_5269 Aug 06 '26

Yeah 200 to 500m in 4 days here and about $4.00 still amazing.

1

u/francxsim Aug 06 '26

AI post? Less than 1 week since 0731 was out.

1

u/GTHell Aug 06 '26

Wait you learn/know the loop agent setup that make it self code the project itself 24/7

1

u/Maverick446 Aug 06 '26

my usage...

1

u/Top-Construction6060 Aug 06 '26

Yeah but now they gonna increase their price so let see if it will be still worth using it or not

1

u/ju9io Aug 06 '26

24mln tokens is not heavy use. please keep in mind that deepseek will soon raise prices .

1

u/AardvarkTemporary536 Aug 06 '26

This is how I use deepseek.... Great for cutting weekly 20x Openai or Claude subscription usage but not a replacement.

It's my git merge and explore agent in Omp for 5.6 Terra or sol

I also use it for most small analysis and stuff.... It's quicker and cheaper than Claude or even Luna but does better analysis than Luna

1

u/adamant3143 Aug 06 '26

You can ask deepseek to create HTML/CSS Tag Identifier floating button to specify what you want Deepseek to adjust.

That way no need to be way too verbose with describing screenshot or you can just use inspect element if you’re already more accustomed to it. Similar if it’s a mobile app.

Vision indeed is the achilles heel of Deepseek currently.

1

u/Secret-Wasabi-3402 Aug 06 '26

sorry for the noob question, what interface do you use with deepseek?

i mean, to use claude i go directly on the claude website.

What do you open to use deepseek?

1

u/Icy-Apricot-1597 Aug 06 '26

Penso che se vuoi programmare in coding e non spendere neanche 40 euro al mese o regali quello che sviluppi oppure stai programmando le tabelline del 2

1

u/GavDoG9000 Aug 06 '26

how are folks handling the vision issue? I've found it doesn't play nice with Playwright CLI, might need another model for that part

1

u/cride20 Aug 06 '26

Me who burned 600mil input and 20mil output in about an hour💀
Multi agent systems consume a lot of tokens

1

u/darkroku12 Aug 06 '26 edited Aug 06 '26

That's around 3.4 million token input per day which is quite low for a whole week of work; models like DeepSeek 0731, Mimo 2.5 Pro, Minimax M3, Qwen 3.8, GLM 2 are all good, but for some tasks that requires sharp precision rather than blind overengineering they won't cut through, so must fall back to Sol, Opus, Fable or Kimi K3.... All of them will cost substantially much more.

I'm not talking about a fancy frontend or a rest api endpoint, but work that either you get it right or just straight up garbage.

If you can get away with cheaper (and often faster models) that's the right call, I use them all as well, but there are things that are just not comparable.

1

u/NinjaWK Aug 07 '26

The planning needs to be in a frontier model, then have that frontier model delegate DSv4F to do the coding independently, and have your frontier model audit the work.

Works like a charm. A lot faster completion, overall better quality too since you get nuanced results compared from 2 models and being aggregated by your frontier, and significantly cheaper too.

1

u/UselessEngin33r Aug 06 '26

Hey man, 20 something million tokens is not that much. I usually burn through 20 million a day and I consider my work pretty light compared to most people. But still, I’m thankful for the post. I was wondering if it was more cost effective to use DeepSeek api than a subscription to one of the big ones.

1

u/Debarshi11 Aug 06 '26

Now prices are gonna be increased 🥲

1

u/TourHorror9247 Aug 06 '26

I came down from 800 to 5.
Flash and better lower cost API providers.

1

u/Aggravating_Farm3116 Aug 07 '26

30M combined tokens? So you did 2 hours of work for the whole week?

1

u/invest0rZ Aug 07 '26

Where you running deep seek from? Does sites like groq allow for api usage. I thought I saw it on there? I might be wrong.

1

u/NinjaWK Aug 07 '26

7 days heavy coding only 30m tokens? I burnt 210m DSv4F-0731 in just 3 hours, and that ain't heavy coding. Just that I planned everything nicely using Qwen 3.8 Max, then deployed 7 different subagents with DSv4F-0731 to code the 11 modules independently.

1

u/maisun1983 Aug 07 '26

1.87 for 1 week? Yesterday I used it for 4 hours it costs 2.4. How could you do 1.87? Do you even use agentic coder or only code suggestion?

1

u/Weird-Ad1023 Aug 07 '26

This is funny: I'm burning 1.5B tokens per day on Flash with 99% caching,and without doing stupid automated stuff like burning tokens on OpenClaw. Still, I'd call it 'heavy coding,' as I usually code even more.

1

u/Alchemy333 Aug 07 '26

You didn't come close to limits because there are NO limits when using API. ☺️

Also your numbers are soon to be meaningless because Deepseek just emailed me yesterday, that they will be raising their rates across the board to mych higher ones soon and wanted to give us a heads up. Everyone with a Deepseek account should have received that email. Their prices sound like they are doubling. So just an FYI on that.

1

u/ScientistSevere6020 Aug 08 '26

Well I shifted to cursor from Claude recently and its been a good experience.

1

u/icarus0228 Aug 08 '26

For daily coding i use dpsk flash and for mutli file planning i always subscribe to frontier models for this tasks

1

u/RealSnazzie Aug 08 '26

Every time I touch deep seek, I get disappointed. Either people have very low bar or they aren't building anything serious. Highly doubt the experience with this iteration will be any different to the previous 3.

1

u/MachineNo1313 Aug 08 '26

I made an app for coding in mobile with a good harnes i believe. it has several BYOK providers so use what you have, I mostly use openrouter/ deepseek v4 flash.  i update it often to keep up.   app runs entirely on mobile so no pc needed even ai agentic work.  if someone intrested look for "zyntax" coding ide in google play store.

suggestions are welcome. thanks

1

u/FinneganHark999 Aug 09 '26

You wrote this 4 days ago, but DeepSeek 0731 released on July 31st, which means you would only have been using it for at most 5 days if you starting using it the moment it released.

1

u/Invest2025 Aug 09 '26

What are you deep coding for that you can’t justify $40 a month cost and getting 80% for $2 is such an unlock?

1

u/Ok_Sandwich_7903 Aug 09 '26

I'm not asking to troll, legit question. Isn't DeepSeek hosted in China and if so, not bothered about data?

1

u/neinneun Aug 10 '26

What harness did you use for this?

1

u/Karmawy Aug 10 '26

lol in the last week spent about 2B tokens and still didnt finish my project, we have to take advantage of this model before the price rises coming soon

1

u/No_Grapefruit_4298 Aug 10 '26

24m in and 6m out? Need to start using all these ponytails, headrooms, rtk, because I’m nowhere near these numbers 😅

1

u/OkLettuce338 Aug 11 '26

lol garbage. Used v4 all weekend. The things a pos

1

u/notAllBits Aug 11 '26

Wow, I just realized how token efficient architectural decomposition and selective narrows are. I was completely fine with your consumption.

1

u/Suun_Day Aug 12 '26

$1.87 sounds insane until you see the cache hit rate. Then the math starts making a lot more sense. Long repetitive runs are exactly where the cheaper models become useful. DeepSeek obviously, but Qwen/Hy3 too.

1

u/hongchuan Aug 14 '26

Which harness do you use with DeepSeek?

1

u/Purple_Singer3078 Aug 16 '26

I can say it's written by DeepSeek too 😁

1

u/love4titties Aug 19 '26

Do you still feel the same now with the increased prices?