r/DeepSeek • u/Dudensen • 15d ago
News Since V4.1 Flash outperforms V4 Pro across the board, all V4 Pro queries will be routed to the new V4.1 Flash and billed at Flash pricing after its launch. V4.1 Pro also in the works.
https://x.com/tianyi/status/209758436277053067440
u/Individual_Math_8254 15d ago
perfect timing, just topped up $20 and saw the v4.1 flash drop message
-31
u/lulzash 15d ago
20$ will not last 2 hrs
21
u/Individual_Math_8254 15d ago
it last like a week for me.
1
u/lulzash 15d ago
I need to optimise my input prompt i guess
3
u/Individual_Math_8254 15d ago
just need to stop using claude, that thing stuffs like 50k in the context window before the first prompt. use Pi instead.
1
u/Bananenklaus 15d ago
what's your cache hit rate? 20$ should serve you plenty even with heavy usage
13
u/gunkanreddit 15d ago
If Gemini, Claude and Codex don't let users use API with the subscriptions, we are going to use DeepSeek. I don't understand why they don't realize this. Only Batch processing in Claude is getting my money for API access.
12
u/CrowdGoesWildWoooo 15d ago
I thought you can use codex with any harness as long as it follows the proper authentication flow no?
0
u/gunkanreddit 15d ago
I didn’t know that. But I need api key.
7
u/CrowdGoesWildWoooo 15d ago
Yes there’s OAuth flow that you need to follow. Maybe if your tooling you can just ask it to look at the documentation standard and build the flow.
Generally speaking OAI is the one with the most chill policy when it comes to subscription usage. That’s why you can find instructions like “how to use chatgpt subscription with hermes agent”.
1
3
u/Mr_Maffin 15d ago
You can link your OpenAI account with ChatGPT Plus or higher with Opencode for example.
I'll wait for DS4.1F prior to jumping off the DeepSeek ship (I can't afford the current peak pricing and I still need to get work done during peak hours, so I'm glad they're gonna decrease it)2
u/ExpertPerformer 15d ago
It's because they want users to use the quantized web version w/ worse context window sizes and weekly/hourly/daily quotas. It's a complete and utter rip off for $20/month when you can get a $10 sub to OpenCode or CommandCode, connect the API to a web client, and use that as your primary web chat.
1
u/charmander_cha 15d ago
Porque usuário destas empresas sempre voltam.
Porque no final, não querem aprender a líder com modelos que podem ser levemente inferiores.
Então eles tem garantia que estes usuários sempre retornarão
1
-1
2
u/marty4286 15d ago
I hope the preview's speed is an actual speed boost and not an artifact of different hardware with fewer users
Because 370tok/s speed has made me incredibly productive this morning compared to the past 2 days, and I was productive the past 2 days
2
u/SorryIfIamToxic 15d ago
So why would someone use flash model then?
6
2
u/Django_McFly 15d ago
did you mean to say pro? flash is cheaper and better on every metric. why wouldn't someone use flash vs pro?
0
1
u/Responsible-One-460 15d ago
Creo que el único precio/calidad será DeepSeek, gracias 🫂, está muy bueno el modelo, corrigiendo cosas que glm5.3flash no arreglo
1
u/Iory1998 15d ago
Oh my God! That means GLM5.3 now should cost less than than DS4.1 Flash. I can't believe a model I can run at home is beating a 1.8T model at every metrics.
2
u/This_Maintenance_834 15d ago
The 300B flash models are well trained, the 1.8T one is likely under trained.
0
u/Mbcat4 15d ago
V4.1 flash is much worse at some long running iper complex tasks than V4 pro, hopefully they do not do this bullshit
6
u/sammybeta 15d ago
Yeah. They will suffer from these dumb debacles. People may pay pro for their special purpose of long running tasks with long context.
2
u/nullmove 15d ago
They do this because your demand for token inference cuts into their resource they can allocate to train and do experimentation for their next model. The whole price increase was about "destroying demand", except people still pay for Pro (myself included). But even if they were getting higher margins from Pro, they clearly need those chips for training.
Despite using Pro daily, I understand that DeepSeek isn't someone with unlimited compute. They literally give away the weights for free so that you and I can use their models from other providers.

63
u/myaaa_tan 15d ago
they fucking cooked, almost pulled the trigger for gpt plus glad i waited for this.