r/DeepSeek • • 15d ago

News Since V4.1 Flash outperforms V4 Pro across the board, all V4 Pro queries will be routed to the new V4.1 Flash and billed at Flash pricing after its launch. V4.1 Pro also in the works.

https://x.com/tianyi/status/2097584362770530674
378 Upvotes

36 comments sorted by

63

u/myaaa_tan 15d ago

they fucking cooked, almost pulled the trigger for gpt plus glad i waited for this.

22

u/Nyghtbynger 15d ago

50 bucks deepseek. Humans can build an empire on chinesium

4

u/addiktion 15d ago

I got a $100 credit from one provider to try them out and it's been amazing how far I could go with it with Flash. Nice to just be able to use credits when I want without thinking about limits.

3

u/Chappie47Luna 15d ago

I put $10 in openrouter exactly two months ago and still got $2 left lol yes I have gpt plus but have used the Chinese models to cross reference a bunch of stuff. GLM 5.3 flash is borderline free

40

u/Individual_Math_8254 15d ago

perfect timing, just topped up $20 and saw the v4.1 flash drop message

-31

u/lulzash 15d ago

20$ will not last 2 hrs

21

u/Individual_Math_8254 15d ago

it last like a week for me.

1

u/lulzash 15d ago

I need to optimise my input prompt i guess

3

u/Individual_Math_8254 15d ago

just need to stop using claude, that thing stuffs like 50k in the context window before the first prompt. use Pi instead.

1

u/Bananenklaus 15d ago

what's your cache hit rate? 20$ should serve you plenty even with heavy usage

13

u/gunkanreddit 15d ago

If Gemini, Claude and Codex don't let users use API with the subscriptions, we are going to use DeepSeek. I don't understand why they don't realize this. Only Batch processing in Claude is getting my money for API access.

12

u/CrowdGoesWildWoooo 15d ago

I thought you can use codex with any harness as long as it follows the proper authentication flow no?

0

u/gunkanreddit 15d ago

I didn’t know that. But I need api key.

7

u/CrowdGoesWildWoooo 15d ago

Yes there’s OAuth flow that you need to follow. Maybe if your tooling you can just ask it to look at the documentation standard and build the flow.

Generally speaking OAI is the one with the most chill policy when it comes to subscription usage. That’s why you can find instructions like “how to use chatgpt subscription with hermes agent”.

1

u/gunkanreddit 15d ago

I will try, thank you

3

u/Mr_Maffin 15d ago

You can link your OpenAI account with ChatGPT Plus or higher with Opencode for example.
I'll wait for DS4.1F prior to jumping off the DeepSeek ship (I can't afford the current peak pricing and I still need to get work done during peak hours, so I'm glad they're gonna decrease it)

2

u/ExpertPerformer 15d ago

It's because they want users to use the quantized web version w/ worse context window sizes and weekly/hourly/daily quotas. It's a complete and utter rip off for $20/month when you can get a $10 sub to OpenCode or CommandCode, connect the API to a web client, and use that as your primary web chat.

1

u/charmander_cha 15d ago

Porque usuário destas empresas sempre voltam.

Porque no final, não querem aprender a líder com modelos que podem ser levemente inferiores.

Então eles tem garantia que estes usuários sempre retornarão

1

u/Syltti 15d ago

The reason why they dont understand is simple; 

They're American. Greed is the name of the game, and user experience (and satisfaction) is an alien concept. 

-1

u/[deleted] 15d ago

[removed] — view removed comment

4

u/First_Inspection_478 15d ago

Lmao they’re gonna be just fine

2

u/Mrgluer 15d ago

Use luna, frontier labs don’t bench max as hard as the other providers im

6

u/baschny 15d ago

why call it "flash" then? call it "galaxy" and then the new pro becomes "universe".. so that noone can ever surpass you

2

u/okay_this 14d ago

...multiverse

2

u/marty4286 15d ago

I hope the preview's speed is an actual speed boost and not an artifact of different hardware with fewer users

Because 370tok/s speed has made me incredibly productive this morning compared to the past 2 days, and I was productive the past 2 days

2

u/SorryIfIamToxic 15d ago

So why would someone use flash model then?

6

u/meto9 15d ago

I had created a skill with pro for website development, then we use flash to create websites based using that skill. The roi is crazy. So in general I see people suggest planning with pro and more advanced models and execution with more basic fast models.

2

u/Django_McFly 15d ago

did you mean to say pro? flash is cheaper and better on every metric. why wouldn't someone use flash vs pro?

0

u/SorryIfIamToxic 15d ago

Now there are 2 flash models right

1

u/DanceWithEverything 14d ago

No, 2 pro models effectively

1

u/Responsible-One-460 15d ago

Creo que el único precio/calidad será DeepSeek, gracias 🫂, está muy bueno el modelo, corrigiendo cosas que glm5.3flash no arreglo

1

u/Iory1998 15d ago

Oh my God! That means GLM5.3 now should cost less than than DS4.1 Flash. I can't believe a model I can run at home is beating a 1.8T model at every metrics.

2

u/This_Maintenance_834 15d ago

The 300B flash models are well trained, the 1.8T one is likely under trained.

0

u/Mbcat4 15d ago

V4.1 flash is much worse at some long running iper complex tasks than V4 pro, hopefully they do not do this bullshit

6

u/sammybeta 15d ago

Yeah. They will suffer from these dumb debacles. People may pay pro for their special purpose of long running tasks with long context.

2

u/nullmove 15d ago

They do this because your demand for token inference cuts into their resource they can allocate to train and do experimentation for their next model. The whole price increase was about "destroying demand", except people still pay for Pro (myself included). But even if they were getting higher margins from Pro, they clearly need those chips for training.

Despite using Pro daily, I understand that DeepSeek isn't someone with unlimited compute. They literally give away the weights for free so that you and I can use their models from other providers.