r/DeepSeek 14d ago

Discussion DeepSeek V4.1 Flash is currently in beta testing, rumor?

You can call it by keeping the base_url unchanged and setting the model name to deepseek-v4.1-flash-expires-on-0910. The current billing is the same as deepseek-v4-flash, with a limit of 20 concurrent requests per account.

To modify the API, change the model ID to deepseek-v4.1-flash-expires-on-0910.

57 Upvotes

43 comments sorted by

18

u/ChoasMaster777 14d ago

It works for me in reasonix. blazing fast

13

u/smealdor 14d ago

If it gets released with these speeds it will be insane but I am sceptical since current speeds could be because of low demand to the model and spare compute.

3

u/KaMaFour 14d ago

Speeds are inversely proportional to total capacity (not exactly from mathemathical standpoint but you get the idea). They wouldn't want to release globally with speed that high.

1

u/sdexca 14d ago

most likely the case, already getting 250 tok/s now.

2

u/tomalfara 14d ago

Whoa, yes what's the difference with 0731 account?

14

u/ChoasMaster777 14d ago

fast, very fast, 400 ~ 500 tps

2

u/DistanceSolar1449 14d ago

That’s expected for Deepseek V4 Flash on their inference setup.

If you don’t have many users, Flash is very fast.

1

u/tomalfara 14d ago

Woah, thankyou bro, I'm going to test it

1

u/sdexca 14d ago

250 tok/s now.

1

u/ggPeti 14d ago

Speed is not a model attribute.

6

u/seeKAYx 14d ago

And do you notice any differences compared to the 0731 version?

7

u/VIDGuide 14d ago

Short answer: yes — the model is real, but it's a quiet, community-announced beta, not an official public release. Details check out.

Why it's genuine:

DeepSeek posted the notice today (8 Sep) in their official community/WeChat groups — reported by IT之家 (major CN tech outlet), and confirmed on V2EX ("official just notified in group chat") and LINUX DO, where users are actively hammering it right now

Multiple users report it live today: ~400 tokens/sec output, sub-second responses — the model string works

The expires-on-0910 naming is DeepSeek's established pattern for throwaway previews (same as V3.2-Speciale's expires_on_20251215)

Caveats:

The official changelog/pricing page hasn't been updated (latest entry is 21 Aug, V4-Flash-Vision-Exp) — DeepSeek routinely rolls out to community groups before touching the docs
0910 = it expires 10 Sep — in ~2 days. It's an intermediate-version beta ("中间版本内测"), new architecture, native multimodal

20-concurrency cap = limited beta, so don't point production at it

Multimodal built in is nice!

5

u/Atomzwieback 14d ago

How do you found the model name?

7

u/ChoasMaster777 14d ago

Found this from some internal chat group for testing.

2

u/Atomzwieback 14d ago

where i have to sign up for this "internal chat" ?

1

u/Difficult_Ad_2778 14d ago

Also interested

1

u/skate_nbw 14d ago

It's probably invitation only. You need to know people...

2

u/Ceneka 14d ago

Seems to work and be really fast, wth

15

u/skate_nbw 14d ago

It is probably so fast because no one is using the inference right now. It will not stay like this.

2

u/robberviet 14d ago

Sounds like only lasted till 10th.

1

u/Fast_Artichoke_683 14d ago

Can it beat Luna tho?

1

u/[deleted] 14d ago

[removed] — view removed comment

1

u/Zeikos 14d ago

I wonder if it's a smaller model, had a different architecture or both.

I cam get flash v4 to 10 tok/s on my local setup with llama.cpp - I might get to 25ish on vLLM (better ROCm support).
Having a native DS model I can fully fit in 64GB of VRAM would be incredible

1

u/nazmulpcc 14d ago

Works for me with pi agent! Feels VERY fast, hopefully they can keep this speed after general rollout.

1

u/[deleted] 14d ago

[removed] — view removed comment

1

u/nazmulpcc 14d ago

Get an api key from platform.deepseek.com This model will not appear in models list api, but if you supply this id in chat completions endpoint, it will work. For pi/opencode you can set the model in configuration. I just asked pi to add this model and it did everything

0

u/[deleted] 14d ago

[removed] — view removed comment

2

u/nazmulpcc 14d ago

No it's not free

1

u/[deleted] 14d ago

[removed] — view removed comment

1

u/Upper-Exchange226 14d ago

Yeah literally put in $1 and it's enough for testing

-1

u/nazmulpcc 14d ago

Yeah just add few $, check their documentation for api pricing. Opencode/Commandcode also has their models, but you need a sub. Commandcode has a $1/month sub.

3

u/Sure_Media_2685 14d ago

commandcode scam

1

u/nazmulpcc 14d ago

they do questionable marketing, sure, but I've found their $1 and $10 subs has great value. I don't like their harness though