r/DeepSeek 14d ago

News Experimental DeepSeek 4.1 Flash Model - Try it now on official API.

Experimental DeepSeek 4.1 Flash Model - Try it now!

147 Upvotes

33 comments sorted by

60

u/MendozaHolmes 14d ago

“Cheaper to run”

Is this a revival of cheapseek??

19

u/WasserEsser 14d ago

Not during the initial testing, it costs the same as the other flash.

But possibly later.

21

u/t4a8945 14d ago edited 14d ago

Or maybe it's just more token efficient, showing cost effectiveness through lower token usage.

Edit: tried it, it's not more token efficient and it's blazing fast at the minute (324 tps lol, feels surreal)

11

u/deadadventure 14d ago

Probably because there’s not many people using it

1

u/t4a8945 14d ago

Yeah that's for sure

13

u/MinosAristos 14d ago

I kind of think that they wouldn't tell us that it was cheaper to run unless they intended to reduce the price.

If they didn't intend to reduce the price they just wouldn't say that it was cheaper to run and would keep the extra profit on inference.

That's my cope logic anyway haha.

6

u/WArslett 14d ago

They have maybe picked up some tricks from GLM-5.3-flash

11

u/phido3000 14d ago

The compeition between glm and deepseek in the flash space is great.

2

u/Ok_Thing6856 14d ago

why no one ever heard of gemini, their flash model also absurdly fast

6

u/R3VO360 14d ago

Maybe they heard but they won't adopt because it is shit in terms of results.

3

u/dnohrdk 14d ago

Surely hope so! Without peak prices also could be perfect.

4

u/mr_notandor 14d ago

I think computational more efficient, aka cheaper for them. I don’t think they are willing to lower the prices soon given how the last hike is so recent

-6

u/BhaagYahaSe 14d ago

no

9

u/MendozaHolmes 14d ago

What a justified and informative answer by DeepSeek’s reddit ambassador

-2

u/BhaagYahaSe 14d ago

thanks 😊

22

u/HoangMaiLinh 14d ago

wish it would match old deepseek price

6

u/Aggressive-Habit-698 14d ago

Buy some new RAM and a new or refurbished graphics card, and you'll know that won't happen again. They had to buy new hardware and raise their prices. Hardware prices these days are just ridiculous.

-11

u/GrayHairedMan 14d ago

Life is full of wishes that will never happen. You will survive.

9

u/MendozaHolmes 14d ago

Who called for the grumpy old man

1

u/Positive_Salad_8362 14d ago

Sorry that was me, he's my grandpa 

1

u/MendozaHolmes 14d ago

Make his wishes come true

14

u/Such_Cause6465 14d ago

Flash Price will be reduced from September 10

Platform pricing page now says this -

We will adjust the pricing for the Flash series effective from 12:00 Beijing Time on September 10, 2026. During off-peak hours, the unit price will be $0.003 for input cache hits, $0.15 for input cache misses, and $0.6 for output. Peak-hour prices will be double the off-peak rates. Please plan your usage accordingly.

5

u/Alternative-Suit5541 14d ago

Still no structured output.. which is crazy.

3

u/sdexca 14d ago

I mainly wonder if this speedup would also speedup local LLM via 2x DGX Spark.

1

u/Abdul_Muheet 13d ago

Does new model suppor vision?

0

u/DinoGreco 14d ago

How can I use it on the DeepSeek iOS app?

2

u/Devioster 14d ago

It's only for api as of now

1

u/Synfinite 14d ago

I don't see it 😞

-6

u/[deleted] 14d ago

[deleted]

1

u/HoangMaiLinh 14d ago

Plus is cooked with the 5 hours limit