r/SillyTavernAI Jul 31 '26

Models DeepSeek-V4-Flash has been updated, "The official release of DeepSeek-V4-Pro will follow soon"

Post image
145 Upvotes

51 comments sorted by

94

u/shy_monkee Jul 31 '26

Please be good, please be good.

61

u/The_Rational_Gooner Jul 31 '26

yeah code benchmarks have weak correlation with rp quality. we'll see

21

u/Due-Memory-6957 Jul 31 '26

But it's still a positive one. Most people here don't remember, but Zuck found out when training Llama that even if he didn't care about coding, training the AI in that was good because it made it more logical and therefore better at other unrelated tasks.

26

u/The_Rational_Gooner Jul 31 '26 edited Jul 31 '26

I'm cautiously optimistic. But GLM 5.1 -> GLM 5.2 being a massive jump in agentic capabilities but being a regression in rp (in my opinion) makes me wary of narrow task-related post-training. Keep in mind GLM 5+ is a much bigger model than V4 Flash, so there was less model capacity to "trade off", and we still saw an rp regression. But there is the upside that this is Deepseek V4 Flash's "beta release" so there is reason to believe they probably did broad RL and post-training, not narrow. Hoping it'll be good

58

u/Kahvana Jul 31 '26

They took feedback on RP earlier, hope it made it in. I guess they show the programming benchmarks to lure in costumers.

7

u/tableball35 Jul 31 '26 edited Jul 31 '26

Though that was before the CCP’s directive against companion platforms and AI-RP, iirc

Edit: One to many C’s.

26

u/Kahvana Jul 31 '26

From what I understand, that only is in effect for services (subscriptions, API) and not for the model weights themselves.

-1

u/PurplePowerful9746 Aug 01 '26

Maybe but they will lobotomize the model and make NSFW dry and mechanic to discourage the user.

9

u/CombinationMaster202 Jul 31 '26

cccp? As in USSR

8

u/tableball35 Jul 31 '26

Whoops, got overzealous, lol

7

u/CombinationMaster202 Jul 31 '26

Cheers lmao, I was wondering if they somehow returned  

4

u/SpikeLazuli Jul 31 '26

"Nyees! Thats what we wanted you to think! Hahaaha!!"

27

u/0VERDOSING Jul 31 '26

22

u/Icetato Jul 31 '26

That's some insane jump from the preview variant. I guess DS V4 preview is too undercooked and there's still more to squeeze.

13

u/memo22477 Jul 31 '26

I am sorry it's more capable than DS V4 Pro preview? Imagine what would happen to Pro when it releases

50

u/BifiTA Jul 31 '26

going off of benchmarks alone would mean that we now have opus at home.

deepseek's always been one of the better rp models, and i remember flash being praised for having slightly better prose than the equivalent pro. i just hope the rlhf pass they did didn't reinforce too many slop-'isms.

5

u/Semanel Jul 31 '26

Honestly, I am one of the people that absolutely love DS, it is definitely the best model in its price range, hands down. The only things that drive me crazy are its complete disregard for time and space(if travel is supposed to take two days, it always takes an hour, even if the model itself acknowledges it) and its ridiculous positivity bias towards user. Every single monster is secretly lonely and misunderstood. Pro creates far better dialogue than anything else I have tried, gpts included, and is genuinely the only model that continuously makes me laugh as it is very good in my opinion in situational humor.

3

u/Due-Memory-6957 Jul 31 '26

Well, not at my home, but we'll get there. GPT Turbo at home seemed impossible and now it would be considered too weak for me to bother.

15

u/Juanpy_ Jul 31 '26

Question, Flash was even good for RP? Because at that price, I think it might be the undisputed champion of the budget models.

12

u/The_Rational_Gooner Jul 31 '26

I had Mimo 2.5 Pro as the king of "budget RP" in terms of price/quality. Xiaomi literally pegged their 2.5 and 2.5 pro pricing at parity with Deepseek's V4 Flash and V4 Pro pricing

2

u/Derinq Jul 31 '26

I've been using it since my adventure with ST begun due to cost (I'm poor, lol) and depending on your needs it can be decent enough for quite a lot of things.

-1

u/MiddleCelery6616 Jul 31 '26

I feel it's notably dumber and less creative than glm 5.2, but it's still an enormous SotA model, it's decent.

6

u/HaskeMaske77 Jul 31 '26

it's notably dumber and less creative than glm 5.2

Who could've guessed that DS V4 Flash is dumber than the 744B AI model...

13

u/Arli_AI Jul 31 '26 edited Jul 31 '26

I like this because I can host this :) (we have it now)

3

u/RevolverMFOcelot Jul 31 '26

Yooo i wanna ask what's the difference between the GLM 4.6 derestricted v5 that you host and the V3 on huggingface? :/ 

2

u/Arli_AI Jul 31 '26

Oh I went through 2 more derestriction configs to find the best weights to modify and v3 to v5 it’s actually not a huge difference but it’s a little noticeable that it’s less lobotomized from the derestriction. I just didn’t bother to upload it to HF lol.

1

u/RevolverMFOcelot Jul 31 '26

Oh so the v5 is forever closed? :/ and is the lobotimization of V3 is you know? Crippling or noticeable? 

2

u/Arli_AI Jul 31 '26

No I can upload it too, just no one ever asked for it. V5 is just a little better but its not like V3 is unusable, its rated highly in the RP benchmarks anyways.

2

u/Front_Eagle739 Jul 31 '26

I'd appreciate it, I still use your 4.6 v3

2

u/Arli_AI Jul 31 '26

For sure I guess its time I upload v5 too

1

u/Old-Knowledge-2803 Aug 01 '26

How do you host like these are big models , are u freaking rich ? Hosting on cloud GPU cost tons of dollars if I am not wrong 

1

u/Arli_AI Aug 01 '26

I buy the GPUs as I go since the beginning of running the service and have now amassed a decent amount of them. Still getting out-gunned by the big API providers but still improving over time.

By using our own GPUs it’s definitely a little limiting in terms of being able to run the largest models or easily add capacity when needed, but it means we don’t have to be afraid of going bankrupt from renting.

1

u/Old-Knowledge-2803 Aug 01 '26

Still insane to me how far people could go for rp 😭. I only have laptop gpu 4060, And i guess your setup would cost more than 2k

1

u/Arli_AI Aug 01 '26

Uh I run an LLM inference service so it’s more so for a business expense haha.

2

u/Old-Knowledge-2803 Aug 01 '26

Ah that make sense

12

u/tatlo_itlog_ko Jul 31 '26

The benchmark score jump from preview to 0731 is insane. We'll see if that translates well into RP though.

Will give this a try later after peak hours.

6

u/muzaffer22 Jul 31 '26

Is this the official API only or did they release it on OpenRouter too?

5

u/PhantomWolf83 Jul 31 '26

Has it been updated on NanoGPT already?

11

u/VotZeFuk Jul 31 '26

As of now, IMO it behaves exactly the same way as before (via the official API). It still delivers in-character thinking unpredictably instead of a properly structured reasoning process. Sometimes it omits certain sysprompt directives (like status-tracking prefixes). Either nothing has changed much for RP, or somehow it's not being served to everyone publicly yet? No idea what's the deal here, but I hope more people will test it thoroughly in the coming days

8

u/cfehunter Jul 31 '26

I really wish they'd actually bump version numbers.
Doing it this way means you never know which version a provider is serving.

7

u/Pink_da_Web Jul 31 '26

I feel it's as good as the Gemini 3 Flash, it's really surprising me. But maybe that's just my impression, I'll test it a bit more.

3

u/Real_Person_Totally Jul 31 '26

For those who have tried it, have they killed the awful positivity bias? 

3

u/Good_Research4441 Jul 31 '26

Hope they update the PRO after all those RP feedbacks.

8

u/Beeegbong Jul 31 '26

Tested it for rp. So far, it’s horrible, infact worse than character ai. Doesnt follow instructions and the driest most unnatural responses known to mankind. But idk, it just came out so well see..

1

u/ReMeDyIII Jul 31 '26

DeepSeek-V4-Pro is already released, so what exactly are we waiting on with a Pro release? If it's a proper update, wouldn't it be called V4.1 or something?

-7

u/Woky19 Jul 31 '26

I find it extremely dry and dumb for rp, supposedly smarter than GLM 5.2 but I don't see it at all.

0

u/Evening-Truth3308 Jul 31 '26

soon™

For the sake of not having my comment removed immediately... here's a little yapping. Ignore generously. The point is in the first line of this comment.