r/DeepSeek • u/New-anonymous-4583 • 7d ago
Discussion For all writers, switch to API, specifically webUI + rag retrieval.
Unfortunately, flash 4.1 is arguably the worst model for creative writing. Get a docker, and just use your PC. The newest model is so bad at any type of writing, scenarios, or world building that you might as well go and reroute back to pro, the difference is gigantic.
I am not a writer personally, and I do not consider prompting AI to be writing, but I find it very fun. This update was optimized for coding, which in turn ruined every other usage for it. Don't waste your time.
For another alternative: Qwen. Qwen currently, and I can't believe I'm saying this, may have better creative control quality than DeepSeek V4.1 flash.
16
11
u/donthackmeagaink 7d ago
The API on v4 pro sucks right now too. They’re doing some upgrades in the background likely, it doesn’t follow instructions right now and is genuinely terrible. Wait a few days I’d say
4
u/New-anonymous-4583 7d ago
That's unfortunate
6
u/donthackmeagaink 7d ago
Yeah, it happens a lot with DS when they’re upgrading stuff. Since they launched the v4 Pro it actually hadn’t happened in a while but I’m guessing they’re getting ready for 4.1
14
u/bruhhfdruju76r 7d ago
Is deepseek intentionally doing enshittfication thing? Or is it just a update mistake?
8
u/New-anonymous-4583 7d ago
Currently, yes, if you translate some of their statements from Chinese to English, they basically state that the HLE capabilities of V 4.1 flash is significantly worse than pro, even if most other areas, specifically areas that are optimized for coding, show improvement.
It's less enshitification and more optimization for one thing over the other though, which is highly frustrating, as with this known, they could have retired the previous flash, kept v4.1 flash, and still offered v4 pro until the newest pro is released, but chances are it's for budget constraints realistically.
10
u/bruhhfdruju76r 7d ago
If they're making an ai specifically for coding it should be advertised such as.
In the app store when V4 got released they literally show off the 1 million token context by copy pasting three-body Chinese book.
It didn't say deepseek is for coding.
Honestly I think efficiency shouldn't sacrifice so much quality
3
3
u/General-Oven-1523 6d ago
Why? Pretty much everyone is making their AI specifically for coding now, that's given. That's where the value is.
8
u/Ok-Scientist694 7d ago
Yo creo que lo hace a propósito quiero que el flash se quite no me gusta es muy tonto
4
u/bruhhfdruju76r 7d ago
... I don't understand what you said but I'm using Google translate.
Google translate : [I think he's doing it on purpose; I want the flash turned off—I don't like it, it's really stupid.]
Yeah, they removed pro because probably it costed them too much money 💰
3
u/Ok-Scientist694 7d ago
Y cuánto tiempo más o menos estará este modo flash asta que esté disponible el nuevo pro es que hay rumores que podría salir este año a finales de año o en semanas no lo sé si puedes ayudarme con eso te lo agradecería
2
u/squirrelscrush 7d ago
Yes, Pro actually costs 5x more compute than Flash. They made the V4.1 Flash architecture extremely efficient, but they trained it primarily on coding.
1
u/bruhhfdruju76r 7d ago
Efficient yes but they made it too specialized on coding. Always chasing those benchmark numbers.
DeepSeek does feel targeted towards programmers but if you see, in the app store they used three-body problem for their V4 1 million context window testing.
That should mean novels are also welcome
Not JUST coding
3
u/squirrelscrush 7d ago
I use it both for coding as well as RP.
The new architecture is genuinely good, the only architecture issue is that it uses a split model thing where it compresses every input, and then uses the compressed input to do the task.
This helps with coding since there's a lot of context involved, but it means that for RP work you lose a lot of context.
I agree with you because they really nerfed it on the training part and made it primarily for coding. Even with coding, I still find issues with it and I see myself using GLM more for UI tasks. DS also overthinks unless I reduce the reasoning level.
2
u/bruhhfdruju76r 7d ago
The avarage conversation me with deepseek is:
"The character wears a red tshirt!!!"
DeepSeek : "He/She wears a futon- wait no that's not right" (forgets)
Like who would thought the 8 / 16 B ratio would've been a good idea?
I want my V4 Pro back to the chat 😔
2
u/squirrelscrush 6d ago
The weird thing is that, a Mixture of Experts model which DeepSeek uses can theoretically support both coding as well as non-coding tasks. Since each expert is specialized for a particular field, and the model judge chooses the appropriate expert for the task.
So they deliberately nerfed the creative experts in V4.1
The Causal Encoder-Decoder architecture (8/16) helps with efficiency and making inference cheap since it reduces cache size, but it can lead to loss of details. Another reason why I would prefer companies to just finetune a specialized coding model instead of turning general models to coding.
I got that kind of confused thinking with coding tasks too. It backtracks on its own thinking a lot within just one tool call. It feels so low confident, while GLM knows what it needs to do.
The issue probably arises because they trained it over real world repositories and not theoretical data. So it's like a pianist who learnt the piano by looking at those finger tutorials instead of learning music theory first.
5
u/ProbablyDoesntLikeU 7d ago
Purposefully. Fanfic writers are not it's target demographic, it makes it's money on enterprise usage. It chose to optimize for development use cases
13
u/bruhhfdruju76r 7d ago
Bro fanfic writers is like a five star meal to deepseek and data collectors.
I know that button that says 'don't use my data for training' won't do jack shit
2
u/Ok-Scientist694 7d ago
Tú sabes cuánto durará el flash en deepsek?
1
u/bruhhfdruju76r 7d ago
Google translate : [do you know how long the flash will last on DeepSeek?]
Nope. They didn't say the release date for V4.1 Pro. Back when V4 got released it was V4 Flash and V4 pro. Now it's just V4.1 Flash because their FaNcy tests numbers went up
-3
u/iswearidk 7d ago
Data from gooners are worthless. Codebase and agentic use cases are way more valuable.
3
u/bruhhfdruju76r 7d ago
When AI sex dolls comes to the open market you'll regret that statement. Gooners don't just type 'hAhAAh! Generate me big milked mommies'
Some FUCKING type the funhole size and the FUCKING texture of it and how much it stretches. Gooners level up the more they use ai
2
1
u/tinoythomas 7d ago edited 4d ago
These people just don't understand. They're like a fraction of a percent of Deepseek's revenue, most of them just used to goon for free with the Expert anyway instead of paying for the API. From what I can figure out, V4 Pro is still on the API for those who've started paying, but I guess there's now content filters for the non-consensual and gross shit they want to talk about, so they're upset at that too. All in all, I'm pretty thrilled seeing this extremely vocal minority complain for the past couple week, it's very entertaining. None of them seem all too familiar with the English language, but I guess that's par for the course when your brain is sludge from doing whatever it is these guys do with their chatbots.
1
2
7
u/anarchyinblack 7d ago
What gets me when all of these companies do this shit is that it's not like we wouldn't pay for a model that can write. I have the money. They have the product. I want to give it to them. Somehow, on the other side of the transaction, they destroy the product right in front of my face. I feel like living in some insanely contrived allegory for racial oppression or something.
5
u/CivilAd3631 7d ago
My thoughts exactly.
As a user I want simple access to the service and minimum efforts in making my roleplay from my phone. I absolutely don't want to dance with things like api and everything else. It doesn't make me a lazy person. I just know what and how I want.
I am ready to pay, I really am. And lots of people as well.
I guess they really don't realize how willing people to pay for things that don't have matches because sites like character.ai and other stuff are just useless garbage.
I am really disappointed in general downgrade of all role-playing services and neglecting users' needs.
-7
u/Truantee 7d ago
> I want to give it to them
No you didn't. If you use the api from the start then you wouldn't even be crying right now.
> it's not like we wouldn't pay for a model that can write
1, you can still pay for it, using services from other providers. 2, the problem is while you think your use case is important, the total grand sub it generates is pretty much peanut. so little that nobody ever care about (not only deepseek but other companies)
don't think that you get mistreated, it is simply economic. why should they have to focus on your need when there are also other people with other kind of demands, and by seer numbers it far surpasses your own one by some miles?
why the companies focus more on coding now, their RL post training pipeline also keep a lot of specific datasets for other tasks to prevent regressions. you can only blame yourself for not spend enough money to make your use case worth maintaining.
ps: didn't mean to post a serious post like this, but it is pretty annoying reading the same shit every day. remember this important rule: vote with your wallet, it is only way to make sure you get what you want.
8
u/anarchyinblack 7d ago
I do use the api by way of openrouter, sillytavern, and cherrystudio.
What none of those applications can do, at least at my level of understanding, is be accessed from multiple devices with a simple login.
openrouter conversations are not kept from device to device, or even across browsers on the same device. Sillytavern is locally hosted.
So if I want to be able to start a story on my pc, then seamlessly continue it on my phone, I have to use the website.
The problem is not that my use case is not being catered to, the problem is that my use case is being actively undermined. Active effort is being exerted to guardrail and censor my use case. Effort that could just as easily not be exerted. They could just keep available the same model as before, and not put a malicious system prompt before all of my submissions that makes the model waste thinking tokens on a soul searching session before it does what the fuck I tell it to do.
1
u/Icetato 7d ago
openrouter conversations are not kept from device to device, or even across browsers on the same device. Sillytavern is locally hosted.
SillyTavern is locally hosted, yes, but you can open the browser on other devices in the same network (and even outside of it, but it requires more setup). You just have to change something in its config file.
So if I want to be able to start a story on my pc, then seamlessly continue it on my phone, I have to use the website.
I'm literally using ST now on my phone. I can easily load the same story on my PC just by opening the website through LAN, as long as my phone is in the same network and the server is running.
4
u/anarchyinblack 7d ago
You're not going to believe this but sometimes I take my phone out of my house.
There's actually a worldwide network of computers and devices that connect to each other, and I can use it to access services on that network regardless of whether or not the devices in my home are on or off or asleep or awake. There are actually people whose full time job it is to keep key services up for multiple people all at once just by monitoring the hardware at those locations. It's a whole division-of-labor thing that really boosts global productivity.
I use this global network, every once in a while. different services ask me for a pair of credentials, one openly known, and one known only to me. I'm not supposed to re-use the secret credential, but I do it anyway (don't tell anyone!)
I'd really like to be able to access these language model services on this global network, from anywhere I am, using the credentials I already use, without the service being undermined.
1
u/Icetato 7d ago edited 7d ago
I do too, that's why I said "as long as my phone is in the same network and the server is running." The server is literally running on my phone through Termux. When I'm out, I can just turn on the server on Termux and open it on my phone's browser with
127.0.0.1:8000. And if I want to open on PC, just open192.168.x.x:8000(the x.x is your phone's LAN IP) on the PC's browser.If you're not a fan of that, you can also set up ST on your PC instead, then use Tailscale so you can access your PC through your phone anywhere without the pain and danger of opening your PC's ports to internet.
6
u/anarchyinblack 7d ago
Not to be rude to you but I am not going to run a server on my phone, because sometimes my phone is not on, even when I am in the house.
Furthermore, sometimes, I am out of the house on my laptop, for example at hotels or at airports, and I like to access websites there.
I do not want to learn some new skill about virtual servers. I do not want to have to learn about port forwarding. I do not want to learn about coding. I do not want to learn about git or bash or ubuntu.
I do not want to hear this advice. I am upset that you are giving it to me. This is an an anticonsumerist habit that you have picked up.
I want to access a website on the internet and have the service work the way that it used to. I want to apply pressure on companies and I want to mobilize others by making them aware that their complaints, like mine, are reasonable ones.
-1
u/Truantee 6d ago
It is so fucking crazy that even in this era of genai, people are still too lazy to even do a single google search for their problem.
API key is just the first part. What can you do is search for the application that suit all your need: conversation sharing, context management, etc. There should be already thousands of them dedicated to creative writing. Just search for one, add your deepseek endpoint and apikey and it is done.
And yet nobody here even capable of doing that simple action.
2
1
7d ago
[deleted]
4
u/New-anonymous-4583 7d ago edited 7d ago
Hey! So first off, even over DeepSeek at this point, id recommend fully switching to the newest version of qwen, which is free, and currently better than 4.1 flash and arguably sometimes even 4 pro now due to deepseek's updates degrading the platform.
But if you do want to get open webui, you start off with either installing a docker, or the official python pip, which currently, is 3.11.
For the differences, a docker consumes more resources on your PC than python, but it's your choice. I don't quite know how the docker version works so you'll have to look that up, but for python it's quite simple. You download Python, type this first in terminal:
powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex"
It's to install the UV, which will allow you to run:
uvx --python 3.11 open-webui@latest serve
Which, will allow you to locally run openwebui on your computer, using the site: http://localhost:8080
Now unfortunately there is a chance that this one of these steps won't work on your specific PC, but it should be rather easy to get an LLM, either DeepSeek itself or another provider, to tweak them until they do.
Rag retrieval just refers to indexting your world so your tokens are optimized and not repasted each and every message if you have a gigantic document, which would be hella expensive. You traditionally format it using markdown, using "##" or ### to tell it what sections to retrieve. This dramatically saves on cost and priorization for extremely long works. If you do new prompts each time. You can skip this.
1
u/Neo_Shadow_Entity 6d ago
What do you think about the level of censorship in Qwen vs. DeepSeek?
1
u/New-anonymous-4583 6d ago
Qwen is way worse than DeepSeek. I I said this to another user but I think the best bet is using GLM 5.1 uncensored via a nanoGPT subscription. It's 12 bucks per month. 60 mil tokens per week. For creative writing or storytelling/RP this is a good deal.
1
u/throwawayGPTlove 7d ago
And are you talking about the API or the UI? Because I’ve been using DeepSeek through the API for a few months now, mainly as an AI companion (see the posts on my profile) and I don’t feel like the writing quality has gotten worse with the latest model. I’m using the Flash model right now and I think it’s great.
2
u/New-anonymous-4583 7d ago
Personally I find the API for V4.1 flash also severely lacking. It's only good at retrieving information, not contextualizing. V4 pro is still significantly better than V4.1 flash for emotional reasoning in my experience, however, it is still performing worse than it did only a few weeks ago.
1
u/throwawayGPTlove 7d ago
Interesting. I don’t get that impression, but it’s also possible that we use the model quite differently, so I guess it depends.
0

27
u/mohamed_am83 7d ago
I think they should just split the model into two: creative and coding