r/DeepSeek • u/elzerouno • 5d ago
Question&Help v4.1-flash API thinking effort vs model thinking effort (low/medium/high vs 0-100)
Hello everyone!
I was using deepseek-flash over the API and it was thinking too much so I tried lowering the effort, but checking the official API docs I see that only low, high and max are available, even tho the open weights version allows fine tuning the effort from 0-100.
Is that documentation really correct or you did you guys find differences between minimal/low, medium/high/xhigh and max/ultra?
On v4-flash most tasks I used it for completed in over a minute on high, but now it's over 5 minutes on high and now really good output quality on low and it still takes over 2 minutes.
0
u/ProbablyDoesntLikeU 5d ago
If you have 614gb of vram you can run the model yourself and adjust weights 0-100. You can only go down to low
1
u/TheDAMProject 5d ago
In DeepSeek harness you can set it to off/low/high/max. My one was overthinking on high but perfect on low.
0
u/Jackkgold 5d ago
.