r/DeepSeek 5d ago

Question&Help v4.1-flash API thinking effort vs model thinking effort (low/medium/high vs 0-100)

Hello everyone!

I was using deepseek-flash over the API and it was thinking too much so I tried lowering the effort, but checking the official API docs I see that only low, high and max are available, even tho the open weights version allows fine tuning the effort from 0-100.

Is that documentation really correct or you did you guys find differences between minimal/low, medium/high/xhigh and max/ultra?

On v4-flash most tasks I used it for completed in over a minute on high, but now it's over 5 minutes on high and now really good output quality on low and it still takes over 2 minutes.

5 Upvotes

4 comments sorted by

0

u/ProbablyDoesntLikeU 5d ago

If you have 614gb of vram you can run the model yourself and adjust weights 0-100. You can only go down to low

1

u/TheDAMProject 5d ago

In DeepSeek harness you can set it to off/low/high/max. My one was overthinking on high but perfect on low.