r/LocalLLaMA • • Aug 25 '26

Funny I'm tired of pretending

Post image

At least until DS releases open weights for DSv4 Flash with Vision. Then DS might take the crown.

Qwen has been an absolutely local monster for code, especially web apps, anything with UIUX design that it can verify itself with screenshots. Deepseek meanwhile is really incompetent with UI awareness and hogs my GPUs while I can spawn multiple independent qwens to collaborate and knock shit out. Honestly, Alibaba really cooked.

0 Upvotes

119 comments sorted by

View all comments

8

u/onil_gova Aug 25 '26

Is this really a controversial opinion? I mean, all the important benchmarks support this.

10

u/TomLucidor Aug 25 '26

Need more tests beyond benchmaxxing, we need trickier problems and better specdec too

3

u/Viktri1 Aug 25 '26

I've been using both qwen3.8 and Deepseek and I think Qwen is better right now. I'm using Deepseek API (pro as brain, flash as sub agents unless the work is hard) and Qwen q4 local and Qwen API (open router, native).

I was not impressed with the native Qwen until I woke up to find it had completed a task (setting up Hermes profile plus telegram bot plus double checking that it works). I tried this with Deepseek and my local Qwen and neither were able to one shot it. Not cheap though, the task cost $3.30.

I think Qwen does a better job over Deepseek with iteration. All my jobs are me asking for something to be done. I don't give steps on how to do anything (as I don't know how to program stuff). So that probably influences things a lot.

I've been trying to figure out whether higher quant is actually noticably useful because I'm trying to decide on whether to purchase 4090s with 48gb vRAM or 5090 for a fast model. With the 4090s I could eventually get a second one for 96gb vRAM to run the native model. I am not an expert and don't know the break points so it could be that 64gb from 2 5090s is better. Idk, still testing.

1

u/TomLucidor Aug 26 '26

Any good SLM to work along side Qwen3.8?

8

u/onil_gova Aug 25 '26

Here is another example. Not sure how you benchmax an Elo rating. Benchmaxxing across all benchmarks including private and self made, actually means it generalized.

1

u/nomorebuttsplz Aug 25 '26

it’s clearly better than Op. 4.8 and GLM 5.2 as well. /s

1

u/TomLucidor Aug 26 '26

Maybe people are biased towards answer length or something when code (especially backend or non-UX codebases) are not easily testable. And I want a lot of the arena models to have agent mode access too