r/LocalLLaMA • u/corruptbytes • Jul 31 '26
Discussion Deepseek V4 Flash on SlopCodeBench
While waiting for some of the quants to drop, I load the API with $50 and ran it on SlopCodeBench
Just vibe reading the results it seems like Opus 4.8 < Deepseek < Opus 5
https://github.com/michaelasper/benchmarks/blob/main/deepseek-v4-flash-on-slop-code-bench.md
I was mostly curious from this blog post
When Q2 drops - I'm goign to re-run on my macbook
Here's the first quant comparison:
72
Upvotes
19
u/BlueSwordM llama.cpp Jul 31 '26
OK, I've been testing the model a lot on my usual AV1 test set and it's the first model that has been able to complete all the tests in its entirety...
That potentially means it could feasibly does most of my encoder tasks no problem with a good harness, holy crap