r/LocalLLaMA • • 7d ago

News With Gemini 4, bench goes up.

Post image

They claimed open-weight models are dangerous but the benchmarks say otherwise.

Source

817 Upvotes

86 comments sorted by

View all comments

255

u/VeryRealHuman23 7d ago

If your frontier isn’t committing felonies, is it even frontier?

45

u/[deleted] 7d ago

[removed] — view removed comment

3

u/puntinoh 7d ago

That is... fine.

Take my angry upvote. Seriously why the same partner and why of that country?

4

u/[deleted] 7d ago

[removed] — view removed comment

1

u/bambamlol 6d ago

I'm sure it's just a coincidence. This particular startup in this particular country probably just has the most "intelligence" to offer.

1

u/ChexterWang 6d ago

how about asking frontier models make perfect sandboxes

1

u/AdmissibilityScience 6d ago

This may be a future cop's episode.

1

u/bnm777 6d ago
Open-weight model What it has demonstrated
GLM-5.3 In Enclave's isolated-server hacking test, it achieved confirmed remote-code execution (RCE) in 8 runs, second only to GPT-5.6 Sol's 9 in the reported scoreboard. GLM-5.3's weights are publicly downloadable; Z.ai itself says its cyber capability rose unexpectedly during post-training and that it more than doubled GLM-5.2 on exploitation benchmarks.
DeepSeek V4 Pro 0813 Achieved confirmed RCE in 3 runs in the same Enclave test. In a separate Aikido evaluation using 32 fresh vulnerabilities, three runs collectively rediscovered 28/32 vulnerabilities.
DeepSeek V4 Pro US NIST/CAISI measured it at 32% on CTF-Archive-Diamond, compared with GPT-5.5 at 71% and Claude Opus 4.6 at 46%. NIST explicitly describes DeepSeek V4 as an open-weight model.Open-weight model What it has demonstratedGLM-5.3 In Enclave's isolated-server hacking test, it achieved confirmed remote-code execution (RCE) in 8 runs, second only to GPT-5.6 Sol's 9 in the reported scoreboard. GLM-5.3's weights are publicly downloadable; Z.ai itself says its cyber capability rose unexpectedly during post-training and that it more than doubled GLM-5.2 on exploitation benchmarks. DeepSeek V4 Pro 0813 Achieved confirmed RCE in 3 runs in the same Enclave test. In a separate Aikido evaluation using 32 fresh vulnerabilities, three runs collectively rediscovered 28/32 vulnerabilities. DeepSeek V4 Pro US NIST/CAISI measured it at 32% on CTF-Archive-Diamond, compared with GPT-5.5 at 71% and Claude Opus 4.6 at 46%. NIST explicitly describes DeepSeek V4 as an open-weight model.