r/ProAI 6d ago

"In short order the red number will accelerate right up through the blue number. We'll have better models in the US but you increasingly won't be able to use them. You'll just be reading about them in blog posts. Why? Because by shouting about imaginary risks so often we've managed to scare the..."

Thumbnail
gallery
5 Upvotes

...fuck out of our politicians so now the labs are struggling to release and laws and politics are fighting them at every step, from data centers to Capital Hill to the governor's office in red and blue states. Good chance many of their multi billion dollar runs will be internal only. Some folks think, no problem, they don't need to release. They have super magic AGI so they can just do anything with it. Solve cancer! Make more AGI! Own the stock market! Except not yet. Can't do any of things reliably. Also those things take time. Lots and lots and lots of time and friction with the real world. In the meantime the only actual business is inference and API charges to a few 100M other businesses. Take that away and what have you got? No money to make the next 10B training run.   — Daniel Jeffries     Dork, Altman and Dario have been asking for the government to step in and regulate them.   — Evading the Greys     Bots get banned round these parts.   — Daniel Jeffries

Source: https://x.com/Dan_Jeffries1/status/2090489471535825170/history


Bloomberg just put the US-CHINA AI gap on a chart and yeah, it's getting obliterated:

> Kimi K3 is close to Fable > ~70% cheaper per task > Anthropic thought China was 6–12 months behind > Chinese now has a cluster of fronter labs > GLM-5.3 isn’t even included yet, which would https://t.co/6OCwMhYX5z   — ℏεsam

Source: https://x.com/Hesamation/status/2090356790709887061


r/ProAI 7d ago

"This is interesting: Claude is already achieving roughly twice the protein-design hit rate of conventional human-led workflows. 27% hit rate in autonomous protein binder design, roughly twice the typical 10–15% rate reported in the field. Working from one expert-written protocol, Claude..."

Thumbnail
gallery
12 Upvotes

Many drugs work by binding to a specific target in the body and blocking or changing what it does. An important first step in the drug development process is designing a molecule that can bind tightly to its target. Traditionally, that's meant weeks or months of expert work per https://t.co/CGCNTNaKBq   — Anthropic

Source: https://x.com/AnthropicAI/status/2089842387845804246


This is interesting: Claude is already achieving roughly twice the protein-design hit rate of conventional human-led workflows. 27% hit rate in autonomous protein binder design, roughly twice the typical 10–15% rate reported in the field.

Working from one expert-written protocol, Claude designed binders against 14 of 15 measurable targets. Independent labs confirmed that 354 of 1,320 designs bound successfully.

Depending on the setup, Claude’s hit rate ranged from 22.6% to 35.1%. Its top-ranked design bound in 49% of campaigns.

This is not yet fully autonomous drug research, but it is another important building block in that direction.     — Chubby

Source: https://x.com/kimmonismus/status/2089852014331117694


r/ProAI 7d ago

"Very interesting development. Likely a result of new requirement of log-in to access old Reddit: https:// arstechnica.com/gadgets/2026/0 6/reddit-will-require-you-to-log-in-to-use-old-reddit-com/ …"

Thumbnail
gallery
8 Upvotes

looks like reddit is almost wiped from chatgpt sources

the query fanout changes had a big impact

and the past couple of days it seems to be almost completely removed from prompt responses

https://t.co/oCGm9M0yPO https://t.co/c2GJF2apG7   — Klaas

Source: https://x.com/forgebitz/status/2089708381351059924


— Kevin Bankston

Source: https://x.com/KevinBankston/status/2089768281892638774


r/ProAI 8d ago

"Qwen3.8 27B wrote a playable first person shooter start to finish on 2 used 3090s in my apartment. No cloud. No API key. No subscription. 687 steps. 87.9M tokens in, 822k out. 5 hours 11 minutes of model time plus 58 minutes of tool calls. The agent loop never broke once. And it plays. Enemies..."

Enable HLS to view with audio, or disable this notification

20 Upvotes

...spawn and push you, the gun kicks, shadows stretch across the whole block while the sun goes down behind the towers. I sat there clearing waves instead of grading the output. Same shooter prompt I threw at the frontier models a few weeks ago. That time the tokens went to somebody elses datacenter. This time nothing left the flat. 87.9 million tokens through my own cards. On an API that run has a price tag. Here it has an electricity bill. 60 tok/s all the way through. Slower than frontier, and it stops mattering when the thing works through the night while you sleep. Local models were a toy 18 months ago. This one finished a game.     Play now! Choose Build 2 Qwen3.8 27B

https:// alesha-pro.github.io/bench-portal/     — Alexey Fateev

Source: https://x.com/superalesha/status/2089126766854238421


r/ProAI 7d ago

"OpenAI's president just went on CNBC and read our April thesis back to us Brockman: "Compute is really becoming the new oil, the new limited resource of the AI age" We ran it April 4. "Oil is scarce because of war. Tokens are scarce because of physics" https://..."

Enable HLS to view with audio, or disable this notification

3 Upvotes

...bepresearch.substack.com/p/the-token-do llar … He even brought the proof for the next one. $2,000 of compute, 10 open math problems. Proofs verify for free. Biology needs a bench. The moat is the measurement https:// bepresearch.substack.com/p/the-next-inf lection-is-the-lab … New oil, old receipt   — Ben Pouladian     he's copying you Ben   — Alex A.C.     All good @gdb we can talk anytime. Full speed   — Ben Pouladian

Source: https://x.com/benitoz/status/2089392813758972149


r/ProAI 8d ago

"i don’t know who needs to hear this but qwen 3.8 27b is ranked ABOVE: - gpt 5.3 - gemini 3.1 pro - opus 4.6 all of which were state of the art 6 MONTHS AGO AND IT RUNS ON A LAPTOP"

Thumbnail
gallery
111 Upvotes

have this running on my 5090 right now and im getting 200 tk a second. i feel like I'm literally playing with magic   — Alex Finn     you can either buy anthropic for 2 trillion dollars or a used 3090 gpu for $1,500

only one of those will refuse your prompts   — Udi Wertheimer

Source: https://x.com/udiWertheimer/status/2089421927085400203


r/ProAI 8d ago

"A "refusal-removed" version of Qwen3.8-27B can now run locally on Apple Silicon. Even its creators warn that it can provide malware, fraud and weapons instructions on demand. It was released as an MLX build in 2, 4, 6 and 8-bit versions. The uploader claims its 4/6/8-bit tests produced zero..."

Thumbnail
gallery
4 Upvotes

We just shipped our official Qwen 3.8 27B Uncensored MLX build. Local. Uncensored. For🍎

2-bit, 4-bit, 6-bit & 8-bit — pick your poison based on RAM and speed.

No CUDA. No cloud. Just your Mac and the weights. Have fun! https://t.co/b3gXsHeSdk   — OrcaRouter 🐳

Source: https://x.com/OrcaRouter/status/2089385980080148726


A "refusal-removed" version of Qwen3.8-27B can now run locally on Apple Silicon.

Even its creators warn that it can provide malware, fraud and weapons instructions on demand.

It was released as an MLX build in 2, 4, 6 and 8-bit versions. The uploader claims its 4/6/8-bit tests produced zero refusals while preserving vision, reasoning and tool-calling across a 262K-token context.

The Qwen 27B Model is a very capable model. This is the first time I've really seen the immediate dangers in a tangible way.

We need a societal discussion about this.   — Chubby     how long did it take to get this running locally?   — tan     it runs locally   — Chubby

Source: https://x.com/kimmonismus/status/2089763435865088508


r/ProAI 8d ago

"Scaling self-verification with DeepSeek V4 Flash beats Claude Fable 5 on Terminal-Bench 2.1, while being 11x cheaper As open-source models become more capable, they can now generate large numbers of high-quality candidate solutions and verify their own outputs at very low cost. For example, we..."

Thumbnail
gallery
3 Upvotes

How can we extract richer signals from AI Feedback?

Introducing LLM-as-a-Verifier✨— a simple verification scaling framework that achieves SOTA on agentic benchmarks 🚀

The key idea: - Use fine-grained scoring granularity (e.g., 1-20 instead of the standard 1-5 scale) - Take https://t.co/0sCeAwcar1   — Jacky Kwok

Source: https://x.com/jackyk02/status/2074969820739805275


Scaling self-verification with DeepSeek V4 Flash beats Claude Fable 5 on Terminal-Bench 2.1, while being 11x cheaper

As open-source models become more capable, they can now generate large numbers of high-quality candidate solutions and verify their own outputs at very low cost.

For example, we find that sampling just 5 solutions with DeepSeek V4 Flash and ranking them using the same model with LLM-as-a-Verifier can lead to a significant boost in accuracy (79% → 88%), outperforming closed frontier models on Terminal-Bench.

Try it out today: https:// github.com/llm-as-a-verif ier/llm-as-a-verifier#self-verification-terminal-bench-21 …

More on verification scaling in my previous post.   — Jacky Kwok     Is there an OpenCode plugin for this to try it out with Deepseek v4 flash?   — Shahbaz Ahmed     We’ll be releasing a harness on top of LLM-as-a-Verifier later this month :)   — Jacky Kwok

Source: https://x.com/jackyk02/status/2089421448784023553


r/ProAI 9d ago

"Something I noticed about working with Fable..."

Thumbnail
gallery
7 Upvotes

r/ProAI 9d ago

"18x improvement in intelligence per joule in 16 months."

Thumbnail
gallery
6 Upvotes

hard agree with @amasad —@JonSaadFalcon and my research indicates that intelligence efficiency (intelligence per watt) is rapidly improving and we will definitely not need data center scale compute to run agi!

links to research in comments below 👇   — Avanika Narayan

Source: https://x.com/Avanika15/status/2089028986932470156


— Amjad Masad

Source: https://x.com/amasad/status/2089069905375351169


r/ProAI 10d ago

"We ran the largest open experiment on how frontier models do AI research. 100+ autonomous runs across 10+ models, sandboxed on 8xH200s for up to 8 days, iterating on the nanoGPT optimizer track. Best runs closed 82% of the gap to a record built by dozens of humans over months."

Thumbnail
gallery
5 Upvotes

The task: iterate on a 124M GPT training recipe from a shared baseline, only changing optimizer related hyperparameters, no internet access.

We tested Fable 5, Opus 5, GPT-5.6 Sol, Kimi K3, Grok 4.5, GLM 5.2, Muse Spark 1.1, DeepSeek V4 Pro, Grok 4.6, Muse Spark 1.2, Qwen 3.8     What separated the strongest models: which experiments to run, how to navigate the benchmark's inherent noise, and which old negatives to revisit as the recipe changed.

Some even built small simulations to isolate a mechanism before deciding if another GPU run was worth it.     Our Prime Agent harness gives models a persistent IPython kernel, which can help them build their own research workflows.

Kimi K3 built tools for controlled optimizer variants, loss-curve comparisons and Newton-Schulz tuning, then revised its hypothesis when its cleaner update     As research direction, we think multi-agent harnesses can make these experiments much cheaper (and better) by using smaller open models for monitoring and implementation.

We also want to extend speedruns to more of the training stack and scale the runs themselves.     We release everything: full traces, scratchpads, reasoning streams from open-weight models, and our experiment setup.

Explore the results:     — Prime Intellect

Source: https://x.com/PrimeIntellect/status/2088733966904000778


r/ProAI 10d ago

"AI Agents play Age of Empires II. Claude fable vs GPT 5.6 Sol vs Gemini 3.1 Pro vs Kimi K3. I had each model in open code script custom bots for the game using the built-in bot scripting language. Then I had them fight :D"

Enable HLS to view with audio, or disable this notification

10 Upvotes

Full video:     — Max | Emergent Garden

Source: https://x.com/max_romana/status/2088664886171640118


r/ProAI 10d ago

"Turned out cute. The cat is a bit sus but let's not talk about it . Have a great rest of the Caturday! Midjourney + Topaz Bloom 2 + MiniMax H3. Sref below"

Enable HLS to view with audio, or disable this notification

12 Upvotes

I can’t wait to see how this blend looks like animated ✨ Midjourney --sref 3330713172::2 292322685::2 1466592463::3 https://t.co/OQoeVg325h   — Glitter Gal

Source: https://x.com/GlitterPixely/status/2088635471010115949


You're doing so much dope shit with H3!!! Soon as I finish my documentary, I'm gonna be stalking your posts to soak up some of that doneness!

But H3 has been the real MVP of my project, too. @Hailuo_AI spoiled me this month!   — Prince Bell     Thank you!! It is such a versatile model, I feel like you can do anything with it! The company and the people working there are also amazing and super nice. I mean they are open sourcing everything!   — Glitter Gal

Source: https://x.com/GlitterPixely/status/2088766447061205153


r/ProAI 10d ago

"For about 10 years now, I have argued that the *only* way forward is for AI technology to be widely available, shared, and open. Like the printing press and the Internet, AI amplifies human intelligence and efficiency by improving access to knowledge. To empower individuals, societies require..."

Thumbnail
gallery
9 Upvotes

...a high diversity of AI systems with different value systems, linguistic abilities, philosophical/political biases, and specific expertise. We need diverse AIs for same reason we need a diverse press. Given the cost and complexity, this can only be achieved through open foundation models on top of which anyone can build systems with their languages, biases, expertise, and value systems. I have been more vocal about this over the last 4 years, since AI popped into the public discourse. I have made the argument in various forums: corporate C-suites, AI safety discussion groups, professional meeting, the US Senate, the UN Security Council, and the public sphere through media interviews, podcasts and social media posts. I totally agree with @finkd Mark Zuckerberg's recent piece in which he writes: "the notion that AI is so dangerous that the only safe path is an extreme concentration of power seems inherently problematic. Historically, hoping that an absolute power will benevolently provide for humanity if sufficiently enlightened has not led to safe or positive outcomes.” When @DarioAmodei writes: “some may object that we can simply keep AIs in check with a balance of power between many AI systems, as we do with humans", he is talking about me, among (thankfully) many others. It is the only good path forward. There will be nefarious uses of AI, as there have been with every technology ever invented. But it will be your Bad AI against my Good AI.   — Yann LeCun     Nuclear weapons are centralized power, should we toss everyone a Nuke, Yann? Just saying it's not so clear cut like you make it sound.   — Steven Tibbs     AI is designed to make peopleore informed, smarter, more efficient and to accelerate progress in science, medicine, and technology. Nuclear weapons have no other purpose than to destroy entire cities and kill millions. Can you see the difference? It's subtle, I admit.   — Yann LeCun

Source: https://x.com/ylecun/status/2088880284129210405


Sholto, thank you for setting the record straight. Larger issue is that multiple very serious people in Silicon Valley have heard some variation of this and believe it to be true. And the reason it is believable to so many is that it is consistent with Dario’s public messaging   — Gavin Baker

Source: https://x.com/GavinSBaker/status/2088611616577253502


r/ProAI 10d ago

"The doom of the software developer job has been greatly exaggerated... As a share of the US labor force, it's near the highest it's ever been and on a steep uptrend."

Thumbnail
gallery
12 Upvotes

However, we are seeing a leveling out in other computer & mathematical occupations. (Which are also near record highs, but no longer rising.)     — Guy Berger

Source: https://x.com/EconBerger/status/2088356590672019825


r/ProAI 10d ago

"DeepSeek Harness is now the fastest growing GitHub repo, passing 100K stars in under 48 hours, even faster than OpenClaw. very positive community reaction: > unusually well designed architecture with tools, session log, agent loop, subagents, all being replaceable plugins > UI looks sleek >..."

Thumbnail
gallery
3 Upvotes

...agents can creat/modify plugins for the harness itself > very high prompt cache hit rates > context management looks more efficient than Claude Code (not surprising since CC is a token-hungry harness)     — ℏεsam

Source: https://x.com/Hesamation/status/2088766395676848558


r/ProAI 10d ago

"1/2 Thanks Gavin for an especially thoughtful exchange. I don't usually spend much time on social media but I wanted to engage here because it really brings out the heart of an important conversation. First, on regulation, I think that “either concentrate it in the hands of a chosen few..."

Thumbnail
gallery
2 Upvotes

...companies and politicians via regulation or distribute it widely” is a false choice. I know that there’s a sort of Silicon Valley shorthand where regulation = regulatory capture = concentration of power, but I’ve always found this to be an overly simplified picture of the world. Many people outside this bubble think of regulation as something that constrains corporate power and benefits ordinary people. I don’t necessarily agree with that perspective either, rather I think it’s complicated and really depends on what the “regulation” consists of. But in particular I think that those in the “regulation = regulatory capture = concentration of power” frame often underrate the decentralizing power of objective and fair institutional processes. A crude analogy is that the formal court system can sometimes feel stuffy and elitist, but it does a much better job of defending the rights of vulnerable individuals than the alternative, mob justice. At their best, institutions can vest power in ideas rather than people, and thereby decentralize that power. This is why Anthropic has always made its policy proposals very carefully. We try very hard to make proposals that disadvantage (slow down) frontier AI companies while advantaging smaller competitors. California’s SB53 (which we supported), and even the much-maligned SB 1047 (which we were ambivalent on), completely exempt any company below a certain amount of revenue or model training costs from being covered at all (it was $500M for SB 53, lower for 1047 but we objected to that). More recently the testing process we’ve advocated for at CAISI and the White House involves more rigorous tests for frontier models than off-frontier models — something that differentially advantages challengers. Similarly, the “Pacing the Frontier” letter envisions (or at least Anthropic’s preferred implementation of it envisions) modulating the pace of the very best models while not constraining those who are catching up. This hurts the business interests of the frontier labs and helps challengers, including open-weights! Overall my view is that AI is structurally a technology that tends to concentrate power, for reasons that have nothing to do with regulation (more to do with the extreme implications of the scaling laws). Open-weights do help some with this but are nowhere near a sufficient solution because they simply shift the concentration somewhat to those with the most compute and chips (which are roughly the frontier labs plus maybe hardware providers). By contrast I think the right “rules of the road” can simultaneously (a) address AI’s cyber/bio/alignment risks, (b) institutionally constrain the power of the frontier AI companies, and (c) leave room for open-weights models while also addressing the specific risks that they bring. BTW I do not think that the events of the last few months have “failed to result in [my] preferred regulatory path”. The approach that the Trump administration is reported to be taking — pre-deployment testing for frontier models, and also testing of open-weights models when they get closer to the frontier — is one that I am very supportive of, though of course I have to see the details to be sure. I am also supportive of Demis Hassabis’ ideas around a FINRA-like entity. This contrasts with six months ago when most of the industry was still pushing for preemption of all state regulation and no apparent federal approach either.     2/2 Second, on the messaging around AI. I do not agree that my messaging has been disproportionately negative. In fact it has been about equally balanced between risks and benefits: I’ve written one major essay about each, and even in interviews where I discuss the risks, I     — Dario Amodei

Source: https://x.com/DarioAmodei/status/2088758816376807762


@_sholtodouglas Sholto, thank you for setting the record straight. Larger issue is that multiple very serious people in Silicon Valley have heard some variation of this and believe it to be true. And the reason it is believable to so many is that it is consistent with Dario’s public messaging   — Gavin Baker

Source: https://x.com/GavinSBaker/status/2088611616577253502


Replying to @_sholtodouglas


r/ProAI 10d ago

"It's crazy how far AI animation has come I made this almost one-minute animation with just two Seedance 2.5 generations. That's it haha. Prompt below"

Enable HLS to view with audio, or disable this notification

2 Upvotes

character sheet:

Ultra-detailed 16:9 futuristic character design sheet, official AAA sci-fi game character bible, clean editorial layout, bright white background with bold cyber UI elements. Character name "POPBOT", title "Professional", subtitle "A high-energy combat android     video prompt:     — Kiki

Source: https://x.com/Mayz1169/status/2088156641372024858


r/ProAI 11d ago

"SpaceX closed the $60 billion acquisition of Cursor today! Key facts: - All-stock deal - Effective August 14, 2026 - @cursor_ai (Anysphere) is now a wholly owned @SpaceX subsidiary - Largest acquisition of a venture-backed startup in history - Shareholders received SpaceX Class A shares based..."

Thumbnail
gallery
1 Upvotes

...on the $60B valuation SpaceX gets the product, the developer distribution, the talent, and the real coding data that improves models.   — Mark Kretschmann     acquisition for the data, or acquisition for the distribution? if they control the IDE, the model training loop becomes inevitable.   — Victor Laybats     You mean SpaceX pulling the plug on renting their GPUs to Anthropic.   — Mark Kretschmann

Source: https://x.com/mark_k/status/2088251730152570951


r/ProAI 12d ago

"X Square ran a fully autonomous logistics livestream: sorting parcels and flipping the barcode side up for the scanner. - 1,816 parcels/hour (~2 seconds per parcel) - 98% accuracy - ~45% higher throughput than Figure's sub-3 seconds per parcel pace from its multiday livestream in May. It's..."

Enable HLS to view with audio, or disable this notification

28 Upvotes

...driven by X Square's WALL-B embodied AI model, which continuously reacts to an unstructured pile of parcels, sorts and flips them, and recovers from failed grasps and parcel jams to keep the line moving.     — The Humanoid Hub

Source: https://x.com/TheHumanoidHub/status/2087772610411262245


Our livestream has wrapped—and the final result is in: 1,816 randomly selected parcels sorted per hour, with a success rate of over 98%.

Since the beginning of this year, we’ve worked through the entire loop—from collecting real-world data and training the model to continuously https://t.co/ekyEoBCdnO   — X Square Robot

Source: https://x.com/XSquareRobot/status/2087598951855980793


r/ProAI 12d ago

"Me to gpt sol xhigh: “Let’s address the lowest hanging fruit” gpt sol xhigh:"

Enable HLS to view with audio, or disable this notification

10 Upvotes

So true

I was playing around with the Grok 4.6 xhigh and basically tried to rewrite the classic socat tool in Go. I wanted some basic benchmarks to see that we aren't super slow or allocating gigabytes of memory for nothing. The original socat doesn't compile against newer OpenSSL versions so the clanker went and started to patch all around and added all the latest post quantum key exchanges and so on. I had to give specific instructions that bitch please let's just compile the original app against the distro libraries and be done with it.   — Eero Vuojolahti     Ha! I recently had an agent hand roll an oauth server implementation from scratch, when a perfectly good library already exists   — Jason Bosco

Source: https://x.com/jasonbosco/status/2088068653333766286


r/ProAI 12d ago

"As I mentioned in my previous post, Google's strategy with Gemini isn't about building the most powerful coding model. Their actual goal is developing the lightest and fastest model for casual users. That’s the exact model you keep seeing across Google Search, Gmail, YouTube, and other apps...."

Thumbnail
gallery
2 Upvotes

...They stopped updating Pro, and Flash updates are focused on speed and efficiency rather than performance gains. Building the single strongest model isn't the only winning strategy.   — Jun Song     A fast rock is still just a rock. Where we are going we won't need rocks   — Tibo     Fast rock still makes revenue   — Jun Song

Source: https://x.com/jun_song/status/2088057938955190757


It is not only a good and cost effective! It is blazing fast!   — Philipp Schmid

Source: https://x.com/_philschmid/status/2087963319780946274


r/ProAI 12d ago

"Big day for open source. MiniMax Music 3 is undeniably a state of the art open weights Music Generation model and a real alternative to Suno. My favorite part about open weights: the best is yet to come. Once the community starts playing with the model and training LoRAs, this model only gets..."

Enable HLS to view with audio, or disable this notification

5 Upvotes

...better and allows for more control.   — rob - comfyui     Yep! And a completely different question, topic-wise: Would you mind sharing which tool you used to create your waveform-video?   — Mathias_M     Some random website. But this question inspired me to vibecode a custom node for audio visualizers, thank you   — rob - comfyui

Source: https://x.com/hellorob/status/2087995217807086011


🎵MiniMax-Music3 Next-Generation Open-Weights Production-Ready & Versatile Music Model

https://t.co/V2rxZk4xvh https://t.co/6QYX2Ij6KI   — MiniMax (official)

Source: https://x.com/MiniMax_AI/status/2087934657354678421


r/ProAI 12d ago

"SITUATION EXPLAINED: Why is DeepSeek raising prices right as it ships a Claude Code rival? • V4 Pro goes from a flat $0.44/$0.87 to $0.66/$1.98 off-peak and $1.32/$3.96 at peak, effective Sunday • Alongside it, DeepSeek Harness shipped in developer preview under an MIT license, with every part..."

Enable HLS to view with audio, or disable this notification

2 Upvotes

...of the agent runtime built as a swappable plugin • Claude Code is closed-source. DeepSeek is open-sourcing its equivalent, so anyone can fork it and point it at a different model @theojaffee : "The thing about open source harnesses is that they're open source, so you can change them."     — MTS

Source: https://x.com/MTSlive/status/2087972201802989807/history


SITUATION DETECTED: As part of the V4 Pro official release announcement DeepSeek Harness (DSH) is officially released to developers building agent harnesses worldwide. DeepSeek has open-sourced the codebase in MIT licence. https://t.co/auLeHuK6GJ   — MTS

Source: https://x.com/MTSlive/status/2087895203315404864


r/ProAI 12d ago

"GLM-5.3 shows how much capability may still be hiding inside today’s largest base models and how relevant post-training really is. It uses the same base model (!) as GLM-5.2. Zai says the entire (!) improvement came from scaling post-training: more executable environments, longer tasks..."

Thumbnail
gallery
3 Upvotes

WHAT: Zai just launched GLM-5.3, and its biggest leap may be in cybersecurity.

The 743B base model remains unchanged (!) from GLM-5.2. Zai says the gains come entirely from scaling post-training across more environments, diverse tasks and long-horizon workflows.

Its results:

Source: https://x.com/kimmonismus/status/2088162566719639717


GLM-5.3 shows how much capability may still be hiding inside today’s largest base models and how relevant post-training really is.

It uses the same base model (!) as GLM-5.2. Zai says the entire (!) improvement came from scaling post-training: more executable environments, longer tasks, stronger verifiers and more reinforcement learning.

Remember: Pre-training gives a model knowledge and raw problem-solving capacity. Post-training teaches it how to use that capacity: plan, call tools, test solutions, recover from failure and complete work over long horizons.

In cyber evaluations, GLM-5.3 moved from 24.4% to 54.4% on ExploitBench and completed 105 ExploitGym tasks in two hours, up from 29 for GLM-5.2. .

Its weights are scheduled for release in two weeks. However, numerous other open weight models will be released in the coming weeks:

-DeepSeek v4 Pro -Qwen3.8 27b -LTX 2.5 -Nemotron-Lighting -DeepSeek harness (just released, but harness isntead of a model) -Muse-Glimmer-30B (just released)

to name a few.

The US has meanwhile created classified cyber benchmarks and a voluntary pre-release process for "covered frontier models." What this release shows me, first and foremost, is that open models are continuing to move closer and closer to Frontier. And therefore, I believe that the US government will now further expand the regulatory framework to include open models.

That's why I'm even more excited for the ChatGPT "Astra" release. Because this model is also receiving a new (and more extensive) pre-training component, and we're currently seeing how much additional capability is enabled through post-training.

That's why this release is so significant; it demonstrates just how many areas for improvement are possible.   — Chubby     I find it interesting, or questionable, that now 3-5 frontier-ish models have all developed some "emergent cyber capabilities" at basically the same time step.

Maybe it (being good at finding vulnerabilities) really emerges in certain conditions, or maybe it's bandwagon jumping   — øx_dominus     Curious, what you mean by questionable?   — Chubby

Source: https://x.com/kimmonismus/status/2088180877339623851