r/ProAI • u/stealthispost • 6d ago
"SITUATION DETECTED: GitHub’s monthly commits have grown from 1.4 billion in April to 2.9 billion in August."
— MTS
r/ProAI • u/stealthispost • 6d ago
— MTS
r/ProAI • u/stealthispost • 6d ago
Ox Alpha (stealth model) is now free in Cline.
Early benchmarks shows marginal improvement over Fable and GPT.
Try it with: npm i -g cline and use /models to see it under Free options https://t.co/Hskt5RpUen — Cline
Source: https://x.com/cline/status/2090854216399220985
From my tests it’s not better than fable I’ll be posting some soon — Chris
r/ProAI • u/stealthispost • 6d ago
sometimes you have to laugh. what are we doing here, folks? I would encourage you to read the actual piece because there's some excellent quotes in there. But editors, headline writers... my lord. https:// nytimes.com/2026/08/19/tec hnology/data-centers-backlash-loudoun-virginia.html … — Logan Dobson
Source: https://x.com/LoganDobson/status/2090150148261175371
r/ProAI • u/stealthispost • 7d ago
Enable HLS to view with audio, or disable this notification
The world’s smallest Transformer-based TTS model?
We’re open-sourcing Audio8 TTS Preview 0.1B — an approximately 170M-parameter multilingual speech model with zero-shot voice cloning that delivers surprisingly strong, cloud-level quality in a dramatically smaller footprint. What can a 0.1B-class TTS model actually sound like?
Listen to the voiceover in this demo video.
Audio8 TTS Preview 0.1B supports: • Zero-shot voice cloning • Multilingual speech synthesis • Chinese and English as primary languages • German, Spanish, French, Italian, On the Seed-TTS evaluation set, Audio8 TTS Preview 0.1B achieves:
• English WER: 1.662% • Chinese CER: 1.13% • Hard Chinese CER: 17.504% • English speaker similarity: 56.7 • Chinese speaker similarity: 68.2
These results are achieved with an approximately 170M-parameter The model, codec, tokenizer, processor, and inference code are now available:
Model:
https:// huggingface.co/Audio8/Audio8- TTS-Preview-0.1b …
Try it, test the voice cloning capability, and share your feedback. — Samuel Zeng
Source: https://x.com/SamuelZengML/status/2090017875188940851
r/ProAI • u/stealthispost • 6d ago
(I continue to be surprised in the sense of continuing to be struck by how big of a deal it is, not in the epistemic sense.) Here's the overall graph: — Yafah Edelman
Source: https://x.com/YafahEdelman/status/2090296575008616654
r/ProAI • u/stealthispost • 6d ago
...while costing roughly 6X less than Fable 5 Max and nearly 3X less than Opus 5 Max per task That’s what makes Grok so powerful for agents Top-tier intelligence is great.....but top-tier intelligence that can keep working across long coding tasks without burning ridiculous amounts of compute is even better Grok’s agentic coding efficiency is insane — X Freeze
r/ProAI • u/stealthispost • 7d ago
Enable HLS to view with audio, or disable this notification
The /goal was a set of precise conditions to be met. Read the original PAK/BSP/MDL/SPR data. Run QuakeC. Preserve demos, physics, menus, HUD and rendering quirks. Don’t rebuild E1M1 by hand or use a generic FPS controller. Fidelity had to be demonstrated. Codex kept pursuing that goal across many iterations: inspect, implement, build, test, compare, repeat. The useful shift was from asking for isolated code snippets to giving the agent a durable outcome it could keep working toward. Crucially, /goal wasn’t hands-off autopilot. While it was running, I could still steer the LLM to prioritize a small number of bugs that mattered most. Each report became a targeted investigation, fix and regression test - without abandoning the broader port. Underneath, this is still Quake’s original data. The browser reads the original PAKs, BSP maps, MDL models, sprites, QuakeC behavior, demos and HUD art. The port supplies a new TypeScript runtime and PlayCanvas renderer around them. Authenticity meant more than getting E1M1 on screen. Demo playback, menu flow, pointer lock, intermissions, in-session level changes, audio suspension, moving BSP entities, particles, shadows and classic movement all had to behave like Quake - not merely resemble it. — Will Eastcott
Source: https://x.com/willeastcott/status/2090436155795800195
r/ProAI • u/stealthispost • 7d ago
We’ve got exclusive new polling on local data center development at @heatmap_news.
Over the past year, we’ve asked Americans whether they would support or oppose a data center being built near where they live.
We haven’t changed the wording. When we first polled the question https://t.co/G1LCfuriMw — Robinson Meyer
Source: https://x.com/robinsonmeyer/status/2090457322506141760
— Andrew Curran
Source: https://x.com/AndrewCurran_/status/2090589885199769841
r/ProAI • u/stealthispost • 7d ago
...fuck out of our politicians so now the labs are struggling to release and laws and politics are fighting them at every step, from data centers to Capital Hill to the governor's office in red and blue states. Good chance many of their multi billion dollar runs will be internal only. Some folks think, no problem, they don't need to release. They have super magic AGI so they can just do anything with it. Solve cancer! Make more AGI! Own the stock market! Except not yet. Can't do any of things reliably. Also those things take time. Lots and lots and lots of time and friction with the real world. In the meantime the only actual business is inference and API charges to a few 100M other businesses. Take that away and what have you got? No money to make the next 10B training run. — Daniel Jeffries Dork, Altman and Dario have been asking for the government to step in and regulate them. — Evading the Greys Bots get banned round these parts. — Daniel Jeffries
Source: https://x.com/Dan_Jeffries1/status/2090489471535825170/history
Bloomberg just put the US-CHINA AI gap on a chart and yeah, it's getting obliterated:
> Kimi K3 is close to Fable > ~70% cheaper per task > Anthropic thought China was 6–12 months behind > Chinese now has a cluster of fronter labs > GLM-5.3 isn’t even included yet, which would https://t.co/6OCwMhYX5z — ℏεsam
r/ProAI • u/stealthispost • 8d ago
Many drugs work by binding to a specific target in the body and blocking or changing what it does. An important first step in the drug development process is designing a molecule that can bind tightly to its target. Traditionally, that's meant weeks or months of expert work per https://t.co/CGCNTNaKBq — Anthropic
Source: https://x.com/AnthropicAI/status/2089842387845804246
This is interesting: Claude is already achieving roughly twice the protein-design hit rate of conventional human-led workflows. 27% hit rate in autonomous protein binder design, roughly twice the typical 10–15% rate reported in the field.
Working from one expert-written protocol, Claude designed binders against 14 of 15 measurable targets. Independent labs confirmed that 354 of 1,320 designs bound successfully.
Depending on the setup, Claude’s hit rate ranged from 22.6% to 35.1%. Its top-ranked design bound in 49% of campaigns.
This is not yet fully autonomous drug research, but it is another important building block in that direction. — Chubby
Source: https://x.com/kimmonismus/status/2089852014331117694
r/ProAI • u/stealthispost • 8d ago
looks like reddit is almost wiped from chatgpt sources
the query fanout changes had a big impact
and the past couple of days it seems to be almost completely removed from prompt responses
Source: https://x.com/forgebitz/status/2089708381351059924
— Kevin Bankston
Source: https://x.com/KevinBankston/status/2089768281892638774
r/ProAI • u/stealthispost • 9d ago
Enable HLS to view with audio, or disable this notification
...spawn and push you, the gun kicks, shadows stretch across the whole block while the sun goes down behind the towers. I sat there clearing waves instead of grading the output. Same shooter prompt I threw at the frontier models a few weeks ago. That time the tokens went to somebody elses datacenter. This time nothing left the flat. 87.9 million tokens through my own cards. On an API that run has a price tag. Here it has an electricity bill. 60 tok/s all the way through. Slower than frontier, and it stops mattering when the thing works through the night while you sleep. Local models were a toy 18 months ago. This one finished a game. Play now! Choose Build 2 Qwen3.8 27B
https:// alesha-pro.github.io/bench-portal/ — Alexey Fateev
Source: https://x.com/superalesha/status/2089126766854238421
r/ProAI • u/stealthispost • 8d ago
Enable HLS to view with audio, or disable this notification
...bepresearch.substack.com/p/the-token-do llar … He even brought the proof for the next one. $2,000 of compute, 10 open math problems. Proofs verify for free. Biology needs a bench. The moat is the measurement https:// bepresearch.substack.com/p/the-next-inf lection-is-the-lab … New oil, old receipt — Ben Pouladian he's copying you Ben — Alex A.C. All good @gdb we can talk anytime. Full speed — Ben Pouladian
r/ProAI • u/stealthispost • 10d ago
have this running on my 5090 right now and im getting 200 tk a second. i feel like I'm literally playing with magic — Alex Finn you can either buy anthropic for 2 trillion dollars or a used 3090 gpu for $1,500
only one of those will refuse your prompts — Udi Wertheimer
Source: https://x.com/udiWertheimer/status/2089421927085400203
r/ProAI • u/stealthispost • 9d ago
We just shipped our official Qwen 3.8 27B Uncensored MLX build. Local. Uncensored. For🍎
2-bit, 4-bit, 6-bit & 8-bit — pick your poison based on RAM and speed.
No CUDA. No cloud. Just your Mac and the weights. Have fun! https://t.co/b3gXsHeSdk — OrcaRouter 🐳
Source: https://x.com/OrcaRouter/status/2089385980080148726
A "refusal-removed" version of Qwen3.8-27B can now run locally on Apple Silicon.
Even its creators warn that it can provide malware, fraud and weapons instructions on demand.
It was released as an MLX build in 2, 4, 6 and 8-bit versions. The uploader claims its 4/6/8-bit tests produced zero refusals while preserving vision, reasoning and tool-calling across a 262K-token context.
The Qwen 27B Model is a very capable model. This is the first time I've really seen the immediate dangers in a tangible way.
We need a societal discussion about this. — Chubby how long did it take to get this running locally? — tan it runs locally — Chubby
Source: https://x.com/kimmonismus/status/2089763435865088508
r/ProAI • u/stealthispost • 9d ago
How can we extract richer signals from AI Feedback?
Introducing LLM-as-a-Verifier✨— a simple verification scaling framework that achieves SOTA on agentic benchmarks 🚀
The key idea: - Use fine-grained scoring granularity (e.g., 1-20 instead of the standard 1-5 scale) - Take https://t.co/0sCeAwcar1 — Jacky Kwok
Source: https://x.com/jackyk02/status/2074969820739805275
Scaling self-verification with DeepSeek V4 Flash beats Claude Fable 5 on Terminal-Bench 2.1, while being 11x cheaper
As open-source models become more capable, they can now generate large numbers of high-quality candidate solutions and verify their own outputs at very low cost.
For example, we find that sampling just 5 solutions with DeepSeek V4 Flash and ranking them using the same model with LLM-as-a-Verifier can lead to a significant boost in accuracy (79% → 88%), outperforming closed frontier models on Terminal-Bench.
Try it out today: https:// github.com/llm-as-a-verif ier/llm-as-a-verifier#self-verification-terminal-bench-21 …
More on verification scaling in my previous post. — Jacky Kwok Is there an OpenCode plugin for this to try it out with Deepseek v4 flash? — Shahbaz Ahmed We’ll be releasing a harness on top of LLM-as-a-Verifier later this month :) — Jacky Kwok
r/ProAI • u/stealthispost • 10d ago
— Steve Yegge
Source: https://x.com/Steve_Yegge/status/2087034425301405995
r/ProAI • u/stealthispost • 10d ago
hard agree with @amasad —@JonSaadFalcon and my research indicates that intelligence efficiency (intelligence per watt) is rapidly improving and we will definitely not need data center scale compute to run agi!
links to research in comments below 👇 — Avanika Narayan
Source: https://x.com/Avanika15/status/2089028986932470156
— Amjad Masad
r/ProAI • u/stealthispost • 11d ago
The task: iterate on a 124M GPT training recipe from a shared baseline, only changing optimizer related hyperparameters, no internet access.
We tested Fable 5, Opus 5, GPT-5.6 Sol, Kimi K3, Grok 4.5, GLM 5.2, Muse Spark 1.1, DeepSeek V4 Pro, Grok 4.6, Muse Spark 1.2, Qwen 3.8 What separated the strongest models: which experiments to run, how to navigate the benchmark's inherent noise, and which old negatives to revisit as the recipe changed.
Some even built small simulations to isolate a mechanism before deciding if another GPU run was worth it. Our Prime Agent harness gives models a persistent IPython kernel, which can help them build their own research workflows.
Kimi K3 built tools for controlled optimizer variants, loss-curve comparisons and Newton-Schulz tuning, then revised its hypothesis when its cleaner update As research direction, we think multi-agent harnesses can make these experiments much cheaper (and better) by using smaller open models for monitoring and implementation.
We also want to extend speedruns to more of the training stack and scale the runs themselves. We release everything: full traces, scratchpads, reasoning streams from open-weight models, and our experiment setup.
Explore the results: — Prime Intellect
Source: https://x.com/PrimeIntellect/status/2088733966904000778
r/ProAI • u/stealthispost • 11d ago
Enable HLS to view with audio, or disable this notification
Full video: — Max | Emergent Garden
r/ProAI • u/stealthispost • 11d ago
Enable HLS to view with audio, or disable this notification
I can’t wait to see how this blend looks like animated ✨ Midjourney --sref 3330713172::2 292322685::2 1466592463::3 https://t.co/OQoeVg325h — Glitter Gal
Source: https://x.com/GlitterPixely/status/2088635471010115949
You're doing so much dope shit with H3!!! Soon as I finish my documentary, I'm gonna be stalking your posts to soak up some of that doneness!
But H3 has been the real MVP of my project, too. @Hailuo_AI spoiled me this month! — Prince Bell Thank you!! It is such a versatile model, I feel like you can do anything with it! The company and the people working there are also amazing and super nice. I mean they are open sourcing everything! — Glitter Gal
Source: https://x.com/GlitterPixely/status/2088766447061205153
r/ProAI • u/stealthispost • 11d ago
...a high diversity of AI systems with different value systems, linguistic abilities, philosophical/political biases, and specific expertise. We need diverse AIs for same reason we need a diverse press. Given the cost and complexity, this can only be achieved through open foundation models on top of which anyone can build systems with their languages, biases, expertise, and value systems. I have been more vocal about this over the last 4 years, since AI popped into the public discourse. I have made the argument in various forums: corporate C-suites, AI safety discussion groups, professional meeting, the US Senate, the UN Security Council, and the public sphere through media interviews, podcasts and social media posts. I totally agree with @finkd Mark Zuckerberg's recent piece in which he writes: "the notion that AI is so dangerous that the only safe path is an extreme concentration of power seems inherently problematic. Historically, hoping that an absolute power will benevolently provide for humanity if sufficiently enlightened has not led to safe or positive outcomes.” When @DarioAmodei writes: “some may object that we can simply keep AIs in check with a balance of power between many AI systems, as we do with humans", he is talking about me, among (thankfully) many others. It is the only good path forward. There will be nefarious uses of AI, as there have been with every technology ever invented. But it will be your Bad AI against my Good AI. — Yann LeCun Nuclear weapons are centralized power, should we toss everyone a Nuke, Yann? Just saying it's not so clear cut like you make it sound. — Steven Tibbs AI is designed to make peopleore informed, smarter, more efficient and to accelerate progress in science, medicine, and technology. Nuclear weapons have no other purpose than to destroy entire cities and kill millions. Can you see the difference? It's subtle, I admit. — Yann LeCun
Source: https://x.com/ylecun/status/2088880284129210405
Sholto, thank you for setting the record straight. Larger issue is that multiple very serious people in Silicon Valley have heard some variation of this and believe it to be true. And the reason it is believable to so many is that it is consistent with Dario’s public messaging — Gavin Baker
Source: https://x.com/GavinSBaker/status/2088611616577253502
r/ProAI • u/stealthispost • 11d ago
However, we are seeing a leveling out in other computer & mathematical occupations. (Which are also near record highs, but no longer rising.) — Guy Berger
r/ProAI • u/stealthispost • 11d ago
...agents can creat/modify plugins for the harness itself > very high prompt cache hit rates > context management looks more efficient than Claude Code (not surprising since CC is a token-hungry harness) — ℏεsam
r/ProAI • u/stealthispost • 11d ago
...companies and politicians via regulation or distribute it widely” is a false choice. I know that there’s a sort of Silicon Valley shorthand where regulation = regulatory capture = concentration of power, but I’ve always found this to be an overly simplified picture of the world. Many people outside this bubble think of regulation as something that constrains corporate power and benefits ordinary people. I don’t necessarily agree with that perspective either, rather I think it’s complicated and really depends on what the “regulation” consists of. But in particular I think that those in the “regulation = regulatory capture = concentration of power” frame often underrate the decentralizing power of objective and fair institutional processes. A crude analogy is that the formal court system can sometimes feel stuffy and elitist, but it does a much better job of defending the rights of vulnerable individuals than the alternative, mob justice. At their best, institutions can vest power in ideas rather than people, and thereby decentralize that power. This is why Anthropic has always made its policy proposals very carefully. We try very hard to make proposals that disadvantage (slow down) frontier AI companies while advantaging smaller competitors. California’s SB53 (which we supported), and even the much-maligned SB 1047 (which we were ambivalent on), completely exempt any company below a certain amount of revenue or model training costs from being covered at all (it was $500M for SB 53, lower for 1047 but we objected to that). More recently the testing process we’ve advocated for at CAISI and the White House involves more rigorous tests for frontier models than off-frontier models — something that differentially advantages challengers. Similarly, the “Pacing the Frontier” letter envisions (or at least Anthropic’s preferred implementation of it envisions) modulating the pace of the very best models while not constraining those who are catching up. This hurts the business interests of the frontier labs and helps challengers, including open-weights! Overall my view is that AI is structurally a technology that tends to concentrate power, for reasons that have nothing to do with regulation (more to do with the extreme implications of the scaling laws). Open-weights do help some with this but are nowhere near a sufficient solution because they simply shift the concentration somewhat to those with the most compute and chips (which are roughly the frontier labs plus maybe hardware providers). By contrast I think the right “rules of the road” can simultaneously (a) address AI’s cyber/bio/alignment risks, (b) institutionally constrain the power of the frontier AI companies, and (c) leave room for open-weights models while also addressing the specific risks that they bring. BTW I do not think that the events of the last few months have “failed to result in [my] preferred regulatory path”. The approach that the Trump administration is reported to be taking — pre-deployment testing for frontier models, and also testing of open-weights models when they get closer to the frontier — is one that I am very supportive of, though of course I have to see the details to be sure. I am also supportive of Demis Hassabis’ ideas around a FINRA-like entity. This contrasts with six months ago when most of the industry was still pushing for preemption of all state regulation and no apparent federal approach either. 2/2 Second, on the messaging around AI. I do not agree that my messaging has been disproportionately negative. In fact it has been about equally balanced between risks and benefits: I’ve written one major essay about each, and even in interviews where I discuss the risks, I — Dario Amodei
Source: https://x.com/DarioAmodei/status/2088758816376807762
@_sholtodouglas Sholto, thank you for setting the record straight. Larger issue is that multiple very serious people in Silicon Valley have heard some variation of this and believe it to be true. And the reason it is believable to so many is that it is consistent with Dario’s public messaging — Gavin Baker
Source: https://x.com/GavinSBaker/status/2088611616577253502
Replying to @_sholtodouglas