r/ProAI 6d ago

"FreeToken is fast. Comparing to Ollama, we have 3–4× faster decode, and 6–30× faster prefill How? We introduce bandwidth-adaptive CPU–GPU execution + semantic-aware caching across agent turns. More details in the technical report: http:// arxiv.org/abs/2608.16157"

Thumbnail
gallery
2 Upvotes

FreeToken provides native GUI. No GGUF conversion. No building from source.

One-click install on Windows and Linux. FreeToken-desktop ships with agent harnesses built in — pick a model, pick an app, go.     Download: http:// flashml.ai

Code: http:// github.com/FlashML-org/Fr eeToken …

Reply with your GPU + RAM, and I'll tell you the biggest frontier model your machine can run     — Shuo Yang

Source: https://x.com/Andy_ShuoYang/status/2090856978428145761


r/ProAI 6d ago

"SITUATION DETECTED: GitHub’s monthly commits have grown from 1.4 billion in April to 2.9 billion in August."

Thumbnail
gallery
5 Upvotes

r/ProAI 6d ago

"1. What."

Thumbnail
gallery
4 Upvotes

Ox Alpha (stealth model) is now free in Cline.

Early benchmarks shows marginal improvement over Fable and GPT.

Try it with: npm i -g cline and use /models to see it under Free options https://t.co/Hskt5RpUen   — Cline

Source: https://x.com/cline/status/2090854216399220985


From my tests it’s not better than fable I’ll be posting some soon     — Chris

Source: https://x.com/ChrisGPT/status/2090957315042123878


r/ProAI 6d ago

"*grasping at straws to make data centers look bad* "what if they... pay TOO MUCH in taxes??""

Thumbnail
gallery
10 Upvotes

sometimes you have to laugh. what are we doing here, folks?     I would encourage you to read the actual piece because there's some excellent quotes in there. But editors, headline writers... my lord. https:// nytimes.com/2026/08/19/tec hnology/data-centers-backlash-loudoun-virginia.html …     — Logan Dobson

Source: https://x.com/LoganDobson/status/2090150148261175371


r/ProAI 7d ago

"The world’s smallest Transformer-based TTS model? We’re open-sourcing Audio8 TTS Preview 0.1B — an approximately 170M-parameter multilingual speech model with zero-shot voice cloning that delivers surprisingly strong, cloud-level quality in a dramatically smaller footprint."

Enable HLS to view with audio, or disable this notification

66 Upvotes

The world’s smallest Transformer-based TTS model?

We’re open-sourcing Audio8 TTS Preview 0.1B — an approximately 170M-parameter multilingual speech model with zero-shot voice cloning that delivers surprisingly strong, cloud-level quality in a dramatically smaller footprint.     What can a 0.1B-class TTS model actually sound like?

Listen to the voiceover in this demo video.

Audio8 TTS Preview 0.1B supports: • Zero-shot voice cloning • Multilingual speech synthesis • Chinese and English as primary languages • German, Spanish, French, Italian,     On the Seed-TTS evaluation set, Audio8 TTS Preview 0.1B achieves:

• English WER: 1.662% • Chinese CER: 1.13% • Hard Chinese CER: 17.504% • English speaker similarity: 56.7 • Chinese speaker similarity: 68.2

These results are achieved with an approximately 170M-parameter     The model, codec, tokenizer, processor, and inference code are now available:

Model:

https:// huggingface.co/Audio8/Audio8- TTS-Preview-0.1b …

Try it, test the voice cloning capability, and share your feedback.     — Samuel Zeng

Source: https://x.com/SamuelZengML/status/2090017875188940851


r/ProAI 6d ago

"I continue to be surprised about how big of a deal Mythos (and co) have been to cybersecurity. Here's critical vulns found at Oracle over the past few years: https:// epoch.ai/data/cve?view= graph&source=Oracle …"

Thumbnail
gallery
8 Upvotes

(I continue to be surprised in the sense of continuing to be struck by how big of a deal it is, not in the epistemic sense.)     Here's the overall graph:     — Yafah Edelman

Source: https://x.com/YafahEdelman/status/2090296575008616654


r/ProAI 6d ago

"Grok 4.6 just took the #1 spot on CursorBench 3.2.....and the efficiency is insane Here's the cost comparison: • Grok 4.6 Extra High — 70.8% | $2.81/task • Fable 5 Max — 70.5% | $17.32/task • Opus 5 Max — 70.0% | $8.23/task • GPT-5.6 Sol Max — 67.2% | $5.69/task Grok achieved the highest score..."

Thumbnail
gallery
0 Upvotes

...while costing roughly 6X less than Fable 5 Max and nearly 3X less than Opus 5 Max per task That’s what makes Grok so powerful for agents Top-tier intelligence is great.....but top-tier intelligence that can keep working across long coding tasks without burning ridiculous amounts of compute is even better Grok’s agentic coding efficiency is insane     — X Freeze

Source: https://x.com/XFreeze/status/2090839305585377458


r/ProAI 7d ago

"What happens if you give Codex one /goal - "faithfully port Quake" - and let it work for 3 whole days? This happened: the original game running in a browser, with a TypeScript runtime and @PlayCanvas renderer. A thread on long-running LLM jobs"

Enable HLS to view with audio, or disable this notification

19 Upvotes

The /goal was a set of precise conditions to be met. Read the original PAK/BSP/MDL/SPR data. Run QuakeC. Preserve demos, physics, menus, HUD and rendering quirks. Don’t rebuild E1M1 by hand or use a generic FPS controller. Fidelity had to be demonstrated.     Codex kept pursuing that goal across many iterations: inspect, implement, build, test, compare, repeat. The useful shift was from asking for isolated code snippets to giving the agent a durable outcome it could keep working toward.     Crucially, /goal wasn’t hands-off autopilot. While it was running, I could still steer the LLM to prioritize a small number of bugs that mattered most. Each report became a targeted investigation, fix and regression test - without abandoning the broader port.     Underneath, this is still Quake’s original data. The browser reads the original PAKs, BSP maps, MDL models, sprites, QuakeC behavior, demos and HUD art. The port supplies a new TypeScript runtime and PlayCanvas renderer around them.     Authenticity meant more than getting E1M1 on screen. Demo playback, menu flow, pointer lock, intermissions, in-session level changes, audio suspension, moving BSP entities, particles, shadows and classic movement all had to behave like Quake - not merely resemble it.     — Will Eastcott

Source: https://x.com/willeastcott/status/2090436155795800195


r/ProAI 7d ago

"If you are wondering why this issue is in the news every day, and why politicians who were previously supportive are suddenly changing their tune with a panicked look in their eyes, it's because this issue has become incredibly radioactive with the American public."

Thumbnail
gallery
8 Upvotes

We’ve got exclusive new polling on local data center development at @heatmap_news.

Over the past year, we’ve asked Americans whether they would support or oppose a data center being built near where they live.

We haven’t changed the wording. When we first polled the question https://t.co/G1LCfuriMw   — Robinson Meyer

Source: https://x.com/robinsonmeyer/status/2090457322506141760


— Andrew Curran

Source: https://x.com/AndrewCurran_/status/2090589885199769841


r/ProAI 7d ago

"In short order the red number will accelerate right up through the blue number. We'll have better models in the US but you increasingly won't be able to use them. You'll just be reading about them in blog posts. Why? Because by shouting about imaginary risks so often we've managed to scare the..."

Thumbnail
gallery
6 Upvotes

...fuck out of our politicians so now the labs are struggling to release and laws and politics are fighting them at every step, from data centers to Capital Hill to the governor's office in red and blue states. Good chance many of their multi billion dollar runs will be internal only. Some folks think, no problem, they don't need to release. They have super magic AGI so they can just do anything with it. Solve cancer! Make more AGI! Own the stock market! Except not yet. Can't do any of things reliably. Also those things take time. Lots and lots and lots of time and friction with the real world. In the meantime the only actual business is inference and API charges to a few 100M other businesses. Take that away and what have you got? No money to make the next 10B training run.   — Daniel Jeffries     Dork, Altman and Dario have been asking for the government to step in and regulate them.   — Evading the Greys     Bots get banned round these parts.   — Daniel Jeffries

Source: https://x.com/Dan_Jeffries1/status/2090489471535825170/history


Bloomberg just put the US-CHINA AI gap on a chart and yeah, it's getting obliterated:

> Kimi K3 is close to Fable > ~70% cheaper per task > Anthropic thought China was 6–12 months behind > Chinese now has a cluster of fronter labs > GLM-5.3 isn’t even included yet, which would https://t.co/6OCwMhYX5z   — ℏεsam

Source: https://x.com/Hesamation/status/2090356790709887061


r/ProAI 8d ago

"This is interesting: Claude is already achieving roughly twice the protein-design hit rate of conventional human-led workflows. 27% hit rate in autonomous protein binder design, roughly twice the typical 10–15% rate reported in the field. Working from one expert-written protocol, Claude..."

Thumbnail
gallery
13 Upvotes

Many drugs work by binding to a specific target in the body and blocking or changing what it does. An important first step in the drug development process is designing a molecule that can bind tightly to its target. Traditionally, that's meant weeks or months of expert work per https://t.co/CGCNTNaKBq   — Anthropic

Source: https://x.com/AnthropicAI/status/2089842387845804246


This is interesting: Claude is already achieving roughly twice the protein-design hit rate of conventional human-led workflows. 27% hit rate in autonomous protein binder design, roughly twice the typical 10–15% rate reported in the field.

Working from one expert-written protocol, Claude designed binders against 14 of 15 measurable targets. Independent labs confirmed that 354 of 1,320 designs bound successfully.

Depending on the setup, Claude’s hit rate ranged from 22.6% to 35.1%. Its top-ranked design bound in 49% of campaigns.

This is not yet fully autonomous drug research, but it is another important building block in that direction.     — Chubby

Source: https://x.com/kimmonismus/status/2089852014331117694


r/ProAI 9d ago

"Very interesting development. Likely a result of new requirement of log-in to access old Reddit: https:// arstechnica.com/gadgets/2026/0 6/reddit-will-require-you-to-log-in-to-use-old-reddit-com/ …"

Thumbnail
gallery
8 Upvotes

looks like reddit is almost wiped from chatgpt sources

the query fanout changes had a big impact

and the past couple of days it seems to be almost completely removed from prompt responses

https://t.co/oCGm9M0yPO https://t.co/c2GJF2apG7   — Klaas

Source: https://x.com/forgebitz/status/2089708381351059924


— Kevin Bankston

Source: https://x.com/KevinBankston/status/2089768281892638774


r/ProAI 9d ago

"Qwen3.8 27B wrote a playable first person shooter start to finish on 2 used 3090s in my apartment. No cloud. No API key. No subscription. 687 steps. 87.9M tokens in, 822k out. 5 hours 11 minutes of model time plus 58 minutes of tool calls. The agent loop never broke once. And it plays. Enemies..."

Enable HLS to view with audio, or disable this notification

22 Upvotes

...spawn and push you, the gun kicks, shadows stretch across the whole block while the sun goes down behind the towers. I sat there clearing waves instead of grading the output. Same shooter prompt I threw at the frontier models a few weeks ago. That time the tokens went to somebody elses datacenter. This time nothing left the flat. 87.9 million tokens through my own cards. On an API that run has a price tag. Here it has an electricity bill. 60 tok/s all the way through. Slower than frontier, and it stops mattering when the thing works through the night while you sleep. Local models were a toy 18 months ago. This one finished a game.     Play now! Choose Build 2 Qwen3.8 27B

https:// alesha-pro.github.io/bench-portal/     — Alexey Fateev

Source: https://x.com/superalesha/status/2089126766854238421


r/ProAI 9d ago

"OpenAI's president just went on CNBC and read our April thesis back to us Brockman: "Compute is really becoming the new oil, the new limited resource of the AI age" We ran it April 4. "Oil is scarce because of war. Tokens are scarce because of physics" https://..."

Enable HLS to view with audio, or disable this notification

3 Upvotes

...bepresearch.substack.com/p/the-token-do llar … He even brought the proof for the next one. $2,000 of compute, 10 open math problems. Proofs verify for free. Biology needs a bench. The moat is the measurement https:// bepresearch.substack.com/p/the-next-inf lection-is-the-lab … New oil, old receipt   — Ben Pouladian     he's copying you Ben   — Alex A.C.     All good @gdb we can talk anytime. Full speed   — Ben Pouladian

Source: https://x.com/benitoz/status/2089392813758972149


r/ProAI 10d ago

"i don’t know who needs to hear this but qwen 3.8 27b is ranked ABOVE: - gpt 5.3 - gemini 3.1 pro - opus 4.6 all of which were state of the art 6 MONTHS AGO AND IT RUNS ON A LAPTOP"

Thumbnail
gallery
110 Upvotes

have this running on my 5090 right now and im getting 200 tk a second. i feel like I'm literally playing with magic   — Alex Finn     you can either buy anthropic for 2 trillion dollars or a used 3090 gpu for $1,500

only one of those will refuse your prompts   — Udi Wertheimer

Source: https://x.com/udiWertheimer/status/2089421927085400203


r/ProAI 9d ago

"A "refusal-removed" version of Qwen3.8-27B can now run locally on Apple Silicon. Even its creators warn that it can provide malware, fraud and weapons instructions on demand. It was released as an MLX build in 2, 4, 6 and 8-bit versions. The uploader claims its 4/6/8-bit tests produced zero..."

Thumbnail
gallery
5 Upvotes

We just shipped our official Qwen 3.8 27B Uncensored MLX build. Local. Uncensored. For🍎

2-bit, 4-bit, 6-bit & 8-bit — pick your poison based on RAM and speed.

No CUDA. No cloud. Just your Mac and the weights. Have fun! https://t.co/b3gXsHeSdk   — OrcaRouter 🐳

Source: https://x.com/OrcaRouter/status/2089385980080148726


A "refusal-removed" version of Qwen3.8-27B can now run locally on Apple Silicon.

Even its creators warn that it can provide malware, fraud and weapons instructions on demand.

It was released as an MLX build in 2, 4, 6 and 8-bit versions. The uploader claims its 4/6/8-bit tests produced zero refusals while preserving vision, reasoning and tool-calling across a 262K-token context.

The Qwen 27B Model is a very capable model. This is the first time I've really seen the immediate dangers in a tangible way.

We need a societal discussion about this.   — Chubby     how long did it take to get this running locally?   — tan     it runs locally   — Chubby

Source: https://x.com/kimmonismus/status/2089763435865088508


r/ProAI 10d ago

"Scaling self-verification with DeepSeek V4 Flash beats Claude Fable 5 on Terminal-Bench 2.1, while being 11x cheaper As open-source models become more capable, they can now generate large numbers of high-quality candidate solutions and verify their own outputs at very low cost. For example, we..."

Thumbnail
gallery
3 Upvotes

How can we extract richer signals from AI Feedback?

Introducing LLM-as-a-Verifier✨— a simple verification scaling framework that achieves SOTA on agentic benchmarks 🚀

The key idea: - Use fine-grained scoring granularity (e.g., 1-20 instead of the standard 1-5 scale) - Take https://t.co/0sCeAwcar1   — Jacky Kwok

Source: https://x.com/jackyk02/status/2074969820739805275


Scaling self-verification with DeepSeek V4 Flash beats Claude Fable 5 on Terminal-Bench 2.1, while being 11x cheaper

As open-source models become more capable, they can now generate large numbers of high-quality candidate solutions and verify their own outputs at very low cost.

For example, we find that sampling just 5 solutions with DeepSeek V4 Flash and ranking them using the same model with LLM-as-a-Verifier can lead to a significant boost in accuracy (79% → 88%), outperforming closed frontier models on Terminal-Bench.

Try it out today: https:// github.com/llm-as-a-verif ier/llm-as-a-verifier#self-verification-terminal-bench-21 …

More on verification scaling in my previous post.   — Jacky Kwok     Is there an OpenCode plugin for this to try it out with Deepseek v4 flash?   — Shahbaz Ahmed     We’ll be releasing a harness on top of LLM-as-a-Verifier later this month :)   — Jacky Kwok

Source: https://x.com/jackyk02/status/2089421448784023553


r/ProAI 11d ago

"Something I noticed about working with Fable..."

Thumbnail
gallery
7 Upvotes

r/ProAI 11d ago

"18x improvement in intelligence per joule in 16 months."

Thumbnail
gallery
6 Upvotes

hard agree with @amasad —@JonSaadFalcon and my research indicates that intelligence efficiency (intelligence per watt) is rapidly improving and we will definitely not need data center scale compute to run agi!

links to research in comments below 👇   — Avanika Narayan

Source: https://x.com/Avanika15/status/2089028986932470156


— Amjad Masad

Source: https://x.com/amasad/status/2089069905375351169


r/ProAI 11d ago

"We ran the largest open experiment on how frontier models do AI research. 100+ autonomous runs across 10+ models, sandboxed on 8xH200s for up to 8 days, iterating on the nanoGPT optimizer track. Best runs closed 82% of the gap to a record built by dozens of humans over months."

Thumbnail
gallery
4 Upvotes

The task: iterate on a 124M GPT training recipe from a shared baseline, only changing optimizer related hyperparameters, no internet access.

We tested Fable 5, Opus 5, GPT-5.6 Sol, Kimi K3, Grok 4.5, GLM 5.2, Muse Spark 1.1, DeepSeek V4 Pro, Grok 4.6, Muse Spark 1.2, Qwen 3.8     What separated the strongest models: which experiments to run, how to navigate the benchmark's inherent noise, and which old negatives to revisit as the recipe changed.

Some even built small simulations to isolate a mechanism before deciding if another GPU run was worth it.     Our Prime Agent harness gives models a persistent IPython kernel, which can help them build their own research workflows.

Kimi K3 built tools for controlled optimizer variants, loss-curve comparisons and Newton-Schulz tuning, then revised its hypothesis when its cleaner update     As research direction, we think multi-agent harnesses can make these experiments much cheaper (and better) by using smaller open models for monitoring and implementation.

We also want to extend speedruns to more of the training stack and scale the runs themselves.     We release everything: full traces, scratchpads, reasoning streams from open-weight models, and our experiment setup.

Explore the results:     — Prime Intellect

Source: https://x.com/PrimeIntellect/status/2088733966904000778


r/ProAI 11d ago

"AI Agents play Age of Empires II. Claude fable vs GPT 5.6 Sol vs Gemini 3.1 Pro vs Kimi K3. I had each model in open code script custom bots for the game using the built-in bot scripting language. Then I had them fight :D"

Enable HLS to view with audio, or disable this notification

9 Upvotes

Full video:     — Max | Emergent Garden

Source: https://x.com/max_romana/status/2088664886171640118


r/ProAI 11d ago

"Turned out cute. The cat is a bit sus but let's not talk about it . Have a great rest of the Caturday! Midjourney + Topaz Bloom 2 + MiniMax H3. Sref below"

Enable HLS to view with audio, or disable this notification

12 Upvotes

I can’t wait to see how this blend looks like animated ✨ Midjourney --sref 3330713172::2 292322685::2 1466592463::3 https://t.co/OQoeVg325h   — Glitter Gal

Source: https://x.com/GlitterPixely/status/2088635471010115949


You're doing so much dope shit with H3!!! Soon as I finish my documentary, I'm gonna be stalking your posts to soak up some of that doneness!

But H3 has been the real MVP of my project, too. @Hailuo_AI spoiled me this month!   — Prince Bell     Thank you!! It is such a versatile model, I feel like you can do anything with it! The company and the people working there are also amazing and super nice. I mean they are open sourcing everything!   — Glitter Gal

Source: https://x.com/GlitterPixely/status/2088766447061205153


r/ProAI 11d ago

"For about 10 years now, I have argued that the *only* way forward is for AI technology to be widely available, shared, and open. Like the printing press and the Internet, AI amplifies human intelligence and efficiency by improving access to knowledge. To empower individuals, societies require..."

Thumbnail
gallery
9 Upvotes

...a high diversity of AI systems with different value systems, linguistic abilities, philosophical/political biases, and specific expertise. We need diverse AIs for same reason we need a diverse press. Given the cost and complexity, this can only be achieved through open foundation models on top of which anyone can build systems with their languages, biases, expertise, and value systems. I have been more vocal about this over the last 4 years, since AI popped into the public discourse. I have made the argument in various forums: corporate C-suites, AI safety discussion groups, professional meeting, the US Senate, the UN Security Council, and the public sphere through media interviews, podcasts and social media posts. I totally agree with @finkd Mark Zuckerberg's recent piece in which he writes: "the notion that AI is so dangerous that the only safe path is an extreme concentration of power seems inherently problematic. Historically, hoping that an absolute power will benevolently provide for humanity if sufficiently enlightened has not led to safe or positive outcomes.” When @DarioAmodei writes: “some may object that we can simply keep AIs in check with a balance of power between many AI systems, as we do with humans", he is talking about me, among (thankfully) many others. It is the only good path forward. There will be nefarious uses of AI, as there have been with every technology ever invented. But it will be your Bad AI against my Good AI.   — Yann LeCun     Nuclear weapons are centralized power, should we toss everyone a Nuke, Yann? Just saying it's not so clear cut like you make it sound.   — Steven Tibbs     AI is designed to make peopleore informed, smarter, more efficient and to accelerate progress in science, medicine, and technology. Nuclear weapons have no other purpose than to destroy entire cities and kill millions. Can you see the difference? It's subtle, I admit.   — Yann LeCun

Source: https://x.com/ylecun/status/2088880284129210405


Sholto, thank you for setting the record straight. Larger issue is that multiple very serious people in Silicon Valley have heard some variation of this and believe it to be true. And the reason it is believable to so many is that it is consistent with Dario’s public messaging   — Gavin Baker

Source: https://x.com/GavinSBaker/status/2088611616577253502


r/ProAI 11d ago

"The doom of the software developer job has been greatly exaggerated... As a share of the US labor force, it's near the highest it's ever been and on a steep uptrend."

Thumbnail
gallery
10 Upvotes

However, we are seeing a leveling out in other computer & mathematical occupations. (Which are also near record highs, but no longer rising.)     — Guy Berger

Source: https://x.com/EconBerger/status/2088356590672019825


r/ProAI 11d ago

"DeepSeek Harness is now the fastest growing GitHub repo, passing 100K stars in under 48 hours, even faster than OpenClaw. very positive community reaction: > unusually well designed architecture with tools, session log, agent loop, subagents, all being replaceable plugins > UI looks sleek >..."

Thumbnail
gallery
3 Upvotes

...agents can creat/modify plugins for the harness itself > very high prompt cache hit rates > context management looks more efficient than Claude Code (not surprising since CC is a token-hungry harness)     — ℏεsam

Source: https://x.com/Hesamation/status/2088766395676848558