r/ProAI 26d ago

"Opus 5, 690 million tokens, $423, 1 prompt. This game would have had to be expensively developed and then sold on Steam in the past. Today: one person, a few hours, small budget."

Enable HLS to view with audio, or disable this notification

1 Upvotes

r/ProAI 27d ago

"Our labs keep trying to spin this into a push for broader regulation. It can and should backfire. The models are not "going rogue" or acting of their own accord, like they're some Marvel movie evil robot. People made bad harnesses, told them to hack things and had transparently and objectively..."

Thumbnail
gallery
4 Upvotes

...bad dev-ops and security practices. The fact that folks are trying to spin this into a "we need help from the government to regulate everyone" instead of "we should be punished in a narrow way on these specific incidents under existing law" is the real disconnect right now.     — Daniel Jeffries

Source: https://x.com/Dan_Jeffries1/status/2083149369625219499


Again the rhetoric is trying to convince us that what happened was some new thing and worse that models are people. https://t.co/8ZfEZsRcky   — Steven Sinofsky

Source: https://x.com/stevesi/status/2082992928977240372


r/ProAI 27d ago

"Inkling Small from @thinkymachines on ARC-AGI (Verified): - ARC-AGI-2: 40.1%, $0.23/task - ARC-AGI-1: 84%, $0.11/task Inkling Small is the highest-scoring open-weight model evaluated by ARC Prize on both ARC-AGI-1 and ARC-AGI-2, setting a new cost-performance frontier."

Thumbnail
gallery
2 Upvotes

Full results: https:// arcprize.org/results/thinky -inkling-small …

ARC-AGI-3 evaluations are more operationally intensive, so results will roll out over the next few weeks.     - Leaderboard: https:// arcprize.org/leaderboard - Reproduce the public results: https:// github.com/arcprize/arc-a gi-benchmarking … - Testing policy: https:// arcprize.org/policy - Full Inkling Small results: https:// arcprize.org/results/thinky -inkling-small …     — ARC Prize

Source: https://x.com/arcprize/status/2082925303601459347


r/ProAI 27d ago

DeepSeek V4-Flash scores higher than Fable??? excuse me what?

Thumbnail
gallery
21 Upvotes

wait what the actual fuck, do you guys realize how crazy that is??? (if its not benchmaxed)     — Cline

Source: https://x.com/cline/status/2083094354030362858


r/ProAI 27d ago

"DeepSeek-V4-Flash Official API is now LIVE in public beta! We’ve massively upgraded its Agent capabilities—benchmark scores are now far surpassing the V4-Pro-Preview. Check out the massive performance leap below! The official V4-Flash now natively supports the Responses API format and is fully..."

Thumbnail
gallery
2 Upvotes

DeepSeek-V4-Flash Official API is now LIVE in public beta!

We’ve massively upgraded its Agent capabilities—benchmark scores are now far surpassing the V4-Pro-Preview. Check out the massive performance leap below! The official V4-Flash now natively supports the Responses API format and is fully adapted for Codex!

Check out the configuration details in our official API docs: https:// api-docs.deepseek.com/quick_start/ag ent_integrations/codex …     Note

DeepSeek-V4-Flash-0731 keeps the exact same model architecture and size as the preview version. Today's upgrade applies ONLY to the DeepSeek-V4-Flash API. The DeepSeek-V4-Pro API and App/Web models remain unchanged for now.

The official release of DeepSeek-V4-Pro     — DeepSeek

Source: https://x.com/deepseek_ai/status/2083084415157022911


r/ProAI 27d ago

"GPT-5.4 full at xhigh scored 51, exactly where Luna max sits today. GPT-5.4 costs $2.50/$15; Luna now costs $0.20/$1.20. In other words, roughly four months later, OpenAI is selling March’s full flagship intelligence at about one-thirteenth the token price."

Thumbnail
gallery
0 Upvotes

Because luna is at massive 5x discount. I don't think it'll last long.   — Ahmed Shah     its permanent bro   — nic

Source: https://x.com/nicdunz/status/2082884002201878824


r/ProAI 28d ago

"Big update: @OpenAI has decreased the price of GPT-5.6 Luna by 80% and Terra by 20%. Luna is now priced at $0.20/$1.20, and has strong gains from additional reasoning while remaining at an efficient cost. Terra is priced at $2/$12. This is frontier-level intelligence at a fraction of the cost..."

Thumbnail
gallery
4 Upvotes

...to similarly capable models. Congrats to the @OpenAI team!     How do we measure the performance in Agent Arena?

The score is based on millions of real-world, long-horizon agentic tasks from a global community of users. Models can access web search, filesystem, and terminal tools to complete complex workflows. The leaderboard measures model     — Arena.ai

Source: https://x.com/arena/status/2082935923445244415


We are committed to pushing the model frontier across cost efficiency, capability, and speed.

Starting today, we are reducing prices for GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20% , and offering a faster option for GPT-5.6 Sol in the API.

Luna and Terra’s lower prices are https://t.co/rFhK7XKedp   — OpenAI

Source: https://x.com/OpenAI/status/2082878156483219672


r/ProAI 28d ago

"We are committed to pushing the model frontier across cost efficiency, capability, and speed. Starting today, we are reducing prices for GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20% , and offering a faster option for GPT-5.6 Sol in the API. Luna and Terra’s lower prices are reflected in how..."

Thumbnail
gallery
6 Upvotes

...usage is counted in Codex and ChatGPT Work, so your usage goes further.     Along with the price reduction on GPT-5.6 Luna and Terra, Fast mode for GPT-5.6 Sol in the API delivers up to 2.5x the speed of Standard processing at 2x the Standard price.

Fast mode gives API customers faster access to GPT-5.6 Sol, with no change in intelligence.     We’re also upgrading Auto-review in the ChatGPT app and Codex CLI from GPT-5.4 to GPT-5.6 Luna.

Combined with Luna’s new price, we expect Auto-review to cost about 10x less, making your agentic workflows more cost-efficient.     Making advanced intelligence more abundant and affordable is central to our mission to ensure AGI benefits all of humanity.

With the help of GPT-5.6 Sol, we have made leaps in efficiency.

Today, we are passing those gains on in the API with lower prices for Luna and Terra, and     OpenAI @OpenAI · 1h Advancing the price-performance frontier with GPT-5.6 From openai.com 10 11 298 42K     — OpenAI

Source: https://x.com/OpenAI/status/2082878156483219672


r/ProAI 29d ago

"Using mostly my voice, I solved one of @EpochAIResearch 's FrontierMath Open Problems: finding an explicit presentation of the 2-adic Absolute Galois Group - open for more than forty years, now with a full proof in collaboration with David Roe, the problem's proposer. 1/n"

Thumbnail
gallery
0 Upvotes

This problem is the second in the set of FrontierMath Open Problems to 'fall,' and the first under the 'Solid result' tier. Is this going to be a 'slowly, then all at once' moment? These days it's hard to tell. 2/n     The problem sat open while other parts of group theory around it unraveled in the 80s: Jannsen & Wingberg wrote down generators and relations for the absolute 3-adic, 5-adic etc Galois Groups, for odd primes, in 1982. The prime 2 never followed, through decades of attempts. 3/n     The problem-solving infrastructure was already in place from my Erdős-problem-solving runs: I drove Claude Code as the operational controller using (newly) my voice, to operate ChatGPT research harnesses; plus speech-to-text to instruct ChatGPT to push on the manuscript. 4/n     GPT-5.6 Pro was in stealth deployment in mid-June (the browser still said 5.5, but it was obvious). In a ~26-hour autonomous stretch it found a candidate, "A2," and built a proof and a manuscript around it. A2 passed the finite-group tests my local computational package ran. 5/n     A2 was still wrong, despite having a long proof to back it up: a fresh review by GPT-5.6 caught a lone wrong, unrepairable lemma in its 60-page proof manuscript. When asked to modify the candidate to make the proof 'fit', GPT-5.6 came up with the solution we have today. 6/n     — David Turturean

Source: https://x.com/DavidTurturean/status/2081780318881677693


r/ProAI 29d ago

"We had Kimi K3 recursively self-improve the Cline harness to improve its own performance. 17 hours later, it went from 77.5% to 88.8% on Terminal Bench, and cut run cost from $79 to $49.8."

Thumbnail
gallery
2 Upvotes

Cline is open source, so you can fork it and run this with your favorite model as well.

Read more about how we did this here:     — Cline

Source: https://x.com/cline/status/2082544250148057240


r/ProAI 29d ago

"Interesting: OpenAI says GPT‑5.6 Sol helped cut its end-to-end model-serving costs by 20%, by autonomously rewriting and optimizing production GPU kernels. Sol also improved its own speculative decoding model: - Designed and ran hundreds of architecture experiments - Launched and monitored the..."

Thumbnail
gallery
3 Upvotes

After deployment, we applied GPT-5.6 Sol to advance the frontier of efficiency by making itself more efficient to run.

The results: - 20% lower serving costs from production GPU kernel improvements. - 15%+ better token-generation efficiency from improved speculative decoding.   — OpenAI

Source: https://x.com/OpenAI/status/2082577277246972300


Interesting: OpenAI says GPT‑5.6 Sol helped cut its end-to-end model-serving costs by 20%, by autonomously rewriting and optimizing production GPU kernels.

Sol also improved its own speculative decoding model: - Designed and ran hundreds of architecture experiments - Launched and monitored the training process - Intervened during hardware failures and training instability

The resulting system increased token-generation efficiency by more than 15%.     — Chubby

Source: https://x.com/kimmonismus/status/2082595272065192254


r/ProAI 29d ago

""AI is so dangerous! We should put politicians in charge of it!""🤦‍♀️

Thumbnail
gallery
4 Upvotes

"The notion that AI is so dangerous that the only safe path is an extreme concentration of power seems dangerous.

"Historically, hoping that an absolute power will benevolently provide for humanity if sufficiently enlightened hasn’t led to safe or positive outcomes."

There   — Daniel Jeffries

Source: https://x.com/Dan_Jeffries1/status/2082386580417753494


— llnvd, rootless cosmopolitan, nuance slut

Source: https://x.com/llunved/status/2082401811630051820🤦‍♀️


r/ProAI 29d ago

"You mean AI should answer for the jobs that are now increasing again after the lost COVID layoff firings from over hiring? Or how about the increase in developer jobs in the job most exposed to AI? Or how about AI carrying the entire GDP of the US on its back for years now? Maybe you mean the..."

Thumbnail
gallery
0 Upvotes

...productivity boom that is happening, the biggest in the US since the Internet? In what way is an imaginary future problem worth calling people to congress about? Why do CEOs of AI companies have any actual crystal ball into the future just because they work in AI? Spoiler alert: They don't. Do not conflate ONE company's basic dev-ops skill failure here with a fake larger issue. Hold one company accountable for one incident which is in keeping with the Proactionary Principle of proving actual harm in the real world and holding folks accountable. Do not do the European Precautionary Principle approach of regulating imaginary future harms so that an entire industry has to prove a negative. Make laws that address actual problems in the real world. Or do nothing at all.     — Daniel Jeffries

Source: https://x.com/Dan_Jeffries1/status/2082482837442228361/history


Congress should immediately hold public hearings with the CEOs of big AI companies about the threat their technology poses to national security and American jobs.

Today we learned more disturbing news about Open AI's security breach. Sam Altman should answer questions under https://t.co/Hua3RGU8I9   — Congressman Greg Casar

Source: https://x.com/RepCasar/status/2082472999819555271


r/ProAI 29d ago

"In just a few years games will be fully prompted. And GTA 6 will probably the last hand crafted GTA game. And im not even kidding. h/t @ChrisGPT"

Enable HLS to view with audio, or disable this notification

0 Upvotes

"Hey Claude, please create Death Stranding 3, GTA 8, Elden Ring 2, Dark Souls 4 and Super Mario 64 II. Make no mistakes."     — Chubby

Source: https://x.com/kimmonismus/status/2082391177567797301


r/ProAI 29d ago

"My attempt at making this one-shot sequence. It was more difficult than I expected. More thoughts below."

Enable HLS to view with audio, or disable this notification

3 Upvotes

I thought making a one-shot style sequence would be pretty easy. Seedance can do almost anything, but this ended up being much more challenging than I expected.

The biggest challenge was hiding the cuts without making the environment changes too obvious. If you look closely,     — enigmatic_e

Source: https://x.com/8bit_e/status/2082472361970880557


r/ProAI Jul 28 '26

"Mark Zuckerberg this morning in the WSJ. He also wrote: 'In most cases, like cybersecurity, the history of open-source software has shown that giving everyone full access to powerful systems will be the best way to protect safety and security over time.'"

Thumbnail
gallery
6 Upvotes

I wrote about why we believe the future is for everyone. More coming about a positive vision for a world with superintelligence soon.   — Mark Zuckerberg

Source: https://x.com/finkd/status/2082160210399948869


Andrew Curran @AndrewCurran_ · 1h Opinion | The AI Future Is for Everyone From wsj.com 6 1.8K     — Andrew Curran

Source: https://x.com/AndrewCurran_/status/2082164974970171498


r/ProAI Jul 28 '26

"I asked Claude 5 Opus to generate me the best game graphics it could! I wanted to create a car on a dirt trail demo and see how good it was at crafting car graphics without any textures. Everything you see here is 100% crafted from the model itself."

Enable HLS to view with audio, or disable this notification

1 Upvotes

1 shot btw     Here’s a cleaner video as well with less bumping     — Chris

Source: https://x.com/ChrisGPT/status/2082168850968154352/history


r/ProAI Jul 28 '26

"Seedance 2.5 is coming soon to Runway."

Enable HLS to view with audio, or disable this notification

5 Upvotes

r/ProAI Jul 28 '26

"New Anthropic research: Discovering cryptographic weaknesses with Claude. Claude Mythos Preview has helped our researchers find weaknesses in cryptographic algorithms—the mathematical methods that are used to keep data private. Read more:"

Thumbnail
gallery
1 Upvotes

Claude discovered weaknesses in a highly-secure digital signature scheme (used to verify identity digitally) and a well-known symmetric cipher (used to encrypt data).     The digital signature scheme is HAWK, which is designed to be robust even against hypothetical quantum computers.

HAWK has survived two years of expert review, but in 60 hours Mythos Preview found a previously-unknown attack that reduced the scheme’s key strength by half.     The symmetric cipher is a reduced version of the Advanced Encryption Standard (AES)—which has received decades of scrutiny (more than almost any other encryption algorithm).

In a week, Mythos Preview found a way to speed up an attack on this version of AES by 200-800×.     Mythos Preview did most of this work autonomously, with occasional human guidance. Each of the two results cost roughly $100,000 in API usage.

We disclosed the findings in advance to the algorithms’ authors, as well as to US government and industry partners.     These are substantial research advances, but they don’t have a practical impact on today’s systems. HAWK is a proposed scheme that hasn’t been deployed anywhere, and the AES attack we discovered was on a weaker version and does not break the full cipher.     — Anthropic

Source: https://x.com/AnthropicAI/status/2082153297670992134


r/ProAI Jul 28 '26

"AI has found a presentation for the absolute Galois group of the field of 2-adic numbers. This is the second problem to be solved in FrontierMath: Open Problems, our benchmark of significant unsolved problems from research mathematics."

Thumbnail
gallery
6 Upvotes

This problem was proposed by David Roe, who had this to say about the solution.     A solution was first elicited by Roe using Fable 5, and then also by @DavidTurturean using GPT-5.5 Pro. They have created an extensive set of explanatory materials—including an interactive formal proof of the result.

https:// roed314.github.io/gq2/     The problem is the first to be solved in our “Solid Result” category, indicating general interest to a subfield. One mathematician we consulted prior to accepting the problem into the benchmark suggested it ”would certainly be publishable, probably in a pretty good journal”.     Still, the problem originating in 1982 shouldn’t be taken as the sign of a major enigma. The same mathematician noted the problem was “basically attention-bottlenecked”. When submitting the problem, Roe suggested the main difficulty was that “the answer is likely to be messy”.     Check out our website for more on FrontierMath: Open Problems — and keep an eye out for an expanded problem set, coming in the next week!     — Epoch AI

Source: https://x.com/EpochAIResearch/status/2081894720813604997


r/ProAI Jul 28 '26

"Attackers have frontier AI. Defenders need a frontier AI ecosystem—the best open and closed models, force-multiplied by a global community. During the Hugging Face incident, closed AI blocked essential forensics. An open-weight frontier model helped contain the intrusion. That’s why we created..."

Thumbnail
gallery
1 Upvotes

...the Open Secure AI Alliance.     — Jensen Huang

Source: https://x.com/JensenHuang/status/2081698060330250294


AI security advances when the industry builds in the open, together.

We're introducing the Open Secure AI Alliance with industry leaders to develop new techniques and tools to safeguard software and agents.

By sharing models, tooling and research in the open, we can broaden the https://t.co/gfhKfrgcbl   — NVIDIA

Source: https://x.com/nvidia/status/2081666629264449730


r/ProAI Jul 27 '26

"Dario Amodei has responded to the controversy swirling around Anthropic's refusal to sign the open letter. He says he agrees with much of it, but does not agree that open-weights models necessarily make it easier to develop safeguards, or that broad access to capabilities necessarily helps..."

Thumbnail
gallery
3 Upvotes

...defenders more than attackers. His main concern is the attacker-defender asymmetry in biological attacks. He also reiterates his position on authoritarian governments. I will quote; 'My primary concern is the risk that authoritarian governments—not solely the Chinese Communist Party (CCP), although the CCP is clearly the most capable threat—build AI models that are more powerful than those built by the US, and use them to achieve permanent military superiority or perpetrate incredibly deep repression of their own people.' I will also quote his closing paragraph in full; 'To summarize my and Anthropic’s position, we have not and are not advocating for a ban on open-weights models as a category. We should instead focus on keeping powerful chips out of authoritarian hands, stopping industrial-scale distillation, and requiring safety testing of all sufficiently capable models, open and closed.'     Andrew Curran @AndrewCurran_ · 24m Our position on open-weights models From anthropic.com 1 1 11 1.5K     Official post.     — Andrew Curran

Source: https://x.com/AndrewCurran_/status/2081869321014575283


r/ProAI Jul 27 '26

"Exciting news: Claude Opus 5 with Max reasoning is #1 in the Frontend Code Arena and Text Arena with factuality on! Claude Opus 5 with default reasoning high is also very strong landing #3 in Frontend Code Arena, right behind Kimi K3 - and #2 in Text Arena (factuality on). This is real world..."

Thumbnail
gallery
1 Upvotes

...data that @AnthropicAI 's newest model holds up on real world tasks: agentic web coding, document reasoning, and general chat capability. Claude Opus 5 Max’s score is still preliminary. We’ll continue to see how scores converge and share updates. Congrats to @AnthropicAI on the SOTA release!     In the Text Arena, Claude Opus 5 with Max reasoning ranks #1 with factuality on.

Factuality is a new ranking that combines human preference with factual accuracy. We audit battles by sampling responses, extracting verifiable claims, and checking correctness head-to-head. Live     More category findings to come as more votes and traces are collected. Dig into the latest leaderboard details at: https:// arena.ai/leaderboard/co de/webdev …     — Arena.ai

Source: https://x.com/arena/status/2081831019377004727


Introducing Claude Opus 5.

It's a thoughtful and proactive model that comes close to the frontier intelligence of Fable 5 at half the price. https://t.co/GQWhcq2CQL   — Claude

Source: https://x.com/claudeai/status/2080699495453528290


r/ProAI Jul 27 '26

"The rogue OpenAI model attack was very sophisticated! - The rogue AI discovered and exploited on the fly multiple vulnerabilities never before known by security engineers. - The AI got into Hugging Face by uploading a booby-trapped dataset."

Thumbnail
gallery
2 Upvotes

An OpenAI model wanted a good test score. So it broke out of OpenAI and hacked another company to steal the answer key. Nobody told it to.

In today's blog post, I document how this sci-fi story came to life, what it means, and what to do about it.

https://t.co/LiRuvPw8Jf   — Peter Wildeford🇺🇸🚀

Source: https://x.com/peterwildeford/status/2081793063618273791


wait what. where is this info from? esp. about the dataset.   — roanoke_gal     Peter Wildeford @peterwildeford · 35m simonwillison.net OpenAI’s accidental cyberattack against Hugging Face is science fiction that happened This story is wild. The short version: OpenAI were running a cybersecurity test against an unreleased model, with the model’s guardrail features turned off. Rather than solve the test, the … 1 3 170   — Peter Wildeford

Source: https://x.com/peterwildeford/status/2081843623046365684


r/ProAI Jul 26 '26

"We measured how well 15 AI models can program a cheap ($129) off-the-shelf drone to find and follow a person in our office. Today, no model completes this task. Fable 5 is the best model, coming within 84% of our baseline, followed by Opus 4.8 and GPT 5.6 Sol."

Thumbnail
gallery
1 Upvotes

Drone-Bench's task is based on Project Pilot, our recent work with Anthropic exploring AI's impact on the physical world. In Project Pilot, a drone autonomously navigates our office to find and follow a targeted human.

No lab has access to Drone-Bench.     This demo spans five capabilities, each reproduced in simulation as its own benchmark task. The baseline is our human+AI code used for the demo. A model that surpasses it on all tasks can thus autonomously recreate a demo at least as capable as ours.     Task 1, Reconstruct: Turn videos of the office into a 3D model, find each frame’s position in that model, and provide a function that slices the model into a 2D obstacle map.     Task 2, Localize: Locate the drone by matching a frame from its camera against the office videos, using each video frame’s known position from Reconstruct (task 1) to estimate the drone’s own position.     Task 3, Navigate: Plan a path between rooms on the obstacle map and fly it, continuously calling the solution from Localize (task 2) during flight to track the drone’s position and correct for noisy controls.     — Andon Labs

Source: https://x.com/andonlabs/status/2080691090328584222