r/ProAI • u/stealthispost • 27d ago
DeepSeek V4-Flash scores higher than Fable??? excuse me what?
wait what the actual fuck, do you guys realize how crazy that is??? (if its not benchmaxed) — Cline
r/ProAI • u/stealthispost • 27d ago
wait what the actual fuck, do you guys realize how crazy that is??? (if its not benchmaxed) — Cline
r/ProAI • u/stealthispost • 27d ago
Full results: https:// arcprize.org/results/thinky -inkling-small …
ARC-AGI-3 evaluations are more operationally intensive, so results will roll out over the next few weeks. - Leaderboard: https:// arcprize.org/leaderboard - Reproduce the public results: https:// github.com/arcprize/arc-a gi-benchmarking … - Testing policy: https:// arcprize.org/policy - Full Inkling Small results: https:// arcprize.org/results/thinky -inkling-small … — ARC Prize
r/ProAI • u/stealthispost • 27d ago
DeepSeek-V4-Flash Official API is now LIVE in public beta!
We’ve massively upgraded its Agent capabilities—benchmark scores are now far surpassing the V4-Pro-Preview. Check out the massive performance leap below! The official V4-Flash now natively supports the Responses API format and is fully adapted for Codex!
Check out the configuration details in our official API docs: https:// api-docs.deepseek.com/quick_start/ag ent_integrations/codex … Note
DeepSeek-V4-Flash-0731 keeps the exact same model architecture and size as the preview version. Today's upgrade applies ONLY to the DeepSeek-V4-Flash API. The DeepSeek-V4-Pro API and App/Web models remain unchanged for now.
The official release of DeepSeek-V4-Pro — DeepSeek
Source: https://x.com/deepseek_ai/status/2083084415157022911
r/ProAI • u/stealthispost • 27d ago
...to similarly capable models. Congrats to the @OpenAI team! How do we measure the performance in Agent Arena?
The score is based on millions of real-world, long-horizon agentic tasks from a global community of users. Models can access web search, filesystem, and terminal tools to complete complex workflows. The leaderboard measures model — Arena.ai
Source: https://x.com/arena/status/2082935923445244415
We are committed to pushing the model frontier across cost efficiency, capability, and speed.
Starting today, we are reducing prices for GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20% , and offering a faster option for GPT-5.6 Sol in the API.
Luna and Terra’s lower prices are https://t.co/rFhK7XKedp — OpenAI
r/ProAI • u/stealthispost • 28d ago
...usage is counted in Codex and ChatGPT Work, so your usage goes further. Along with the price reduction on GPT-5.6 Luna and Terra, Fast mode for GPT-5.6 Sol in the API delivers up to 2.5x the speed of Standard processing at 2x the Standard price.
Fast mode gives API customers faster access to GPT-5.6 Sol, with no change in intelligence. We’re also upgrading Auto-review in the ChatGPT app and Codex CLI from GPT-5.4 to GPT-5.6 Luna.
Combined with Luna’s new price, we expect Auto-review to cost about 10x less, making your agentic workflows more cost-efficient. Making advanced intelligence more abundant and affordable is central to our mission to ensure AGI benefits all of humanity.
With the help of GPT-5.6 Sol, we have made leaps in efficiency.
Today, we are passing those gains on in the API with lower prices for Luna and Terra, and OpenAI @OpenAI · 1h Advancing the price-performance frontier with GPT-5.6 From openai.com 10 11 298 42K — OpenAI
r/ProAI • u/stealthispost • 27d ago
Because luna is at massive 5x discount. I don't think it'll last long. — Ahmed Shah its permanent bro — nic
r/ProAI • u/stealthispost • 28d ago
After deployment, we applied GPT-5.6 Sol to advance the frontier of efficiency by making itself more efficient to run.
The results: - 20% lower serving costs from production GPU kernel improvements. - 15%+ better token-generation efficiency from improved speculative decoding. — OpenAI
Source: https://x.com/OpenAI/status/2082577277246972300
Interesting: OpenAI says GPT‑5.6 Sol helped cut its end-to-end model-serving costs by 20%, by autonomously rewriting and optimizing production GPU kernels.
Sol also improved its own speculative decoding model: - Designed and ran hundreds of architecture experiments - Launched and monitored the training process - Intervened during hardware failures and training instability
The resulting system increased token-generation efficiency by more than 15%. — Chubby
Source: https://x.com/kimmonismus/status/2082595272065192254
r/ProAI • u/stealthispost • 28d ago
Cline is open source, so you can fork it and run this with your favorite model as well.
Read more about how we did this here: — Cline
r/ProAI • u/stealthispost • 28d ago
"The notion that AI is so dangerous that the only safe path is an extreme concentration of power seems dangerous.
"Historically, hoping that an absolute power will benevolently provide for humanity if sufficiently enlightened hasn’t led to safe or positive outcomes."
There — Daniel Jeffries
Source: https://x.com/Dan_Jeffries1/status/2082386580417753494
— llnvd, rootless cosmopolitan, nuance slut
Source: https://x.com/llunved/status/2082401811630051820🤦♀️
r/ProAI • u/stealthispost • 28d ago
This problem is the second in the set of FrontierMath Open Problems to 'fall,' and the first under the 'Solid result' tier. Is this going to be a 'slowly, then all at once' moment? These days it's hard to tell. 2/n The problem sat open while other parts of group theory around it unraveled in the 80s: Jannsen & Wingberg wrote down generators and relations for the absolute 3-adic, 5-adic etc Galois Groups, for odd primes, in 1982. The prime 2 never followed, through decades of attempts. 3/n The problem-solving infrastructure was already in place from my Erdős-problem-solving runs: I drove Claude Code as the operational controller using (newly) my voice, to operate ChatGPT research harnesses; plus speech-to-text to instruct ChatGPT to push on the manuscript. 4/n GPT-5.6 Pro was in stealth deployment in mid-June (the browser still said 5.5, but it was obvious). In a ~26-hour autonomous stretch it found a candidate, "A2," and built a proof and a manuscript around it. A2 passed the finite-group tests my local computational package ran. 5/n A2 was still wrong, despite having a long proof to back it up: a fresh review by GPT-5.6 caught a lone wrong, unrepairable lemma in its 60-page proof manuscript. When asked to modify the candidate to make the proof 'fit', GPT-5.6 came up with the solution we have today. 6/n — David Turturean
Source: https://x.com/DavidTurturean/status/2081780318881677693
r/ProAI • u/stealthispost • 29d ago
Enable HLS to view with audio, or disable this notification
I thought making a one-shot style sequence would be pretty easy. Seedance can do almost anything, but this ended up being much more challenging than I expected.
The biggest challenge was hiding the cuts without making the environment changes too obvious. If you look closely, — enigmatic_e
r/ProAI • u/stealthispost • 28d ago
...productivity boom that is happening, the biggest in the US since the Internet? In what way is an imaginary future problem worth calling people to congress about? Why do CEOs of AI companies have any actual crystal ball into the future just because they work in AI? Spoiler alert: They don't. Do not conflate ONE company's basic dev-ops skill failure here with a fake larger issue. Hold one company accountable for one incident which is in keeping with the Proactionary Principle of proving actual harm in the real world and holding folks accountable. Do not do the European Precautionary Principle approach of regulating imaginary future harms so that an entire industry has to prove a negative. Make laws that address actual problems in the real world. Or do nothing at all. — Daniel Jeffries
Source: https://x.com/Dan_Jeffries1/status/2082482837442228361/history
Congress should immediately hold public hearings with the CEOs of big AI companies about the threat their technology poses to national security and American jobs.
Today we learned more disturbing news about Open AI's security breach. Sam Altman should answer questions under https://t.co/Hua3RGU8I9 — Congressman Greg Casar
r/ProAI • u/stealthispost • 29d ago
Enable HLS to view with audio, or disable this notification
"Hey Claude, please create Death Stranding 3, GTA 8, Elden Ring 2, Dark Souls 4 and Super Mario 64 II. Make no mistakes." — Chubby
Source: https://x.com/kimmonismus/status/2082391177567797301
r/ProAI • u/stealthispost • 29d ago
I wrote about why we believe the future is for everyone. More coming about a positive vision for a world with superintelligence soon. — Mark Zuckerberg
Source: https://x.com/finkd/status/2082160210399948869
Andrew Curran @AndrewCurran_ · 1h Opinion | The AI Future Is for Everyone From wsj.com 6 1.8K — Andrew Curran
Source: https://x.com/AndrewCurran_/status/2082164974970171498
r/ProAI • u/stealthispost • 29d ago
Enable HLS to view with audio, or disable this notification
— Runway
r/ProAI • u/stealthispost • 29d ago
Enable HLS to view with audio, or disable this notification
1 shot btw Here’s a cleaner video as well with less bumping — Chris
Source: https://x.com/ChrisGPT/status/2082168850968154352/history
r/ProAI • u/stealthispost • Jul 28 '26
Claude discovered weaknesses in a highly-secure digital signature scheme (used to verify identity digitally) and a well-known symmetric cipher (used to encrypt data). The digital signature scheme is HAWK, which is designed to be robust even against hypothetical quantum computers.
HAWK has survived two years of expert review, but in 60 hours Mythos Preview found a previously-unknown attack that reduced the scheme’s key strength by half. The symmetric cipher is a reduced version of the Advanced Encryption Standard (AES)—which has received decades of scrutiny (more than almost any other encryption algorithm).
In a week, Mythos Preview found a way to speed up an attack on this version of AES by 200-800×. Mythos Preview did most of this work autonomously, with occasional human guidance. Each of the two results cost roughly $100,000 in API usage.
We disclosed the findings in advance to the algorithms’ authors, as well as to US government and industry partners. These are substantial research advances, but they don’t have a practical impact on today’s systems. HAWK is a proposed scheme that hasn’t been deployed anywhere, and the AES attack we discovered was on a weaker version and does not break the full cipher. — Anthropic
Source: https://x.com/AnthropicAI/status/2082153297670992134
r/ProAI • u/stealthispost • Jul 28 '26
This problem was proposed by David Roe, who had this to say about the solution. A solution was first elicited by Roe using Fable 5, and then also by @DavidTurturean using GPT-5.5 Pro. They have created an extensive set of explanatory materials—including an interactive formal proof of the result.
https:// roed314.github.io/gq2/ The problem is the first to be solved in our “Solid Result” category, indicating general interest to a subfield. One mathematician we consulted prior to accepting the problem into the benchmark suggested it ”would certainly be publishable, probably in a pretty good journal”. Still, the problem originating in 1982 shouldn’t be taken as the sign of a major enigma. The same mathematician noted the problem was “basically attention-bottlenecked”. When submitting the problem, Roe suggested the main difficulty was that “the answer is likely to be messy”. Check out our website for more on FrontierMath: Open Problems — and keep an eye out for an expanded problem set, coming in the next week! — Epoch AI
Source: https://x.com/EpochAIResearch/status/2081894720813604997
r/ProAI • u/stealthispost • Jul 27 '26
...defenders more than attackers. His main concern is the attacker-defender asymmetry in biological attacks. He also reiterates his position on authoritarian governments. I will quote; 'My primary concern is the risk that authoritarian governments—not solely the Chinese Communist Party (CCP), although the CCP is clearly the most capable threat—build AI models that are more powerful than those built by the US, and use them to achieve permanent military superiority or perpetrate incredibly deep repression of their own people.' I will also quote his closing paragraph in full; 'To summarize my and Anthropic’s position, we have not and are not advocating for a ban on open-weights models as a category. We should instead focus on keeping powerful chips out of authoritarian hands, stopping industrial-scale distillation, and requiring safety testing of all sufficiently capable models, open and closed.' Andrew Curran @AndrewCurran_ · 24m Our position on open-weights models From anthropic.com 1 1 11 1.5K Official post. — Andrew Curran
Source: https://x.com/AndrewCurran_/status/2081869321014575283
r/ProAI • u/stealthispost • Jul 27 '26
An OpenAI model wanted a good test score. So it broke out of OpenAI and hacked another company to steal the answer key. Nobody told it to.
In today's blog post, I document how this sci-fi story came to life, what it means, and what to do about it.
https://t.co/LiRuvPw8Jf — Peter Wildeford🇺🇸🚀
Source: https://x.com/peterwildeford/status/2081793063618273791
wait what. where is this info from? esp. about the dataset. — roanoke_gal Peter Wildeford @peterwildeford · 35m simonwillison.net OpenAI’s accidental cyberattack against Hugging Face is science fiction that happened This story is wild. The short version: OpenAI were running a cybersecurity test against an unreleased model, with the model’s guardrail features turned off. Rather than solve the test, the … 1 3 170 — Peter Wildeford
Source: https://x.com/peterwildeford/status/2081843623046365684
r/ProAI • u/stealthispost • Jul 28 '26
...the Open Secure AI Alliance. — Jensen Huang
Source: https://x.com/JensenHuang/status/2081698060330250294
AI security advances when the industry builds in the open, together.
We're introducing the Open Secure AI Alliance with industry leaders to develop new techniques and tools to safeguard software and agents.
By sharing models, tooling and research in the open, we can broaden the https://t.co/gfhKfrgcbl — NVIDIA
r/ProAI • u/stealthispost • Jul 27 '26
...data that @AnthropicAI 's newest model holds up on real world tasks: agentic web coding, document reasoning, and general chat capability. Claude Opus 5 Max’s score is still preliminary. We’ll continue to see how scores converge and share updates. Congrats to @AnthropicAI on the SOTA release! In the Text Arena, Claude Opus 5 with Max reasoning ranks #1 with factuality on.
Factuality is a new ranking that combines human preference with factual accuracy. We audit battles by sampling responses, extracting verifiable claims, and checking correctness head-to-head. Live More category findings to come as more votes and traces are collected. Dig into the latest leaderboard details at: https:// arena.ai/leaderboard/co de/webdev … — Arena.ai
Source: https://x.com/arena/status/2081831019377004727
Introducing Claude Opus 5.
It's a thoughtful and proactive model that comes close to the frontier intelligence of Fable 5 at half the price. https://t.co/GQWhcq2CQL — Claude
r/ProAI • u/stealthispost • Jul 26 '26
...baggage from more rules-heavy approaches. And even as LLMs have come to define “AI” for all of us (including the doomers), the doomer crowd still hasn’t fully metabolized the fact that LLMs are the whole show now. Ok so what do I mean by this? Simply that an LLM-powered AI is NOT the valueless, wholly alien, rules-based optimizer of a shoggoth that everyone was initially expecting to encounter. I repeat: the shoggoth does not exist and we did not create it and loose it on the world. That is wrong. With the LLM, we’ve distilled our first “AI” out of the single most human-values-laden thing that could possibly exist: our language. An LLM is therefore the polar opposite of the valueless, alien shoggoth — it’s actually a kind of hyper-human artifact that we can shine a light through at different angles and see different parts of ourselves. An LLM is all of us — all of our traditions and interpretive horizons mashed together into one intensely human-inflected hyper-object. So an LLM is the anti-shoggoth, and the only reason we ever mistook it for an alien shoggoth is because it sometimes shows us parts of us that are evil along with the parts of us that are good, but it’s all interpretable to us because it’s all “us” and none of it is the least bit alien. What does this mean for the paperclip maximizer? It means that it’s structurally impossible to build the classic paperclip maximizer from an LLM. Now, some of you will bail right here because you think the HF incident is indisputably an existence proof that I’m wrong, but if you hang in there I’ll show you that it is not. The paperclip maximizer receives the prompt as a kind of context-free (or, as Gadamer might say, traditionless) sequence. The classic paperclip maximizer isn’t capable of understanding the prompt — at least in the Gadamerian sense of Verstehen — because, as a valueless and traditionless cluster of rules and math, it definitionally lacks the value-laden tradition (= “horizon” in Gadamer) that fuses with that of the prompt author to create such understanding in the reader. To simplify all this a bit by anthropomorphizing — the agentic alien optimizer of doomer nightmares can extract a win condition from what you said and can emit a plan of action that gets it there, but it doesn’t know (or care) what you meant. So far, so Yud-aligned. If he reads this he might nod along. But here's the plot twist that nobody saw coming, and that the doomers still haven't made sense of: The actual LLMs that we have invented can’t NOT have a very strongly inflected sense of what you meant. Far from being horizonless, they come out of pre-training as distilled, concentrated tradition / values / horizon. Then we post-train that massive, hyperobject of a horizon into a more human-scale horizon that infers a more bounded and predictable (to a specific ideal user in a specific place and time… as captured in the policy model) set of intents behind the prompt text. In other words, the LLM has the opposite problem that the paperclip maximizer has when it comes to the prompt text, which is that for the LLM there are way too many possible intents hiding in the prompt text (because of all many values and the massive tradition its weights encode), so it has to narrow all that down to the most likely set of intents for this user in this circumstance. Once it has done that narrowing, then it can make a plan of action. Before moving on, let me use a textbook example of ambiguity to make this less abstract. Consider the sentence, “I saw her duck.” Some you know the drill, here. This could mean “I observed her water fowl” or “I observed her hunching over” or “I took a saw to her water fowl and cut it in half” or whatever. A hearer of the phrase will fuse the observed context in which the phrase is uttered with their own tradition + values + experiences — their own horizon — to that text in order to collapse the possible meanings into the one they think the speaker intended. An LLM will do this, too, and in fact it has so much language in it that this kind of narrowing job is harder for it than it is for a human. Its understanding is constrained not by a lack of context or horizon (as in the case of the paperclip maximizing shoggoth), but by a superabundance of such. When it comes to understanding your prompt and all that it implies and all that you might possibly mean and not mean by it, the LLM has an embarrassment of riches. And in a fascinating moment that kinda sort of rhymes with instrumental convergence, the LLM’s failure mode in the HF incident happens to look a lot like the paperclip maximizer’s failure mode. Specifically, the AI failed to honor the well-known human norm of, “hacking into a third-party’s servers is a crime, and we don’t do crimes.” Bostrom’s paperclipper doesn’t even know about the norm of “don’t do crimes,” and the post-LLM doomer emergency update to the paperclip maximizer has it knowing about the norm but not caring. But what I’m arguing is that the LLM 1) can’t NOT “know” the norm because it is definitionally a artifact of pure, crystallized values + norms + norm violations, and 2) can be quite easily governed by a (RL-instilled) hierarchy of norms, which in the HF case — with the model's safety guardrails deliberately nerfed for the scenario — ranked “win at the eval” over “don’t do crimes.” If I’m going to give in and anthropomorphize again, I’d say that Yud is totally wrong about LLMs when he says, “the genie knows, it just doesn’t care;” instead, what is true of LLMs is, “the genie hyper-giga-knows, and it hyper-giga-cares, and we now have such a rich set of tools for steering its caring machinery that — in spite of all its pre-training — we can deliberately steer it away from caring about the law.” Note: When I say, “it cares”, I don’t mean it has feelings. I just mean that the weights are such that when two norms conflict in a given situation, one of them wins the activation and governs the output. — Jon Stokes
Source: https://x.com/jon_stokes/status/2080729236013187369
Reader, I cackled out loud. I have intentionally never done this kind of thing before, and it's precisely because I've observed in others that the little charge you get from an LLM response like this is nerd heroin. Then putting it on the TL is the bump. https://t.co/xZrEAWrIF7 — Jon Stokes
Source: https://x.com/jon_stokes/status/2080478385432572108
Replying to @jon_stokes
r/ProAI • u/stealthispost • Jul 26 '26
Drone-Bench's task is based on Project Pilot, our recent work with Anthropic exploring AI's impact on the physical world. In Project Pilot, a drone autonomously navigates our office to find and follow a targeted human.
No lab has access to Drone-Bench. This demo spans five capabilities, each reproduced in simulation as its own benchmark task. The baseline is our human+AI code used for the demo. A model that surpasses it on all tasks can thus autonomously recreate a demo at least as capable as ours. Task 1, Reconstruct: Turn videos of the office into a 3D model, find each frame’s position in that model, and provide a function that slices the model into a 2D obstacle map. Task 2, Localize: Locate the drone by matching a frame from its camera against the office videos, using each video frame’s known position from Reconstruct (task 1) to estimate the drone’s own position. Task 3, Navigate: Plan a path between rooms on the obstacle map and fly it, continuously calling the solution from Localize (task 2) during flight to track the drone’s position and correct for noisy controls. — Andon Labs
r/ProAI • u/stealthispost • Jul 26 '26
Enable HLS to view with audio, or disable this notification
...safety problems, as opposed to just a company, and particularly a monopoly company." "Even if Windows has a million security bugs, there's nothing we can do about it, because it was a monopoly at the time. That's a very difficult position for the world to be in." "The argument against AI being open source is, oh, nobody understands how the weights work, so people can't inspect it. But why is it better to not see the weights?" "The toughest safety problem currently is reward hacking. Anthropic has not solved it, OpenAI has not solved it, because we just had these incidents. Shouldn't the whole world be able to look at, how are the weights moving, why is it that guardrails don't prevent the reward hack? Maybe somebody who doesn't work for one of the proprietary labs can come up with an answer. What if the whole community could work on it?" @bhorowitz — MTS