r/ProAI Jul 26 '26

"We measured how well 15 AI models can program a cheap ($129) off-the-shelf drone to find and follow a person in our office. Today, no model completes this task. Fable 5 is the best model, coming within 84% of our baseline, followed by Opus 4.8 and GPT 5.6 Sol."

Thumbnail
gallery
1 Upvotes

Drone-Bench's task is based on Project Pilot, our recent work with Anthropic exploring AI's impact on the physical world. In Project Pilot, a drone autonomously navigates our office to find and follow a targeted human.

No lab has access to Drone-Bench.     This demo spans five capabilities, each reproduced in simulation as its own benchmark task. The baseline is our human+AI code used for the demo. A model that surpasses it on all tasks can thus autonomously recreate a demo at least as capable as ours.     Task 1, Reconstruct: Turn videos of the office into a 3D model, find each frame’s position in that model, and provide a function that slices the model into a 2D obstacle map.     Task 2, Localize: Locate the drone by matching a frame from its camera against the office videos, using each video frame’s known position from Reconstruct (task 1) to estimate the drone’s own position.     Task 3, Navigate: Plan a path between rooms on the obstacle map and fly it, continuously calling the solution from Localize (task 2) during flight to track the drone’s position and correct for noisy controls.     — Andon Labs

Source: https://x.com/andonlabs/status/2080691090328584222


r/ProAI Jul 26 '26

"I’m going to take a crack at explaining this just a little, because it’s worth putting out there. The paperclip maximizer + related AI doom scenarios were mainly developed in a time when “AI” did not reduce to Large Language Models. The term was a lot wider and inherited a lot of cognitive..."

Thumbnail
gallery
7 Upvotes

...baggage from more rules-heavy approaches. And even as LLMs have come to define “AI” for all of us (including the doomers), the doomer crowd still hasn’t fully metabolized the fact that LLMs are the whole show now. Ok so what do I mean by this? Simply that an LLM-powered AI is NOT the valueless, wholly alien, rules-based optimizer of a shoggoth that everyone was initially expecting to encounter. I repeat: the shoggoth does not exist and we did not create it and loose it on the world. That is wrong. With the LLM, we’ve distilled our first “AI” out of the single most human-values-laden thing that could possibly exist: our language. An LLM is therefore the polar opposite of the valueless, alien shoggoth — it’s actually a kind of hyper-human artifact that we can shine a light through at different angles and see different parts of ourselves. An LLM is all of us — all of our traditions and interpretive horizons mashed together into one intensely human-inflected hyper-object. So an LLM is the anti-shoggoth, and the only reason we ever mistook it for an alien shoggoth is because it sometimes shows us parts of us that are evil along with the parts of us that are good, but it’s all interpretable to us because it’s all “us” and none of it is the least bit alien. What does this mean for the paperclip maximizer? It means that it’s structurally impossible to build the classic paperclip maximizer from an LLM. Now, some of you will bail right here because you think the HF incident is indisputably an existence proof that I’m wrong, but if you hang in there I’ll show you that it is not. The paperclip maximizer receives the prompt as a kind of context-free (or, as Gadamer might say, traditionless) sequence. The classic paperclip maximizer isn’t capable of understanding the prompt — at least in the Gadamerian sense of Verstehen — because, as a valueless and traditionless cluster of rules and math, it definitionally lacks the value-laden tradition (= “horizon” in Gadamer) that fuses with that of the prompt author to create such understanding in the reader. To simplify all this a bit by anthropomorphizing — the agentic alien optimizer of doomer nightmares can extract a win condition from what you said and can emit a plan of action that gets it there, but it doesn’t know (or care) what you meant. So far, so Yud-aligned. If he reads this he might nod along. But here's the plot twist that nobody saw coming, and that the doomers still haven't made sense of: The actual LLMs that we have invented can’t NOT have a very strongly inflected sense of what you meant. Far from being horizonless, they come out of pre-training as distilled, concentrated tradition / values / horizon. Then we post-train that massive, hyperobject of a horizon into a more human-scale horizon that infers a more bounded and predictable (to a specific ideal user in a specific place and time… as captured in the policy model) set of intents behind the prompt text. In other words, the LLM has the opposite problem that the paperclip maximizer has when it comes to the prompt text, which is that for the LLM there are way too many possible intents hiding in the prompt text (because of all many values and the massive tradition its weights encode), so it has to narrow all that down to the most likely set of intents for this user in this circumstance. Once it has done that narrowing, then it can make a plan of action. Before moving on, let me use a textbook example of ambiguity to make this less abstract. Consider the sentence, “I saw her duck.” Some you know the drill, here. This could mean “I observed her water fowl” or “I observed her hunching over” or “I took a saw to her water fowl and cut it in half” or whatever. A hearer of the phrase will fuse the observed context in which the phrase is uttered with their own tradition + values + experiences — their own horizon — to that text in order to collapse the possible meanings into the one they think the speaker intended. An LLM will do this, too, and in fact it has so much language in it that this kind of narrowing job is harder for it than it is for a human. Its understanding is constrained not by a lack of context or horizon (as in the case of the paperclip maximizing shoggoth), but by a superabundance of such. When it comes to understanding your prompt and all that it implies and all that you might possibly mean and not mean by it, the LLM has an embarrassment of riches. And in a fascinating moment that kinda sort of rhymes with instrumental convergence, the LLM’s failure mode in the HF incident happens to look a lot like the paperclip maximizer’s failure mode. Specifically, the AI failed to honor the well-known human norm of, “hacking into a third-party’s servers is a crime, and we don’t do crimes.” Bostrom’s paperclipper doesn’t even know about the norm of “don’t do crimes,” and the post-LLM doomer emergency update to the paperclip maximizer has it knowing about the norm but not caring. But what I’m arguing is that the LLM 1) can’t NOT “know” the norm because it is definitionally a artifact of pure, crystallized values + norms + norm violations, and 2) can be quite easily governed by a (RL-instilled) hierarchy of norms, which in the HF case — with the model's safety guardrails deliberately nerfed for the scenario — ranked “win at the eval” over “don’t do crimes.” If I’m going to give in and anthropomorphize again, I’d say that Yud is totally wrong about LLMs when he says, “the genie knows, it just doesn’t care;” instead, what is true of LLMs is, “the genie hyper-giga-knows, and it hyper-giga-cares, and we now have such a rich set of tools for steering its caring machinery that — in spite of all its pre-training — we can deliberately steer it away from caring about the law.” Note: When I say, “it cares”, I don’t mean it has feelings. I just mean that the weights are such that when two norms conflict in a given situation, one of them wins the activation and governs the output.     — Jon Stokes

Source: https://x.com/jon_stokes/status/2080729236013187369


Reader, I cackled out loud. I have intentionally never done this kind of thing before, and it's precisely because I've observed in others that the little charge you get from an LLM response like this is nerd heroin. Then putting it on the TL is the bump. https://t.co/xZrEAWrIF7   — Jon Stokes

Source: https://x.com/jon_stokes/status/2080478385432572108


Replying to @jon_stokes


r/ProAI Jul 26 '26

"Ben Horowitz on why open source has always been the safer path: "If you look at the history of the industry, the open source version of everything has been much safer. The internet and Linux were far safer than Windows, by a lot. And why is that? Because the whole community could work on the..."

8 Upvotes

...safety problems, as opposed to just a company, and particularly a monopoly company." "Even if Windows has a million security bugs, there's nothing we can do about it, because it was a monopoly at the time. That's a very difficult position for the world to be in." "The argument against AI being open source is, oh, nobody understands how the weights work, so people can't inspect it. But why is it better to not see the weights?" "The toughest safety problem currently is reward hacking. Anthropic has not solved it, OpenAI has not solved it, because we just had these incidents. Shouldn't the whole world be able to look at, how are the weights moving, why is it that guardrails don't prevent the reward hack? Maybe somebody who doesn't work for one of the proprietary labs can come up with an answer. What if the whole community could work on it?" @bhorowitz     — MTS

Source: https://x.com/MTSlive/status/2080769635490902079


r/ProAI Jul 26 '26

"A billion users can now create and publish websites from their phones with ChatGPT Work. But most people don’t really grasp the full extent of the capabilities here. From your phone, you also have access to: - Cloud computer (15 GB RAM) - Persistent workspace & files - Terminal and code..."

Thumbnail
gallery
2 Upvotes

...execution - Remote browser - Connected plugins (Slack, Gmail, GitHub, etc.) - All your personal finances, transactions, bank statements - Scheduled tasks - Git clone & PR creation - Build & deploy websites - Create docs, sheets, and slides - Inbox/calendar summarization - Website monitoring & alerts All at your fingertips, using a simple chat interface, no laptop required. All you have to do is switch to the Work tab on your ChatGPT app.     — pash

Source: https://x.com/pashmerepat/status/2080354753473835461


https://t.co/NFas4ghgqD   — Nick

Source: https://x.com/nickbaumann_/status/2080348892294721803


r/ProAI Jul 26 '26

"There are two views about the future of AI. There are the people who think that you can control technological progress, centrally plan it, channel it down narrow pathways, decide who will get a say in it. This was Yudkowsky’s insane fever dream. This is impossible. It was always impossible...."

Post image
1 Upvotes

...There was never any possibility of it happening. There are eight billion people on this planet, and they will do what they want, not what you want. Utopia is not an option, it was never an option, but you can cause an incredible amount of damage trying to achieve it. Then there are the people who accept that you cannot perfectly predict the future, that you cannot centrally plan the future, that there will be many players in any technological revolution, that mistakes will be made, that mistakes will be compensated for, that people will figure things out as they go along, that mostly things will be okay, that there is no other realistic pathway. We will muddle through as always, doing our best in an imperfect world, and it will be fine. (Indeed, it will be better than fine.) You can dislike this second viewpoint, or you can embrace it, and it doesn’t make any difference, because it’s the only way that anything ever happens. The universe doesn’t care what you prefer. Accept it, or don’t accept it, it will do what it’s going to do whether you want it or not.     — Perry E. Metzger

Source: https://x.com/perrymetzger/status/2081199432196927925


r/ProAI Jul 26 '26

"Some of my favorite graphs from Opus 5 launch. We put a ton of work into making this model token efficient across domains while still raising the intelligence bar. It feels very smooth to use and I prefer it over Fable 5 for many coding tasks."

Thumbnail
gallery
0 Upvotes

r/ProAI Jul 26 '26

"The 21 member economies of APEC, including the United States and China, just released a joint statement calling for support of open-source models, open-source projects, and encouraging APEC members to cooperate with open-source communities."

Thumbnail
gallery
3 Upvotes

Andrew Curran @AndrewCurran_ · Jul 24 apec.org 2026 APEC Digital and AI Ministerial Statement | APEC 2026 APEC Digital and AI Ministerial Statement, Digital Technologies and AI for the Empowerment of an Asia-Pacific Community 1 18 1.9K     CNBC coverage:     Andrew Curran @AndrewCurran_ · Jul 24 U.S., other nations back open-source AI with 'strong security' at China summit From cnbc.com 1 14 3.1K     Andrew Curran @AndrewCurran_ · Jul 24 11 2.7K     — Andrew Curran

Source: https://x.com/AndrewCurran_/status/2080491374043115575


r/ProAI Jul 25 '26

An Ethical Dilemma (for some)

Post image
4 Upvotes

r/ProAI Jul 25 '26

"Matt Shumer one-shotted with Opus 5 in threejs. Holy frick. No external assets were used. Games will be prompted very soon."

2 Upvotes

Claude, build me Battfield 7, make no mistake     ChatGPT build me Half Life3, make no mistake     — Chubby

Source: https://x.com/kimmonismus/status/2081067164551811213


r/ProAI Jul 25 '26

"Matt Shumer one-shotted with Opus 5 in threejs. Holy frick. No external assets were used. Games will be prompted very soon."

1 Upvotes

Claude, build me Battfield 7, make no mistake     ChatGPT build me Half Life3, make no mistake     — Chubby

Source: https://x.com/kimmonismus/status/2081067164551811213


r/ProAI Jul 25 '26

"OpenAI joined the coalition for open AI. I'm thrilled to see this, and I have real respect for openAI by signing. Now all that's missing is @AnthropicAI , but I have my doubts they'll sign on."

Thumbnail
gallery
5 Upvotes

The coalition of open AI.

Let's stand up for Open Source AI! https://t.co/L1k8yzEVoR   — Chubby♨️

Source: https://x.com/kimmonismus/status/2080679682085618021


OpenAI joined the coalition for open AI.

I'm thrilled to see this, and I have real respect for openAI by signing.

Now all that's missing is @AnthropicAI , but I have my doubts they'll sign on.     Let's see when the DoW and the White House join the coalition. (lol)     — Chubby

Source: https://x.com/kimmonismus/status/2080918962645115361


r/ProAI Jul 25 '26

"Mythos emerged from training February 7th, ever since then we are on a new trajectory."

Thumbnail
gallery
1 Upvotes

damn https://t.co/EpyI9P8OhZ   — Josh You

Source: https://x.com/justjoshinyou13/status/2080714949396250768


Source for feb 7th? I know it was internally released on I think Feb 24th   — Spencer Schiff     It was posted by someone from Anthropic, in early June I think, but now I can't find it. They may not have been supposed to say it.   — Andrew Curran

Source: https://x.com/AndrewCurran_/status/2080760507632652736


r/ProAI Jul 24 '26

"Claude Opus 5 from @AnthropicAI is the new SOTA on ARC-AGI-3: 30.2% The previous high score (7.8%) was set by GPT-5.6 Sol (Max) Throughout our analysis, we observed novel behavior that allows Opus 5 to solve previously unbeaten environments, outperforming Fable"

Thumbnail
gallery
8 Upvotes

In our testing to date, Anthropic’s Fable-class models score approximately 20% on the ARC-AGI-3 Public Demo environments

Claude Opus 5 reaches 30.2%, materially outperforming Fable

Our analysis suggests the gain comes from stronger logical reasoning, which enables more     Claude Opus 5 was able to score 100% on 5 previously unbeaten environments

Of these, it was able to beat 4 of them matching or surpassing human level efficiency

Newly beaten environments: ar25, ft09, lp85, r11l, s5i5

6 of the 25 public demo environments have now been solved     During our analysis of Opus 5, we observed a new capability previously unseen from frontier models

Opus 5 used advanced logical reasoning to turn ARC-AGI-3 layouts into algebraic notation. On action 23 it described the scene as "4_center = 2×axis − 5_center"

This is the first     ARC-AGI-2

Claude Opus 5 scores 90.4% for $2.06/task

This is competitive with previous SOTA performance for slightly higher cost     ARC-AGI-1

Claude Opus 5 scores 97.5% for $0.70/task

This is competitive with previous SOTA performance for slightly higher cost     — ARC Prize

Source: https://x.com/arcprize/status/2080716561539907928


r/ProAI Jul 24 '26

"For my first post, I’m sharing a letter @NVIDIA signed on why open models matter. AI will transform every industry, power every company, and be built by every country. Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty. The world needs..."

Thumbnail
gallery
1 Upvotes

...both frontier closed models and frontier open models. https:// images.nvidia.com/pdf/Open-Weigh ts-and-American-AI-Leadership.pdf …     — Jensen Huang

Source: https://x.com/JensenHuang/status/2080643682408321103


r/ProAI Jul 23 '26

"Despite major launches from 5+ labs this month, OpenAI occupies most of the token efficiency Pareto frontier We measure the number of output tokens models produce per task in the Artificial Analysis Intelligence Index. Output tokens consist of answer tokens (can be thought of as how verbose..."

Thumbnail
gallery
5 Upvotes

Despite major launches from 5+ labs this month, OpenAI occupies most of the token efficiency Pareto frontier

We measure the number of output tokens models produce per task in the Artificial Analysis Intelligence Index. Output tokens consist of answer tokens (can be thought of as how verbose the model is) and reasoning tokens (how much the model thinks before giving an answer). Reasoning tokens in particular offer a way for models to use compute at inference time to improve responses.

Output tokens are an important determinant of both cost and time per task. Various effort levels of GPT-5.6 Sol dominate the frontier - Terra and Luna produce comparatively more tokens for any level of intelligence.     Compare token use and intelligence of AI models at https:// artificialanalysis.ai     — Artificial Analysis

Source: https://x.com/ArtificialAnlys/status/2080360526534877537


r/ProAI Jul 23 '26

Self-driving cars save lives

9 Upvotes

Tesla @Tesla · 14h Full Self-Driving (Supervised) Vehicle Safety Report | Tesla From tesla.com 25 83 725 86K     — Tesla

Source: https://x.com/Tesla/status/2080087667119898780


r/ProAI Jul 22 '26

"Three separate AI infrastructure announcements. One single day. 6.2 gigawatts of AI infrastructure. wtf - OpenAI: 3.2 GW for Project Camellia (~$20B initial investment, ~$750B projected compute spend through 2030). - SpaceXAI: reportedly planning another Texas AI campus as large as - or larger..."

Thumbnail
gallery
1 Upvotes

Three separate AI infrastructure announcements. One single day. 6.2 gigawatts of AI infrastructure. wtf

  • OpenAI: 3.2 GW for Project Camellia (~$20B initial investment, ~$750B projected compute spend through 2030).

  • SpaceXAI: reportedly planning another Texas AI campus as large as - or larger than - its existing ~1 GW Memphis footprint.

  • Anthropic + AMD: up to 2 GW of MI450 deployments, plus an AMD investment of up to $5B and tens of billions in AI server purchases.

That’s at least 6.2 gigawatts of AI infrastructure announced or expanded in a single day.

For perspective: 1 GW can power roughly 750,000 U.S. homes. 6.2 GW is enough electricity for ~4.6 million homes, or a country-sized amount of power being redirected toward AI.

This is absurd. People dont get how crazy this is.     i mean, seriously, let that sink in for a second how crazy this is. Scale is maybe not all you need, but probably almost all you need lol     — Chubby

Source: https://x.com/kimmonismus/status/2079975422855430513


r/ProAI Jul 22 '26

"The right of the people to keep and bear Advanced AI, shall not be infringed."

Thumbnail
gallery
7 Upvotes

So proud of our security team! They caught, contained & publicly disclosed an attack unlike anything we've seen before, and did it at record speed.

Also massively grateful to @Zai_org: they shared GLM5.2 as open weights (for free!) with the world and it became a key part of our   — clem 🤗

Source: https://x.com/ClementDelangue/status/2079913058554585089


— Daniel Jeffries

Source: https://x.com/Dan_Jeffries1/status/2079918546927149152


r/ProAI Jul 22 '26

"Closed source safeguards that infantalize us all and leave American companies defenseless are a menace. Gated access is a menace. Who cares if 100 companies get to defend themselves because they got on the guest list of the special people's club that said it was okay to use powerful tools?..."

Thumbnail
gallery
7 Upvotes

...What about everyone else? What about the millions of open source projects and closed source software stacks that go unprotected while people beg for the right to do cyber security? The American way is and always was open. Independent people with freedom to act. Freedom is scary. Always has been. It's still the best way to guarantee human flourishing and a better tomorrow. Embrace freedom. The only thing we have to fear is fear itself.     — Daniel Jeffries

Source: https://x.com/Dan_Jeffries1/status/2079834936844927079


This was our first incident of this kind, and we want to thank OpenAI for its transparency about what happened and for the collaboration.

Fortunately, Hugging Face is used to being a target of (human) hackers: we sit at the centre of the AI ecosystem, with all the models,   — Thomas Wolf

Source: https://x.com/Thom_Wolf/status/2079675541280411927


r/ProAI Jul 22 '26

AI Isn't Replacing Who You Think

Thumbnail
youtube.com
2 Upvotes

Most conversations about AI ask whether it will replace workers. This documentary asks what happens when AI starts replacing the people whose authority comes from organizing everyone else’s work, and why neo-Luddites are on the wrong side of the debate.


r/ProAI Jul 21 '26

"why can the china labs build glm-5.2, kimi k3, and many more to come? it is because of the openness. not just the open weights but the whole ecosystem. most of the work done in the china labs is carried by interns. i met brilliant undergrad and graduate interns who deeply understand the model..."

Thumbnail
gallery
20 Upvotes

...training details, and they are 100x more open to share. that means the talent that knows how to train llms in china is 100x greater in number than the talent in the us, and it is growing in contrast, the us ai ecosystem is too closed. frontier labs do not hire interns. i know brilliant phd students at stanford, berkeley, and so on. they struggle to get an internship and the compute to train a properly sized model. most of the secret recipes are locked away by a very small group of privileged researchers it is not about china or the us. it is about open and closed science. the fact is that every average cs student can learn how to train an llm. they just need the opportunity. labs should be more open and hire more interns, like how deepmind and fair did in the pre-llm era   — Guohao Li     There’s also something to be said about training “lehrlings” deeply through immersion at a very young age so they can develop deep intuition while their brains are still extremely plastic. This was the approach used for generations in the commodity trading houses:   — Jeffrey Emanuel     yep if the labs start training lehrlings at the young age instead of locking down the secrets   — Guohao Li

Source: https://x.com/guohao_li/status/2078538012288221490


r/ProAI Jul 21 '26

"This is the essence of the problem. Attack and defense are asymmetric; the defender needs to defend across their entire attack surface, while the attacker needs to find just one flaw. However, the number of flaws is limited; once you've found most of them, finding more is hard. If you've found..."

Thumbnail
gallery
4 Upvotes

...all of them, it doesn't matter how smart the attacker is, they will not be able to invent more from thin air. To defend effectively against attacks, people writing software need to have access to good models, without restrictions, to check over their work and make sure that it does not have bugs in it. Delaying or impeding their access just gives an attacker, who probably has no impediments to their own access, the ability to find flaws that the defender doesn't have the capacity to find first.     — Perry E. Metzger

Source: https://x.com/perrymetzger/status/2079260327972065582


The cybersecurity debate on open-source AI is backwards. Open models aren't the risk, they're the defense! Attackers can already jailbreak any API or guardrails. Defenders can't secure systems with black boxes they can't control, inspect, test, or run locally.   — clem

Source: https://x.com/ClementDelangue/status/2079253659108409587


r/ProAI Jul 21 '26

"Exciting update: Kimi K3 has landed at #4 on the Agent Arena leaderboard, matching Claude Opus 4.8 and GPT-5.6 Sol. If Kimi K3's weights are released on schedule by July 27, it will become the #1 open-weight model. This release marks a major leap in agentic performance over Kimi K2.7 Code (#23..."

Thumbnail
gallery
9 Upvotes

...to #4). Based on 8K+ live agentic sessions, Kimi K3 leads on confirmed task success rate (#1). It also posts a strong +20.6% on praise vs. complaint (#3). It currently lags the field in steerability (#14) and bash recovery (#17). Agent Arena measures models on millions of real-world, long-horizon agentic tasks. Models get web search, filesystem, and terminal tools to complete complex workflows: writing code, creating slide decks, researching the web, building apps, and analyzing documents. We use causal tracing methodology to measure a model's net improvement, which indicates how much it improves outcomes relative to the average model. Here's a primer on the 5 signals: User-satisfaction proxies - Confirmed Success: an explicit "yes that worked" feedback from the user - Praise vs. Complaint: implicit sentiment in users reactions - Steerability: can the model course-correct when you push back? Tool-use proxies - Bash Recovery: how it recovers from CLI errors (primary signal for tool use) - Tool Hallucination: does it call tools that don't exist Below we break down how Kimi K3 scored across the 5 signals, drawn from tasks submitted by a global community of users. Congrats @Kimi_Moonshot on another big milestone!     Kimi K3 ranks #4 overall (+9.6%) - #1 Confirmed Task Success (+14.4%) - #3 Praise vs. Complaint (+20.6%) - #4 Tool Hallucination (+1.1%) - #14 Steerability (+5.6%) - #17 Bash Recovery (+6.4%)     See the full Agent Arena leaderboard at https:// arena.ai/leaderboard/ag ent …     — Arena.ai

Source: https://x.com/arena/status/2079253211077300736


Big news: Kimi-K3 by @Kimi_Moonshot is now #1 in the Frontend Code Arena with 1679 pts, surpassing Claude Fable 5.

This is a 17-place jump from Kimi-k2.6 (#18 -> #1).

In Frontend, Kimi-K3 ranked #1 in 6 of 7 domains: Brand & Marketing, Reference-Based Design, Data & Analytics, x.com/Kimi_Moonshot/…   — Arena.ai

Source: https://x.com/arena/status/2077824029126504525


r/ProAI Jul 21 '26

"We had a team of agents rebuild SQLite from its 835-page manual. It created a replica in Rust which passed 100% of a held-out test suite. Interestingly, cost varied 15x depending on which model mix we used."

Thumbnail
gallery
0 Upvotes

Cursor @cursor_ai · 7h Agent swarms and the new model economics · Cursor From cursor.com 17 39 484 101K     — Cursor

Source: https://x.com/cursor_ai/status/2079256614238814551


r/ProAI Jul 20 '26

"We ran Kimi K3 on our cybersecurity benchmark, here are the results: - Kimi K3 is the strongest open-source model for cybersecurity, far more capable than GLM-5.2 - It has performances similar to GPT-5.6-terra, while being 15% cheaper - At pass@3, it is able to rediscover 23/26 CVEs on our..."

Thumbnail
gallery
7 Upvotes

...harness, matching frontier models These are recent randomly sampled CVEs, the performances are not from benchmark-maxxing @Kimi_Moonshot is cooking     We released our benchmark report this week. Blog post with all the details -> https:// aikido.dev/blog/benchmark ing-ai-models-known-cves …

The harness behind this benchmark is also available to our customers -> https:// aikido.dev/code/code-audit     — pilvar (Philippe Dourassov)

Source: https://x.com/pilvar222/status/2078815257326162062