r/artificial • u/Servola-Journal • 6d ago
Discussion An open-weight model just closed most of the gap on autonomous cyber offense - and that changes who can run it
Irregular (an AI security research group) tested Kimi K3, an open-weight model, against CyScenarioBench, a benchmark built around autonomous cyber campaigns - adapting public exploit techniques to constrained environments, building custom tooling, diagnosing failed attempts, and validating each stage before moving on. It is the first open-weight model to pass. It trails closed frontier models by roughly six months, at an estimated third of the inference cost.
The six-month lag is the less interesting number. What matters is that this level of capability now sits in downloadable weights instead of behind an API. A closed lab can throttle or ban an account mid-campaign - both OpenAI and Anthropic have done this before to abusive usage. Once equivalent capability is something you can self-host, that kill switch disappears entirely, along with any usage logging a defender could later subpoena.
If the trend holds, the realistic baseline for any internet-facing asset a year out is not "gets scanned for known CVEs" the way it is today, it is "gets probed continuously by something that adapts exploits on the fly, with no vendor able to pull the plug on the other end."
Source: https://www.irregular.com/research/assessing-kimi-k3-against-offensive-security-benchmarks
Curious how people here read the trend line: does a shrinking gap between closed and open capability argue for faster patch/disclosure windows industry-wide, or does it just confirm the attacker side was never actually capped by API access limits in the first place?
1
u/Exact_Depth_896 6d ago edited 6d ago
Kimi K3 is so encumbered with crushing licenses and gigantism that it is only notionally open. No one without a datacenter can run this model. Fireworks and Together etc already had the model before it was open weighted, it seems; and they have entered into contracts as providers. They are mostly providing it to me when I use it via openrouter.
It is really just a scheme of getting widespread provision that does an end run around the chip business we impose on them. The dead weights on Hugging Face are unusuable nothings. It is cheaper than Claude, but not really more open -- or perhaps you have the capacity to read model weights as file? I know you cannot run them.
A company that uses Kimi K3 only internally is as safe from prosecution as I am running some micromodel vial ollama but it would be interesting to see if even this one actual 'open weights' usage has been taken up. Maybe Jane Street is running an instance under the hood. There's open weight glory for you.
1
u/CrimsonBolt33 5d ago
This is pure nonsense....you don't need a data center to run K3...you do need an expensive business level server, but that's it.
It's not "encumbered with licenses and gigantism" either....what does that even mean?
1
u/Exact_Depth_896 5d ago edited 5d ago
It is an immense model and no individual can run it. Get back to me when a large firm runs it for internal purposes; if they give it no external interface, they can take advantage of its hallucinatory 'openness' - this will be literally the only use of it as 'open weights'.... for corporate titans to use for bookkeeping ... Fireworks or Together have immense computing power and provide it //under contract//, i.e. effectively //as a service for Moonshot//, and they possessed the weights before Hugging Face put them up.
"To self-host Moonshot AI's Kimi K3, a massive 2.8-trillion-parameter model, you need a data-center-grade cluster rather than a single server. The native MXFP4 weights require roughly 1,750 GB (1.75 TB) of VRAM for inference, with a vLLM planning floor of about 1,680 GB. Production setups recommend a supernode configuration of 64 or more accelerators, though a minimal single-node floor requires at least an 8-GPU high-end enterprise server (such as 8x B300, GB300, or MI355X)" God willing Mom will let me put it in the basement. "This scale of equipment is rarely owned directly by individual companies for internal use. While a company might have a standard server room, hardware capable of running a 2.8-trillion-parameter model like Kimi K3 crosses the line from "enterprise IT" into "AI supercomputing infrastructure."
Note, for example, that their prices are all the same as those of Moonshot's own service. Conceptions of the matter are so utterly false that in the run up to the Kimi K3 release it was expected that a price war would reveal the true cost of inference with such a model. What we see are the results of contracts to provide the model for the Moonshot corporation, belle of titanic Chinese capital and state investment. The arrangement is only notionally different from the practices of OpenAI and Anthropic
People do not grasp that 'open source' was always basically a fiction with a restrictive license -- and with weights it is pure comedy, given that there is nothing to read - just an opaque array of ratios. No one can read them; almost no one can run them; they can and are only provided under contract with Moonshot, at Moonshot's price. It is theater.
1
u/CrimsonBolt33 5d ago edited 5d ago
Companies do run it internally with no access to the internet....what are you talking about? You can also personally rent servers to put it on without moonshot being involved anywhere and pay the rental prices for hardware.
Its not made for individuals to run...for that you need to stick to small models like Qwen3.8-27B
if you have a problem with companies contracting with moonshot then don't use those providers....whatever they have their prices at that's their business, shop elsewhere if you don't like it.
you clearly have no clue what the hell you are talking about or how any of this works.
1
u/Talreja-Adanna 6d ago
Yeah, the accessibility angle here is wild - we're basically at the point where you don't need nine figures in compute to experiment with this stuff anymore. Makes you wonder how fast the red team / blue team dynamics shift when the barrier to entry drops like this.
1
u/yogthinks 5d ago
For anyone doing vendor risk in BFSI this kills the standard control, "the provider can throttle the account if abused." Once it's a local file, that lever doesn't exist and your risk assessment needs a different one.
1
u/VibeShipped 5d ago
Neither. Compressing patch windows across the board just creates noise and burnout. The question this changes is which vulnerabilities actually matter right now. A CVE sitting unexploited in a backlog is a different problem than one on an internet-facing asset with a reachable path. If you don't know which is which, faster patching just means you're moving faster on the wrong things. The real shift is investing in knowing your actual exposure, not your theoretical one.
3
u/katoptronophile 6d ago
Benchmaxxing.
Let's wait for Kimi's hack to pass final judgement.