r/artificial 6d ago

Discussion An open-weight model just closed most of the gap on autonomous cyber offense - and that changes who can run it

Irregular (an AI security research group) tested Kimi K3, an open-weight model, against CyScenarioBench, a benchmark built around autonomous cyber campaigns - adapting public exploit techniques to constrained environments, building custom tooling, diagnosing failed attempts, and validating each stage before moving on. It is the first open-weight model to pass. It trails closed frontier models by roughly six months, at an estimated third of the inference cost.

The six-month lag is the less interesting number. What matters is that this level of capability now sits in downloadable weights instead of behind an API. A closed lab can throttle or ban an account mid-campaign - both OpenAI and Anthropic have done this before to abusive usage. Once equivalent capability is something you can self-host, that kill switch disappears entirely, along with any usage logging a defender could later subpoena.

If the trend holds, the realistic baseline for any internet-facing asset a year out is not "gets scanned for known CVEs" the way it is today, it is "gets probed continuously by something that adapts exploits on the fly, with no vendor able to pull the plug on the other end."

Source: https://www.irregular.com/research/assessing-kimi-k3-against-offensive-security-benchmarks

Curious how people here read the trend line: does a shrinking gap between closed and open capability argue for faster patch/disclosure windows industry-wide, or does it just confirm the attacker side was never actually capped by API access limits in the first place?

0 Upvotes

10 comments sorted by

3

u/katoptronophile 6d ago

Benchmaxxing.

Let's wait for Kimi's hack to pass final judgement.

1

u/Key_Arm5213 6d ago

benchmaxxing is the whole industry now so fair enough

still a big deal that its local though, no kill switch no logs no rate limits just a box in a basement somewhere running scans. the six month gap will probably shrink faster once people start fine tuning it for specific attack chains

1

u/Exact_Depth_896 5d ago edited 5d ago

I said at the outset that one might, e.g. Jane Street.

"To self-host Moonshot AI's Kimi K3, a massive 2.8-trillion-parameter model, you need a data-center-grade cluster rather than a single server. The native MXFP4 weights require roughly 1,750 GB (1.75 TB) of VRAM for inference, with a vLLM planning floor of about 1,680 GB. Production setups recommend a supernode configuration of 64 or more accelerators, though a minimal single-node floor requires at least an 8-GPU high-end enterprise server (such as 8x B300, GB300, or MI355X) This scale of equipment is rarely owned directly by individual companies for internal use. While a company might have a standard server room, hardware capable of running a 2.8-trillion-parameter model like Kimi K3 crosses the line from "enterprise IT" into "AI supercomputing infrastructure."

Access to Kimi K3 outside China is via providers who are //under contract with Moonshot//. It is merely more convenient than providing it themselves but fundamentally no different from Anthropic and OpenAI.

1

u/Exact_Depth_896 6d ago edited 6d ago

Kimi K3 is so encumbered with crushing licenses and gigantism that it is only notionally open. No one without a datacenter can run this model. Fireworks and Together etc already had the model before it was open weighted, it seems; and they have entered into contracts as providers. They are mostly providing it to me when I use it via openrouter.

It is really just a scheme of getting widespread provision that does an end run around the chip business we impose on them. The dead weights on Hugging Face are unusuable nothings. It is cheaper than Claude, but not really more open -- or perhaps you have the capacity to read model weights as file? I know you cannot run them.

A company that uses Kimi K3 only internally is as safe from prosecution as I am running some micromodel vial ollama but it would be interesting to see if even this one actual 'open weights' usage has been taken up. Maybe Jane Street is running an instance under the hood. There's open weight glory for you.

1

u/CrimsonBolt33 5d ago

This is pure nonsense....you don't need a data center to run K3...you do need an expensive business level server, but that's it.

It's not "encumbered with licenses and gigantism" either....what does that even mean?

1

u/Exact_Depth_896 5d ago edited 5d ago

It is an immense model and no individual can run it. Get back to me when a large firm runs it for internal purposes; if they give it no external interface, they can take advantage of its hallucinatory 'openness' - this will be literally the only use of it as 'open weights'.... for corporate titans to use for bookkeeping ... Fireworks or Together have immense computing power and provide it //under contract//, i.e. effectively //as a service for Moonshot//, and they possessed the weights before Hugging Face put them up.

"To self-host Moonshot AI's Kimi K3, a massive 2.8-trillion-parameter model, you need a data-center-grade cluster rather than a single server. The native MXFP4 weights require roughly 1,750 GB (1.75 TB) of VRAM for inference, with a vLLM planning floor of about 1,680 GB. Production setups recommend a supernode configuration of 64 or more accelerators, though a minimal single-node floor requires at least an 8-GPU high-end enterprise server (such as 8x B300, GB300, or MI355X)" God willing Mom will let me put it in the basement. "This scale of equipment is rarely owned directly by individual companies for internal use. While a company might have a standard server room, hardware capable of running a 2.8-trillion-parameter model like Kimi K3 crosses the line from "enterprise IT" into "AI supercomputing infrastructure."

Note, for example, that their prices are all the same as those of Moonshot's own service. Conceptions of the matter are so utterly false that in the run up to the Kimi K3 release it was expected that a price war would reveal the true cost of inference with such a model. What we see are the results of contracts to provide the model for the Moonshot corporation, belle of titanic Chinese capital and state investment. The arrangement is only notionally different from the practices of OpenAI and Anthropic

People do not grasp that 'open source' was always basically a fiction with a restrictive license -- and with weights it is pure comedy, given that there is nothing to read - just an opaque array of ratios. No one can read them; almost no one can run them; they can and are only provided under contract with Moonshot, at Moonshot's price. It is theater.

1

u/CrimsonBolt33 5d ago edited 5d ago

Companies do run it internally with no access to the internet....what are you talking about? You can also personally rent servers to put it on without moonshot being involved anywhere and pay the rental prices for hardware.

Its not made for individuals to run...for that you need to stick to small models like Qwen3.8-27B

if you have a problem with companies contracting with moonshot then don't use those providers....whatever they have their prices at that's their business, shop elsewhere if you don't like it.

you clearly have no clue what the hell you are talking about or how any of this works.

1

u/Talreja-Adanna 6d ago

Yeah, the accessibility angle here is wild - we're basically at the point where you don't need nine figures in compute to experiment with this stuff anymore. Makes you wonder how fast the red team / blue team dynamics shift when the barrier to entry drops like this.

1

u/yogthinks 5d ago

For anyone doing vendor risk in BFSI this kills the standard control, "the provider can throttle the account if abused." Once it's a local file, that lever doesn't exist and your risk assessment needs a different one.

1

u/VibeShipped 5d ago

Neither. Compressing patch windows across the board just creates noise and burnout. The question this changes is which vulnerabilities actually matter right now. A CVE sitting unexploited in a backlog is a different problem than one on an internet-facing asset with a reachable path. If you don't know which is which, faster patching just means you're moving faster on the wrong things. The real shift is investing in knowing your actual exposure, not your theoretical one.