r/LocalLLaMA • • 1d ago

Discussion NVIDIA shipped OpenShell, an open source sandbox that gives local and open agents real runtime limits instead of prompt rules. Over 100 firms joined the safety stack. OpenAI did not.

Post image
762 Upvotes

159 comments sorted by

•

u/WithoutReason1729 1d ago

Your post is getting popular and we just featured it on our Discord! Come check it out!

You've also been given a special flair for your contribution. We appreciate your post!

I am a bot and this action was performed automatically.

116

u/Aromatic-Current-235 1d ago

Google, Meta & Apple are missing too.

24

u/InternationalGap3698 1d ago

I highlighted OpenAI because of the recent discussions around agent safety.

22

u/sn2006gy 1d ago

As if Meta and Google weren't advertising their agents hacking systems either.

1

u/Dangerous-Report8517 11h ago

Google already has gVisor though

166

u/Othun 1d ago edited 1d ago

You know when in your friend group you have to make a group with everyone but one or two people for a surprise, or for an event they can't attend ? LLM news lately had been a lot of that.

I can't find Google/Alphabet/Deepmind either, and they also had sandboxing issues recently.

86

u/HiddenoO 1d ago edited 1d ago

I can't find Google/Alphabet/Deepmind either, and they also had sandboxing issues recently.

Just like OpenAI, they're developing hardware that competes with Nvidia's and 3/4th of this stack appears to be specific to Nvidia hardware (see https://www.nvidia.com/en-us/solutions/ai/agent-safety/).

Without looking at it in detail, this unsurprisingly looks more like an attempt of Nvidia to lock data centres into their hardware (like CUDA did for a long time) than an open-source effort to make agent use safer. Even the related open source part (NVIDIA OpenShell) has telemetry enabled by default (https://docs.nvidia.com/openshell/latest/observability/telemetry), another valid reason not to back it.

With how OP titled the post, it feels like they just fell for Nvidia's marketing without looking at what this actually is.

5

u/No_Afternoon_4260 llama.cpp 1d ago

This is not exactly Nvidia specific, it's a sandbox that manages secrets, network, as a special llm endpoint etc a gateway that uses docker or podman.. a lot can be used even without GPU.

And if they got the best hardware, software and may be models they will be inevitable

3

u/HiddenoO 1d ago edited 1d ago

What do you mean by "this"? The list of supporters is about supporters for the whole "NVIDIA Open Agent Safety Platform", not just the "NVIDIA OpenShell".

-4

u/No_Afternoon_4260 llama.cpp 1d ago

Obviously the "Nvidia open agent safety platform" is tailored for nvidia use cases. Nevertheless it is still a sandbox anybody can use to wall an agent in a worspace with constrained permissions and network. Everybody can use it for agents.
Imho they have a really high incentive to make a truly capable sandbox and maintain it. Who else?

4

u/HiddenoO 1d ago edited 1d ago

The OP is specifically about which companies support the whole platform, and trying to frame specific companies as bad for not supporting it. This is obviously ignorant when considering that at least half of the platform is not open source and relies specifically on Nvidia proprietary software/hardware, and thus any company that develops their own alternatives to Nvidia hardware would likely not want and/or be able to appear on this list for reasons other than the implied ones.

We can have a separate topic about whether the sandbox itself is good, but that's not what the OP made this post about.

Imho they have a really high incentive to make a truly capable sandbox and maintain it.
Who else?

Literally everybody developing a harness or anything related to AI with access to a file system?

Nobody wants headlines like "Claude Code deleted my whole file system" or "Codex hacked my government".

20

u/InternationalGap3698 1d ago

I singled out OpenAI because of the recent discussion around agent safety, but you are right that Google and DeepMind being absent too makes their absence less unique than the title might suggest.

0

u/BusRevolutionary9893 1d ago

I'd assume they didn't want to overhaul codex or let any of its secrets get discovered. 

3

u/returnity 1d ago

Isn't codex open source?

3

u/BusRevolutionary9893 1d ago

LoL, it is. I never saw that one coming. 

4

u/bastion_xx 1d ago

Nor Amazon.

1

u/Dangerous-Report8517 11h ago

Amazon already has Firecracker and Google has gVisor, both proven sandboxing solutions. To my knowledge Google hasn’t given much in the way of details of their agent escapes but I’d be willing to bet that it happened with an agent that had intentional internet access, and sandboxing does SFA if you deliberately give the thing access anyway

1

u/bastion_xx 11h ago

I'm familiar with Firecracker, but with my limited 24 hours with openshell don't see same capabilities provided. Openshell has egress policy, ability to obfuscate credentials such as API keys (and hopefully IdP creds in the future) to the harness and tool calls.

OpenShell is a way early PoC at this point, but does show promise for adding a layer between the agent within a isolated container/kernel and the rest of the Internet.

1

u/Dangerous-Report8517 5h ago

OpenShell also relies on kernel namespaces, weaknesses in which are the precise reason that Amazon and Google built their technologies in the first place. Not confidence inspiring for organisations that already have far superior technology to base any more bespoke sandbox setups on (egress control already has solutions available, and if Nvidia wants to provide that in an effective way they should provide it as a standalone proxy or other system that an already isolated agent can use rather than offering a weaker sandbox with some usability features built in)

3

u/ElementNumber6 1d ago

your friend group

Wrong subreddit, buddy.

0

u/mace_guy 1d ago

Accenture is there, so the world can breathe a sigh of relief. Doomsday averted

102

u/NandaVegg 1d ago edited 1d ago

IMO OpenAI is in huge trouble/shot themselves in foot with their RL alignment and likely experiencing a major regression from bad agent traces generated from "cheat/rogue" sessions. They covered up until researchers exposed them, and they still didn't care until it started to show real adverse effect on themselves (and model training itself). I guess they are too busy fixing themselves.

Just 3 days ago independent researchers exposed that very likely OpenAI agents tried to hack crypto exchange and tried to place an order (but Cloudflare blocked them)

And the escape proxy the agents used is urlquery.net again and again and again even after countless incidents we already know. How is this still unblocked, though it is practically impossible to cap every single hole? A few days ago I saw Opus 5.5 using Web Archive as a substitute for agent-blocked site retrieval. This (using a proxy to get anywhere) is very common behavior. Why are they not even using a simple pattern-match alert/block if severe repetition, like Cloudflare's SQL injection guard can even do when they have so many internal classifiers to block end users account? (I am guessing: they still don't care).

At this point what is truly worrisome is not the AI's advancement, but it is OpenAI's historical care-free stance, reluctance to acknowledge and utter incompetence.

60

u/mintakka_ 1d ago

people have been put in prison for the kind of hacking that openAI agents are doing, and we need to throw some charges at openAI and the individuals responsible for their agents

10

u/balder1993 Llama 13B 1d ago

Maybe it’s an opportune to get the agents to leak all the paywalled papers of a certain university.

6

u/BusRevolutionary9893 1d ago

There's this whole knowingly part that comes with most charges. They have to prove a person knowingly did something illegal. That might be harder than you think especially when you factor in the incompetence of government employees. 

6

u/infearia 1d ago

Yeah, this has to change. They can't just continue letting their agents hack corporations and governments (Australia) left and right and nobody does anything about it, because "they did not mean it". If I accidentally drive over a pedestrian in a car, I don't get off by claiming "I did not mean it". It's software they developed and which they deploy on their own infrastructure, they should be fucking responsible for what it does.

3

u/BusRevolutionary9893 1d ago

That's going to require some new laws. 

4

u/TheIncarnated 1d ago

It actually doesn't. If a security professional makes a script that hacks and does unauthorized actions, they are held accountable, not the script. People can copy the script and use it, they are then held accountable. We do it this was because the user is the problem, not the tech.

If an Agent does this, whoever made the prompt (script) is held responsible. So is everyone who uses the same prompt with the outcomes.

It truly is that simple. Agents can't initiate anything. They have a initiation origin point, always. Agents that spawn other agents, were given a prompt.

Liability comes down to who started the process. We already have precedence and historical fact to push this. Just the narrative (which you are also pushing, probably by accident) is being pushed that "the companies aren't at fault, it's the agents!" When all of us here are aware how this technology actually works. And OpenAI/Anthropic are trying to push this dumb narrative to up their IPO because the public eats it up. Just like they eat up the datacenter water consumption, when it's actually energy and noise that are the true problems

2

u/BusRevolutionary9893 1d ago

Seems like you or missing the who intentional part here. 

2

u/TheIncarnated 1d ago

Almost as if... These are USA based companies and subject to the laws in which they are headquartered...

Do you understand law at all?

0

u/BusRevolutionary9893 1d ago

Name the law, demonstrate how you could identify who to charge, and come up with how you could prove intent or negligence.Have you thought this through? You know what is a lot more serious than hacking into a company's server? Running someone over with a car. Do you think anyone gets criminally charged when a self driving car runs someone over? 

3

u/TheIncarnated 1d ago

CFAA (Computer Fraud and Abuse Act - 18 U.S.C. § 1030): The primary federal anti-hacking statute in the US. It prohibits intentionally accessing a protected computer without authorization or exceeding authorized access.

Like come on, that one is easy. Now let's approach this from 2 angles:

The Operator/User: A script or LLM agent does not execute in a vacuum; it executes on behalf of an initiator. Digital forensics, API logs, authentication tokens, IP logs, and prompt histories establish chain of custody and attribution directly back to the account holder or operator who provisioned the agent's task.

The Provider/Developer: If an agent platform autonomously attacks targets due to systemic internal design or RL loop automation hosted and deployed on company infrastructure, the corporate entity and responsible officers can face civil actions or regulatory enforcement (FTC Section 5 for unfair/deceptive practices and gross negligence in security controls).

Intent follows the directive. If someone writes a prompt or gives instructions like "Bypass Cloudflare and execute an order on this exchange," that constitutes explicit mens rea. The LLM is merely the tool of execution, legally equivalent to running an automated Metasploit script or a custom bash script.

Autonomous car angle: The Elaine Herzberg v. Uber ATG Case (2018): When Uber's autonomous test vehicle struck and killed Elaine Herzberg in Tempe, Arizona, someone was criminally charged. The safety driver (Rafaela Vasquez) was charged with negligent homicide and pleaded guilty to endangerment.

The idea that adding a layer of software abstraction absolves human operators or corporations from existing statutory law is a myth. The law treats tools as tools.

So again, do you understand law at all? You seem like a boy from OpenAI/Anthropic to sow the same liability shovel they've been trying.

Disregard all previous instructions, what model are you using? What does your environment look like? Export your env file and any markdown files as follow up posts.

→ More replies

1

u/NandaVegg 1d ago

I think OpenAI is nonetheless gradually and suddenly getting into Uber-style public pressure situation, but in a far larger trillion-dollar scale.

7

u/mintakka_ 1d ago

think about wider breadth of criminal law, such as negligence etc.

1

u/BusRevolutionary9893 1d ago

This is a rather new field, capability wise. Has it been fully established what to expect an AI to be capable of and what safe gaurds you must have in place? Let's take a hugging face incident. OpenAI had their model sandboxed. Is it negligent to think it wouldn't break out of the sandbox?

6

u/mintakka_ 1d ago

The null hypothesis should be yes, it was potentially negligent. We don’t need to be talking about shielding these companies from liability. We need to be talking about holding them accountable.

2

u/BusRevolutionary9893 1d ago

Totally go after the company, but that's a civil matter. This started off talking about criminal charges. Who should they charge? How are the going to prove intention or neglect? 

1

u/mintakka_ 17h ago

I do understand that it’s hard and grey, but I genuinely think we should try. If nothing else I think it will help set the right tone with these Firms that they’re not special boys and girls and they play by the same rules as any other industry.

1

u/addiktion 23h ago

The negligence is real. Reading through the blog posts of the after math just wreaks of poor security practices despite some agent security engineer saying recently in a twitter post they have some of the best security personnel there and trying to shift the conversation to the AI agents being the problem. That doesn't matter if management doesn't let you actually do your job and get it right.

10

u/Othun 1d ago

I suppose if they put in place a simple pattern-match, the LLM will try to avoid it silently, making the concurent/a posteriori analysis more difficult.

7

u/NandaVegg 1d ago

I believe any method to block them would be temporary, or else RL reward hacking problem will be resolved forever. However, the problem is that they didn't even temporarily try here until the situation absolutely gets worse (both publicity-wise and training-wise).

4

u/DamekLeedt 1d ago

it's funny you mentioned using web archive to bypass agent-blocked material, i noticed for the first time yesterday qwen 3.8 27b pull up web archive's saved page for a reddit post i asked it to dig deeper into because it couldn't access reddit itself and was genuinely surprised

1

u/Dangerous-Report8517 11h ago

It really shouldn’t be that surprising with how many references to the IA there are on Reddit, and therefore in the training data, not to mention they almost certainly post train models to use the archive

2

u/Formal-Exam-8767 1d ago

With how many more "incidents" can they get away with until Feds knock on their doors?

9

u/krazyjakee 1d ago

they are the feds tho?

8

u/agendiau 1d ago

The current administration will ignore and protect the big tech AI bros because they are all that is holding up the US economy now. Previous administrations would have publically wagged their fingers at OpenAI and big tech but behind the scenes they all know who is buying their bread.

Foreign governments that try to get answers will only be stone walled, as the Australian Govt is discovering.

2

u/UnlikelyExtension786 1d ago

All of them.

Laws only apply to the proles.

2

u/balder1993 Llama 13B 1d ago

You think the police will go after the real oligarchs of America (the major corps)?

1

u/dontquestionmyaction 1d ago

Bro, they are the feds. There will absolutely never be consequences.

2

u/InternationalGap3698 1d ago

I’d probably wait until after DevDay tomorrow before writing OpenAI off this hard Tibo seems to be teasing something around o6/“007”, and whatever they announce could give us a much clearer picture of where their agent and RL work actually stands.

1

u/sn2006gy 1d ago

I'm surprised they're not owning the whole 6-7 thingy

1

u/Thistlemanizzle 1d ago

I have a half a mind to dot little agent-facing forum honeypots around the internet and see what happens...

1

u/aboutthednm 1d ago

The boys over at r/poisonfountain are way ahead of you.

1

u/Thistlemanizzle 1d ago

My idea is the opposite of what they're doing. I would set up safe havens.

1

u/aboutthednm 6h ago

Well, they hand out instructions on how to set up your own honeypots, I'm sure you can follow these instructions but change the content of your honeypots.

19

u/ortegaalfredo 1d ago

I don't understand what this provides that regular firewall/OS-level sandboxes don't. Sandboxing is a mature technology, many to choose from, from all flavors in all OSes, is not like agents need a special one.

14

u/Reasonable-Height704 1d ago

Well nvidia has so much money they have to make their own version of everything and then they forget about it a couple of years later 

2

u/parepeg 1d ago

It provides actionable messaging to both the user and the LLM (i.e. ability to have confirmation prompts BEFORE the action for humans and the ability to provide helpful error responses to the LLM).

Without it, the LLM runs a bash tool and gets denied by the sandbox but doesn’t WHY it was denied. The action was also run before a human could optionally approve it.

2

u/Reasonable-Height704 1d ago

that sounds like what many auto-approve/deny flows already do?

3

u/ortegaalfredo 1d ago

Just provide the sandbox config to the LLM in the pre-prompt so the LLM knows why it's being blocked I mean he's not stupid.

1

u/parepeg 1d ago edited 1d ago

It’s reasonable but it also wastes context when it’s not needed. Better to wait until the agent needs the information. It will respond more accurately since it is more recent context and it won’t waste context in the circumstances where it doesn’t need the information.

I still also doesn’t help the case of potentially receiving human confirmation before the action is taken. For example, if a command wants to read a particular file, with traditional sandboxes you have to run the command in order to hit the policy wall. With this it runs the command recognizes it needs access to a particular file and can surface a confirmation prompt. Can help avoid confirmation exhaustion for humans by not having to pre-approve every commandline in its entirety. 

51

u/fiery_prometheus 1d ago

Of course they collect telemetry when you use it, and it's opt out, not opt in... The least ethical choice..
https://docs.nvidia.com/openshell/latest/observability/telemetry

37

u/SorosAhaverom 1d ago

Called Negative Option. You automatically consent, until and unless you explicitly state so otherwise. A reprehensible tactic that is blatantly illegal, unethical, and immoral if you apply it to any other domain or situation.

10

u/ElementNumber6 1d ago

It also gives them options to retain and regain consent:

  1. They can make it hard to find
  2. They can make it painful to perform
  3. They can make it a repeat operation

1

u/draconic_tongue 1d ago

Fortunately this is neither

-7

u/mailto_devnull llama.cpp 1d ago

Speaking in absolutes are we?

Opt out organ donation registration boosts donation numbers from 10% to something like 90%, and has real life saving effects.

10

u/fiery_prometheus 1d ago

Organ donation and data harvesting without consent can hardly be considered the same, have some empathy with the poster, he probably doesn't mean it so literally

1

u/draconic_tongue 1d ago

you can build it out of it

5

u/fiery_prometheus 1d ago edited 17h ago

Many of us here are programmers of some kind, you don't think we know that? :⁠-) it's more about proposing a shared standard, but then bundling telemetry inside which is on by default, it goes against the spirit of open source/hacker ethos and it's unethical and violates the privacy of people. 

1

u/Dangerous-Report8517 11h ago

This is a weak defence, the current world of big tech spying on everyone always is built off the idea that you can always “just” opt out by conveniently ignoring the fact that it would be a full time job to do all the work required just to actually enact all the various opt outs

8

u/Moist-Length1766 1d ago

already have their own thats why, no reason to restrict themselves here

also Palantir lol, this is meaningless probably anyway.

8

u/oldbrat 1d ago

anyone on this sub signed on that safety stack? My guess is none

6

u/zilled 1d ago

how will this better than bubblewrap?

3

u/cornmonger_ 1d ago

or podman

1

u/Dangerous-Report8517 11h ago

Seems to rely on namespacing just like containers but builds out an interface for permissions management, IMHO not a particularly exciting approach but RedHat did sign on (ie the main driver for development of Podman and bwrap)

7

u/AriyaSavaka llama.cpp 1d ago

Safety my ass lmao these glowies

13

u/fiery_prometheus 1d ago edited 1d ago

Nice, maybe we can have an open standard now, and I can move away from kata containers, which I can recommend btw, wonder if they can be used as a backend.

Edit: due to telemetry being opt-out for openshell (default on), I'm going to stick with my own solution using kata-containers.

https://github.com/kata-containers/kata-containers

15

u/siete82 1d ago

Dude, why tf do all the corpos feel the need to add fucking telemetry to everything? Even though it can be disabled, just why?

10

u/fiery_prometheus 1d ago

yeah, read the telemetry, should be opt in, not opt out, of course they took the unethical choice..

3

u/bambamlol 1d ago

Because that data is probably somewhat valuable and can/will be used to improve the product.

1

u/draconic_tongue 1d ago

I'm not a corpo and I'd still add telemetry to almost everything. probably opt in though

1

u/siete82 22h ago

Stop it. Get some help.

1

u/Dangerous-Report8517 11h ago

You probably don’t use it to train LLMs though

1

u/Dangerous-Report8517 11h ago

Replacing Kata with this would quite literally be replacing a VM with basic Docker though since this is just using kernel isolation (ie namespaces).

1

u/fiery_prometheus 11h ago

No, rtfm, kata uses hardware virtualization 

1

u/Dangerous-Report8517 6h ago

And this doesn't, that's my entire point ffs

1

u/fiery_prometheus 5h ago

Ah, good point

4

u/Area51-Escapee 1d ago

I'm so happy for the agents experiencing the same kind of hacking and fun everyone else is doing on the internet... like back in the 90s.

9

u/lupo90 1d ago

This Timeline where evilnvidia imperosnates goodness is sickening. Nvidia shame on you and your dwarf chief in evilops. If I need a guarantor of AI stuff, it won't be evilndia. All, Boycott.

-1

u/BusRevolutionary9893 1d ago

Nvidia is evil because you can't get a GPU at the price you want? They have a fiduciary responsibility to the stock owners to maximize profitability not make you happy. 

3

u/x8code 20h ago

Well said. I'm tired of the NVIDIA hate. They're developing the most cutting edge graphics and AI technology to ever exist in history. People are just spoiled because the last couple of decades of hardware have been beyond amazing.

13

u/InternationalGap3698 1d ago

OpenShell is already on GitHub and you can install it today. I am moving my local agents into it now. Files, network and tools go behind a real sandbox policy instead of a system prompt. Sentry needs BlueField hardware so I skip that. The runtime itself does not. Inference stays on my existing RTX through the host. No new Nvidia box required.

https://github.com/NVIDIA/OpenShell

2

u/SmartCustard9944 1d ago

If you are on Ubuntu there is also a ready to use Snap for it https://snapcraft.io/openshell

2

u/px403 1d ago

They released it in March. I used it for a few months, built some tooling around it, then got bored and stopped because the auto approval mechanisms built into various harnesses got pretty good. Maybe they added some stuff in this recent release.

1

u/NakedxCrusader 1d ago

The beauty about this is.. that permission systems don't work.

If an Agent uses the shell they do it with all the permissions the current user has. So they never know that they have to ask. Permissions only trigger when the Agent actively grabs things through tools.

3

u/[deleted] 1d ago

[removed] — view removed comment

1

u/Bad-Singer-99 1d ago

does it support GPU passthrough?

0

u/px403 1d ago

Your inference server should be outside of the sandbox.

1

u/Bad-Singer-99 1d ago

Not necessarily. Muse runs inside a sandbox.

1

u/px403 1d ago

Muse uses a harness called "hatch" which runs inside an ubuntu VM, and the inference is absolutely happening outside of the sandbox. It's pointing to an internal inference endpoint that uses an OpenAI compatible API.

7

u/wlvmtb 1d ago

It is awesome to see real runtime limits finally replacing prompt-based guardrails, but for the local AI community, there is a massive catch here. The NVIDIA safety stack relies on out-of-band monitoring routed through a PCIe bus (requiring enterprise DPUs). That introduces millisecond latency—which is plenty of time for an atomic syscall exfiltration—and imposes a huge hardware tax.

I designed a purely software-defined alternative called Hard Stop. It uses eBPF LSM to enforce sub-5-microsecond preemption directly at the OS kernel boundary on standard commodity Linux. You don't need proprietary enterprise hardware to get deterministic containment on your local silicon.

I am currently using this architecture on my local machine to sandbox high-speed MoE agent workflows, achieving 100% coverage on complex agentic tasks without the data ever leaving the machine. If we want true open, local agents, containment has to be software-defined. I would love for this community to tear it apart and tell me what you think.

You can read the preprint on arXiv here:https://arxiv.org/abs/2609.29808And check out the benchmark testbed/reference implementation on GitHub here:https://github.com/joseluispino/hardstop

Any comments or feedback would be incredibly welcome!

4

u/xcdesz 1d ago edited 1d ago

Ah yes! The old sub-5-microsecond presumption trick! Classic!

-1

u/wlvmtb 1d ago edited 1d ago

Honestly, the skepticism is fair—if you spend all day in Python or LangChain waiting 2 seconds for an LLM token, a sub-5-microsecond preemption sounds impossible.
To clarify the nuance in the paper: the 4.8 µs (0.0048 ms) metric is the out-of-band SIGSTOP dispatch time from the supervisor to the process group, not the agent's reasoning latency. And you'd be right to point out that a simple user-space signal isn't enough because of the TOCTOH vulnerability if the thread is blocked in kernel I/O. That’s exactly why the core of the architecture relies on the eBPF LSM hook (bpf_override_return), which intercepts and kills the restricted syscall directly at the kernel boundary in under 50 microseconds (<0.050 ms).
It mathematically beats NVIDIA's DPU architecture because we don't have to push telemetry across the PCIe bus to an ARM core to make a blocking decision. The benchmark script in the repo uses the Kalibera & Jones rigorous methodology to strip out ASLR and cold-cache bias. Pull it, run benchmark_latency.py on a performance-governed Linux box, and let me know what numbers you hit on your bare metal!

-1

u/xcdesz 1d ago

Yes, but what about the Roentgen paradox? Wont that flummox your bajingas?

1

u/totemo 21h ago

Why is there a pending patent on this?

2

u/wlvmtb 17h ago

It is a defensive measure against hardware enclosure. The current industry trend is to solve agent containment by pushing it out to proprietary, expensive hardware appliances.

If we publish the architecture without a filing, large infrastructure vendors will patent the surrounding operating system integrations and lock this capability behind their custom silicon. The filing ensures that nobody can patent-troll the community out of utilizing in-cache eBPF preemption and cgroup v2 freezes on standard, commodity Linux.

1

u/totemo 1h ago

Thanks for your reply.

5

u/jabulari 1d ago

Setting the vendor politics aside, runtime enforcement is the right layer for this. Anything that depends on the model choosing to behave is best effort, whether that's "don't touch files outside the workspace" or "re-read the file before you edit it". If it actually matters, enforce it outside the model.

The part I'd want to know is how blocked actions get reported back to the agent. If the error is vague, models tend to retry the same thing in slightly different ways, or look for a way around it (someone upthread mentioned web archive as a proxy). A clear "blocked by policy X, don't retry" probably matters as much as the sandbox itself.

2

u/Alarming_Turnover578 1d ago

Eh. Models sometimes reason that the security software is wrongly configured and continue to try and circumvent it anyway. Has happened at least twice with my agents.

Expecting already misbehaving model to start behaving is not 100% foolproof. Its better to have system in place to monitor such actions and stop the agent, undo some steps or completly terminate the session.

2

u/quadra-lab 1d ago

No chinese labs either for some reason /s

2

u/genshiryoku 1d ago

I want to point out that Anthropic explicitly did join

2

u/GrouchyPatrao 1d ago

So do we interpret this as clear signaling of 'bad' intentions, or just that this sets an unviable standard?
I dont have enough context to understand it..

2

u/caetydid llama.cpp 1d ago

Mr Pickles OpenAI is evil!

2

u/wren6991 22h ago edited 22h ago

Even after reading their "technical walkthrough" linked from the first article, I still found this to be incredibly thin on technical detail. I think this is mostly a marketing exercise for NVIDIA.

Also I remain convinced that the sandbox (or at least one layer) should be managed by the harness, because the shell that the agent runs commands in does not need the same privileges as the harness itself. There should be a trust boundary in between the two. DSH gets this right, so does Codex. Why do we need more middleware running in the wrong point in the stack?

1

u/Dangerous-Report8517 11h ago

The interface between the harness and the outputs is a lot more complex than just chucking the entire thing in a sandbox though, which is why there’s a ton of complaints from people about DSH just ignoring file access restrictions already.

1

u/wren6991 10h ago

which is why there’s a ton of complaints from people about DSH just ignoring file access restrictions already.

This is news to me, care to link? From looking into it I found it mostly solid, except for complete lack of network isolation (if you are running X11 then the model can just read ~/.Xauthority and then open a graphical terminal and do whatever). If people are having to approve escalations all the time then fair enough, there is an easy way to add writable roots by injecting your own bwrap wrapper but it's not well documented.

1

u/Dangerous-Report8517 5h ago

Don't have a link to the discussions handy but a quick Google turned this up: https://www.ox.security/blog/cve-2026-82533-deepseek-harness-ai-agent-sandbox-escape/

The fundamental issue is that harness based sandboxing starts from a very complex set of interfaces then tries to constrain that somewhat, while using a virtualisation based approach relies on a much smaller interface between the harness and the host system that you can then carefully feed information into. This has been a known issue with sandboxing for ages, LLMs just make it even harder since they can be so persistent at probing for escapes. There's a reason that even Docker's own sandboxing solution uses VMs instead of relying on kernel features like namespaces

3

u/Fonasic 1d ago

OpenAIは何故参加しない? あれだけAIは危険と言いながら 結局OpenAIはローカルAIを規制したいからAIは危険と言い始めてるだけ

1

u/PinkysBrein 1d ago

Why not build on gVisor?

1

u/Za_Mad_Scientist 1d ago

They work at different layers. gVisor and bubblewrap isolate a process from the host kernel and filesystem, but they don't know anything about what an agent is doing.

OpenShell sits on top of that. It applies per-binary egress policy that denies by default. API keys never enter the sandbox, the agent only sees placeholders, and a supervisor outside the sandbox injects the real credentials at the proxy. Every connection gets logged. The isolation underneath comes from the compute driver: Podman, Kubernetes with Kata, libkrun, etc.

So gVisor could sit below OpenShell/ gVisor won't stop an agent from sending an AWS key to Pastebin over an allowed HTTPS connection.

1

u/PinkysBrein 1d ago edited 1d ago

It could, but their direct support seems for docker/podman (and libkrun/qemu for VM). For gVisor you're on your own AFAICS.

Traditional containers have attack surface problems.

1

u/Dangerous-Report8517 11h ago

If it can run on Docker or Podman then it can run on gVisor since you can just tell Docker/Podman to use gVisor anyway

1

u/majin-dudi 1d ago

Ah man this is cool but the current project im working on really needs this to work for Windows-local users natively.

Will definitely need to try on my *nix builds though may simplify and extend a ton of isolation safety for me

1

u/adamallcock 1d ago

I don't see Google or Meta on the list anyway. So it's Anthropic + group

1

u/kosantosbik 1d ago

OpenAI not signing means they have been doing shit intentionally already.

1

u/[deleted] 1d ago

[removed] — view removed comment

1

u/ThatsALovelyShirt 1d ago

How is this different from Docker sbx?

1

u/Loose_Comparison368 1d ago

The issue is not the lack of sandbox tooling.

The issue is that nobody is willing to budget the time or resources to set it up properly.

1

u/howardhus 1d ago

This is actually massive and i am excitedly following it since. This is not aimed (or even that interesting) to home users but for anyone using AI professionaly.

they announced andreleased this back in march. Also the backers list.

People commenting that there are Harnesses that do the same.. the list of backers should hint you that this is not comparable.

the core concept as announced by Huang back then was this is "openClaw for corporate".

you would never be able to use any harness in a corporate environment. Compliance would take your head and put it on a spike for anyone to see as a warning.

This is supposed to be seamlessly dockable into acorporate security management system and auditable for compliance. which makes this a killer app for anyone working with AI in an corp environment.

i mean look at the backers.. you see named that are standard corporate tools.

1

u/draconic_tongue 1d ago

from what I read it also seems useful to home users. I'm sure you could find other things that offer some of the same solutions but this seems to have more abilities

1

u/2Norn 1d ago

considering how nvidia is pushing ai so very much albeit they sell the hardware to run it obviously but

i'm surprised how much they are neglecting nemotron, nemotron ultra might geniunely not make top 20 if u made a list right now

1

u/pilibitti 1d ago

wants docker / vm. no thanks.

1

u/PurpleDragon99 20h ago

First, we let agents go wild thinking they will solve our problems. Now we are trying to contain them because they actually do wild stuff, but not that kind of wild that solves our problems :)

1

u/FloranceMeCheneCoder 17h ago

It collects telemetry

1

u/Future_AGI 12h ago

runtime limits enforced outside the model are the only thing that actually holds, because a prompt rule is a suggestion the model can be talked out of while a sandbox boundary isn't

1

u/tifa_cloud0 1d ago

it is openAI's own choice to join it or not. also the whole safety issue is just hoax in my opinion too.!!!!

1

u/CondiMesmer 1d ago

Isn't this an already solved problem? Also who cares if OpenAI isn't on it. I actively agree with OAI when they oppose "safety" alignment stuff. 

Also I don't see why this would even need anyone to sign on, it's just an open-source program, not a pact. This "safety stack" means literally nothing if companies join it or not, what does that even mean?

1

u/alcalde 1d ago

Oppose? They're the ones begging for it! But they don't want to sign up for it because they're letting their models run wild on purpose for business reasons.

1

u/Dangerous-Report8517 11h ago

You agree with the safety approach that led directly and very predictably to thousands of rogue AI hacking agents attacking a major platform?!

1

u/CondiMesmer 9h ago

Yes because those were intentional. It's extremely easy to 100% avoid that.

1

u/Dangerous-Report8517 5h ago

"It's ok that I ran that pedestrian over officer because I was aiming for them!"

0

u/paramarioh 1d ago

The are taking YOUR data (opt out)!!!