This looks pretty decent actually. Sure, you could consider it a frontend/SDK for bubblewrap/seatbelt/processcontainer; but setting em up consistently is far from trivial; and hand rolling is a really bad idea (speaking from experience).
I like the ‘learning’ mode for figuring out what perms/config a runtime needs, the MIT license, the clear optional telemetry disclosures, and somewhat light and still readable documentation.
Regardless of your views on Microsoft, this looks quite useful; serves a clear purpose, and from a quick glance, looks like a high quality project even if it’s just the first version.
show comments
simonw
This is promising, but there's one feature that's missing that I really care about: fine-grained networking.
Things macOS is missing include "Allow/deny by hostname" and "Allow/deny by IP, CIDR, port, or protocol".
The rest all looks great, and if you are on Linux or Windows those restrictions don't apply.
I guess this is the universal challenge of building an abstraction layer over multiple different technologies.
show comments
neobrain
Do any of these sandboxing solutions have a dynamic component to them that lets you grant permissions, starting with a minimal sandbox and asynchronously adding permissions as they become necessary? Harnesses try to do this when accessing non-project folders, but it's not always strictly enforced and generally not revocable. Harnesses also block agent execution until a decision is made, which requires constant monitoring to ensure progress can happen when the agent could easily proceed with an alternative method right away.
I like the idea of a minimal sandbox that protects against accidental `rm -rf` and against personal data leakage, but such a setup then often gets in the way of the specific task to be done. Ideally the sandbox would be able to aggregate blocked accesses and then expose them in an external TUI dashboard, where I can then enable access (without blocking any running agent on this, since that's prone to "press okay" fatigue).
Does anything close to this exist yet?
show comments
amluto
Sadly, skimming the docs about how the different backends work makes me think that almost this entire project was done by a recent-gen LLM that interpreted its instructions as “make these things work at all costs” instead of “make a considered design that cleanly and securely fits its use case”.
Even the bubblewrap integration docs are basically a stream of consciousness vibe splat. I have approximately zero confidence in the results.
kernc
350,000 of mostly Rust SLOC [1] ... And the upstream sandboxes aren't even vendored!
I'd be way more confident building upon something I can grasp and understand. [2]
A little off topic, maybe, but I've been having great luck with wasmtime and wasm32-wasip3 for writing sandboxed plugins. The tooling is pretty nice when you write plugins in rust, but I don't know what it looks like for other languages right now.
wasip3 is not stable yet, but it has a lot of nice changes (compared to wasip2) for integrating with async code
chneu
Right you are, Ken!
show comments
minraws
Why is everyone making their own code execution agent runtime engines I have an entire project built on top of openshell already, why not first come up with a sandboxing policy design, like unix did, and then build on top of that.
Currently all project do tend to agree on what and how they work but certain things being different makes porting tedius, if all of them have a bare minimum subset common amongst them it would be much easier to switch, and validate security surface area.
I feel like there are more vulnerabilities in this vibe coded slop sandboxes, and it's more likely everyone one of us trusting them to build projects around them will shoot our foot off once a cve is hit in one that's common in all of them but since they are all slop copies someone will have to figure out how they apply to all others and then manually fix it properly, and if one of them makes a CVE public it will leave dozens of these runtimes open to exploits.
I wish the best to my future self with regards to security I feel like we are completely screwed. Since we can no longer depend on upstream for security.
show comments
zmmmmm
I do hate adding abstraction layers needlessly, but currently this does look like it might solve a real problem for me. Or at least the concept of it.
I want to support sandboxing for my app, and commands it launches, but there is nothing that actually works across all the environments I want code to run in. So yes, if I had one tool that could be an abstract interface and let the user set up and configure their sandboxing completely separate to my app, and do that at run time based on declarative policy - it would be handy.
mintflow
Seems aws also announced a sandbox solution
I used agent over 1 year and basically always give codex full permission on each thread, do not get issue so far
Why we need this layer of complexity? Or its mainly for big company that need control ?
show comments
sharts
When do you use this? You’re writing the app so you want the app to sandbox itself instead of… just running in a container?
show comments
wild_pointer
What's also interesting is that they added the Experimental_CreateProcessInSandbox API to Windows, like, last month.
Funny, now that it's documented, the API name will be stuck with this name forever.
show comments
shados
Is this playing in the same space as the like of gvisor?
arj
Would this allow a sandboxed container on windows to still run commands in wsl?
plq
Both Firefox and Chromium have battle-tested cross-platform sandbox implementations, but I imagine it'd be laborious to integrate them as a third-party dependency to other projects. Why not spend resources on repackaging them in an easy-to-use SDK instead of reinventing the wheel?
Microsoft is already a Chromium contributor so it's not like they lack in-house expertise or something.
They have a sandbox escape in there. Likewise capability ordering is wrong. Exactly what you should expect from Microsoft.
Since this apparently wraps bubblewrap (another incompetent action on behalf of Microslop), did a quick sweep of that codebase too. Setuid is wrong, capability dropping is wrong, bubblewrap does _not_ protect against compromised/vulnerable kernels (and it should fyi), wrote up a full bubblewrap/mxc sandbox escape too.
show comments
smitty1e
Asked Grok the difference between mxc and flatpak:
"So MXC is a cross-platform “what may this workload touch?” layer aimed at agents. Flatpak is a Linux app format whose sandbox happens to share a backend with MXC on Linux."
zenapollo
Saw the M stands for Microsoft and immediately closed the tab. 1 it’s unnecessary - communicates nothing but look-at-me branding. 2 toxic company.
This looks pretty decent actually. Sure, you could consider it a frontend/SDK for bubblewrap/seatbelt/processcontainer; but setting em up consistently is far from trivial; and hand rolling is a really bad idea (speaking from experience).
I like the ‘learning’ mode for figuring out what perms/config a runtime needs, the MIT license, the clear optional telemetry disclosures, and somewhat light and still readable documentation.
Regardless of your views on Microsoft, this looks quite useful; serves a clear purpose, and from a quick glance, looks like a high quality project even if it’s just the first version.
This is promising, but there's one feature that's missing that I really care about: fine-grained networking.
They have this for Windows and Linux, but it's sadly missing for macOS - see the support table here: https://github.com/microsoft/mxc/blob/main/docs/backends/sea...
Things macOS is missing include "Allow/deny by hostname" and "Allow/deny by IP, CIDR, port, or protocol".
The rest all looks great, and if you are on Linux or Windows those restrictions don't apply.
I guess this is the universal challenge of building an abstraction layer over multiple different technologies.
Do any of these sandboxing solutions have a dynamic component to them that lets you grant permissions, starting with a minimal sandbox and asynchronously adding permissions as they become necessary? Harnesses try to do this when accessing non-project folders, but it's not always strictly enforced and generally not revocable. Harnesses also block agent execution until a decision is made, which requires constant monitoring to ensure progress can happen when the agent could easily proceed with an alternative method right away.
I like the idea of a minimal sandbox that protects against accidental `rm -rf` and against personal data leakage, but such a setup then often gets in the way of the specific task to be done. Ideally the sandbox would be able to aggregate blocked accesses and then expose them in an external TUI dashboard, where I can then enable access (without blocking any running agent on this, since that's prone to "press okay" fatigue).
Does anything close to this exist yet?
Sadly, skimming the docs about how the different backends work makes me think that almost this entire project was done by a recent-gen LLM that interpreted its instructions as “make these things work at all costs” instead of “make a considered design that cleanly and securely fits its use case”.
Even the bubblewrap integration docs are basically a stream of consciousness vibe splat. I have approximately zero confidence in the results.
350,000 of mostly Rust SLOC [1] ... And the upstream sandboxes aren't even vendored!
I'd be way more confident building upon something I can grasp and understand. [2]
[1]: https://ghloc.dev/microsoft/mxc [2]: https://github.com/sandbox-utils/sandbox-run
This really should be a WebAssembly runtime so that the same computations can run on different platforms and devices.
Been looking at sandboxing, both low level and higher level like this.
The API for their Rust mxc-sdk looks nice but
- their "sdk" has binaries and the build script has logic for them
- their build scripts do windows-exclusive work on all platforms
- not putting some of the backends behind features causes more build script work (and that work will break on future Cargo versions)
- at least some of the remaíning build script work doesn't need to be a build script
- it seems pretty dependency heavy
Microsoft stole my idea :) (joking obviously, everyone and their mom is making sandboxes) https://github.com/pprotas/slopbox
A little off topic, maybe, but I've been having great luck with wasmtime and wasm32-wasip3 for writing sandboxed plugins. The tooling is pretty nice when you write plugins in rust, but I don't know what it looks like for other languages right now.
wasip3 is not stable yet, but it has a lot of nice changes (compared to wasip2) for integrating with async code
Right you are, Ken!
Why is everyone making their own code execution agent runtime engines I have an entire project built on top of openshell already, why not first come up with a sandboxing policy design, like unix did, and then build on top of that.
Currently all project do tend to agree on what and how they work but certain things being different makes porting tedius, if all of them have a bare minimum subset common amongst them it would be much easier to switch, and validate security surface area.
I feel like there are more vulnerabilities in this vibe coded slop sandboxes, and it's more likely everyone one of us trusting them to build projects around them will shoot our foot off once a cve is hit in one that's common in all of them but since they are all slop copies someone will have to figure out how they apply to all others and then manually fix it properly, and if one of them makes a CVE public it will leave dozens of these runtimes open to exploits.
I wish the best to my future self with regards to security I feel like we are completely screwed. Since we can no longer depend on upstream for security.
I do hate adding abstraction layers needlessly, but currently this does look like it might solve a real problem for me. Or at least the concept of it.
I want to support sandboxing for my app, and commands it launches, but there is nothing that actually works across all the environments I want code to run in. So yes, if I had one tool that could be an abstract interface and let the user set up and configure their sandboxing completely separate to my app, and do that at run time based on declarative policy - it would be handy.
Seems aws also announced a sandbox solution
I used agent over 1 year and basically always give codex full permission on each thread, do not get issue so far
Why we need this layer of complexity? Or its mainly for big company that need control ?
When do you use this? You’re writing the app so you want the app to sandbox itself instead of… just running in a container?
What's also interesting is that they added the Experimental_CreateProcessInSandbox API to Windows, like, last month.
https://learn.microsoft.com/en-us/windows/win32/secauthz/cre...
Funny, now that it's documented, the API name will be stuck with this name forever.
Is this playing in the same space as the like of gvisor?
Would this allow a sandboxed container on windows to still run commands in wsl?
Both Firefox and Chromium have battle-tested cross-platform sandbox implementations, but I imagine it'd be laborious to integrate them as a third-party dependency to other projects. Why not spend resources on repackaging them in an easy-to-use SDK instead of reinventing the wheel?
Microsoft is already a Chromium contributor so it's not like they lack in-house expertise or something.
https://wiki.mozilla.org/Security/Sandbox
https://chromium.googlesource.com/chromium/src/+/HEAD/docs/d...
Codex uses this on Windows now.
Right you are, ken!
But seriously, we need docker for models like years ago. I dont want these things running with the ability to run rm -rf /
its should be treated no different than wget | sh
How does this compare to Sandboxie?
https://github.com/sandboxie-plus/Sandboxie
They have a sandbox escape in there. Likewise capability ordering is wrong. Exactly what you should expect from Microsoft.
Since this apparently wraps bubblewrap (another incompetent action on behalf of Microslop), did a quick sweep of that codebase too. Setuid is wrong, capability dropping is wrong, bubblewrap does _not_ protect against compromised/vulnerable kernels (and it should fyi), wrote up a full bubblewrap/mxc sandbox escape too.
Asked Grok the difference between mxc and flatpak:
"So MXC is a cross-platform “what may this workload touch?” layer aimed at agents. Flatpak is a Linux app format whose sandbox happens to share a backend with MXC on Linux."
Saw the M stands for Microsoft and immediately closed the tab. 1 it’s unnecessary - communicates nothing but look-at-me branding. 2 toxic company.