Linux containers in 500 lines of code (2016)

141 points42 commentsa day ago
smashed

Coincidentally I mis-prompted claude code the other day while working on a toy project and failed to specify the project should be built on top of docker and not "like docker".

It went on to waste all my tokens creating a specialized docker clone. Cool I guess.

js2

(2016). Previous submissions w/comments:

https://news.ycombinator.com/item?id=30623372 (250 points | March 10, 2022 | 27 comments)

https://news.ycombinator.com/item?id=22232705 (267 points | Feb 4, 2020 | 29 comments)

https://news.ycombinator.com/item?id=15608435 (440 points | Nov 2, 2017 | 53 comments)

setheron

I have written https://fzakaria.com/2020/05/31/containers-from-first-princi... a while ago in similar vein.

abidinberkay

This was written in 2016. What would be different if you wrote it today? For example would cgroups v2 or newer seccomp features change that much?

show comments
zoobab

I discovered proot-distro build yesterday, which does not require a special kernel with network and pid namespaces, it could run on older machines that do not have those features, or as a unix user that don't have those permissions.

https://pypi.org/project/proot-distro/

TZubiri

The takeaway from this is that (the most popular impmementation of) containers have been birthed from linux primitives, and to the extent they weren't primitives were developed to fulfill those needs.

So, like many linux userspace applications, containers in linux are just a thin wrapper over kernel functionality.

ranger_danger

> I wanted specifically to find a minimal set of restrictions to run untrusted code.

I don't think we should consider containers to be a security boundary. Even full VMs can be escaped, and have been, many times.

The fact that this is possible in the first place makes me think we need a much better approach.

show comments
kragen

Last night I was looking for how to run Graphviz on untrusted input in a secure way, because recent versions of Graphviz give untrusted input to a whole insane rat's nest of code: Harfbuzz, Pango, libfribidi, libthai, libgraphite2, and its own internal format parser, each of which has a rap sheet of CVEs that makes Charlie Manson look like a petty shoplifter. And apparently Pango is even multithreaded, so we can expect nondeterminism. (Most of this doesn't show up in a simple ldd check; Graphviz sneakily waits until runtime to dlopen graphviz/libgvplugin_pango.so.6!) So, naturally, I wanted to sandbox it so that the worst thing a malicious attacker could do would be to make it draw Dickbutt or something. What I ended up with was less than 500 lines of code using Claude's suggestion of Bubblewrap http://canonical.org/~kragen/sw/dev3/wrapdot:

    #!/bin/sh
    # Confine dot in bubblewrap, taking input from stdin and writing PNG
    # output to stdout.

    # 64 megs seems to be enough, 21 megs isn’t.
    address_space=64001000

    # With zero --fsize, we can’t write the output file on stdout if it's
    # redirected to a file, but you can pipe it to `cat`.
    file_size=0

    cpu_seconds=5

    # Apparently Pango or fontconfig is multithreaded now‽
    # (process:2): GLib-ERROR **: 00:37:18.600: creating thread '[pango] FcInit': Error creating thread: Resource temporarily unavailable
    processes=4

    # We’re using --unshare-user, etc., explicitly, because --unshare-all
    # uses the wimpy --unshare-user-try and --unshare-cgroup-try options.
    # --remount-ro / prevents malicious code from filling the root
    # filesystem with empty files.

    exec bwrap \
          --ro-bind /bin /bin \
          --ro-bind /lib /lib \
          --ro-bind /lib64 /lib64 \
          --ro-bind /sbin /sbin \
          --ro-bind /usr/lib /usr/lib \
          --ro-bind /usr/share/fonts /usr/share/fonts \
          --ro-bind /var/cache/fontconfig /var/cache/fontconfig \
          --ro-bind /etc/fonts /etc/fonts \
          --remount-ro / \
          --unshare-user --unshare-ipc --unshare-pid --unshare-net --unshare-uts \
          --unshare-cgroup --die-with-parent --new-session --cap-drop ALL \
          --clearenv --setenv PATH /bin \
          prlimit --as="$address_space" --fsize="$file_size" \
                  --cpu="$cpu_seconds" --nproc="$processes" \
                      dot -Tpng -Gdpi=192

    # For testing, to verify that network access is indeed blocked:
    #                 nc.traditional -v -v 127.0.0.1 8000
Still, this is enough code that I'm not sure I haven't left something out. Still pending: run ImageMagick or netpbm inside the sandbox to convert the PNG file into a PPM or BMP — there have been CVEs in libpng in the past, and of course it's potentially vulnerable to zip bombs.

Of course this is still exposing most of the Linux kernel system call interface, although fortunately not /proc and /dev. And it's still fucking insane that drawing a node-link graph with three nodes requires more virtual memory than my first Linux machine had in total, in which it ran web browsers and recompiled the kernel. But that's a little further down the line.

show comments