Linux containers in 500 lines of code (2016)

9 years ago (blog.lizzie.io)

Since others are sharing lightweight container implementations, I'll throw mine in. This one is written in Scheme: http://git.savannah.gnu.org/cgit/guix.git/tree/gnu/build/lin...

I like to think that there are sufficient code comments and docstrings to help demystify what is going on under the hood with containers.

The root capabilities section is worrying. Instead of trying to exhaustively list every capability that needs to be dropped, shouldn't the code instead just list the capabilities that are allowed, and clear all the others?

This would seem to better guard against accidentally missing existing capabilities, and also protect against newer capabilities that might be added in a future kernel.

  • I fully agree that whitelisting is safer than blacklisting.

    But on the top of the article, the author explicitly recognzied this and explained why they went down that route:

    > I wanted specifically to find a minimal set of restrictions to run untrusted code. This isn't how you should approach containers on anything with any exposure: you should restrict everything you can. But I think it's important to know which permissions are categorically unsafe!

    • Fair point. And the article does make a good read, with explanations about why each capability in particular should be disallowed.

  • Ditto for the seccomp syscall (and args) blacklist, although I wouldn't describe it as "worrying" considering the disclaimer at the top. Jess Frazelle's contained.af uses a seccomp whitelist[1], which doesn't need to be that long to allow reasonable programs to execute.

    This is a very good piece of writing with extensive references and I'll definitely find this useful to share in the future. She documents a ton of tradeoffs she made and resources she chose not to constrain (including the aforementioned syscalls and capabilities), which is important in this type of design and something that I wish I saw more of.

    [1]: https://github.com/jessfraz/contained.af/blob/master/seccomp...

  • That's true.

    I agree, you should iterate over all your existing capabilities and drop those that are not in the white-list. (I have implemented this functionality in one of my projects this way.)

    BUT:

    Maybe the reason some people do it otherwise, is that capabilities API have only a drop function for the bounding set, and people just don't think they should use it in reverse mode - Sapir-Whorf Hypothesis in operation! ;)

    edit: seems like the author have a different reason though

What a fine piece of literate programming!

I've just printed it out, and it literally contains 100 page of explanation and context for that 500 lines of code. Great work!

  • Just curious, why print it out?

    • Printed out works are less distracting, you can also write on them with a pencil or highlight things you like. Also if the page ever disappears / becomes inaccessible you have a backup of it. Not something I do, but I can see the uses. A coworker did it to analyze someone else's code. Also it doesn't drain your eyes as much. I sometimes wish I could have a Kindle that just showed me code (with some syntax highlighting) so I could look at code on an e-reader type of display but never drain my eyes doing so.

      2 replies →

    • Some people like to read the dead-tree version. Nostalgia, maybe paper is better for the eyes, paper books are generally lighter than the equivalent sized tablets or ebooks, etc.

      7 replies →

    • Because viewing things on a LCD is hard on the eyes after a while. E-ink is great, but paper is easier for most.

      This is one thing I hate about Apple these days. They once had pdf versions of all the developer documents but now they only have web pages and Xcode. For a company so concerned about users, they sure don’t seem to care about developer’s eye strain.

      1 reply →

    • I ask the same question from my wife! Some people only read on paper. We have two of each book because I read on Kindle and she reads it on the actual physical book.

      3 replies →

She mentions five Linux kernel mechanisms – "namespaces", "capabilities", "cgroups", and "setrlimit". Is any of those what I should use if I want to run an application inside some kind of container that lets me intercept file system calls (for example for the purpose of creating a file on the fly as it is accessed)?

  • Seccomp with ptrace is the way I'd do this. You can setup the rules to signal the ptracing process to intercept the syscall. I've not done it before but it should be possible. Id also look at doing it in a mount namespace with overlayfs on top of everything the process can see, so that you can manipulate anything you want or need filewise without destroying the original system. Then you can copy out any changed files later if you want to preserve them.

  • Not really. You probably need to use FUSE or similar to pull something like that off.

Is it only me or both firefox and chromium refuse to connect with something like ERR_SSL_PROTOCOL_ERROR?

off-topic, but that's a pretty cool e-mail address she's got :-) I wonder how many clients fail on it.

  • Huh. That is interesting. _ makes me think of a placeholder in Scala. It is a totally valid e-mail address .. makes me want to add it to my own domain and use it as a junk address.

that's a pretty cool e-mail address

  • Are you the same person who wrote the other comments about the e-mail address? (both heavily downvoted, one flagged and disappeared).

    What do you want to achieve by repeating that ever and ever again?

    • There is some weird stuff going on with new accounts. Look around the bottom of recent threads. The new accounts copy a part of a comment in the same thread and repost it. Seems like someone is getting their feet wet in botting.