hn.today

Linux containers in 500 lines of code

blog.lizzie.io111 points22 comments
Screenshot of Linux containers in 500 lines of code

Lizzie implements a compact container runtime in C (roughly 500 lines) that demonstrates how Linux namespaces, capabilities, cgroups, setrlimit and seccomp can be combined to run untrusted code with constrained privileges. The project is presented as literate source (GPLv3) with runnable examples and a usage demo showing pivot_root, mount namespacing, UID/GID mapping, capability dropping, syscall filtering, and cleanup of cgroups. A recurring theme is finding a minimal set of permissions that are categorically unsafe while documenting tradeoffs: user namespaces can simplify isolation but introduce complex, historically buggy semantics, so they’re used only when available and handled carefully.

Concrete implementation details matter: the runtime creates a socketpair for parent-child coordination, allocates a temporary stack for clone(2) and starts a child with CLONE_NEWNS|CLONE_NEWCGROUP|CLONE_NEWPID|CLONE_NEWIPC|CLONE_NEWNET|CLONE_NEWUTS, then the parent writes /proc/<pid>/{uid_map,gid_map} when appropriate. The code validates kernel version/architecture, prepares cgroups and resource limits, generates container hostnames, and explicitly drops capabilities and blacklists syscalls. The writeup includes code snippets, rationale for design choices, and pointers to source and related hardening material for readers who want to inspect or extend the runtime.

Read on blog.lizzie.io22 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in Security

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.