For the last couple of weeks, I’ve been working on garbage collection
for composefs repositories. A composefs repository is a
content-addressed object store, and it needs a garbage collector to
reclaim space from images that are no longer referenced. The hard part
is deciding what “referenced” actually means when an image has no ref
in the repository but is mounted somewhere, possibly from another
process, possibly in a different mount namespace, possibly by the
running system itself. On
composefs-rs#346
was suggested to use flock() on the EROFS backing file. Mounters
take a shared lock, the GC probes with an exclusive one, and if the
exclusive lock fails, the image is still in use.
Posts for: #Erofs
Image sealing with composefs
Composefs achieves whole-filesystem integrity verification through image sealing: a single cryptographic digest authenticates an entire filesystem, covering both file contents and metadata (directory structure, permissions, ownership, symlinks, and xattrs). This goes further than fs-verity alone, which can only verify individual file contents, and avoids the fixed-partition requirement of dm-verity. The mechanism combines an EROFS image for metadata, a content-addressed object store for file data, and overlayfs with verity=require to enforce integrity checks on every file access at the kernel level.
Sharing EROFS superblocks across container mounts
For a while now we’ve had composefs support in the container-libs. For each OCI layer in an image, we build a read-only EROFS blob that can be verified with fs-verity and that the kernel mounts directly. It works well if you look at a single container, but the moment you start thinking about composefs as the storage for a real container host where many containers are running, we have a problem. Many of them share the same image layers but with composefs every one of those containers mounts the same EROFS blobs, and each mount is a separate world as far as the kernel is concerned. The bytes on disk are shared, but the memory used to cache them is not. That is one of the reasons composefs is not yet the default way to store and run containers, and it is what I want to talk about here.