feat(guests): let a guest nest OCI containers (task 0009)

Add a `nesting` placement field to the Host-side guest interface, a bool
off by default. On, it grants the guest's container the prerequisites its
interior needs to run Podman and other OCI containers: the `CAP_NET_ADMIN`
capability an OCI runtime uses to build its bridges and firewall rules,
and the `/dev/net/tun` and `/dev/fuse` device nodes it reaches for to
network those containers and back their overlay storage. Off, both the
capability and device lists are empty, so a non-nesting guest is untouched.

cgroup delegation, the other nested prerequisite, the NixOS container
backend already grants every container unconditionally, so the Skeleton
records it with an absence pointer rather than re-emitting it.

Add a nesting-sample guest whose interior defines an `oci-containers`
workload on Podman, and enable it on neogaia with `nesting` on, so the
path builds end to end through the Host's `nix flake check` — which pulls
in podman and the generated container unit for the nested system.
This commit was merged in pull request #35.
This commit is contained in:
2026-07-25 22:10:46 -04:00
parent 0b7d409fbc
commit 969737b6b5
4 changed files with 90 additions and 0 deletions

32
lib.nix
View File

@@ -279,6 +279,19 @@ let
'';
};
};
nesting = lib.mkOption {
type = lib.types.bool;
default = false;
description = ''
Grant the guest's interior the prerequisites to run Podman or other
OCI containers of its own. Off by default, so a guest cannot nest
containers. On, the guest's container gains the network-administration
capability its container runtime uses to build bridges and firewall
rules, along with the tun and fuse device nodes such a runtime reaches
for, so the interior's `virtualisation.oci-containers` works with
Podman as its default runtime.
'';
};
autoStart = lib.mkOption {
type = lib.types.bool;
default = true;
@@ -335,6 +348,25 @@ let
# A private-user mapping would shift the ids and reintroduce those errors, so it stays off.
privateUsers = lib.mkDefault "no";
# A nesting guest runs Podman or other OCI containers in its interior.
# The network-administration capability lets that runtime build its
# bridges and firewall rules.
# The tun and fuse device nodes are what it reaches for to network
# those containers and back their overlay storage.
# The remaining prerequisite, a delegated cgroup subtree for the
# runtime to manage, the container backend already grants every guest.
additionalCapabilities = lib.optionals cfg.nesting [ "CAP_NET_ADMIN" ];
allowedDevices = lib.optionals cfg.nesting [
{
node = "/dev/net/tun";
modifier = "rwm";
}
{
node = "/dev/fuse";
modifier = "rwm";
}
];
bindMounts = userMounts // secretMounts;
inherit specialArgs;