test(guests): boot a guest end to end in a VM (task 0010)

Add a NixOS VM integration test to the flake's checks, so nix flake check
boots the network foundation and one guest on a virtual L2 segment and
asserts the three behaviors a VM can honestly reproduce: the guest presents
its own MAC distinct from the host's, gets its own IP on a tagged VLAN
across the segment, and writes to a bind mount owned by the shared storage
group. A tagged router node serving DHCP only on VLAN 10 makes the address
and reverse ping prove 802.1Q tagging end to end, not plain reachability.

Booting a networked guest surfaced a latent defect: the guest's own networkd
default-enables systemd-resolved, which conflicts with the nested-container
default of inheriting the host's resolv.conf, failing the guest toplevel
build. Fix it in the guest networking realization so a networked guest keeps
its own resolver.
This commit is contained in:
2026-07-25 23:02:58 -04:00
parent 969737b6b5
commit f1d6df7d51
4 changed files with 246 additions and 2 deletions

View File

@@ -0,0 +1,40 @@
---
spec: guests
blocked-by: [0004-guest-networking-placement, 0006-guest-storage-placement]
---
## What to build
One NixOS VM integration test, added to the flake's `checks` so `nix flake check` runs it, exercising the external observable behavior of a Guest and its foundations rather than the internal shape of the generated nested-container config.
A single harness boots the `modules.network` foundation and one sample Guest and asserts the three hard requirements together: the Guest presents its own MAC, gets its own IP on a tagged VLAN across a virtual L2 segment, and can write to a bind-mounted directory owned by the shared `storage` group.
The upstream NixOS test suite's nested-container networking cases (macvlan, extra-veth) are the model.
The two behaviors a VM cannot honestly reproduce — real 802.1Q against the physical switch and real ZFS identity-mapped writes on the pool — are out of this test and verified manually on the target Host.
## Acceptance criteria
- [x] A NixOS VM test is added to the flake's `checks` and runs as part of `nix flake check`.
- [x] The test boots `modules.network` and one sample Guest on a virtual L2 segment.
- [x] The test asserts the Guest presents its own MAC distinct from the Host's.
- [x] The test asserts the Guest gets its own IP on the correct tagged VLAN across the virtual segment.
- [x] The test asserts a guest service can write to a bind-mounted directory owned by the shared `storage` group.
## Implementation Notes
The test lives in `tests/guest-integration.nix` and is merged into `checks.x86_64-linux` as `guest-integration`, built through `pkgs.testers.runNixOSTest`.
The flake builds its nodes with the same `my`/`inputs` special arguments every configuration gets, passed through `node.specialArgs`, so the host node imports the real `modules/network.nix` and the real `my.guest` builder rather than a hand-rolled stand-in.
The harness is two nodes on one test-framework segment, the "virtual L2 segment".
The host runs `modules.network` with `trunk = "eth1"` and `vlans = [10]`, and one guest placed on VLAN 10.
A second `router` node speaks VLAN 10 only on a tagged `eth1.10` sub-interface and serves DHCP there, so the guest getting a `10.0.10.x` lease and the reverse `router → guest` ping succeed only when 802.1Q tagging works end to end across the segment.
This exercises the tagged path honestly rather than plain co-segment reachability.
The MAC assertion checks the guest's `eth0` equals its derived placement MAC and differs from the host trunk, and the storage assertion relies on the identity map (`privateUsers = "no"`) carrying gid 10000 through unshifted, so `stat -c %G` reading `storage` on the host is the "no permission errors" mechanism under test.
Booting a networked guest surfaced a latent defect in the guest networking foundation: enabling the guest's own networkd default-enables `systemd-resolved`, which conflicts with the nested-container default of inheriting the host's `resolv.conf`, and the guest's toplevel failed to build with "Using host resolv.conf is not supported with systemd-resolved".
Task 0004 never hit this because it only evaluated derived values, never built a networked guest's toplevel, and no committed host places a networked guest.
The fix is one line in `guestNet` (`networking.useHostResolvConf = false`), so any networked guest keeps its own resolver.
The guest interior here carries a storage-writing service as test scaffolding, since a real guest seals its own interior and the sample guest carries no service.
The host node also orders `container@sample` after the bridge's device unit, since the container enslaves its veth to the bridge at start and the upstream containers module orders only after `network.target`, not after the specific `hostBridge`.
Following the pattern of tasks 0003, 0004, and 0006, no committed host places a networked or pool-mounted guest, since the repo's only host is a wifi laptop with no bridge or pool.
The behaviors a VM cannot honestly reproduce — real 802.1Q against the physical switch and real ZFS identity-mapped writes on the pool — stay out of this test and are verified manually on the target Host, as the spec's testing decisions direct.