fcvm podman prepare checks that the snapshot filesystem has room for the full snapshot only at the moment it takes the snapshot. For a workload with a long warm-up that is the last step of a long run, so a host that is a few GiB short loses the whole run.
What happened
A prepare of a 128 GiB guest whose readiness takes about 104 minutes (image import, application warm-up, then the health check turns ready):
07:56:44 Starting fcvm podman lifecycle lifecycle=Prepare(PrepareOptions { tag: Some("warm-tested"), force: true, ready_timeout: 10800s })
...
09:40:52 Creating prepared startup snapshot (VM healthy) snapshot_key=warm-tested
09:40:52 cleaning up resources
09:40:52 SIGKILL to Firecracker process
09:41:29 ERROR fcvm: Error: creating prepared startup snapshot warm-tested: Not enough disk space for full snapshot: need 131072 MiB, have 80791 MiB free on /mnt/fcvm-btrfs/snapshots/warm-tested. Use --mem to limit VM memory, or increase btrfs_size in rootfs-config.toml.
The filesystem had 133 GiB free when the prepare started. The VM's own disk growth during the boot, with the previous snapshot of the same tag still on disk (--force replaces it only after the new one is written), took it to 78.9 GiB by the time the snapshot was due. The VM is disposable in this flow, so it was killed and 6,283 s of work were gone.
Proposal
- Run the same free-space check before the VM boots. The memory size is known from the command line, so the requirement is known up front:
need <mem> MiB for the snapshot, have <free> MiB. A prepare that cannot fit its snapshot at the start should refuse to start.
- Keep the check at snapshot time as it is, because the disk can fill during the boot. When it fails there, say how much the free space fell since the start, which points at the VM's own disk growth or at another writer.
- With
--force and an existing snapshot of the same tag, say in the refusal that the old snapshot still occupies space until the new one is installed, and how much.
A unit test can pin the first point without a VM: a prepare against a filesystem with less free space than --mem fails before any VM process is spawned.
fcvm podman preparechecks that the snapshot filesystem has room for the full snapshot only at the moment it takes the snapshot. For a workload with a long warm-up that is the last step of a long run, so a host that is a few GiB short loses the whole run.What happened
A prepare of a 128 GiB guest whose readiness takes about 104 minutes (image import, application warm-up, then the health check turns ready):
The filesystem had 133 GiB free when the prepare started. The VM's own disk growth during the boot, with the previous snapshot of the same tag still on disk (
--forcereplaces it only after the new one is written), took it to 78.9 GiB by the time the snapshot was due. The VM is disposable in this flow, so it was killed and 6,283 s of work were gone.Proposal
need <mem> MiB for the snapshot, have <free> MiB. A prepare that cannot fit its snapshot at the start should refuse to start.--forceand an existing snapshot of the same tag, say in the refusal that the old snapshot still occupies space until the new one is installed, and how much.A unit test can pin the first point without a VM: a prepare against a filesystem with less free space than
--memfails before any VM process is spawned.