Skip to content

Restore distro-kernel parity for long-lived VMs: swap+zram, MGLRU, vmstat counters #145

Description

@jmagder

libkrunfw kernels have CONFIG_SWAP=n and CONFIG_ZRAM=n, with CONFIG_MODULES=n, so a guest whose live working set exceeds its krun.ram_mib ceiling can only OOM-kill — even when the host has plenty of memory. There is no other swap path: no block devices are exposed through podman+krun, and virtiofs cannot host swapfiles. zram is the only possible in-guest backend and the config forbids it.

Context: config-libkrunfw_x86_64 had CONFIG_SWAP=y until it was dropped in b085fa0 ("Drop swap and unused memory features", 2025-01). The sev/tdx and aarch64/riscv64 variants still carry SWAP=y (plus ZSWAP/ZSMALLOC on the non-x86 ones), so this is restore-to-parity rather than new surface. The same slimming pass also dropped CONFIG_VM_EVENT_COUNTERS.

Suggested lines for config-libkrunfw_x86_64 (the zstd selectors materialise ZSMALLOC=y and CONFIG_ZRAM_DEF_COMP="zstd" via make olddefconfig; ZSTD_COMPRESS/ZSTD_DECOMPRESS are already =y):

CONFIG_SWAP=y
CONFIG_ZRAM=y
CONFIG_ZRAM_BACKEND_ZSTD=y
CONFIG_ZRAM_DEF_COMP_ZSTD=y
CONFIG_LRU_GEN=y
CONFIG_VM_EVENT_COUNTERS=y
CONFIG_PSI=y

All of these are defaults on mainstream distro kernels, and all are inert until guest userspace opts in: built-in zram creates exactly one zram0 with disksize 0 and no swap exists until mkswap/swapon; MGLRU is strictly-better reclaim; VM_EVENT_COUNTERS/PSI are telemetry-only.

Size cost: in our rebuild the resulting libkrunfw.so came out the same size as the stock artifact (21,432,000 B) — the added code fits entirely within the kernel image's existing inter-segment alignment padding, so the shipped library doesn't grow. Content-wise the kernel image is ~1% larger (~68 KiB xz-compressed), i.e. a few hundred KiB of resident kernel text; and zram allocates nothing until guest userspace initialises the device (disksize starts at 0).

Motivating cases, both field-observed on 4 GiB podman+krun VMs:

  1. Three node processes resume heavy work simultaneously at boot — guest OOM at t=16s with ~3.8 GiB nearly-all-inactive anon, zero file cache, and the host idle.
  2. With swap enabled, the same resume OOM'd via the classic LRU declaring all_unreclaimable with 3.4 GiB of swap still free — CONFIG_LRU_GEN (MGLRU) fixed it.

End result with all of the above: 11 node sessions + dev servers on a 4 GiB ceiling, 2 GiB compressed at ~4:1 in zram, zero OOMs across resume bursts and repeated restarts.

Happy to send a PR against the configs if the maintainers are open to it.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions