Skip to content

Set the pods' keepalive sysctls so gen_rpc client sockets get a bound (T-614) - #448

Merged
chasers merged 2 commits into
mainfrom
t-614-keepalive-sysctls
Oct 2, 2026
Merged

chasers merged 2 commits into
mainfrom
t-614-keepalive-sysctls

Conversation

@chasers

@chasers chasers commented Oct 2, 2026 •

Copy link
Copy Markdown
Owner

TL;DR: Every smolquery pod now sets the keepalive sysctls (5 s / 5 s / 2). A gen_rpc client socket that is idle when its peer dies closes in about 15 s, not 2 hours.

Tracker: T-614. Follows #447.

Why

  • gen_rpc 3.6.1 turns SO_KEEPALIVE on for the sockets it dials.
  • It does not let an application set the keepalive timers or TCP_USER_TIMEOUT on them.
  • So those sockets used Linux's defaults: first probe after 2 hours.
  • On 2026-10-02 the quiet buffer-1 -> buffer-0 channel waited the full 15 minutes.

What changed

  • deploy/base/api.yaml, buffer.yaml, storage.yaml and deploy/overlays/kind-symmetric/server.yaml set these in the pod securityContext:
    • net.ipv4.tcp_keepalive_time = 5
    • net.ipv4.tcp_keepalive_intvl = 5
    • net.ipv4.tcp_keepalive_probes = 2
  • These are Kubernetes safe sysctls from kubelet 1.29.
  • They set the defaults for every socket in the pod that has SO_KEEPALIVE.
  • A new kind test reads the three values from an api, a buffer and a storage pod.
  • docs/deployment.md updates the gen_rpc row and adds an upgrade note.

What it covers, and what it does not

The test ran in a network namespace with the peer blackholed, using the same timers set per socket:

socket after the peer dies result
idle (nothing sent) ✅ closed after 15.1 s (etimedout)
sent data first ⚠️ still open after 70 s
  • Keepalive does not run while sent data is unacknowledged. That socket is on the retransmission path (tcp_retries2, about 15 min).
  • So this bounds a gen_rpc channel that was quiet when its peer died.
  • gen_rpc's own ping only goes out after 60 s of silence, so the keepalive probes come first.
  • A channel that sends to the dead peer first still depends on T-613: the drop on :nodedown, and the timeout probe.
  • HotClient sockets set their own timers per socket (Bound sockets to a peer node with keepalive and TCP_USER_TIMEOUT (T-614) #447), so these defaults do not change them.

Tests

  • ✅ New kind test: every role's pod runs with the keepalive sysctls (T-614). It passed in the Cluster CI job, so the pods start with the sysctls.
  • ✅ mix precommit: 3,078 tests pass
  • ✅ mix ci

Review

A Fable review found nothing serious. It confirmed the Kubernetes claims (safe from 1.29), the YAML placement, the overlays, and the Linux semantics. Fixed in Review of T-614 sysctls: ...:

  • a broken Markdown link in the upgrade note
  • the admission caveats and the affected sockets, now in the docs
  • the kind test reads only the last 3 lines of kubectl exec output

Watch out

  • ⚠️ kubelet 1.29 or newer is required. An older kubelet rejects the pod with SysctlForbidden, and the pod does not start. Check the sandbox and production node versions before you roll this out. The kind CI cluster (kind v0.32.0) is new enough.
  • ⚠️ The sysctls apply to every keepalive socket in the pod:
    • Finch pooled connections to object storage
    • the libpq catalog connections that DuckDB's postgres and ducklake extensions open
    • A peer that answers nothing at the TCP level for 15 s now closes those idle connections too. They reconnect on next use.
    • Erlang distribution and Postgrex set no keepalive, so they keep their behaviour.
  • ⚠️ Admission can also refuse the sysctls:
    • Pod Security admission with enforce-version pinned to v1.28 or older
    • a policy engine (Gatekeeper, Kyverno) with its own sysctl allowlist
    • a hostNetwork: true pod (it cannot set net.* sysctls)
  • ⚠️ Outside Kubernetes, set the same values with sysctl.

🤖 Generated with Claude Code

Chase Granberry added 2 commits October 2, 2026 22:05
… (T-614)

gen_rpc 3.6.1 turns SO_KEEPALIVE on for its client sockets but lets an
application set neither the keepalive timers nor TCP_USER_TIMEOUT, so a
socket to a dead peer used Linux's defaults: first probe after 2 h.

Every smolquery StatefulSet in deploy/ now sets the pod sysctls
net.ipv4.tcp_keepalive_time 5, net.ipv4.tcp_keepalive_intvl 5 and
net.ipv4.tcp_keepalive_probes 2 in its securityContext. They are
Kubernetes safe sysctls from kubelet 1.29, and they set the defaults for
every socket in the pod that has SO_KEEPALIVE.

Measured in a network namespace with the peer blackholed, using the same
timers per socket: an idle socket closed after 15.1 s (etimedout); a
socket that sent data after the peer died was still open after 70 s,
because keepalive does not run while data is unacknowledged. So this
bounds a gen_rpc channel that was quiet when its peer died, which is the
case that waited the full 15 minutes on 2026-10-02 (buffer-1 to
buffer-0); a channel that sends first stays with T-613's membership drop
and timeout probe.

A kind test reads the three values from an api, a buffer and a storage
pod. docs/deployment.md has the row and an upgrade note with the kubelet
requirement.
…r test

Fable review of PR 448.

- The upgrade note's link to the socket-bounds section was split across
  two lines and did not render as a link.
- The rollout note now names the other ways the sysctls are refused: Pod
  Security admission pinned to enforce-version v1.28 or older, a policy
  engine with its own sysctl allowlist, and hostNetwork pods.
- It also names the other sockets that take the 5 s timers: Finch's
  pooled connections to object storage and the libpq catalog connections
  DuckDB's postgres and ducklake extensions open. Distribution and
  Postgrex set no keepalive.
- The kind test reads only the last three lines of kubectl exec's output,
  so a kubectl warning on stderr cannot fail it.
@chasers
chasers merged commit e91ff50 into main Oct 2, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant