Set the pods' keepalive sysctls so gen_rpc client sockets get a bound (T-614) - #448
Merged
Merged
Conversation
added 2 commits
October 2, 2026 22:05
… (T-614) gen_rpc 3.6.1 turns SO_KEEPALIVE on for its client sockets but lets an application set neither the keepalive timers nor TCP_USER_TIMEOUT, so a socket to a dead peer used Linux's defaults: first probe after 2 h. Every smolquery StatefulSet in deploy/ now sets the pod sysctls net.ipv4.tcp_keepalive_time 5, net.ipv4.tcp_keepalive_intvl 5 and net.ipv4.tcp_keepalive_probes 2 in its securityContext. They are Kubernetes safe sysctls from kubelet 1.29, and they set the defaults for every socket in the pod that has SO_KEEPALIVE. Measured in a network namespace with the peer blackholed, using the same timers per socket: an idle socket closed after 15.1 s (etimedout); a socket that sent data after the peer died was still open after 70 s, because keepalive does not run while data is unacknowledged. So this bounds a gen_rpc channel that was quiet when its peer died, which is the case that waited the full 15 minutes on 2026-10-02 (buffer-1 to buffer-0); a channel that sends first stays with T-613's membership drop and timeout probe. A kind test reads the three values from an api, a buffer and a storage pod. docs/deployment.md has the row and an upgrade note with the kubelet requirement.
…r test Fable review of PR 448. - The upgrade note's link to the socket-bounds section was split across two lines and did not render as a link. - The rollout note now names the other ways the sysctls are refused: Pod Security admission pinned to enforce-version v1.28 or older, a policy engine with its own sysctl allowlist, and hostNetwork pods. - It also names the other sockets that take the 5 s timers: Finch's pooled connections to object storage and the libpq catalog connections DuckDB's postgres and ducklake extensions open. Distribution and Postgrex set no keepalive. - The kind test reads only the last three lines of kubectl exec's output, so a kubectl warning on stderr cannot fail it.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
TL;DR: Every smolquery pod now sets the keepalive sysctls (5 s / 5 s / 2). A gen_rpc client socket that is idle when its peer dies closes in about 15 s, not 2 hours.
Tracker: T-614. Follows #447.
Why
SO_KEEPALIVEon for the sockets it dials.TCP_USER_TIMEOUTon them.buffer-1 -> buffer-0channel waited the full 15 minutes.What changed
deploy/base/api.yaml,buffer.yaml,storage.yamlanddeploy/overlays/kind-symmetric/server.yamlset these in the podsecurityContext:net.ipv4.tcp_keepalive_time= 5net.ipv4.tcp_keepalive_intvl= 5net.ipv4.tcp_keepalive_probes= 2SO_KEEPALIVE.docs/deployment.mdupdates the gen_rpc row and adds an upgrade note.What it covers, and what it does not
The test ran in a network namespace with the peer blackholed, using the same timers set per socket:
etimedout)tcp_retries2, about 15 min).:nodedown, and the timeout probe.HotClientsockets set their own timers per socket (Bound sockets to a peer node with keepalive and TCP_USER_TIMEOUT (T-614) #447), so these defaults do not change them.Tests
every role's pod runs with the keepalive sysctls (T-614). It passed in the Cluster CI job, so the pods start with the sysctls.mix precommit: 3,078 tests passmix ciReview
A Fable review found nothing serious. It confirmed the Kubernetes claims (safe from 1.29), the YAML placement, the overlays, and the Linux semantics. Fixed in
Review of T-614 sysctls: ...:kubectl execoutputWatch out
SysctlForbidden, and the pod does not start. Check the sandbox and production node versions before you roll this out. The kind CI cluster (kind v0.32.0) is new enough.postgresandducklakeextensions openenforce-versionpinned tov1.28or olderhostNetwork: truepod (it cannot setnet.*sysctls)sysctl.🤖 Generated with Claude Code