Bound sockets to a peer node with keepalive and TCP_USER_TIMEOUT (T-614) - #447
Merged
Merged
Conversation
added 2 commits
October 2, 2026 20:33
A pod killed without closing its sockets leaves every pooled HTTP connection to it pointed at its old IP. Linux keeps such a socket until it gives up retransmitting, about 15 minutes, and an idle one with no keepalive is never checked. After T-613 the sandbox still failed a seal about once a minute after a buffer pod was replaced: each stale pooled HotClient connection waited out the full 30 s receive timeout once. Smolquery.PeerSocket gives the TCP options for a socket to a peer: keepalive (idle 5 s, interval 5 s, 2 probes, gen_rpc's acceptor timers) and a 20 s TCP_USER_TIMEOUT, as Linux raw options. HotClient passes them as Req connect_options, so its pool's sockets carry them. The kernel then closes an idle dead socket in about 15 s and Finch drops it from the pool; a request already in flight on one fails after 20 s. Measured in an unprivileged network namespace with iptables: a host name remapped to a new IP while the old one is blackholed, four pooled connections, a request every 5 s. Before: every request failed after 30 s for the whole 75 s run. After: one in-flight request failed after 20.5 s, and every request succeeded from 28.5 s on. SMOLQUERY_PEER_TCP_USER_TIMEOUT_MS and SMOLQUERY_PEER_TCP_KEEPALIVE_IDLE_S tune it. docs/deployment.md lists every inter-node connection's bound; gen_rpc client sockets still have none at the TCP level, since gen_rpc 3.6.1 does not let an application set those options.
…live Fable review of PR 447. - On Linux TCP_USER_TIMEOUT overrides the keepalive probe count: an idle dead socket closes at the first unanswered probe at or after the user timeout. With 20 s and probes every 5 s the idle bound was 20 s, not the 15 s the docs said, and keepalive_count did nothing. The user timeout is now 15 s, the bound for both an in-flight request and an idle socket; keepalive_count is gone and the docs state the rule. Measured in a network namespace: all four stale pooled connections were gone by 16 s, and the first request after the blackhole succeeded. - A keepalive idle value Linux rejects (over 32767) was ignored silently, leaving the 2 h default. SMOLQUERY_PEER_TCP_KEEPALIVE_IDLE_S is now limited to 1..32767. - Finch and gen_rpc already turn SO_KEEPALIVE on, with Linux's 2 h timers. The docs say so, and note that the pod's tcp_keepalive_* sysctls would bound gen_rpc's client sockets without a gen_rpc fork. - The socket tests assert every raw option, not only the user timeout, and skip off Linux.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
TL;DR: Sockets from
HotClientto a buffer node now close 15 s after the peer stops answering, idle or in flight. A stale pooled connection to a replaced pod no longer fails one seal every minute.Tracker: T-614. Follows T-613 (#444, #445).
Why
{:manifest_unreachable, %Req.TransportError{reason: :timeout}}about once a minute for 4+ minutes after buffer-0 was killed.HotClient's pooled sockets use Linux's default keepalive timers. Finch turnsSO_KEEPALIVEon, but the first probe comes after 2 hours.tcp_retries2, about 15.4 min.What changed
Smolquery.PeerSocket.tcp_options/0. It returns:gen_tcpoptions for a socket to a peer node:TCP_USER_TIMEOUT15 skeepalive: trueis set.HotClientpasses them to Req asconnect_options: [transport_opts: ...]. Req keeps one pool for one option set.SMOLQUERY_PEER_TCP_USER_TIMEOUT_MSandSMOLQUERY_PEER_TCP_KEEPALIVE_IDLE_S(1 to 32767).docs/deployment.mdhas a new section. It lists every inter-node connection and what bounds it. It also has an upgrade note for T-613 and T-614.How it works
TCP_USER_TIMEOUTdecides when keepalive gives up, not the probe count. The socket closes at the first unanswered probe at or after 15 s.Evidence
The test ran in an unprivileged network namespace (
unshare -rn) withiptables:peer.testpointed at127.0.0.2, and 4 pooled connections were opened.127.0.0.2was dropped, andpeer.testwas remapped to127.0.0.3.Tests
PeerSocketTest: the option values on Linux, keepalive-only elsewhere, config overrides, and a real socket that reads back all three raw values.HotClientTest: after a manifest read, the pooled socket toHotServerreads back keepalive plus all three raw values. This test fails if the options are removed.iptables.Review
A Fable review found no bug in the mechanism. It verified that Mint keeps all
:rawoptions over HTTP and HTTPS, that Req reuses one pool, and that Finch drops a closed idle connection. These fixes are inReview of T-614: ...:TCP_USER_TIMEOUToverrides the keepalive probe count on Linux. So the idle bound was 20 s, not 15 s, andkeepalive_countdid nothing. Now there is one 15 s bound,keepalive_countis removed, and the docs state the rule.SO_KEEPALIVEwas already on (Finch, gen_rpc) with 2 h timers.Checks
mix precommit: 3,078 tests passmix cimix dialyzerWatch out
connect. The T-613 drop on:nodedownand the timeout probe cover gen_rpc. The pod sysctlsnet.ipv4.tcp_keepalive_*would also bound them, with no fork.httpfsreads hot segments over its own HTTP client. This PR does not bound it.🤖 Generated with Claude Code