Skip to content

Bound sockets to a peer node with keepalive and TCP_USER_TIMEOUT (T-614) - #447

Merged
chasers merged 2 commits into
mainfrom
t-614-bound-peer-sockets
Oct 2, 2026
Merged

chasers merged 2 commits into
mainfrom
t-614-bound-peer-sockets

Conversation

@chasers

@chasers chasers commented Oct 2, 2026 •

Copy link
Copy Markdown
Owner

TL;DR: Sockets from HotClient to a buffer node now close 15 s after the peer stops answering, idle or in flight. A stale pooled connection to a replaced pod no longer fails one seal every minute.

Tracker: T-614. Follows T-613 (#444, #445).

Why

  • After T-613, the sandbox rerun (2026-10-02 20:08 UTC) showed gen_rpc recover in seconds.
  • The buffer manifest HTTP path did not. storage-2 logged {:manifest_unreachable, %Req.TransportError{reason: :timeout}} about once a minute for 4+ minutes after buffer-0 was killed.
  • Cause: HotClient's pooled sockets use Linux's default keepalive timers. Finch turns SO_KEEPALIVE on, but the first probe comes after 2 hours.
  • A pooled connection to the old pod IP stays in the pool.
  • Each request that draws a stale connection waits the full 30 s receive timeout. Then the pool drops that one connection.
  • Linux would close the socket itself only after tcp_retries2, about 15.4 min.

What changed

  • New Smolquery.PeerSocket.tcp_options/0. It returns :gen_tcp options for a socket to a peer node:
    • TCP_USER_TIMEOUT 15 s
    • keepalive probes after 5 s idle, then every 5 s
  • These are Linux raw options. On another OS only keepalive: true is set.
  • HotClient passes them to Req as connect_options: [transport_opts: ...]. Req keeps one pool for one option set.
  • New env vars: SMOLQUERY_PEER_TCP_USER_TIMEOUT_MS and SMOLQUERY_PEER_TCP_KEEPALIVE_IDLE_S (1 to 32767).
  • docs/deployment.md has a new section. It lists every inter-node connection and what bounds it. It also has an upgrade note for T-613 and T-614.

How it works

  1. A buffer pod dies hard. Its sockets stay open on the storage and query nodes.
  2. Idle socket: after 5 s idle, the kernel sends a keepalive probe every 5 s.
  3. On Linux, TCP_USER_TIMEOUT decides when keepalive gives up, not the probe count. The socket closes at the first unanswered probe at or after 15 s.
  4. Socket in flight: a request on a dead socket fails when its data is unacknowledged for 15 s, not after the 30 s receive timeout.
  5. Finch gets the socket error for an idle connection and removes it from the pool.
  6. The next request opens a new connection. DNS gives the new pod's IP.

Evidence

The test ran in an unprivileged network namespace (unshare -rn) with iptables:

  • peer.test pointed at 127.0.0.2, and 4 pooled connections were opened.
  • Then all traffic to 127.0.0.2 was dropped, and peer.test was remapped to 127.0.0.3.
  • Requests ran every 5 s.
first request at result
before (Req defaults) 3 s every request failed after 30 s, for all 75 s
first version (20 s user timeout) 3 s 1 in-flight failure at 20.5 s, all ok from 28.5 s
this version (15 s) 16 s all 4 stale connections already gone; every request ok

Tests

  • ✅ PeerSocketTest: the option values on Linux, keepalive-only elsewhere, config overrides, and a real socket that reads back all three raw values.
  • ✅ HotClientTest: after a manifest read, the pooled socket to HotServer reads back keepalive plus all three raw values. This test fails if the options are removed.
  • Both socket tests skip off Linux.
  • ⚠️ The blackhole run is not in CI. It needs a network namespace and iptables.

Review

A Fable review found no bug in the mechanism. It verified that Mint keeps all :raw options over HTTP and HTTPS, that Req reuses one pool, and that Finch drops a closed idle connection. These fixes are in Review of T-614: ...:

  • TCP_USER_TIMEOUT overrides the keepalive probe count on Linux. So the idle bound was 20 s, not 15 s, and keepalive_count did nothing. Now there is one 15 s bound, keepalive_count is removed, and the docs state the rule.
  • A keepalive idle value over 32767 was ignored silently (the 2 h default stayed). The env var now must be 1 to 32767.
  • The docs now say SO_KEEPALIVE was already on (Finch, gen_rpc) with 2 h timers.
  • The tests assert every raw option.

Checks

  • ✅ mix precommit: 3,078 tests pass
  • ✅ mix ci
  • ✅ mix dialyzer
  • ✅ Cluster tests (kind): one failure on the first run was a known flake. It answered 400 during a buffer kill, which also happened on 2026-09-12. It passed on rerun. Tracked as T-623.

Watch out

  • ⚠️ gen_rpc client sockets still have no TCP-level bound from smolquery. gen_rpc 3.6.1 only passes buffer sizes to connect. The T-613 drop on :nodedown and the timeout probe cover gen_rpc. The pod sysctls net.ipv4.tcp_keepalive_* would also bound them, with no fork.
  • ⚠️ DuckDB httpfs reads hot segments over its own HTTP client. This PR does not bound it.
  • ⚠️ The bound acts only on a peer that answers nothing at the TCP level. A slow peer still ACKs, so a slow response is never cut short.
  • ⚠️ T-614's last acceptance item is open: repeat the kill on the sandbox after deploy.

🤖 Generated with Claude Code

Chase Granberry added 2 commits October 2, 2026 20:33
A pod killed without closing its sockets leaves every pooled HTTP
connection to it pointed at its old IP. Linux keeps such a socket until it
gives up retransmitting, about 15 minutes, and an idle one with no
keepalive is never checked. After T-613 the sandbox still failed a seal
about once a minute after a buffer pod was replaced: each stale pooled
HotClient connection waited out the full 30 s receive timeout once.

Smolquery.PeerSocket gives the TCP options for a socket to a peer:
keepalive (idle 5 s, interval 5 s, 2 probes, gen_rpc's acceptor timers)
and a 20 s TCP_USER_TIMEOUT, as Linux raw options. HotClient passes them
as Req connect_options, so its pool's sockets carry them. The kernel then
closes an idle dead socket in about 15 s and Finch drops it from the pool;
a request already in flight on one fails after 20 s.

Measured in an unprivileged network namespace with iptables: a host name
remapped to a new IP while the old one is blackholed, four pooled
connections, a request every 5 s. Before: every request failed after
30 s for the whole 75 s run. After: one in-flight request failed after
20.5 s, and every request succeeded from 28.5 s on.

SMOLQUERY_PEER_TCP_USER_TIMEOUT_MS and SMOLQUERY_PEER_TCP_KEEPALIVE_IDLE_S
tune it. docs/deployment.md lists every inter-node connection's bound;
gen_rpc client sockets still have none at the TCP level, since gen_rpc
3.6.1 does not let an application set those options.
…live

Fable review of PR 447.

- On Linux TCP_USER_TIMEOUT overrides the keepalive probe count: an idle
  dead socket closes at the first unanswered probe at or after the user
  timeout. With 20 s and probes every 5 s the idle bound was 20 s, not the
  15 s the docs said, and keepalive_count did nothing. The user timeout is
  now 15 s, the bound for both an in-flight request and an idle socket;
  keepalive_count is gone and the docs state the rule. Measured in a
  network namespace: all four stale pooled connections were gone by 16 s,
  and the first request after the blackhole succeeded.
- A keepalive idle value Linux rejects (over 32767) was ignored silently,
  leaving the 2 h default. SMOLQUERY_PEER_TCP_KEEPALIVE_IDLE_S is now
  limited to 1..32767.
- Finch and gen_rpc already turn SO_KEEPALIVE on, with Linux's 2 h
  timers. The docs say so, and note that the pod's tcp_keepalive_* sysctls
  would bound gen_rpc's client sockets without a gen_rpc fork.
- The socket tests assert every raw option, not only the user timeout, and
  skip off Linux.
@chasers
chasers merged commit 00f3a18 into main Oct 2, 2026
11 of 12 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant