Skip to content

Probe a gen_rpc channel after a call timeout, drop it if silent (T-613) - #445

Merged
chasers merged 1 commit into
t-613-drop-stale-gen-rpc-clientsfrom
t-613-probe-stale-channel-on-timeout
Oct 2, 2026
Merged

chasers merged 1 commit into
t-613-drop-stale-gen-rpc-clientsfrom
t-613-probe-stale-channel-on-timeout

Conversation

@chasers

@chasers chasers commented Oct 2, 2026 •

Copy link
Copy Markdown
Owner

TL;DR: After a gen_rpc call times out, probe that channel's client. If the probe gets no answer, kill that client.

Tracker: T-613. Stacked on #444 (T-613).

Why

  • Drop a peer's gen_rpc clients when it leaves the cluster (T-613) #444 reacts only to :nodedown.
  • A call can also time out to a node that is still in Node.list/0, for example during a partition distribution has not seen yet.
  • In that case the client keeps its dead socket until the kernel kills it (about 15 min).
  • A timeout alone does not prove the socket is dead. The remote operation can be slow.

What changed

  • New Smolquery.Cluster.RpcProbe.pong/0. It returns :pong. Nothing else.
  • RpcProbe is added to gen_rpc's rpc_module_list.
  • New RpcClients.check/2. It returns at once and probes in the background.
  • Transport.GenRpc calls check/2 on {:badrpc, :timeout}.
  • QueryService.WorkerTransport calls check/2 on {:badrpc, :timeout}.

How it works

  1. A call to {node, key} returns {:badrpc, :timeout}.
  2. The transport casts the destination to RpcClients.
  3. RpcClients spawns one probe process for that destination.
  4. The probe looks up the destination's current client pid. No client: stop. A check never dials.
  5. The probe calls RpcProbe.pong/0 on that destination, with a 10 s timeout.
  6. No answer: kill the pid from step 4, and only that pid. The next call dials fresh.
  7. Any answer (an error too): keep the client.

Why 10 s

  • The probe waits in the client's mailbox behind earlier sends.
  • A send to a slow reader blocks for up to gen_rpc's send_timeout (5 s). Then the client fails and stops by itself.
  • So the probe waits longer than 5 s. It only catches a socket that takes bytes but sends nothing back.

Why a slow call keeps its socket

  • gen_rpc runs each call in its own process on the remote side.
  • So a slow operation does not block the probe.
  • The probe answers, and the client stays.

Tests

  • ✅ Frozen peer: check/2 kills the :control client and keeps :bulk and :scatter.
  • ✅ Slow call: a 3 s remote :timer.sleep runs on :control. The probe runs during it. The client stays.
  • ✅ No client: check/2 does not start one.
  • ✅ One probe per destination at a time.
  • ✅ Mutation checks: "always kill" fails the slow-call test. "Never kill" fails the frozen-peer test. "Dial without a client" fails the no-client test.

Review

  • A Fable review found 2 problems in the first version. Both are fixed in this commit:
    • The probe could start a new client, then kill it. Now it probes only an existing client and kills only that pid.
    • A 2 s probe could kill a healthy channel behind a slow send. Now the timeout is 10 s.

Checks

  • ✅ mix precommit: 3073 tests pass
  • ✅ mix ci
  • ✅ mix dialyzer
  • ✅ Transport peer tests (gen_rpc_test, worker_transport_peer_test, multi_node_ring_test) pass

Watch out

  • ⚠️ The task asked to check "no bytes received since the call was sent". This PR uses a probe instead. gen_rpc keeps the socket inside the client's private state, so recv_oct is not reachable without reading gen_rpc internals.
  • ⚠️ No test checks that a transport timeout calls check/2. That path is two lines in each transport.
  • ⚠️ A peer on an older build does not have RpcProbe on its allowlist. It answers {:badrpc, :unauthorized}. That counts as an answer, so the client stays.
  • ⚠️ Callers on a killed client get {:badrpc, {:unknown_error, _}}, which holds the request. Normalising it is T-616.

🤖 Generated with Claude Code

@chasers
chasers added this pull request to stack #446 October 2, 2026 13:45
A call can time out to a node distribution still lists, for a partition
distribution has not noticed yet. Layer 1 only reacts to nodedown, so
that client would keep its dead socket until the kernel gives up. A
timeout alone does not prove the socket is dead, though: the remote
operation may just be slow.

On {:badrpc, :timeout}, the buffer transport and the scatter transport now
hand the destination to RpcClients.check/2. It takes the destination's
current client and calls Smolquery.Cluster.RpcProbe.pong/0 on it. gen_rpc
runs each call in its own process on the remote side, so a slow operation
does not hold the probe up. No answer within probe_timeout_ms kills that
client, the one probed; any answer, an error included, keeps it. A
destination with no client is not probed, so a check never dials. One
probe runs per destination at a time, and the caller never waits for it.

probe_timeout_ms defaults to 10 s, above gen_rpc's 5 s send_timeout: the
probe queues behind sends on the client, and a send stuck on a slow
reader fails the client by itself first.

RpcProbe joins gen_rpc's rpc_module_list. It only answers :pong.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@chasers
chasers force-pushed the t-613-probe-stale-channel-on-timeout branch from 3b134e9 to 821b092 Compare October 2, 2026 14:11
@chasers
chasers merged commit 4426373 into main Oct 2, 2026
13 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant