Skip to content

feat: injectable redis client and idle-connection handling for state persistence - #1209

Open
maartenbreddels wants to merge 1 commit into
masterfrom
feat/state-redis-client-factory
Open

maartenbreddels wants to merge 1 commit into
masterfrom
feat/state-redis-client-factory

Conversation

@maartenbreddels

@maartenbreddels maartenbreddels commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Managed Redis servers (and NATs in between) drop connections that sat idle for a few minutes. The Redis state backend built its client with from_url(...) and no health check. With from_url, redis-py gives each connection zero retries. So the next takeover on such a connection failed with ConnectionError: Error while reading from ...: Connection reset by peer, and the user got a fresh session instead of the restored one.

Built-in client

  • health_check_interval=30: a connection that sat idle is pinged before it is reused.
  • Retry(NoBackoff(), 1, supported_errors=(redis.ConnectionError,)): one immediate retry on a connection error.
  • Timeouts are never retried. A takeover that completes after connect_timeout deletes the stored state (the claim-or-delete path in kernel_context), so a retry must not outlive the deadline. supported_errors is set explicitly because redis-py's default includes TimeoutError. retry_on_error is set as well because redis-py 4.x ignores retry without it.

SOLARA_STATE_REDIS_CLIENT_FACTORY

A dotted path to a callable(settings) -> redis.Redis. It replaces the built-in client, for deployments that need their own pool, keepalive, TLS or Sentinel setup. validate_state_settings() imports it at startup, so a typo stops the server instead of failing every restore. The docs describe the contract: synchronous, thread-safe, decode_responses=False, bounded by the connect deadline, and dedicated to Solara.

No set_backend() setter: the backend is a lazy process-wide singleton, and swapping it while kernels hold the old one would leave two backends with separate breaker state.

Known and accepted

The retry covers every command, so a lost reply also re-runs a delete. A fenced delete then returns False although the first attempt deleted the key; that value only goes to a log line. For a retried delete to remove new state, another instance would have to recreate the same kernel's key inside the millisecond gap between the lost reply and the retry. Deletes happen on a tab close and on a poisoned state for that same kernel, so I think this is acceptable.

Tests

tests/unit/state_redis_client_test.py puts a small TCP proxy between the real redis-py client and a fakeredis TCP server:

  • reset idle connection: the proxy answers the next request with a RST, as after a silent idle drop. The takeover succeeds and the script ran once. Without the fix, this reproduces the production error exactly (Error while reading from 127.0.0.1:...: (54, 'Connection reset by peer')).
  • lost reply: the server ran the script but the reply is lost. The retry runs it again, and the caller holds the generation the store has.
  • timeout: the request is swallowed. It raises TimeoutError, with no reconnect and no second send.
  • The factory setting is used. Startup validation rejects a missing, unimportable or non-callable factory, and ignores the setting when a non-Redis backend is selected.

Note: a connection closed with a FIN before the next command is already handled by redis-py itself (can_read() check on checkout). The failing case is a drop the client never saw, which is why the proxy sends the RST in reply to the request.

I also checked end to end against a local Redis, with an app-side factory loaded through the env var: a takeover after CLIENT KILL restores the state, and the generation goes up by exactly one.

🤖 Generated with Claude Code

…persistence

Managed Redis servers and NATs drop connections that sat idle for a few
minutes. The built-in client had no health check and no retry, so the next
takeover on such a connection failed with "connection reset by peer" and the
user came back to a fresh session instead of the restored one.

The built-in client now pings a connection that sat idle and retries once,
immediately, on a connection error. It never retries a timeout: a takeover
that completes after the connect deadline deletes the stored state, so a
retry must not outlive that deadline.

Deployments that need their own pool, keepalive or TLS setup can now point
SOLARA_STATE_REDIS_CLIENT_FACTORY at a function that builds the client,
instead of overriding a private method. A bad path stops the server at
startup rather than failing every restore.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

This branch had an error being deployed

1 failed deployment
feat/state-redis-client-factory - solara-stable PR #1209 — a9b5d08b Deployed Sep 29, 2026 by maartenbreddels
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant