Skip to content

[Bug] Forwarded Proxy Admin calls have no timeout when the caller omits a deadline #11228

Description

@3219378872

Before Creating the Bug Report

Runtime platform environment

Ubuntu 24.04 / WSL2, x86_64. Reproduced with a real loopback Netty gRPC peer and the production ProxyAdminForwarder, using a stubbed consumer directory to route a remote client. No running broker is needed for this forwarding-path reproduction.

RocketMQ version

develop at 78b96bc5e21216cd7896efae08f90c5cde4cae53, including the pinned rocketmq-apis submodule at 3e60073c6dab3430feaec443824802e23ecda681.

JDK Version

Temurin 8u504-b01, Maven 3.9.9.

Describe the Bug

ProxyAdminForwarder.invokeOnPeer takes a cached AdminGrpc.AdminStub, attaches metadata and invokes the peer without setting a deadline. If the caller does not supply a gRPC deadline, the outgoing call has none. A peer that accepts the call but never responds can retain the forwarded request indefinitely, even though the proxy has grpcAdminServerRequestTimeoutMillis configured.

This affects the client-directed admin RPCs forwarded to the owning proxy: GetConsumerRunningInfo, PrintThreadStackTrace and VerifyMessage. Local broker/telemetry timeouts do not bound a stalled network hop to the peer.

Steps to Reproduce

  1. Enable Proxy Admin and set its request timeout (the regression fixture uses 1000 ms).
  2. Register a remote client in the consumer directory, pointing to a loopback peer running the normal Admin gRPC service.
  3. Have that peer accept GetConsumerRunningInfo but return no response.
  4. Call forwardIfRemote without an incoming deadline and inspect the peer's gRPC context.

The peer receives no deadline. In the new regression test, the baseline fails:

Tests run: 2, Failures: 1, Errors: 0, Skipped: 0
java.lang.AssertionError: the peer call must carry the configured deadline

The shorter-caller-deadline control passes on the baseline, confirming that propagation works when the caller supplies one; it is the proxy's own bound that is missing.

What Did You Expect to See?

Apply a fresh configured timeout to each forwarded call. An unresponsive peer should terminate with DEADLINE_EXCEEDED; a shorter incoming deadline should still win. The cached channel/stub must remain usable for subsequent calls after an earlier call times out.

Proposed fix and validation

Apply withDeadlineAfter on the per-invocation stub using a dedicated grpcAdminServerForwardTimeoutMillis budget (proposed default 15000 ms), and add real-transport regression tests for the stalled peer, cached-channel reuse and shorter incoming deadline.

The existing broker timeout (3000 ms by default) must keep its per-broker meaning: a forwarded VerifyMessage may first query a broker and then wait for the telemetry relay (5000 ms plus its timeout slack by default). The peer hop therefore needs its own whole-request budget rather than prematurely reusing the broker timeout.

The reproduction is scoped to forwarding and cancellation; it is not a full multi-proxy/broker deployment benchmark. I am preparing the fix and tests with AI assistance and will include the exact validation commands in the PR.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions