Skip to content

Unreachable MCP server hangs the runner forever instead of failing #292

Description

@anticomputer

A remote MCP server that cannot be connected to does not fail the run. It deadlocks it, and the agent produces no output until an external timeout kills it.

Cause

mcp_lifecycle.mcp_session_task connects every server inside one try, and sets the connected event only after the whole loop succeeds:

try:
    for entry in entries:
        ...
        await entry.server.connect()

    connected.set()          # never reached if any connect() raises
    await cleanup.wait()
    ...
except RuntimeError:
    logging.exception("RuntimeError in mcp session task")
except asyncio.CancelledError:
    logging.exception("Timeout on main session task")
finally:
    entries.clear()

The handlers swallow the failure and return without setting connected and without recording an error. Meanwhile runner.py:745 waits on that event with no timeout:

mcp_sessions = asyncio.create_task(mcp_session_task(entries, servers_connected, start_cleanup))
await servers_connected.wait()          # blocks forever

So the task finishes, nothing sets the event, and the awaiting coroutine is parked permanently.

The except asyncio.CancelledError arm makes it worse than it looks. A failed connect() surfaces as CancelledError from the anyio cancel scope rather than as a transport error, so this is the normal failure path, not an edge case. Verified by executing it:

import asyncio
from agents.mcp import MCPServerStreamableHttp

async def main():
    s = MCPServerStreamableHttp(
        name="x", params={"url": "", "headers": None, "timeout": 10},
        client_session_timeout_seconds=10,
    )
    try:
        await s.connect()
    except BaseException as e:
        print(type(e).__module__ + "." + type(e).__name__, "|", e)

asyncio.run(main())
# asyncio.exceptions.CancelledError | Cancelled by cancel scope 10d9b1400

An unreachable host or a refused port reaches the same handler.

Impact

Any misconfigured or transiently-down remote toolbox turns a run that should fail in seconds into one that burns its full time budget and yields nothing. It is also hard to diagnose from the outside: the process looks alive and the log line is "Timeout on main session task", which points at a timeout rather than at a connect failure.

We hit this with a kind: streamable toolbox whose URL was empty because its sidecar was not running. Six taskflows requested that toolbox, and every audit hung.

Suggested fix

Signal failure rather than swallowing it. Roughly:

  • Record the exception on the entry or a shared holder, and set connected in a finally so the waiter always wakes.
  • Have the waiter re-raise, or skip the servers that failed, depending on whether a missing toolbox should be fatal.
  • Consider await asyncio.wait_for(servers_connected.wait(), timeout=...) in runner.py as a backstop, so a future non-signalling path degrades into a timeout instead of a deadlock.
  • Narrowing except asyncio.CancelledError around cleanup.wait() only, so connect failures are not misreported as timeouts, would also make the logs honest.

Happy to send a PR if you have a preference on fatal vs skip.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions