A remote MCP server that cannot be connected to does not fail the run. It deadlocks it, and the agent produces no output until an external timeout kills it.
Cause
mcp_lifecycle.mcp_session_task connects every server inside one try, and sets the connected event only after the whole loop succeeds:
try:
for entry in entries:
...
await entry.server.connect()
connected.set() # never reached if any connect() raises
await cleanup.wait()
...
except RuntimeError:
logging.exception("RuntimeError in mcp session task")
except asyncio.CancelledError:
logging.exception("Timeout on main session task")
finally:
entries.clear()
The handlers swallow the failure and return without setting connected and without recording an error. Meanwhile runner.py:745 waits on that event with no timeout:
mcp_sessions = asyncio.create_task(mcp_session_task(entries, servers_connected, start_cleanup))
await servers_connected.wait() # blocks forever
So the task finishes, nothing sets the event, and the awaiting coroutine is parked permanently.
The except asyncio.CancelledError arm makes it worse than it looks. A failed connect() surfaces as CancelledError from the anyio cancel scope rather than as a transport error, so this is the normal failure path, not an edge case. Verified by executing it:
import asyncio
from agents.mcp import MCPServerStreamableHttp
async def main():
s = MCPServerStreamableHttp(
name="x", params={"url": "", "headers": None, "timeout": 10},
client_session_timeout_seconds=10,
)
try:
await s.connect()
except BaseException as e:
print(type(e).__module__ + "." + type(e).__name__, "|", e)
asyncio.run(main())
# asyncio.exceptions.CancelledError | Cancelled by cancel scope 10d9b1400
An unreachable host or a refused port reaches the same handler.
Impact
Any misconfigured or transiently-down remote toolbox turns a run that should fail in seconds into one that burns its full time budget and yields nothing. It is also hard to diagnose from the outside: the process looks alive and the log line is "Timeout on main session task", which points at a timeout rather than at a connect failure.
We hit this with a kind: streamable toolbox whose URL was empty because its sidecar was not running. Six taskflows requested that toolbox, and every audit hung.
Suggested fix
Signal failure rather than swallowing it. Roughly:
- Record the exception on the entry or a shared holder, and set
connected in a finally so the waiter always wakes.
- Have the waiter re-raise, or skip the servers that failed, depending on whether a missing toolbox should be fatal.
- Consider
await asyncio.wait_for(servers_connected.wait(), timeout=...) in runner.py as a backstop, so a future non-signalling path degrades into a timeout instead of a deadlock.
- Narrowing
except asyncio.CancelledError around cleanup.wait() only, so connect failures are not misreported as timeouts, would also make the logs honest.
Happy to send a PR if you have a preference on fatal vs skip.
A remote MCP server that cannot be connected to does not fail the run. It deadlocks it, and the agent produces no output until an external timeout kills it.
Cause
mcp_lifecycle.mcp_session_taskconnects every server inside onetry, and sets theconnectedevent only after the whole loop succeeds:The handlers swallow the failure and return without setting
connectedand without recording an error. Meanwhilerunner.py:745waits on that event with no timeout:So the task finishes, nothing sets the event, and the awaiting coroutine is parked permanently.
The
except asyncio.CancelledErrorarm makes it worse than it looks. A failedconnect()surfaces asCancelledErrorfrom the anyio cancel scope rather than as a transport error, so this is the normal failure path, not an edge case. Verified by executing it:An unreachable host or a refused port reaches the same handler.
Impact
Any misconfigured or transiently-down remote toolbox turns a run that should fail in seconds into one that burns its full time budget and yields nothing. It is also hard to diagnose from the outside: the process looks alive and the log line is
"Timeout on main session task", which points at a timeout rather than at a connect failure.We hit this with a
kind: streamabletoolbox whose URL was empty because its sidecar was not running. Six taskflows requested that toolbox, and every audit hung.Suggested fix
Signal failure rather than swallowing it. Roughly:
connectedin afinallyso the waiter always wakes.await asyncio.wait_for(servers_connected.wait(), timeout=...)inrunner.pyas a backstop, so a future non-signalling path degrades into a timeout instead of a deadlock.except asyncio.CancelledErroraroundcleanup.wait()only, so connect failures are not misreported as timeouts, would also make the logs honest.Happy to send a PR if you have a preference on fatal vs skip.