Skip to content

fix(loop): make runaway recovery reach Anthropic and stop treating it as no_tool - #77

Merged
zhanghanduo merged 3 commits into
mainfrom
fix/opus-runaway-effort-and-loop-exit
Oct 9, 2026
Merged

zhanghanduo merged 3 commits into
mainfrom
fix/opus-runaway-effort-and-loop-exit

Conversation

@zhanghanduo

@zhanghanduo zhanghanduo commented Oct 9, 2026 •

Copy link
Copy Markdown
Collaborator

Problem

On ApodexHarness's 2026-10-08 GDPval full run (claude-opus-5-5, effort=max), three tasks ended stopped_by=no_tool with no deliverable after 3–7 turns (40a99a31, 8079e27d, 99ac6944). All three were tasks that ask for live external data (S&P 500 closing prices, a priced equipment list, a hardware bill of materials) in a sandbox with no network. The model tried to recall that data in private reasoning, and the turn went:

runaway_retry=1/3 next_cap=8192 next_thinking_mode=reduced
runaway_retry=2/3 next_cap=4096 next_thinking_mode=disabled
runaway_retry=3/3 next_cap=2048 next_thinking_mode=disabled
LLM reasoning runaway persisted after 4 resamples; returning empty response for loop-level nudge handling
[ReAct] done stopped_by=no_tool

Two independent defects:

  1. The Anthropic adapter ignored the ladder's override. ThinkingRetryOverride is read only by providers/openai_chat.py. AnthropicClient._build_kwargs always sent the construction-time thinking and effort; with the override active the request body was byte-identical except max_tokens. Every rung re-asked the model for effort=max in less room.
  2. The loop had no "loop-level nudge handling". A runaway response (cap reached, no text, no tool call) fell into if not parsed_calls, and under no_tool_behavior="stop" ended the run as an ordinary finish.

Evidence that effort is the right lever

  • thinking={"type":"disabled"} on claude-opus-5-5 is rejected by the upstream (Bedrock 400: "thinking.type.disabled" is not supported for this model. Use "thinking.type.adaptive" and "output_config.effort" to control thinking behavior). So the disabled rung cannot be expressed through thinking.
  • Same runaway-shaped prompt, max_tokens=8192:
effort stop output tokens text chars
max max_tokens 8192 0
high end_turn 1241 2365
medium end_turn 1319 2311
low end_turn 953 2157

Change

  • AnthropicClient sends override.reasoning_effort as output_config.effort when a retry override is active. Like the OpenAI adapter it only replaces an effort the profile opted into, and it never touches thinking.
  • The loop recognises a runaway response before the no_tool branch (public is_runaway_response), appends RUNAWAY_LOOP_RECOVERY_GUIDANCE and continues, up to LoopConfig.runaway_max_loop_recoveries (default 1, reset by any tool-call turn). When those run out it stops with stop_reason="reasoning_runaway", registered in AgentBus fan-in as an incomplete report.

Change fragment: changes/77.feature.md (MINOR: new config field and stop reason).

Tests

tests/test_runaway_loop_exit.py (10 tests): effort replaced for every ladder mode with thinking unchanged; no effort added to a client without one; runaway turn recovered under no_tool_behavior="stop"; persistent runaway named reasoning_runaway; allowance resets after progress; zero allowance still names the failure.

Mutation check: reverting the provider change fails 3 tests; disabling the loop branch fails 4. Full suite: 2322 passed, 2 skipped.

🤖 Generated with Claude Code

zhanghanduo and others added 3 commits October 9, 2026 21:08
… as no_tool

Two defects turned a reasoning runaway on claude-opus-5-5 into a lost
task (GDPval 2026-10-08: 40a99a31 / 8079e27d / 99ac6944, all asked to
look up live data offline):

- providers/anthropic.py never read the ladder's ThinkingRetryOverride,
  so 8192 -> 4096 -> 2048 went out at effort=max every time. It now maps
  the override's reasoning_effort onto output_config.effort. Thinking
  itself is left alone: opus-5-5 400s on thinking.type=disabled.
  Measured on a runaway prompt at an 8192 cap: max -> 0 text at the cap;
  high/medium/low -> full answers.
- call_llm returned the runaway "for loop-level nudge handling", but the
  loop had no such branch, so no text + no tool call reached no_tool and
  ended the run under no_tool_behavior="stop". The loop now recovers once
  (LoopConfig.runaway_max_loop_recoveries) with explicit guidance, then
  stops with stop_reason="reasoning_runaway".

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@zhanghanduo
zhanghanduo merged commit 588bc40 into main Oct 9, 2026
5 checks passed
@zhanghanduo
zhanghanduo deleted the fix/opus-runaway-effort-and-loop-exit branch October 9, 2026 13:34
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant