Repository navigation
fix(loop): make runaway recovery reach Anthropic and stop treating it as no_tool - #77
Merged
Merged
Conversation
… as no_tool Two defects turned a reasoning runaway on claude-opus-5-5 into a lost task (GDPval 2026-10-08: 40a99a31 / 8079e27d / 99ac6944, all asked to look up live data offline): - providers/anthropic.py never read the ladder's ThinkingRetryOverride, so 8192 -> 4096 -> 2048 went out at effort=max every time. It now maps the override's reasoning_effort onto output_config.effort. Thinking itself is left alone: opus-5-5 400s on thinking.type=disabled. Measured on a runaway prompt at an 8192 cap: max -> 0 text at the cap; high/medium/low -> full answers. - call_llm returned the runaway "for loop-level nudge handling", but the loop had no such branch, so no text + no tool call reached no_tool and ended the run under no_tool_behavior="stop". The loop now recovers once (LoopConfig.runaway_max_loop_recoveries) with explicit guidance, then stops with stop_reason="reasoning_runaway". Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
On ApodexHarness's 2026-10-08 GDPval full run (claude-opus-5-5,
effort=max), three tasks endedstopped_by=no_toolwith no deliverable after 3–7 turns (40a99a31,8079e27d,99ac6944). All three were tasks that ask for live external data (S&P 500 closing prices, a priced equipment list, a hardware bill of materials) in a sandbox with no network. The model tried to recall that data in private reasoning, and the turn went:Two independent defects:
ThinkingRetryOverrideis read only byproviders/openai_chat.py.AnthropicClient._build_kwargsalways sent the construction-timethinkingandeffort; with the override active the request body was byte-identical exceptmax_tokens. Every rung re-asked the model foreffort=maxin less room.if not parsed_calls, and underno_tool_behavior="stop"ended the run as an ordinary finish.Evidence that effort is the right lever
thinking={"type":"disabled"}on claude-opus-5-5 is rejected by the upstream (Bedrock 400: "thinking.type.disabled" is not supported for this model. Use "thinking.type.adaptive" and "output_config.effort" to control thinking behavior). So thedisabledrung cannot be expressed throughthinking.max_tokens=8192:Change
AnthropicClientsendsoverride.reasoning_effortasoutput_config.effortwhen a retry override is active. Like the OpenAI adapter it only replaces an effort the profile opted into, and it never touchesthinking.no_toolbranch (publicis_runaway_response), appendsRUNAWAY_LOOP_RECOVERY_GUIDANCEand continues, up toLoopConfig.runaway_max_loop_recoveries(default 1, reset by any tool-call turn). When those run out it stops withstop_reason="reasoning_runaway", registered in AgentBus fan-in as an incomplete report.Change fragment:
changes/77.feature.md(MINOR: new config field and stop reason).Tests
tests/test_runaway_loop_exit.py(10 tests): effort replaced for every ladder mode withthinkingunchanged; no effort added to a client without one; runaway turn recovered underno_tool_behavior="stop"; persistent runaway namedreasoning_runaway; allowance resets after progress; zero allowance still names the failure.Mutation check: reverting the provider change fails 3 tests; disabling the loop branch fails 4. Full suite: 2322 passed, 2 skipped.
🤖 Generated with Claude Code