Repository navigation
fix(llm): unify empty completion recovery and provider signals - #70
Conversation
Some OpenAI-compatible gateways send `refusal: ""` alongside ordinary content or tool calls. Treating any non-None refusal as a decline made agent_loop end the run as `refusal` and drop the tool calls. An empty refusal now counts only when the turn carries nothing else; the streaming path holds that marker until the finishing chunk. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
LLMProxy.stream handed after_llm an LLMResponse carrying only text and reasoning, so streamed calls looked usage-free to RateLimitMiddleware and TokenAccountingMiddleware, and refusal-aware middleware never saw stop_details. Fold usage, usage_source, finish_reason, model, provider, refusal, stop_details and stop_reason the same way the loop's stream assembler does. Transport heartbeats are still ignored. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Review follow-ups: two fixes pushed to this branch
|
Blank responses after PR #68 still recover differently in chat and streaming modes: retries can reset the logical deadline, native fallback chains cannot receive the assembler's late error, and rejected blanks are recorded as accepted attempts. Synthetic usage and dropped refusal fields can also hide or misclassify the response.
This change makes
call_llmown blank recovery in both modes. Same-leg resamples, native chain advances and attempt accounting share one logical deadline; per-call cursors preserve bindings, middleware and concurrency isolation. Provider adapters retain usage provenance and OpenAI refusal signals. Direct fallback streams buffer candidate-empty fragments so discarded tool arguments cannot contaminate the next leg. Failed tool-argument replays get independent attempt identities while the original response remains deliverable. Empty refusal placeholders are evaluated against text, tool calls and both reasoning channels; streaming inference waits for normal EOF even when finish_reason is omitted, without changing error or cancellation semantics.Compatibility details are documented in
docs/llm-runtime-boundary.mdand a breaking release fragment:empty_completion_max_retriesapplies to both transports independently of the generic retry allowance, per serving leg; matching native fallback triggers determine advancement.empty_completion; explicit refusals and filters stop asrefusal/content_filter, including under a nudge policy. Accompanying tools are not executed and history remains replayable.estimated: trueor setusage_source="estimated". Unmarked legacy usage remains conservative.for_logical_callandadvance_empty_completion; dynamic attribute forwarding alone never removes the wrapper.Validation:
Middleware consumer corrections: reported usage is distinguished from estimates before cost/budget/event/aggregation effects; tracing retains estimate provenance. Canonical cache read/write fields preserve zeros and cache-only reports. Rate buckets correct zero estimates and reported zero usage against per-instance capped reservations, with legacy-context compatibility and no refund for unknown/estimated usage. Proxy and runtime share indexed tool assembly so real loop detection sees native streamed calls; failed/consumer-closed proposals stay out of history while authentic usage is still accounted. These behaviors are covered by 50 additional real-consumer integration cases using CostSink, BudgetState, usage aggregation and quota buckets. Streaming billing/budget changes and downstream recalibration are documented in the existing breaking fragment.
Verified against current main's stream-terminator guards: protocol validation precedes EOF refusal inference, and truncated Chat/Responses refusal streams still raise rather than becoming accepted declines. Combined full suite: 2,284 passed, 1 skipped.