Skip to content

fix(server): deliver screenshot images to the model on custom providers - #2768

Merged
Dani Akash (DaniAkash) merged 2 commits into
mainfrom
fix/agent-screenshot-image-transport
Sep 28, 2026
Merged

Dani Akash (DaniAkash) merged 2 commits into
mainfrom
fix/agent-screenshot-image-transport

Conversation

@DaniAkash

@DaniAkash Dani Akash (DaniAkash) commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Fixes #2722

Summary

The built-in agent could not see screenshot images on custom / OpenAI-compatible providers (reproduced on OpenRouter with a vision-capable model). The image arrived on providers that keep media inside tool results, but on every other provider the tool result was replaced with [Tool content omitted] and the picture never reached the model, so the agent could only reason from the accessibility tree.

Root cause

For providers that cannot carry media inside a tool result, the agent strips the media out of the tool result and re-attaches it as a following user message (normalizeMessagesForModel). That extractor only recognized the deprecated image-data / file-data content-part shapes. The AI SDK MCP client actually emits the current canonical v7 shape:

{ type: 'file', mediaType, data: { type: 'data', data: <base64> } }

A file part fell through to the default case, so nothing was extracted and the image was dropped. The same shape gap in the compaction helpers is why the placeholder read [Tool content omitted] and why a screenshot was invisible to the binary-content budget. Providers that keep media in tool results (Anthropic, OpenAI, Azure, Bedrock, Gemini 3) were unaffected because they skip normalization.

Fix

  • message-normalization.ts: handle the canonical file part when re-attaching tool-result media. Only the inline data variant of the tagged data carries bytes, so a url / reference variant is left as-is.
  • compaction/content.ts: recognize the file part in the three helpers that classify, render, and budget binary tool-result content, so the placeholder reads [Image] / [File] and images are counted for the compaction budget.

No change to the providers that already worked; they continue to keep media inside the tool result.

Tests

Added message-normalization.test.ts covering the canonical file image re-attachment, a non-image file, the still-supported deprecated image-data shape, a non-inline (url) file, native-media pass-through, and the no-vision case. Existing compaction suites still pass.

The AI SDK MCP client emits an image tool result as the canonical v7 `file`
content part ({ type: 'file', data: { type: 'data', data }, mediaType }), but
the tool-result media extractor only handled the deprecated image-data and
file-data shapes. For any provider that cannot carry media inside a tool result
(OpenRouter and every other custom provider), the image was stripped to
"[Tool content omitted]" and never re-attached, so the agent could not see
screenshots. Providers that keep media in tool results (Anthropic, OpenAI, and
similar) were unaffected because they skip normalization entirely.

Handle the canonical `file` part when re-attaching tool-result media as a
following user message, and in the compaction helpers that classify, render,
and budget binary content, so a screenshot reaches vision-capable models on
every provider.

Fixes #2722
@github-actions github-actions Bot added the fix label Sep 28, 2026
@github-actions

github-actions Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

✅ Tests passed: 1370/1373

Ran 8 of 16 suites (8 not affected by this change).

Suite Passed Failed Skipped
✅ server-agent 216/216 0 0
✅ server-api 279/279 0 0
✅ server-tools 248/248 0 0
✅ server-browser 10/10 0 0
✅ server-integration 10/10 0 0
✅ server-lib 197/197 0 0
✅ server-root 38/41 0 3
✅ agent 372/372 0 0
⏩ claw-app n/a n/a not affected
⏩ claw-onboard n/a n/a not affected
⏩ app-onboard n/a n/a not affected
⏩ build n/a n/a not affected
⏩ release n/a n/a not affected
⏩ claw-server-rust n/a n/a not affected
⏩ claw-server-rust-quality n/a n/a not affected
⏩ claw-mcp n/a n/a not affected

View workflow run

@greptile-apps

greptile-apps Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 4/5

[Medium risk] Handles screenshot images for custom AI model providers.

The PR appears safe to merge, though the new regression tests should be included in the server test suite.

Summary

The PR recognizes canonical AI SDK file parts so screenshot images can be re-attached for providers that cannot carry media in tool results. It also updates compaction’s file-part classification and adds normalization tests.

  • The new tests are outside the server test runner’s selected directories.

Diagram

%%{init: {'theme': 'neutral'}}%%
flowchart LR
  A[MCP screenshot result] --> B{Provider supports media in tool results?}
  B -->|Yes| C[Keep tool result intact]
  B -->|No| D[Replace file part with text placeholder]
  D --> E[Attach inline image in following user message]
  E --> F[Compaction and model request]
Loading

Reviews (1) · Last reviewed commit: "fix(server): deliver screenshot images t..."

@DaniAkash
Dani Akash (DaniAkash) merged commit 1143158 into main Sep 28, 2026
19 checks passed
@DaniAkash
Dani Akash (DaniAkash) deleted the fix/agent-screenshot-image-transport branch September 28, 2026 06:52
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Agent never receives screenshot tool images — result arrives as "[Tool content omitted]"

1 participant