Skip to content

feat(claude): report workflow runs as structured phases and agents - #26

Merged
dviejokfs merged 3 commits into
mainfrom
feat/claude-workflow-progress
Oct 2, 2026
Merged

dviejokfs merged 3 commits into
mainfrom
feat/claude-workflow-progress

Conversation

@dviejokfs

Copy link
Copy Markdown
Contributor

Summary

A run of Claude Code's Workflow tool reaches consumers as one workflow task plus a TaskActivity of kind Progress on every tick. Each of those activities has the same description and the most recently active agent's label in last_tool_name. A UI built on that shows hundreds of near-identical rows and no structure.

Claude already reports the structure. Each workflow task_progress carries a workflow_progress list: phases, agents with their state, tokens, tool calls, timings and previews, and script logs. The Workflow tool's result also names the run's transcript directory. This PR surfaces that.

Workflow snapshot

AgentTask::workflow: Option<AgentWorkflow>, replaced on every update:

  • The run's name, run_id and transcript_dir, taken from task_started.workflow_name and the Workflow tool's tool_use_result.
  • phases: index, title and kind.
  • agents: index, label, phase, normalized state, the raw native_state, agent_id, model, type, isolation, attempt, cached, tokens, tool calls, duration, queued/started/last-progress times, last_tool_name/last_tool_summary, prompt and result previews, and error. Normalized states are queued, running, completed, failed and blocked.
  • logs (the last 10) and omitted_agents.

Every tick resends the whole workflow, so it is bounded: 100 agents, 32 phases, and 240 characters per text field.

Activity per agent state change, not per tick

For a workflow task, TaskActivity is emitted only when an agent appears or changes state. The changed agent is in the new AgentTaskActivity::workflow_agent (boxed), and summary reads like scan:read completed. A real 3-agent run now produces 7 activities, one per agent state change.

Per-agent activity

Claude does not stream a workflow agent's own tool calls; it writes them to <transcriptDir>/agent-<agentId>.jsonl.

  • AgentWorkflow::agent_transcript_path(agent) returns that path. It does so only for plain agent IDs ([A-Za-z0-9_-]), so a value reported by the provider cannot point outside the directory.
  • Claude::transcript_activity(transcript, max_entries) replays the transcript through the adapter's live parser. It returns timestamped TextDelta/ReasoningDelta/ToolCall events with the same bounds and redaction as a live turn, skips a line that is still being written, and keeps the most recent entries.

Also

  • The result of a background shell's Bash call now carries the shell task's ID, so it groups under that task. Fleet had been carrying this as a local patch on its vendored copy; this upstreams it with its test.

Compatibility

AgentTask and AgentTaskActivity gain fields, so code that builds them literally must set workflow/workflow_agent (this is noted in the CHANGELOG). Both fields are #[serde(default)], so previously stored tasks and activities still deserialize; a test covers this.

Testing

  • Unit tests on anonymized fixtures captured from a real Workflow run (paths replaced, attachments dropped). They cover:
    • the snapshot's phases, agents and transcript path;
    • activity records state changes only;
    • state normalization;
    • bounds;
    • transcript paths cannot escape the directory;
    • transcript-to-activity, including a partial last line and max_entries;
    • deserializing older payloads;
    • the local_bash attribution.
  • Live, against Claude CLI 2.1.287 (haiku): cargo run --example claude_workflow_smoke printed PASS.
    • The snapshot showed phases Scan/Report and 3 completed agents, with tokens, tool counts and latest tool.
    • 7 agent state changes were recorded.
    • Each agent's transcript read back with its real Read/Bash calls.
  • CI gate, run locally: fmt, clippy -D warnings (Rust 1.99), the feature matrix, package checks, cargo doc, cargo package, fullstack-example clippy, and cargo test --all-features (415 tests).

🤖 Generated with Claude Code

David Viejo and others added 2 commits October 2, 2026 18:17
A Claude Workflow run is a `workflow` task. Its AgentTask::workflow now
carries the run's phases, agents (label, phase, normalized state, model,
tokens, tool calls, timings, latest tool, prompt and result previews) and
recent logs, parsed from Claude's workflow_progress and replaced on every
update, plus the run's name, identifier and transcript directory from the
Workflow tool's result. Collections and text are bounded because every tick
carries the whole workflow.

Workflow TaskActivity now records agent state changes, with the changed
agent in AgentTaskActivity::workflow_agent, instead of one Progress per tick.

Claude does not stream a workflow agent's own tool calls, so
AgentWorkflow::agent_transcript_path (plain agent identifiers only) and
Claude::transcript_activity rebuild an agent's text and tool calls from its
transcript with the live parser.

Also attributes a background shell's Bash result to its shell task.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Every workflow tick carries the whole snapshot, and consumers commonly
persist each event. A tick that only moves counters (tokens, tool calls,
durations, latest tool) is now re-emitted once per five such ticks, while
any state, result, error, phase or log change is emitted at once. The held
snapshot is still updated on every tick, so the next emission is current.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@greptile-apps

greptile-apps Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

RetriggerConfidence Score: 5/5

[Medium risk] Adds structured workflow task reporting to Claude provider.

The PR appears safe to merge; all three previous findings are addressed.

What we checked:

  • Final counters stay current: Every tick saves the latest workflow. Task completion sends that saved snapshot even when the next scheduled counter update has not arrived.

Summary

This PR adds structured Claude workflow snapshots and records each agent’s state changes instead of logging every progress tick.

  • Exposes phases, agents, recent logs, and transcript paths.
  • Replays recorded text, reasoning, and tool calls from agent transcripts.
  • Suppresses unchanged snapshots and slows counter-only updates.
  • Links background shell results to their task.
  • Addresses all three previous findings.
Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart TD
    A[Claude workflow progress] --> B[Save phases and agents]
    B --> C{What changed?}
    C -->|Agent state| D[TaskActivity with workflow_agent]
    C -->|Structure or task details| E[TasksChanged immediately]
    C -->|Counters only| F[TasksChanged every fifth tick]
    C -->|Nothing| G[No snapshot]
    H[Agent transcript] --> I[transcript_activity]
    I --> J[Text, reasoning, and tool events]
Loading

Reviews (2) · Last reviewed commit: "fix(claude): replay transcript reasoning..."

Comment thread src/providers/claude.rs
Comment thread src/providers/claude.rs Outdated
Comment thread src/providers/claude.rs Outdated
… change nothing

Transcripts record reasoning only as thinking blocks, so transcript_activity
now returns them as ReasoningDelta. Workflow ticks without workflow_progress
or with an identical snapshot no longer resend the whole workflow; counter
changes still go out at most every fifth tick. Docs no longer imply that
transcript entries are redacted.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@dviejokfs

Copy link
Copy Markdown
Contributor Author

@greptile-apps review

@dviejokfs
dviejokfs merged commit 4ee1a3d into main Oct 2, 2026
7 checks passed
@dviejokfs
dviejokfs deleted the feat/claude-workflow-progress branch October 2, 2026 20:41
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant