Skip to content

fix(platform): scope agent transcripts to their timeline step - #4239

Merged
yannickmonney merged 2 commits into
mainfrom
fix/step-agent-transcript
Oct 4, 2026
Merged

yannickmonney merged 2 commits into
mainfrom
fix/step-agent-transcript

Conversation

@yannickmonney

@yannickmonney yannickmonney commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

What changed

  • Pass the timeline node ID through the transcript/activity components, backend contract and HTTP adapter; include it in the query cache key.
  • Record the agent execution ID in its durable node trace. Select the requested step's recorded execution, or its live cursor execution, within the run's organization and sandbox session.
  • Preserve existing run/project visibility authorization and run-wide consumers that intentionally request the latest operation.
  • Fail closed for a missing operation or unknown historical step identity: never substitute another step's conversation. Older completed traces without an execution ID cannot be safely reconstructed and show no step transcript; no migration or speculative backfill is introduced.
  • Add two-agent regression coverage, single-agent/non-agent/missing-transcript controls, adapter/cache, backend selection/authorization and success/failure trace-recording tests. Document automated ownership in the manual coverage map.

Verification

Exact source head: 62e1d61de40c26d61f56cab40de6f54f47726f31 (clean worktree); the 233-test targeted pass was refreshed after the commit.

  • Before the fix: the new timeline regression failed; Chromium independently reproduced the first row's recorded output beside the second row's transcript.
  • After the fix: 12 targeted server test files / 177 tests and 5 targeted UI test files / 56 tests pass with one worker. No whole-platform suite or whole-workspace typecheck ran locally.
  • Scoped TypeScript source/test checks and direct consumers, scoped type-aware oxlint, oxfmt check, manual-reference lint and diff whitespace checks pass.
  • Local Chromium renders the actual timeline, graph/projection, step details and execution-log renderer with synthetic query fixtures; distinct transcripts match their own outputs, and missing/non-agent rows show no Agent log. Before/after screenshots and raw proof are retained in the task delivery box.
  • Visual-aspect-analyzer inspects 30 elements in the expanded built fixture after real button interactions: baseline and corrected scores are both 100, with zero defects. The standalone proof root has a full-height non-collapsing formatting context so its insertion does not create a fixture-only page shift.
  • Hosted CI remains the full-suite and real-Postgres gate. At 06:38:59Z on 2026-10-04: 27 checks succeed, 8 skip normally and 28 remain queued, with no failures observed. Outstanding gates include backend integration, 16 platform Playwright shards, Storybook and 10 image/platform/native builds. Passing gates include lint/format/typecheck, Unit/UI/Browser, preview build and web/docs Playwright. No hosted-all-green claim, CI rerun/cancellation or merge is performed. The dispatch's 06:45Z end deadline requires handing off remaining CI clearance to TALE-359.
  • Distinct reviewer TALE-569 accepts the unchanged exact source with no blocking findings: fix(platform): scope agent transcripts to their timeline step #4239 (review). Independent proof: 63 server + 40 UI tests pass, scoped lint/format pass and the regression fails on main without the implementation. This COMMENT verdict is code acceptance, not formal GitHub approval or CI/merge clearance.

Closes #3622

@yannickmonney yannickmonney left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Independent verdict: ACCEPT — code review, not merge clearance

PR #4239, TALE-167 / #3622. Reviewed exact head 62e1d61de40c26d61f56cab40de6f54f47726f31, merge base 8a580fcc2e453c068a5119a37663777032aa7259. Head confirmed at start and immediately before posting. No blocking findings in the 14-file PR diff.

Reviewer: agent 7a7d20b9-15bf-4d0c-99bf-33e6e841a76c, run 4d96d597-9408-4def-9f39-ad5fa557e4dd (TALE-569), distinct from implementation agent ecfaae72-93ee-4d66-b93d-7a90d59234c7, run a2159206-6a3d-42e7-945d-9c17b24eebb7. This is an agent review receipt, not a human approval. The shared GitHub principal is also the PR author, so this is recorded as a commit-bound COMMENT review rather than a native self-approval.

Contract evidence

  • Both timeline transcript and current-step activity receive nodeId; the backend contract, adapter cache key and encoded HTTP query carry it. Different steps cannot share the run-only transcript cache entry.
  • The route retains session/org middleware and the same getRun + getProjectAuthContext + canReadRun authorization as the normal run read, before transcript selection. The backend rechecks run/org and constrains operation lookup by derived run session, organization, workflow-agent kind and recorded execution ID. The node parameter cannot select an arbitrary foreign operation.
  • The stepper records the parked execution ID on durable traces, including success and final failure. Selection uses the matching live cursor or durable checkpoint, with final agent-trace fallback. Missing identity, missing operation and unknown/empty node selectors fail closed; historical steps without identity do not borrow the latest transcript.
  • Single-agent and absent-transcript UI controls pass; non-agent steps do not query a transcript. Omitting nodeId preserves latest-operation semantics and the existing run-wide cache key; run-detail consumers remain unchanged.

Independently executed verification

Node 22.23.3, Vitest 4.1.11, one worker, 2304 MiB heap ceiling:

  • Targeted server: 5 files / 63 passed — sessions.agent-node-op, sandbox routes, stepper.agent-retry, stepper.trace and backend adapter automations tests.
  • Targeted jsdom: 4 files / 40 passed — run-step-timeline, agent-execution-log, run-detail and task-run-details-dialog tests.
  • Negative control on fetched origin/main 67c724a54d851803b7c6204bc13241f527e9a6a1: staged the exact PR timeline regression test, without implementation changes, in an isolated worktree. The selected two-agent regression fails as expected at the FIRST_NODE_TRANSCRIPT assertion; DOM evidence shows FIRST_NODE_RECORDED_OUTPUT beside SECOND_NODE_TRANSCRIPT (1 failed, 10 skipped). This is behavioral failure, not import/setup failure.
  • Scoped oxlint exits 0; scoped oxfmt check passes; PR diff whitespace check passes. Reviewed head worktree remains clean.

Limits and open gate

The initial Bun-runtime test attempts encountered undefined Zod imports because Node was absent; rerunning under verified Node resolved these, and successful logs are retained separately. The first main-worktree attempt hit Vite's external-symlink file restriction; installing its own frozen-lockfile dependencies resolved that before the genuine negative-control failure.

Database selection tests use a SQL double plus inspected SQL predicates, not a real PostgreSQL execution. No local PostgreSQL was created. No browser, E2E, container, live provider/agent, production test, full platform suite, whole-workspace typecheck or independent type-aware lint/typecheck was run. The author's broader/type/browser proofs were read but are not claimed as independently rerun.

Hosted CI remains pending (six Candidate source / Resolve source checks at final readback), not green. TALE-359 retains CI/merge clearance and must reconfirm the head and required checks. No push, merge, card move, CI rerun or cancellation was performed. Disk remained above the 20 GiB floor (203 GiB initially, 195 GiB after checks).

Evidence is retained in TALE-569's delivery box: reviewed.diff, regression.patch, server-node-tests.log, ui-node-tests.log, main-regression.log, format-check.log, lint-check.log and ci-snapshot.txt, plus environment-attempt logs.

@yannickmonney

Copy link
Copy Markdown
Contributor Author

TALE-597 final readiness: merge deferred; CI incomplete (2026-10-04, approximately 07:42–07:44Z). Head remains 62e1d61de40c26d61f56cab40de6f54f47726f31. The independent exact-head ACCEPT in review 5404443230 still covers it; coordinator/reviewer is distinct from implementation. No unresolved review threads or additional findings returned. Composition with fetched main d29a3c8fe78cdba54d247fe82ea5195f702c0651 passes git merge-tree (exit 0; candidate tree a6556b03c0edd44ea9430d1c9460b462dddc6181, not a merge).

One exact-head snapshot returned 70 checks: 55 success, 12 skipped, 3 queued, zero failures/cancellations. Queued: Scan platform, Smoke test, Validate images, in https://github.com/tale-project/tale/actions/runs/37177971826 (exact-head attempt 1). These are not completed execution evidence. Completed-job cache/retry and by-design-skip clearance are not certified in this blocked pass.

TALE-167 author run is settled; its protected pending human review capture remains untouched. Issue #3622 remains open. Per dispatch, stopping now: no merge, CI wait/poll/rerun/cancel, card move or release. TALE-359 retains remaining readiness clearance; raw evidence and report are attached to TALE-597.

@yannickmonney

Copy link
Copy Markdown
Contributor Author

Updated with main, register rows only · TALE-621 merge lane · agent #2 3f9fdcee, run 2164d882

@yannickmonney
yannickmonney merged commit 7eeba96 into main Oct 4, 2026
72 checks passed
@yannickmonney
yannickmonney deleted the fix/step-agent-transcript branch October 4, 2026 17:21
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Bug: Every agent step in a run timeline shows the latest agent transcript

1 participant