Repository navigation
feat(platform): keep, read and replay a step-by-step record of every automation run - #4647
Merged
Merged
Conversation
An action declares an English title and per-locale overrides (i18n.<locale>.title), and a connector named with ordinary words its localized display names (i18n.<locale>.displayName), on the locale grammar every declared text already uses.
Each of the 121 shipped connector actions gets a sentence-case title with German and French overrides, and the seven connectors named with ordinary words (Tasks, Documents, Conversations, …) their German and French display names. A guard fails a shipped action without its titles, a sloppy one, a duplicate, and a new connector that is neither a listed brand nor translated.
GET /automations/catalog/node-types answers, per connector action, its connector and its title with the German and French overrides, and once per connector its display name, overrides and shipped icon as a data URL. The registry carries the display data on the engine's connector surface; the app parses the answer with one zod schema instead of casting it, and the editor and run page read its nodeTypes.
The connector contract pages in English, German and French list the action's title and its i18n overrides, and say when a connector carries translated display names and when a brand keeps its own.
Read-only code (chat, docs, the guides) used Shiki's min-light and
min-dark themes, whose comments, parameters, constants and string
expressions fell below 4.5:1. The --code-* variables in globals.css now
hold one palette, corrected to AA on every code surface in both themes,
and Shiki reads it through its css-variables theme: one highlight serves
light and dark, so a theme switch no longer tokenizes again.
Template-aware grammars (tale-template, json-template, yaml-template,
markdown-template) highlight {{ js }} expressions as JavaScript, and
only under their own roots, so plain YAML keeps Helm or Jinja braces as
text. code-palette.test.ts measures every token on each surface.
A widget that uses Escape itself, such as a code editor closing its completion list, sat inside dialogs, sheets and popovers whose layer closed on the same key: Radix hears Escape first, in the capture phase. A widget now marks itself with data-claims-escape while it needs the key, and every @tale/ui overlay passes its onEscapeKeyDown through respectEscapeClaims, which keeps the layer open for a claimed Escape.
A code editor's text is a contenteditable element, which <label for> cannot name or focus. Field now gives its label an id and names a child that declares fieldLabelling = 'labelledby' with aria-labelledby; Label focuses a target that is not labelable when it is clicked. A go-to request for a part of a registered anchor (/nodes/0/input/to on /nodes/0/input) now reaches a custom target with that part: the rest of the path and the range, which is offsets into the part, so a JSON or YAML editor can select it. Element targets ignore it, as before.
@tale/ui/code-editor is the one control for code-ish fields, on
CodeMirror 6 behind a lazy chunk: the light module draws the frame and a
same-size placeholder, registers with IssueFocus and queues focus until
the editor arrives, so a page without a code field never loads it.
- Languages: a script, one expression, JSON, YAML, Markdown (from
@lezer/markdown, so no HTML or CSS language comes along), templates and
text. {{ js }} templates parse as JavaScript inside text, JSON and YAML
strings and Markdown, drawn as one tinted chip; typing {{ opens a pair.
- Colours: the --code-* palette Shiki reads; a parity test compares the
role of every character with the read-only highlight.
- Keyboard: Tab indents, Esc then Tab leaves, Mod-Enter submits; a
one-line editor never takes Tab. A claimed Escape is handled before any
overlay, so the first Escape in a sheet never closes it.
- Field and IssueFocus: named by its label, described by the field's
problems, and a go-to lands on the exact range, also inside a JSON or
YAML value (locate*Pointer).
- Every CodeMirror phrase is translated (en, de, fr, de-CH); guards keep
one copy of @codemirror/state and view and refuse the language bundles.
- Problems: the host's (a check's results, placed on the text it saw
and mapped through later edits), a host's instant lint provider, and the
editor's own syntax marks (JSON or YAML that does not parse, an unclosed
{{, unreadable JavaScript), merged in one state field. Underlines say the
severity by style, the gutter by glyph; the tooltip shows the message,
the host's detail, the code and fix buttons. F8 / Shift-F8 walk them and
read each aloud; Mod-. applies a fix; a summary describes the field
unless a Field already lists its problems.
- Completion from the host's provider with the member chain, a JSON or
YAML pointer and the region the cursor is in; names that are not
identifiers go in as ["a b"]; Tab or Enter accepts; the list keeps
inside a sheet and Escape closes it before the sheet.
- Hover types after 200 ms, and on Mod-K Mod-I, read aloud.
- Expand opens the field in a large dialog and brings the caret back.
- Every new string in en, de, fr and de-CH.
SchemaTree lists the fields a value has the way a form reads: each field's name, its kind in words (text, a number, list of objects, one of “open” or “closed”), whether it is required, a host's tag (from the trigger) and whether it may be empty; compact for a summary, comfortable with nested fields and descriptions, and the same shape as TypeScript behind a disclosure. schemaKindLabel gives the kind alone. French carries the preposition in the plural kinds, so a list of objects elides. VendorIcon moves unchanged from the platform's credential settings to @tale/ui/vendor-icon, so automation nodes can show a connector's icon.
Shiki stops tokenizing a line after 500 ms and leaves the rest of it plain. The first highlight in a language compiles its grammar, which on a loaded machine took that long, so the palette test saw a comment as plain text and failed. Both highlight tests now pay for the grammars on other snippets first.
A page that shows no code field, or only imports locate, template-scan or providers, must never download CodeMirror. The guard walks the static imports from each public code-editor module through the package's own files and refuses any CodeMirror or Lezer module, or the editor view itself, which loads only through a dynamic import. A walk from the view proves that it would catch one.
Read-only code now colours through the --code-* variables, so one highlight serves light and dark. The document and skill asset previews still asked again for the min-* theme on every switch, blanked to plain text until the answer came back, and their tests expected the old min-* class. They now highlight without a theme and keep it across a switch; the tests check the one palette in both themes.
ui.tale.dev gains a Code editor guide (fields, languages, templates, problems, completion and types, the keyboard, read-only code, go-to and every prop) with six live demos, and a Schema tree guide with one. The Dialog guide says how a widget claims Escape, Colors lists the --code-* roles and the contrast they keep, and Icons shows VendorIcon with its fallback. Stories cover the three components; the package README lists them.
CSS reaches most animation through motion-safe and the global reduced-motion rule; a React Flow viewport ease, a Web Animations call or a caret's blink runs from script and must ask itself. The hook and its one-off reader name the query once, and the code editor reads its caret blink through it.
A FlowGraph is plain data a host builds from its document: Start, steps, gates, End, the edges between them and a frame round each iterating step. layoutFlowGraph hands it to ELK (layered, orthogonal, model order, one port per edge end) and answers absolute boxes, frame headers, axis-aligned routes, label boxes and rows. A graph already on screen is laid out again with position hints, so no row changes order. ELK runs in elkjs's own same-origin worker, falls back to the bundled build with one warning when the worker cannot start, and to a single column when ELK fails or takes five seconds. Sizes are a pure function of the graph; fresh layouts are cached for the session. Measured on the fixtures (shipped Triage, review and inbox automations, a when/elseOf/else-if/repeat graph, a gate on the input, 40 nodes, a cycle): no edge crosses a box or a frame header, no label overlaps, five scripted edits keep every row in order. Pinning Start and End to their own layers cost crossings (2 against 0 on 40 nodes) for no guarantee the graph does not already give, so they are not pinned.
…inds WorkflowCanvas renders a FlowGraph from data on its own layout: Start with its triggers and inputs, steps with type, returns, chips and what they read, conditions as pills with Yes and No, End with what a run returns and how it ends, dashed frames round iterating steps. Every edge follows its ELK route with rounded corners and a tinted arrow; none is React Flow's handle-to-handle curve any more. The chart is one Tab stop: arrows follow the lines and the rows, Home and End jump, Enter opens. Each node is a real button named and described in words (position, neighbours, conditions, reads); the List view (FlowStepList) says the same as text, and FlowLegend explains the marks. FlowNodeStatus is the one run-state vocabulary. FlowCanvas loses the unused AI toggle, gains the auto fit policy, top corners and corner actions, and eases viewport moves out (quint) over the duration tokens, instantly under reduced motion. The No branch takes --flow-edge-negative (amber-700 in light, 2.2:1 amber failed). Package strings live under flow in en, de and fr, with de-CH for ss.
Knip found constants and helpers the flow modules exported but nothing outside them reads: the box and gate size tables, the frame header id, the row grouping, the node ring and lift classes, the chip pill and the contrast helper's internals. They stay module-private; an unused frame header height and end-direction helper go.
Three 40-node graphs laid out in the worker in Chromium, after one warm-up layout, must each finish under 1.5 s (an order-of-magnitude guard for a loaded runner; locally 45 to 84 ms). The times are logged.
React Flow turns pointer events off on an edge nothing selects; the routed edge's wide hit path takes them back. Hit-testing the middle of Open issues → Report finds the path whose title carries its detail.
ui.tale.dev gains a Workflow canvas guide after Flow node issues: the graph a host builds, the routed lines and their kinds, frames, Start and End, problem markers, the List view, the keyboard, fit and touch, layoutFlowGraph and flowNodeSize, and every prop. Six live demos show a laid-out workflow, a route round a frame, frames, Start and End, problems and the List view; Colors explains the lines' tokens, including --flow-edge-negative. Stories mirror the demos on the shared fixtures, and the package README names the new imports.
An arrow key moved focus with focus(), whose scroll-into-view shifted the clipped canvas frame under the chart before the canvas could pan. Focus now moves without scrolling and the canvas pans the box in with the least move. A browser test walks to End and back to Start through revealId on a frame too short for the graph.
The list opens on "nodes" while the test types; under a loaded full browser run the update after the dot landed after the test read the options and it failed once (it passes alone). The test now waits for the list to show what follows the dot.
A host that controls the view puts its own switch in topStart; the List view dropped it, so the reader could not switch back. The list's top row now holds topStart on the left and topEnd on the right. onLayout fires once per layout, whatever the callback's identity.
A workflow's possible paths are plain data now (@tale/ui/flow/paths): which nodes run on each and how each condition decided. Pure helpers turn them into the highlight a canvas draws — every path, the paths through one branch of a condition (or, when the host could not list them, what lies below the branch), a set of nodes in the error tone, a node's own lines — Start and End on every path, a Yes or No line only where its condition decided that way. FlowPathList is the list a reader explores them with: pointing at or focusing a row previews its path, Enter, Space or a click pins it and says so once, Escape or Show all unpins. Rows that are not paths (the nodes that end a run when they fail) activate instead. One Tab stop across every section, arrows and Home/End between rows, and while a path is pinned the list claims Escape from a sheet around it.
@tale/ui/flow/playback carries a run onto a workflow chart: an overlay (where each node ended, no time) or a playback timeline (spans of work, values travelling the lines, wait and failure marks) at the host's moment t. flowStateAt and flowStateFromOverlay derive the one frame a chart draws — each node's state, items and passes, each condition's decision, which lines were travelled or not taken, and the first failure with the way the run took to it. buildPlaybackTimeline compresses a recorded run piece by piece (a step plays at least 240 ms, a gap at most 1.2 s, a long wait becomes a mark with its real length), gives each value time to reach its target, and maps playback time back to real time both ways. usePlaybackClock drives it: it opens on the end of the run, plays at 0.5x to 4x, follows a live run's end, and under reduced motion steps from event to event instead of sweeping. FlowPlaybackBar is the replay's control row: play or pause, previous and next event, a named scrubber that says where it is in words, event ticks with failures in the error red and waits hatched, the time, the speed (a menu on a phone) and Live with Follow live. Fixtures of a failed Triage run and a branching run with a wait ship in @tale/ui/testing/flow.
A graph that changes on screen (same layoutKey) now glides to its new layout: what leaves shrinks out at its old place (150 ms), what stays moves on its node wrapper's transform (300 ms, out-quint), what joins grows in after 150 ms and new or re-routed lines fade in last at 250 ms, settled in 450 ms. The canvas draws the graph its layout was computed for until the next one lands, so nothing vanishes before it can leave. A live refit eases with it; the open node is brought back into view if it moved out. Nodes changed outside this tab (changed) ring once after their relayout. A new layoutKey, the first layout and reduced motion play none of it. overlay and playback put a run on the chart: each node framed and worded by its state (a running node's top bar sweeps, a failed one gets a red edge and its error line), conditions show how they decided, travelled lines take the emphasis colour, lines not taken step back, a frame counts items or passes, and a dot rides each line a value is on (a paused Web Animation set from the host's moment; none under reduced motion). A failed run brings the way to its first failure forward; strips say "The run stopped here" without an error line. paths, highlight and onHighlightChange bring parts of the chart forward: a pointer resting on a condition or a Yes/No label lights its paths, a node under the pointer or the keyboard its own lines, and a host highlight wins and is announced once. Outside a highlight boxes are dashed on a muted surface with their reason in the strip and lines thin to 1 px; colour and width never animate, two stacked lines crossfade. The List view says the same: state glyphs, run words first, quiet rows dashed with their reason.
Guides for the possible-paths list and run playback, the canvas guide's overlay and relayout sections, four live demos (static overlay, possible paths, playback with its bar, live relayout) and the matching stories.
Escape only unpinned from a path row. With focus on Show all, the list's Escape claim kept a sheet around it open and nothing unpinned, so the key did nothing. Escape now unpins from Show all as well and hands focus back to the row that was pinned. Unpinning, with Escape or Show all, also ends the preview, so the chart shows every path again. Before, focus landed back on the pinned row and previewed its path, so Show all went on showing that one path. The row keeps the focus without previewing; the next arrow key previews as usual.
A replayed step showed its recorded detail, such as its 1.2 s duration, while it was still running at the moment shown. A span's detail now shows once the replay reaches the span's end. A span still open on a live run shows its detail at once, because there the host's words describe the present. A wait's hatched band on the scrubber ran only to the next event, so work going on during the wait cut it short. buildPlaybackTimeline now gives a wait mark its end, and the bar draws the band to it. A wait still open has no end, and its band runs to the end of the timeline. Two values travelling one line at once, or two marks at the same moment, no longer share a React key, so each keeps its own dot or tick.
a && (b || c) and (a && b) || c both read "A and B or C" on the canvas, in every language. A group of the other kind inside a condition now reads in brackets, so the two read apart.
Two of the recorder's credential shapes, the JSON web token and the URL with a password, scanned a long run of token characters once for every place a match could start inside it: 128 KB of 'a-' or 'eyJ-' took about nine seconds per scan. They run on whatever a run receives, a webhook body included, on the run-start path and again for every condition and recorded value, long enough to outlast the run's lease. A web token is now only looked for where a run of token characters starts, and a URL scheme is at most 32 characters: the same text reads in milliseconds, and the shapes still find what they found.
A failure's callee was cut at 80 characters with a plain slice, which could keep half of an emoji; Postgres refuses a lone surrogate (and a NUL character) in jsonb, so the run's finishing write failed, the run stayed running, and every re-claim failed at the same write. A program's own output could do the same to the step records. One shared helper (lib/shared/utils/storable-text) now cuts on a character boundary and makes text storable; the record's failure parameters and message use it (the message is also capped at a run detail's 4,096 characters), the step rows are made storable before they are written, and the write itself runs in a savepoint, so a batch the database still refuses is logged and dropped and the run's progress commits. The record core's three private copies of the cut are gone.
A connector input that missed its schema was refused with a sentence that pasted the whole resolved input, so a key the step's record withholds came back in the failure every reader of the run sees. The sentence now names the action and what is wrong; the record shows the input, its secrets withheld.
…gone An erasure found the runs that replay the person's runs only through the replay's link to its source, and deleting the source (retention, a person, an earlier erasure) nulls that link: a colleague's replay kept the person's input, and the receipt said the erasure was complete. A replay now records the starters of its whole lineage when it starts (migration 0192), the erasure matches on them as well as on the link, and held runs are counted across the same lineage. ERASE-R10 says so, and the node-run lane proves the deleted-source case on Postgres.
A probed evaluation that node-vm stopped at its deadline (it kills and replaces its process) failed the step as a broken runner, though probing must never fail what the plain evaluation passes: such a stop now carries timedOut, and the unit is evaluated plainly on the fresh process, as one that timed out inside it already was. The failure's trace kept the evaluator's sentence whole, and that sentence may quote what the expression read — a key passed to JSON.parse, cut by the engine to its start. The trace's copy now leaves out every secret of the unit's scope and the longest start of one it quotes, is cut to a run detail's 4,096 characters, and is storable.
…ns and turns Every step a subautomation walked for an item was kept as a row of its own, so rows multiplied with the items that walked them (100 orders × 100 lines × 5 steps is about 60,000 rows, not the 1,000 a run keeps), and every read of the record loaded them all. And the cap counted from zero every turn. A step of a subautomation walk is now kept only while the run's rows last, like items and passes (the calling step's counts still say what happened), and each turn starts from the rows the run already stores.
…t view A condition's problems were counted on its own box, which the List view has no row for: no row showed them and no screen reader heard them, though the List view is the chart's text alternative and the default on narrow phones. The step a condition guards now carries its marker and says its problems in its description.
The canvas's problem counts and its select handler took a new identity on every keystroke in an inspector field, which rebuilt the graph's words and re-rendered every node view, though the layout itself waits for a pause. The counts are now read against the document the canvas draws, which settles at that pause, and the handler reads the typed document through a ref.
"Go to" a problem selected the check's raw offsets though the reader had typed since the check settled, so the next keystroke replaced the wrong characters. The range is now mapped from the checked text to the text shown, as the underlines are; a range inside what changed selects nothing.
A JSON field's problems were placed against the two-space JSON the check read; in a value the author had formatted another way, a problem inside the reformatted part lost its underline and its one-click fix. The field now places them in its own text whenever that text holds the value the check read, through a diagnosticsAt function the inspectors pass.
The undo history slot is read in the view's creation effect and in the value effect; it never changes, but the hook rules want every value an effect reads in its dependency list.
yannickmonney
enabled auto-merge
October 9, 2026 22:32
Failures in words, replays, playback, the Steps view with items, comparison, step data and the runs table — the run view squashed onto the run debugger.
Main and this branch each gave the dialog, the responsive dialog and the popover an onEscapeKeyDown — one keeps Escape inside an IME composition, the other lets a field that claims Escape (an editor's completion list) keep it — and the merge kept both props, so the later one silently replaced the first. The claim wrapper now carries the composition check: a claimed Escape and an Escape inside a composition both leave the layer open.
Nothing outside the module calls it any more, so it is no longer exported, as on main.
yannickmonney
force-pushed
the
feat/automation-run-debugger
branch
from
October 9, 2026 22:34
7dfa886 to
82b5c81
Compare
The run page's Run again and Edit input and run reported a refused replay through mutate's onError, which react-query drops when another call starts or the page unmounts, while the replay write itself stays quiet — so a refusal could pass without a word. Each call now awaits its own promise and reports from it, as the one-failure-one-toast guard requires.
The ValueTree marks demo built its map from mixed kinds of marks, which the type checker read as the first entry's kind alone; the map now names the mark type it holds. The End shape test declared a doc that shadowed the file's own document helper.
…tion The automation run view added Änderungen as the shipped German word for Changes, so the German changelog description now uses it instead of the English loanword.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Every automation run now keeps a step-by-step record of what it did, and every door reads it: the app, REST v1 and MCP. A record says why each step ran, was skipped or failed, with the values its conditions read. A run can also be run again, whole or from one of its steps. This is the backend and API half of the run debugger; the run view UI comes in follow-up PRs.
What changes
The record (migration
0190,AUTO-R38,AUTO-R39)__start), each step, each item of a step that runs per item, each pass of a step that repeats, and the output (__end).reasonfrom a fixed list, with itsparams.Why a condition held
explainConditionturns that into a tree of the expression's parts with their values; parts it cannot vouch for read unknown.{{ }}unit's text landed.Readings: one pure read model,
lib/engine/core/record/read.ts, shared by every host:unknown, neverchanged. The comparison names the first step where the runs diverged.Doors (
AUTO-R40: each reads the run as the run is read; a hidden run answers like a missing one)/runs/:runId/record(also?since=andinclude=travels),/record/node,/record/items,/compare/:otherRunId.GET …/runs/{runId}/record,…/record/node,…/record/items(the house keyset page, signed cursor),…/compare/{otherRunId}. The OpenAPI schemas are checked against answers built from a real run.get_runtakesinclude: ["record"|"travels"];get_run_nodeandcompare_runsare new read tools. The server instructions and the failed-run prompt point an agent atfailure.reasonand its explanation.NODE_RUN_NOT_FOUND(404) andRUN_COMPARE_MISMATCH(400).truncatednames what was left out.Run a run again (migration
0191,AUTO-R41,ERASE-R10)again(its own input),edited(a new input) orfroma step. A fork is born with the steps the run finished outside that step and what it feeds; they are marked reused, with their results and record copied and never their effects. It runs the rest.automation.run.replayed).writesAgain. Doors:GET/POST /runs/:runId/replay;…/replay, withIdempotency-Key;replay_run, withdryRun.@tale/sharedserves all three doors. SixREPLAY_*refusal codes.Contract 3.29.0. #4619 holds 3.28.0. Whichever lands second renumbers; this branch's entry already assumes it comes second.
Checklist
bun run checkscope run locally: platformtscclean; oxlint clean on every changed file; the whole server vitest project passes (18,906 tests). The one red after merging main is main's ownTASK_REVIEW_FORBIDDENregistry gap, which fix: restore shared CI contract checks #4625 fixes. knip: only main'sFINISH_PATHS/decideMergeGroup.bun run lint:sast— the staged-file Opengrep hook ran on every commit.docs/{en,de,fr}/develop/api-reference.md("Read a run step by step") anddocs/{en,de,fr}/develop/mcp-endpoint.md(the tools and the fix-a-failed-run loop). The manual reference rows and the error-code rows are updated.Test plan
bunx vitest run lib/engine/core/record backend/domains/automations scripts/openapi backend/rest lib/mcp backend/domains/mcpITEST_LANES=checkAutomationNodeRuns bun run backend:integration(86/86), including replays (a fork's reused steps and copied record, a double submit, refusals, the lineage surviving the source's deletion) and ERASE-R10. It covers:get_run {runId, include: ["record"]}on a failed run names the failed step'sfailure.reasonand explains the expression it failed on.Review fixes (adversarial review, 2026-10-09)
An independent review of the run record found one high, three medium and four low findings; seven are fixed here, each with a test:
lib/shared/utils/storable-text) cuts on character boundaries and makes text storable; the step rows are written storable and in a savepoint, so the record can never fail a run's progress.Left: record reads still load a run's rows whole (now bounded by the cap), and binding a replay's start to the confirmed plan is done by the run page sending the plan's numeric version (#4658).