Problem
Visual snapshot testing over the emitted Storybook has to be invented by every customer, yet the CLI already knows every story id it emitted. The run hand-built the whole loop in ~70 lines (Playwright shoot of every *--default story via iframe.html?id=… against a static build, pixelmatch diff) and it worked remarkably well: 142/142 stories, 1m55s per full run, and zero pixel flake across three full runs (unchanged stories diffed at exactly 0 pixels — the emitted output + static build is fully deterministic; the only hygiene needed was an animation/transition freeze and caret hide).
Since that run, a fuller harness has been built and exercised against four library fixtures. It is not a sketch any more — it measures Figma variant renders against Storybook renders of emitted output, scores pass/fail per variant, ranks by dependency depth, and publishes a status page. This issue is about moving that into the product, not inventing it.
Evidence it works
Four fixtures, one run each, same day:
| Fixture |
Components |
Pairs |
Pass |
Fail |
Pass rate |
| A (large commercial library) |
69 |
2,513 |
513 |
33 |
94.0% |
| B (mid-size library) |
32 |
377 |
334 |
12 |
96.5% |
| C (mid-size library) |
29 |
420 |
323 |
0 |
100% |
| D (adversarial fixture set) |
133 |
421 |
72 |
59 |
55.0% |
Fixture D's low rate is the point — those fixtures are deliberately hostile, and a failure there is usually the fixture doing its job.
The hard parts are solved and written down, not waiting to be discovered:
- Width pinning. A render is pinned to the variant's authored width, because Figma constrains a FILL root by its parent frame and Storybook has no equivalent. Narrowing the pin to FILL roots only reads as the more principled rule and was measured to be worse.
- White-union flatten before diff. Figma exports carry alpha, Playwright shots are transparent; raw RGBA comparison flags every white-vs-transparent pixel. Both sides flatten onto an opaque-white union canvas first.
- Enum-casing normalization. The spec carries Figma's value casing; the emitters normalize it. Values are matched case-insensitively against the emitted contract's string-literal union and replaced with the contract's spelling, so a manifest stays correct across output from different emitter versions.
- Interaction classification. A prop declared in the spec but absent from the emitted contract does not exist at render time (hover/pressed are CSS pseudo-classes). Variants driving such a prop are marked deferred — not shot, not failed.
- The noise valve. A per-library ignore file carries
passPct, dimTolerancePx, threshold, skip, skipVariants, sampleVariants, pinWidth. Every entry requires a note: saying why, so a permanent measurement fact stays distinguishable from a temporary allowance.
Two baselines, not one
The original framing here and the harness as built are different products. Both diff the same shots against a different left-hand image, so they share everything except where the baseline comes from.
| Mode |
Baseline |
Answers |
Needs |
| Fidelity |
the Figma export |
"does the code match the design?" |
FIGMA_TOKEN, the variant↔story join, the manifest |
| Regression |
the last accepted render |
"did my change break anything?" |
nothing but a running Storybook |
Fidelity is the differentiated thing and carries all the complexity. Regression is cheap, offline, and is what the zero-flake evidence above actually measured.
One consequence worth stating: in fidelity mode there is no "accept the new render." Percy and Chromatic are built around approving a changed snapshot; here the baseline is the design, and a render that disagrees is wrong by definition. accept exists only in regression mode, and is named so nobody reaches for it expecting the other thing.
On-disk layout
<workspace>/testing/visual/
figma/ <kind>/<key>/<nodeId>.png design truth, captured overtly
accepted/ <kind>/<key>/<nodeId>.png accepted renders (regression mode)
render/ <kind>/<key>/<nodeId>.png this run
diff/ triptychs and masks
report/ visual-report.{json,md}
manifest.json
visual-ignore.yaml
Not under storybook/. That directory is generated wholesale by specs storybook and rewritten by init --force — a hand-tuned ignore file kept there is a file waiting to be overwritten.
Commands
specs testing visual, with the testing namespace left open for the other suites that already exist as scripts (parity, emitted-tree typecheck, render round-trip, schema compliance).
The subcommand split is not cosmetic. Each stage has a different cost, and the point of splitting is to pay only the one you need — editing a tolerance and re-scoring must cost seconds, not minutes:
| Subcommand |
Cost |
Network |
Browser |
Flags |
init |
instant |
— |
— |
--force |
status |
instant |
— |
— |
--components, --json |
manifest |
seconds |
— |
— |
--components, --check |
baseline |
slow |
Figma REST |
— |
--components, --force |
shoot |
minutes |
— |
Playwright |
--components, --workers N, --port, --target react|webcomponents |
diff |
seconds |
— |
— |
--components, --against figma|accepted |
report |
instant |
— |
— |
--format json|md |
accept |
instant |
— |
— |
--components, --all |
| (bare) |
minutes |
— |
yes |
shoot → diff → report |
Universal: --config <path>, --kinds components,compositions, --verbose.
baseline is never implicit. Not in the bare command, not on a watch, not on a schedule. It is slow, it is rate-limited, and it is the one stage that costs someone else's API quota. The bare command prints how many shootable variants have no baseline and stops there.
init scaffolds, the customer installs. The same contract as specs storybook init: it writes testing/visual/package.json declaring Playwright, pixelmatch and pngjs, then prints the one install command to run. We never ship, vendor, or install them — shoot resolves Playwright from that directory at runtime and fails with a pointer to init when it is absent. Everything except shoot works with nothing installed.
manifest --check is a dry run reporting unmapped Figma props. Drive it to zero before trusting a library's numbers — an unmapped prop defers every variant that uses it.
diff is the only stage that reads visual-ignore.yaml. sampleVariants and pinWidth are manifest-time keys and need a rebuild plus a re-shoot; everything else re-scores in seconds with no browser.
- Scoring stays in the ignore file, not in flags. A tolerance is a durable fact about a library that needs a
note: explaining it. A flag is something passed once to make red go away, and nobody can tell afterwards whether it hid a bug.
Staleness is provenance, not scheduling
An earlier draft keyed baseline staleness off the Figma file's lastModified. That is the wrong signal: the timestamp moves on any edit anywhere in the file, so a run recaptures the whole catalogue to discover nothing changed.
With capture overt and nothing automatic, there is no trigger to compute, so v1 computes none. What it does instead:
status reports present / missing per shootable variant. Missing is knowable with no model at all.
- The report records each baseline's
capturedAt beside the payload's own fetch time, so a diff cannot silently blame the emitter for a baseline that predates the design.
Making baseline self-scoping — so an overt run recaptures the three components that changed rather than all eighty — is a later optimization, tracked separately. The cheap signal for it is a hash of the component's Figma node subtree out of the already-fetched payload: offline, free, exact per component, and unmoved by edits elsewhere in the file. Deliberately not the spec version from the ledger: the spec is lossy by design, so that signal goes blind exactly where visual testing earns its keep.
Not in build or run
Out of scope, and not deferred-pending-design — ruled out for v1 on the merits.
Every candidate trigger fires far more often than the thing it is meant to detect:
| Trigger |
Why it fails |
| Figma file changed |
moves constantly; few or no components actually change |
| Spec regenerated |
a spec-first authoring session rewrites specs every few seconds |
| Emitted code changed |
every save, and a shoot is minutes with a browser |
Until there is a signal that identifies what changed at component granularity, any automatic trigger is a browser launch tax on an unrelated edit. There is also a mechanical mismatch: steps.ts steps are in-process function calls over files, and shooting needs an external HTTP server listening on a port — unlike every existing step.
A component-scoped recapture driven by a narrow watch is imaginable later. It is not this issue.
Why this cannot stay outside the CLI
The harness discovers specs with a flat directory read, keeping any child that holds an api.yaml. Under the current layout (ADR-096) the children of specs/ are components/ and compositions/, neither of which holds an api.yaml — so discovery returns nothing. The harness still runs against exactly one fixture, and only because that fixture is the last one in the pre-ADR-096 flat layout.
Contract resolution has the matching problem: it hard-codes react/src/components/<Pascal>, which no composition will ever be under.
Both are one fix — discovery and contract resolution go through the layout seam — and that seam exists only in the CLI, deliberately: it is documented as the only thing that knows the layout, after two divergent filters preceded it and a third kind had nowhere to go. Repairing the harness in place means reimplementing it, which is what the seam exists to prevent.
Compositions
Visual testing targets both kinds. A composition's spec settles what that means: no props, a plain FRAME source node, and exactly one emitted story — the emitter writes the reason into the file ("A composition has no variant axes — no sticker sheet").
| Concern |
Component |
Composition |
| Figma node |
COMPONENT_SET → enumerate variant children |
single FRAME → one node |
| Variant name parsing |
Prop=Value, …, can fail |
none |
| Prop mapping |
the main source of deferred pairs |
nothing to map |
| Variant↔story join |
exact / overlay / no-story heuristic |
exact by construction |
| Pairs per spec |
1–150 |
exactly 1 |
The manifest side is nearly free — the builder already takes a non-COMPONENT_SET branch yielding one variant with an empty config. The cost is downstream:
- Baseline keying becomes kind-aware. A component and a composition sharing a name is legal and is handled that way in the version assembler.
figma/<key>/ therefore collides; it needs figma/<kind>/<key>/, and the same split in the render and diff trees, the report, and the ignore file.
- Crop rule — decided. Compositions are frames placed directly on a page node, so every one of them is fixed-width. Pin width unconditionally, and pin height only when the frame's vertical sizing is FIXED; a hugging frame keeps its natural height. This extends the existing width pin rather than introducing a second rule, and turns vertical overflow into visible pixels instead of a dimension mismatch. The height half is the only part that needs measuring, and it can be measured when the first hugging composition appears rather than up front.
- Scoring is advisory — decided. Compositions are captured, diffed, and reported with pixel counts, but their pass/fail does not decide the run's exit status. A composition's diff is dominated by its children, so one leaf defect lights up every composition containing it; the existing depth-ascending ranking already says to fix leaves first, and this makes the report agree with it.
- Pro gating. Composition output is Pro (ADR-097). Composition visual testing filters on a free run the same way
transform does, reporting the count once.
- Images and fonts. Canonical page compositions are where image sources and missing licensed fonts concentrate. The noise valve needs the kind dimension from point 1 before it can absorb that separately from component noise.
What ports, and where it lands
~3,000 lines across 15 files in the existing harness.
| Harness module |
LOC |
Destination |
| workspace resolution |
128 |
Delete. storybook/workspace.ts already resolves root, specs, data sources, emitted trees, scaffold state, and the dev-server port |
| spec index |
114 |
Port, shrinks — the CLI knows its own layout |
| variant↔story join |
270 |
Port as-is. The hardest and most valuable code |
| PNG utilities |
144 |
Port as-is |
| manifest builder |
450 |
Port, shrinks — reads payloads through SectionedFile instead of hand-walking JSON |
| baseline status |
118 |
Port, shrinks — no staleness model |
| Figma baseline capture |
118 |
Port as-is |
| shoot |
276 |
Port as-is, plus the composition pin rule |
| diff + report |
435 |
Port as-is, plus kind-aware keying and advisory scoring |
| orchestrator |
148 |
Becomes the bare command |
| status page |
298 |
Becomes a visualtesting Storybook concern |
The concern registry (storybook/concerns/registry.ts) is exactly what the report page needs: detect() = does a report directory exist, build() = emit the page. The three hand edits to .storybook/main.ts that currently mount it disappear — and have to, because init --force now owns that file. The scaffolded host already reserves the mount point, serving a visual-testing baselines directory at /baselines when one exists; that line is currently the only mention of visual testing anywhere in the CLI.
No new CLI dependencies. Playwright, pixelmatch and pngjs are declared in the scaffold init writes and installed by the customer in their own testing/visual/ directory — the same contract as the Storybook host, where we never ship or vendor Storybook (ADR A). A customer who never does visual testing never downloads a browser, and the published CLI stays the nine small dependencies it is today.
Build order
In sequence. Each step is independently verifiable. Fixture A is the muse throughout steps 1–7; the others widen coverage in step 8, once the suite is complete.
| # |
Step |
Validated by |
Size |
| #711 |
Layout-seam discovery and kind-aware baseline keying |
both kinds enumerate where zero do today |
m |
| #712 |
Command surface, manifest, status |
unmapped props at zero; no browser, no network |
l |
| #713 |
Overt Figma baseline capture |
scoped capture writes only what was named |
m |
| #714 |
Shoot, variant↔story join, pin rules |
three full shoots diff render-against-render at 0 pixels |
l |
| #715 |
Diff, scoring, ignore resolution, report |
a tolerance change re-scores in seconds with no browser |
l |
| #716 |
Report page as a Storybook concern |
survives storybook init --force |
m |
| #717 |
Regression mode against accepted renders |
runs end to end with no Figma token |
m |
| #718 |
Broaden to the remaining fixtures, and document |
all three complete a full cycle |
l |
Follow-on, not required to ship:
Steps 1–5 are a straight line. #716 and #717 both depend only on #715 and are independent of each other.
Acceptance criteria
Non-goals
- Any automatic trigger, and any integration with
build or run.
- A staleness model.
status reports present/missing only.
- Drop-shadow bounds. Figma exports include them; a tight render crop does not. Needs a shadow-aware rule before elevated variants are scoreable.
- Hover and pressed states. Driving them with Playwright is deliberately out of scope.
- Figma variable modes — everything assumes mode defaults.
- Fixing what the diff finds. This measures; the fixing workflow lives elsewhere.
Case data
- Workspace: atlassian — Atlassian Design System (ADS) community files; sealed first-customer onboarding simulation, 2026-09-27
- Territory: cli
- Size: l
Notes
Priority: P3. Related: #489, #490 (VT content policies), #405, #406 (existing VT concerns). Determinism evidence above is the strongest argument this belongs in the product: the hard part (stability) is already solved by the output's nature.
Implementation details are tracked internally.
Problem
Visual snapshot testing over the emitted Storybook has to be invented by every customer, yet the CLI already knows every story id it emitted. The run hand-built the whole loop in ~70 lines (Playwright shoot of every
*--defaultstory viaiframe.html?id=…against a static build, pixelmatch diff) and it worked remarkably well: 142/142 stories, 1m55s per full run, and zero pixel flake across three full runs (unchanged stories diffed at exactly 0 pixels — the emitted output + static build is fully deterministic; the only hygiene needed was an animation/transition freeze and caret hide).Since that run, a fuller harness has been built and exercised against four library fixtures. It is not a sketch any more — it measures Figma variant renders against Storybook renders of emitted output, scores pass/fail per variant, ranks by dependency depth, and publishes a status page. This issue is about moving that into the product, not inventing it.
Evidence it works
Four fixtures, one run each, same day:
Fixture D's low rate is the point — those fixtures are deliberately hostile, and a failure there is usually the fixture doing its job.
The hard parts are solved and written down, not waiting to be discovered:
passPct,dimTolerancePx,threshold,skip,skipVariants,sampleVariants,pinWidth. Every entry requires anote:saying why, so a permanent measurement fact stays distinguishable from a temporary allowance.Two baselines, not one
The original framing here and the harness as built are different products. Both diff the same shots against a different left-hand image, so they share everything except where the baseline comes from.
FIGMA_TOKEN, the variant↔story join, the manifestFidelity is the differentiated thing and carries all the complexity. Regression is cheap, offline, and is what the zero-flake evidence above actually measured.
One consequence worth stating: in fidelity mode there is no "accept the new render." Percy and Chromatic are built around approving a changed snapshot; here the baseline is the design, and a render that disagrees is wrong by definition.
acceptexists only in regression mode, and is named so nobody reaches for it expecting the other thing.On-disk layout
Not under
storybook/. That directory is generated wholesale byspecs storybookand rewritten byinit --force— a hand-tuned ignore file kept there is a file waiting to be overwritten.Commands
specs testing visual, with thetestingnamespace left open for the other suites that already exist as scripts (parity, emitted-tree typecheck, render round-trip, schema compliance).The subcommand split is not cosmetic. Each stage has a different cost, and the point of splitting is to pay only the one you need — editing a tolerance and re-scoring must cost seconds, not minutes:
init--forcestatus--components,--jsonmanifest--components,--checkbaseline--components,--forceshoot--components,--workers N,--port,--target react|webcomponentsdiff--components,--against figma|acceptedreport--format json|mdaccept--components,--allUniversal:
--config <path>,--kinds components,compositions,--verbose.baselineis never implicit. Not in the bare command, not on a watch, not on a schedule. It is slow, it is rate-limited, and it is the one stage that costs someone else's API quota. The bare command prints how many shootable variants have no baseline and stops there.initscaffolds, the customer installs. The same contract asspecs storybook init: it writestesting/visual/package.jsondeclaring Playwright, pixelmatch and pngjs, then prints the one install command to run. We never ship, vendor, or install them —shootresolves Playwright from that directory at runtime and fails with a pointer toinitwhen it is absent. Everything exceptshootworks with nothing installed.manifest --checkis a dry run reporting unmapped Figma props. Drive it to zero before trusting a library's numbers — an unmapped prop defers every variant that uses it.diffis the only stage that readsvisual-ignore.yaml.sampleVariantsandpinWidthare manifest-time keys and need a rebuild plus a re-shoot; everything else re-scores in seconds with no browser.note:explaining it. A flag is something passed once to make red go away, and nobody can tell afterwards whether it hid a bug.Staleness is provenance, not scheduling
An earlier draft keyed baseline staleness off the Figma file's
lastModified. That is the wrong signal: the timestamp moves on any edit anywhere in the file, so a run recaptures the whole catalogue to discover nothing changed.With capture overt and nothing automatic, there is no trigger to compute, so v1 computes none. What it does instead:
statusreports present / missing per shootable variant. Missing is knowable with no model at all.capturedAtbeside the payload's own fetch time, so a diff cannot silently blame the emitter for a baseline that predates the design.Making
baselineself-scoping — so an overt run recaptures the three components that changed rather than all eighty — is a later optimization, tracked separately. The cheap signal for it is a hash of the component's Figma node subtree out of the already-fetched payload: offline, free, exact per component, and unmoved by edits elsewhere in the file. Deliberately not the spec version from the ledger: the spec is lossy by design, so that signal goes blind exactly where visual testing earns its keep.Not in
buildorrunOut of scope, and not deferred-pending-design — ruled out for v1 on the merits.
Every candidate trigger fires far more often than the thing it is meant to detect:
Until there is a signal that identifies what changed at component granularity, any automatic trigger is a browser launch tax on an unrelated edit. There is also a mechanical mismatch:
steps.tssteps are in-process function calls over files, and shooting needs an external HTTP server listening on a port — unlike every existing step.A component-scoped recapture driven by a narrow watch is imaginable later. It is not this issue.
Why this cannot stay outside the CLI
The harness discovers specs with a flat directory read, keeping any child that holds an
api.yaml. Under the current layout (ADR-096) the children ofspecs/arecomponents/andcompositions/, neither of which holds anapi.yaml— so discovery returns nothing. The harness still runs against exactly one fixture, and only because that fixture is the last one in the pre-ADR-096 flat layout.Contract resolution has the matching problem: it hard-codes
react/src/components/<Pascal>, which no composition will ever be under.Both are one fix — discovery and contract resolution go through the layout seam — and that seam exists only in the CLI, deliberately: it is documented as the only thing that knows the layout, after two divergent filters preceded it and a third kind had nowhere to go. Repairing the harness in place means reimplementing it, which is what the seam exists to prevent.
Compositions
Visual testing targets both kinds. A composition's spec settles what that means: no props, a plain
FRAMEsource node, and exactly one emitted story — the emitter writes the reason into the file ("A composition has no variant axes — no sticker sheet").COMPONENT_SET→ enumerate variant childrenFRAME→ one nodeProp=Value, …, can failno-storyheuristicThe manifest side is nearly free — the builder already takes a non-
COMPONENT_SETbranch yielding one variant with an empty config. The cost is downstream:figma/<key>/therefore collides; it needsfigma/<kind>/<key>/, and the same split in the render and diff trees, the report, and the ignore file.transformdoes, reporting the count once.What ports, and where it lands
~3,000 lines across 15 files in the existing harness.
storybook/workspace.tsalready resolves root, specs, data sources, emitted trees, scaffold state, and the dev-server portSectionedFileinstead of hand-walking JSONvisualtestingStorybook concernThe concern registry (
storybook/concerns/registry.ts) is exactly what the report page needs:detect()= does a report directory exist,build()= emit the page. The three hand edits to.storybook/main.tsthat currently mount it disappear — and have to, becauseinit --forcenow owns that file. The scaffolded host already reserves the mount point, serving a visual-testing baselines directory at/baselineswhen one exists; that line is currently the only mention of visual testing anywhere in the CLI.No new CLI dependencies. Playwright, pixelmatch and pngjs are declared in the scaffold
initwrites and installed by the customer in their owntesting/visual/directory — the same contract as the Storybook host, where we never ship or vendor Storybook (ADR A). A customer who never does visual testing never downloads a browser, and the published CLI stays the nine small dependencies it is today.Build order
In sequence. Each step is independently verifiable. Fixture A is the muse throughout steps 1–7; the others widen coverage in step 8, once the suite is complete.
manifest,statusstorybook init --forceFollow-on, not required to ship:
baselineself-scoping via node-subtree hashing, so an overt capture takes the three components that changed rather than all eighty.Steps 1–5 are a straight line. #716 and #717 both depend only on #715 and are independent of each other.
Acceptance criteria
Non-goals
buildorrun.statusreports present/missing only.Case data
Notes
Priority: P3. Related: #489, #490 (VT content policies), #405, #406 (existing VT concerns). Determinism evidence above is the strongest argument this belongs in the product: the hard part (stability) is already solved by the output's nature.