Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
47 changes: 45 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -35,6 +35,11 @@ patchstory render ./pr-walkthrough.json --out ./site
`fetch` required). Zip-friendly.
- **AI optional.** A heuristic generator always works with no API key. AI
*improves* the story; it is never required.
- **Narrated play mode.** Press **▶ Play** to turn the walkthrough into a
self-playing screencast: each chapter becomes a scene that pans the actual
diff and spotlights the lines it references, narrated aloud via the browser's
built-in speech synthesis (captions included). No ffmpeg, no API key, no
network — the same single `.html`, just playing itself.

---

Expand Down Expand Up @@ -172,6 +177,7 @@ Commands
file <path.diff> raw unified diff file
github <pr-url> GitHub PR (uses `gh` if available, else public .diff)
render <walkthrough> render an existing pr-walkthrough.json
video <walkthrough> render a narrated .mp4 screencast of the walkthrough
serve [dir|file] serve an output folder/file on your LAN
schema print the pr-walkthrough.json JSON Schema

Expand All @@ -190,8 +196,13 @@ Options
--serve serve the result on your LAN after generating
--open open the result in a browser
--port <n> port for --serve / serve (default 8137)
--diff <file> (render only) raw diff to fill the diff explorer
--diff <file> (render/video) raw diff to fill the diff explorer
--zip also write <out>.zip
--tts <engine> (video) auto | elevenlabs | espeak-ng | flite | say | none
--voice <id> (video) voice id (elevenlabs) or name (espeak-ng/say)
--chrome <path> (video) Chrome/Chromium used to rasterize scenes
--fps <n> (video) frames per second (default 30)
--keep (video) keep the intermediate working dir
-h, --help show help
--version show version
```
Expand All @@ -207,6 +218,36 @@ produces output.
others can open. `--redact` masks secrets (token shapes, `KEY=value`, private
keys) in the diff before it's embedded *or* sent to an AI generator.

### Narrated video (opt-in MP4)

`patchstory video <walkthrough.json> --diff <pr.diff> -o walkthrough.mp4` renders the
walkthrough into a real, shareable `.mp4`: a title card, one **animated** scene per
chapter (the diff reveals line-by-line and the referenced lines light up as they're
narrated), and an outro. Unlike everything else here, this shells out to **system
tools** — it adds no npm runtime deps, and they're only touched when you ask for a video.

Two engines (`--engine`):

- **`hyperframes`** (default) — generates a [HyperFrames](https://hyperframes.heygen.com)
composition (HTML + GSAP) and renders it frame-by-frame in headless Chrome via
`npx hyperframes`. This is the animated one. Needs network for `npx` on first use.
- **`pan`** — a fully local fallback: rasterizes each scene with Chromium and pans it
with **ffmpeg**. No `npx`/network; lower production value.

**Text-to-speech** (`--tts`, default `auto`): `elevenlabs` (`ELEVENLABS_API_KEY`, best
quality), `kokoro` (local neural TTS via HyperFrames — no key, the keyless default),
local `espeak-ng` / `flite` / macOS `say`, or `none` (silent; captions still shown).

**ffmpeg/ffprobe** are resolved from `PATH`, then `/usr/bin`, then
`PATCHSTORY_FFMPEG` / `PATCHSTORY_FFPROBE` — each validated by actually running it, so a
broken or shadowing PATH entry is skipped (and the working one is handed to HyperFrames).

Narration is the audio track — nothing is burned into or captioned over the frame, so the
code and motion graphics stay unobstructed.

It's slower and heavier than the HTML — the in-page play mode is the local-first default;
the MP4 is for when you need a file to drop in Slack or a release thread.

### Interactive UI

Syntax-highlighted diffs (highlight.js, bundled at build time — lazily applied
Expand All @@ -217,7 +258,9 @@ dark, copy-summary, a "Start here" guide and recurring-theme detection on the
overview, related commits per chapter, and a footer build stamp.

Keyboard: `j`/`k` next/prev chapter · `/` search · `e`/`c` expand/collapse all ·
`r` toggle reviewed · `t` theme · `?` shortcuts · `Esc` close.
`r` toggle reviewed · `p` play narrated walkthrough · `t` theme · `?` shortcuts ·
`Esc` close. In play mode: `space` play/pause · `←`/`→` prev/next scene · `m`
mute (captions only) · `Esc` close.

---

Expand Down
20 changes: 18 additions & 2 deletions integrations/claude-code/patchstory/skills/patchstory/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,9 @@ description: >
open PR if there is one, otherwise the branch vs its default base — or takes an explicit
PR number / PR URL / git range. The agent authors the narrative itself (chapters with
intent, risk, reviewer questions, and verification steps), then patchstory renders it as
one self-contained .html with secrets redacted and opens it in the browser.
one self-contained .html with secrets redacted and opens it in the browser. The rendered
page also has a narrated "play" mode — a self-playing screencast that pans the diff while
reading each chapter aloud — so author a short spoken `narration` per chapter.
Triggers: "/patchstory", "patchstory this", "make a walkthrough of this PR",
"tell the story of this PR", "PR walkthrough", "explain this PR for a human",
"patchstory #123", "patchstory the current branch".
Expand Down Expand Up @@ -89,6 +91,10 @@ who has never seen the change:
one chapter when they tell one sub-story. Each chapter:
- **`intent`** — *why* this exists / what problem it solves (the most valuable field).
- **`summary`** — what the diff in this chapter does.
- **`narration`** — 1–4 sentences of *spoken* prose for the page's "play" mode: plain and
conversational, what you'd say out loud while walking someone through this chapter. Avoid
symbols/paths that sound bad read aloud. Optional, but author it — without it, play mode
falls back to reading `intent` then `summary`.
- **`risk_level`** — `low|medium|high`. Raise for auth, payments, migrations, money math,
deletions, or anything externally observable.
- **`review_notes`** — sharp reviewer questions.
Expand Down Expand Up @@ -120,7 +126,8 @@ Authoritative copy: `patchstory schema`. Required: `version`, `title`, `summary`
(+ `source.type` ∈ `github_pr|git_diff|commit_range|diff_file`), `stats` (`files_changed`,
`additions`, `deletions` — numbers), `chapters`. Each chapter needs a **unique** `id`, `title`,
`summary`, `risk_level` (`low|medium|high`), and `files`. `diff_hunks` items need `file`,
`start_line`, `end_line` (line numbers in the **new** file). Everything else is optional.
`start_line`, `end_line` (line numbers in the **new** file). Everything else is optional,
including the chapter's `narration` (spoken script for "play" mode).

```jsonc
{
Expand All @@ -139,6 +146,7 @@ Authoritative copy: `patchstory schema`. Required: `version`, `title`, `summary`
"title": "Detect multiple faces in uploaded media",
"summary": "Adds metadata and detection logic for multi-face media.",
"intent": "Determine whether creator approval is needed before publishing.",
"narration": "When media is uploaded, we now count the faces in it. If there's more than one person, the upload can't auto-publish — it routes to the creator for approval first. This chapter adds the detection logic and the fields that track that state.",
"risk_level": "medium",
"files": ["app/models/media.rb", "app/services/face_detection_service.rb"],
"diff_hunks": [
Expand Down Expand Up @@ -166,5 +174,13 @@ Authoritative copy: `patchstory schema`. Required: `version`, `title`, `summary`
"$WORK/site"`, drop `--single-file`) and then `patchstory serve "$WORK/site"` in the
background — it binds `0.0.0.0` and prints a URL other devices can open. The `--port` is a
starting hint; if taken, serve picks the next free port and prints the real one.
- **Want a shareable video instead of HTML?** `patchstory video "$WORK/pr-walkthrough.json"
--diff "$WORK/pr.diff" --redact -o "$WORK/walkthrough.mp4"` renders a narrated, **animated**
`.mp4` from the same JSON (a title card + one scene per chapter where the diff reveals and
the referenced lines light up as they're narrated — this is why authoring `narration` is
worth it). Default engine `hyperframes` (animated, via `npx hyperframes`; needs network);
`--engine pan` is a local ffmpeg fallback. TTS `--tts auto` picks ElevenLabs if
`ELEVENLABS_API_KEY` is set, else local `kokoro` (no key). Slower/heavier than the HTML —
only reach for it when a video file is the deliverable.
- **Private PRs** need `gh` (authenticated). The public `.diff` fallback is public-repos-only.
- The work dir under `~/.cache/patchstory/` persists; old runs can be deleted freely.
2 changes: 1 addition & 1 deletion packages/cli/src/args.ts
Original file line number Diff line number Diff line change
Expand Up @@ -9,7 +9,7 @@ export interface ParsedArgs {
/** Flags that take no value (presence = true). */
const BOOLEAN_FLAGS = new Set([
"zip", "help", "version", "no-open", "single-file", "serve", "open", "redact",
"scaffold",
"scaffold", "keep",
]);

export function parseArgs(argv: string[]): ParsedArgs {
Expand Down
44 changes: 42 additions & 2 deletions packages/cli/src/cli.ts
Original file line number Diff line number Diff line change
Expand Up @@ -23,7 +23,8 @@ import type {
ResolvedSource,
WalkthroughBundle,
} from "@patchstory/core";
import { renderWalkthrough, renderSingleFile } from "@patchstory/renderer";
import { renderWalkthrough, renderSingleFile, renderVideo } from "@patchstory/renderer";
import type { TtsProvider } from "@patchstory/renderer";
import { parseArgs, flagStr } from "./args.ts";
import { zipDirectory } from "./zip.ts";
import { serveStatic, findFreePort, openBrowser, lanIp } from "./serve.ts";
Expand All @@ -45,6 +46,7 @@ Commands:
file <path.diff> Walkthrough of a raw unified diff file
github <pr-url> Walkthrough of a GitHub pull request
render <walkthrough> Render an existing pr-walkthrough.json
video <walkthrough> Render a narrated .mp4 screencast of the walkthrough
serve [dir|file] Serve an output folder/file on your LAN
schema Print the pr-walkthrough.json JSON Schema

Expand All @@ -64,8 +66,14 @@ Options:
--serve Serve the result on your LAN after generating
--open Open the result in a browser
--port <n> Port for --serve / serve (default: 8137)
--diff <file> (render only) raw diff to populate the diff explorer
--diff <file> (render/video) raw diff to populate the diff explorer
--zip Also write <out>.zip
--engine <name> (video) hyperframes (animated, default) | pan (static)
--tts <engine> (video) auto | elevenlabs | kokoro | espeak-ng | flite | say | none
--voice <id> (video) voice id (elevenlabs) or name (kokoro/espeak-ng/say)
--chrome <path> (video) Chrome/Chromium binary for the pan engine
--fps <n> (video) frames per second (default: 30)
--keep (video) keep the intermediate working dir
-h, --help Show this help
--version Show version

Expand Down Expand Up @@ -136,6 +144,38 @@ async function main() {
return;
}

// `video` renders a narrated .mp4 from an existing walkthrough JSON (+ diff).
if (command === "video") {
const redactV = !!flags.redact;
const result = await renderCommand(positionals, flags, redactV);
const bundle: WalkthroughBundle = { walkthrough: result.walkthrough, diff: result.diff };
const outFlag = flagStr(flags, "out") ?? "./walkthrough.mp4";
const outFile = /\.mp4$/i.test(outFlag) ? outFlag : `${outFlag}.mp4`;
process.stdout.write(`\nRendering video → ${resolve(outFile)}\n`);
try {
const res = await renderVideo(bundle, {
out: outFile,
engine: (flagStr(flags, "engine") as "hyperframes" | "pan" | undefined),
tts: flagStr(flags, "tts") as TtsProvider | undefined,
voice: flagStr(flags, "voice"),
chrome: flagStr(flags, "chrome"),
fps: flags.fps ? Number(flagStr(flags, "fps")) : undefined,
keep: !!flags.keep,
onProgress: (m) => process.stdout.write(` ${m}\n`),
});
process.stdout.write(
`\n✓ Video written to ${res.file}\n` +
` ${res.sceneCount} scenes · ~${Math.round(res.durationSec)}s · tts: ${res.ttsProvider}\n`,
);
if (redactV) {
process.stdout.write("🛈 Redaction on: secrets masked in the diff shown on-screen.\n");
}
} catch (err) {
fail(err instanceof Error ? err.message : String(err));
}
return;
}

const generator = flagStr(flags, "generator") ?? "none";
const model = flagStr(flags, "model");
const redact = !!flags.redact;
Expand Down
1 change: 1 addition & 0 deletions packages/core/src/schema.ts
Original file line number Diff line number Diff line change
Expand Up @@ -56,6 +56,7 @@ export const WALKTHROUGH_JSON_SCHEMA = {
title: { type: "string" },
summary: { type: "string" },
intent: { type: "string" },
narration: { type: "string" },
risk_level: { type: "string", enum: ["low", "medium", "high"] },
files: { type: "array", items: { type: "string" } },
diff_hunks: {
Expand Down
8 changes: 8 additions & 0 deletions packages/core/src/types.ts
Original file line number Diff line number Diff line change
Expand Up @@ -54,6 +54,14 @@ export interface Chapter {
summary: string;
/** Why this part of the change exists. */
intent?: string;
/**
* Spoken narration for this chapter, used by the in-page "play" mode (a
* narrated, auto-advancing screencast). One to four sentences of plain,
* conversational prose — what a reviewer would say out loud while walking
* someone through this change. Optional: when absent, play mode falls back
* to `intent` then `summary`.
*/
narration?: string;
risk_level: RiskLevel;
files: string[];
diff_hunks: DiffHunkRef[];
Expand Down
4 changes: 4 additions & 0 deletions packages/renderer/src/index.ts
Original file line number Diff line number Diff line change
Expand Up @@ -16,6 +16,10 @@ import { dirname, join } from "node:path";
import type { WalkthroughBundle } from "@patchstory/core";
import { WEB_CSS, WEB_HTML, WEB_JS } from "./assets.generated.ts";

// Opt-in MP4 export (drives system ffmpeg + headless Chromium + a TTS engine).
export { renderVideo } from "./video/index.ts";
export type { VideoOptions, VideoResult, TtsProvider } from "./video/index.ts";

export interface RenderOptions {
/** ISO timestamp stamped into the document. */
generatedAt?: string;
Expand Down
Loading