A build spec you can hand to a coding agent (or follow yourself) to transcribe a whole library of audio — a local share, a NAS, or a cloud bucket — accurately, resumably, and with structured output you can actually build on.
It distills the design and the non-obvious lessons from a community transcription pipeline that processed ~2,300 hours of long-form talk audio. The goal isn't "run Whisper" — that's one line. It's everything around Whisper that makes a large batch finish cleanly, read accurately, and produce output a downstream tool can use.
Stack-agnostic. A reference implementation exists in Python (single file, MIT), but the ideas port directly to
whisper.cpp(great for a C++/FFmpeg codebase), faster-whisper, or any Whisper binding. Build it however fits.
my-recording.mp3
│
▼ (transcribe)
├── my-recording.txt plain transcript
├── my-recording.srt standard subtitles (segment-level)
└── my-recording.segments.json the useful one ▼
Three artifacts, one pass. The JSON is what makes this more than a subtitle generator: per-word timings let you cut clips or karaoke captions to the exact word, and per-segment confidence lets you later flag the handful of shaky segments for review instead of re-reading the whole thing.
┌─────────────────────── BATCH LOOP (model loaded once) ───────────────────────┐
│ │
folder/ │ for each media file: │
bucket ─┼─▶ already done? ──yes──▶ skip (resumable) │
(mount) │ │ no │
│ ▼ │
│ ┌─────────┐ ┌──────────────┐ ┌───────────────┐ ┌──────────────────┐ │
│ │ FFmpeg │──▶│ Whisper │──▶│ glossary + │──▶│ write txt / srt / │ │
│ │ decode │ │ (tuned flags) │ │ QC gate │ │ segments.json │ │
│ │ 16k mono│ │ word+conf │ │ (optional) │ │ │ │
│ └─────────┘ └──────────────┘ └───────────────┘ └──────────────────┘ │
│ │ on error: log + continue (one bad file can't kill the run) │
└───────────────────────────────────────────────────────────────────────────────┘
Engine. Whisper large-v3 for quality.
| Platform | Recommended backend | Notes |
|---|---|---|
| Apple Silicon | mlx-whisper |
~8× faster than CPU; the sweet spot on a Mac |
| NVIDIA GPU | faster-whisper (CTranslate2) or openai-whisper + CUDA | big speedup |
| CPU-only / C++ | whisper.cpp with a quantized model |
no Python, runs anywhere |
| CPU-only / Python | openai-whisper turbo |
fine default, slower |
Decode with FFmpeg. Whisper wants 16 kHz mono. Let the library handle it, or
normalize up front: ffmpeg -i in -ac 1 -ar 16000 out.wav. FFmpeg also means you
accept anything — mp3, wav, m4a, flac, opus, and video containers.
Four settings that matter for long-form audio:
condition_on_previous_text = False. With it on, each chunk is prompted with the previous chunk's text, so a single accidental repeated phrase self-reinforces into minutes of looping output. Off = no loops. This is the single most important flag for talks, sermons, podcasts, meetings.- Clamp the temperature fallback to ~
(0.0, 0.2, 0.4). The default ladder climbs to 1.0, where retries on noisy audio emit foreign-script and nonsense tokens rather than a usable transcript. Capping it keeps failures graceful. word_timestamps = True. Cheap, and the only way to get the per-word timing that makes clipping and precise captions possible.- Pin the language (e.g.
language="en") when you know it — prevents spurious mid-file language switches.
Write all three artifacts next to each other, named off the source file's stem (see the JSON shape above). Keep the sidecar data (words, confidence) index- aligned to the segments so anything downstream can re-join them without guessing. One self-describing JSON per file beats a pile of parallel files.
transcribe --batch /mnt/recordings --recursive --out ./transcripts
- Scan the directory (optionally recursive) for audio/video extensions.
- Load the model once, then loop. Model load is the slow part — never pay it per file.
- Resumability is non-negotiable at scale. Before transcribing, skip any file whose output already exists. A 500-file run will be interrupted — a reboot, a closed laptop, a Ctrl-C. Make stop-and-rerun free and idempotent.
- Isolate failures. Wrap each file in try/except and continue; one corrupt or
zero-length file shouldn't end a 12-hour run. Tally
done / skipped / failedat the end and log which files failed. - Handle SIGINT cleanly so Ctrl-C stops between files and a rerun resumes.
Skip per-provider SDKs. Mount the bucket as a filesystem and point the batch scanner at the mountpoint:
| Provider | Mount with |
|---|---|
| Amazon S3 | s3fs or rclone mount |
| Google Cloud Storage | gcsfuse or rclone mount |
| Azure Blob | blobfuse2 or rclone mount |
| Anything else | rclone mount |
gcsfuse my-church-recordings /mnt/recordings
transcribe --batch /mnt/recordings --recursive --out ./transcripts
This keeps the tool storage-agnostic and dependency-free; writes stream back to the bucket if your output dir is on the mount too. (If you prefer to pull objects directly, a thin "list + download to temp dir" adapter works — but mounting is simpler and handles streaming and retries for you.)
- Domain glossary. ASR reliably garbles domain-specific terms — product names, place names, people's names, jargon. A small find/replace pass against a glossary of your known terms (correct spelling + the common mis-hearings) dramatically improves perceived accuracy for almost no effort. Keep personal names in a private list if the output is published.
- QC gate instead of blind trust. Score each transcript with cheap heuristics
— words-per-minute in a sane band, subtitle coverage vs. audio duration, runs
of repeated characters, foreign-script ratio — and flag
suspect/failfor a human glance. This catches the ~2% that went wrong without reviewing the 98% that didn't. - Multi-engine cross-check (advanced). Run two different engines on the same file and compare; where they disagree is a likely error worth flagging, and a term-fidelity vote can pick the better word. Worth it only with diverse engines — the same engine run twice mostly agrees with itself, mistakes included. If you do this, decide the winner before overwriting anything.
A single machine chewing a folder overnight covers most needs. Add this only when volunteers genuinely want to pitch in — it's a real step up in complexity. The minimal shape that works:
┌──────────────┐ claim (lease) ┌─────────────────────┐
│ │ ◀───────────────────────────│ Volunteer worker A │
│ Job queue │ │ download→transcribe│
│ + tiny API │ ────────── submit ─────────▶│ →submit, loop │
│ │ └─────────────────────┘
│ status: │ claim (lease) ┌─────────────────────┐
│ available │ ◀───────────────────────────│ Volunteer worker B │
│ checked_out │ ────────── submit ─────────▶└─────────────────────┘
│ done | failed│
└──────────────┘ lease expires → back to "available" (nobody strands a file)
- Job queue: a DB table (even SQLite), one row per file, with
status (available | checked_out | done | failed),lease_expires_at, and anattemptscounter. - Two API endpoints:
claim(atomically hand out the oldestavailablejob- a short lease) and
submit(accept the result, run the QC gate, markdone). The server never touches the audio or the CPU — volunteers do the work locally and send back only the transcript.
- a short lease) and
- A tiny worker each volunteer runs:
claim → download → transcribe → submit, loop. Version-gate claims so you can hold back outdated installs when you ship a quality fix. - Leases + requeue: if a volunteer quits mid-file, its lease expires and the
job returns to
available. Capattemptsso a poison file fails out instead of looping forever. - Consent gate: only process recordings that are cleared to publish. Gate on an explicit "is this allowed" flag at claim time — not on a filename guess — so a private recording can never slip into a public transcript.
- Phase 1 — one-file transcribe with the tuned flags + the three artifacts.
- Batch — folder scan, load-once, resumable skip, failure isolation.
- Mount a bucket and point Phase 2's scanner at it — done for most use cases.
- Phase 2 (queue + workers) only when there's real demand for distribution.
Start small, get Phase 1 right, and most of the value is already yours.
MIT licensed — use it, fork it, build on it.
{ "title": "my-recording", "engine": "whisper-large-v3", "word_count": 8421, "segments": [ { "start": 12.34, "end": 15.90, "text": "welcome everybody to the show", "words": [ // per-word timing → clip/caption trimming { "w": "welcome", "start": 12.34, "end": 12.80 }, { "w": "everybody", "start": 12.81, "end": 13.40 } ], "avg_logprob": -0.21, // confidence → flag likely-wrong spans "no_speech_prob": 0.01, "compression_ratio": 1.34 } ] }