Conversation
Released gateways pin the digest of each whisper runtime archive, and a rebuilt archive differs byte for byte, so republishing a release must leave the archives it holds alone. The release job still creates a missing release whole, not marked latest. For an existing release it uploads only the archives the release lacks, then replaces the checksum list, keeping its old lines verbatim and appending the new ones. It refuses a release whose archives and checksum lines disagree, the state a failed upload leaves, and a run with nothing new uploads nothing. The publish logic reads only its token, the repository name, the whisper tag, and the archive directory, so it runs unchanged outside CI. - `softprops/action-gh-release`: The SHA-pinned release action and the separate `Assemble SHA256SUMS` step give way to one inline bash script that drives `gh`, whose version the workflow does not pin. - `gh release upload`: New archives upload without `--clobber`, so a name the release already holds fails the publish instead of replacing an archive that gateways pin. `SHA256SUMS` is the one asset uploaded with `--clobber`, after the new archives land. - `gh release view`: A release counts as missing only when this command's stderr holds the exact line `release not found`; any other failure prints gh's error and stops the publish. - `SHA256SUMS`: For an existing release, the script requires its `.zip` assets and checksum lines to name the same archives before any upload, and otherwise exits naming the assets without a line and the lines without an asset. Nothing in the workflow repairs that state, so a publish that stops after an archive lands but before the checksum upload blocks every later publish until someone fixes the release. - `github.token`: The script authenticates with the job's own token, and the job's `contents: write` permission and dispatch-only gate are unchanged. - `Publish the release add-only`: The change adds no automated test for the create, append, refuse, or nothing-new paths, and the step still runs only on a manual dispatch. Design: new surface-growth @ .github/workflows/whisper-lib.yml::Publish the release add-only boundary: persisted Deferred: Verifying the five existing archives and their SHA256SUMS lines unchanged waits for the release dispatch. Plan: vibe/2026-09-28-2-whisper-cuda-backend.md
The whisper runtime workflow gains a Windows x86-64 row that builds whisper for the CPU on a hosted runner, so Windows hosts without an NVIDIA GPU get a build that does not depend on CUDA. The row configures like the Windows CUDA row without the CUDA backend, and it builds in the same push and manual runs as the other rows. A manual run adds its archive to the release beside the archives already there. The Windows packaging and load check now serve both Windows rows, and only the CUDA row bundles the CUDA runtime libraries. The Unix packaging and load check skip every Windows row. - `startsWith(matrix.platform, 'windows-')`: The Windows packaging and smoke-load steps, and the negated guards of the two Unix steps, now select rows by the platform name's `windows-` prefix instead of by equality with `windows-x86_64-cuda`. Any later row whose name starts with `windows-` gets the Windows packaging and smoke-load, and every other row gets the Unix ones. - `Package Windows runtime`: One step packages both Windows rows and branches inside its script on `windows-x86_64-cuda`, so only the CUDA row locates the CUDA toolkit and copies `cudart64_*.dll`, `cublas64_*.dll`, and `cublasLt64_*.dll`. - `windows-x86_64`: The row runs on the hosted `windows-2022` image, and the add-only publish of a manual run adds its archive and a `SHA256SUMS` line to a release that lacks them. The publish job needs every build row, so this row's failure also blocks a release publish. - `Configure Windows CPU`: The step passes the `Configure Windows CUDA` flags without `-DGGML_CUDA=ON`, an x64 shared-library build with `-DGGML_NATIVE=OFF` and with examples and tests off. - `Get-ChildItem stage -Recurse -Filter *.dll`: The CPU archive takes every DLL under the install stage, as the CUDA archive does, while the Unix packaging keeps only `libwhisper` and `libggml` libraries. - `Smoke-load Windows runtime`: No automated test covers the new row, and the workflow has no pull-request trigger, so the row first runs on a push to master or a manual run. Its only checks are this load and the packaging's `whisper.dll` and `ggml*.dll` presence checks, and none asserts that the CPU archive holds no CUDA DLL. Design: new surface-growth @ .github/workflows/whisper-lib.yml::windows-x86_64 boundary: persisted Design: new dispatch-on-tag @ .github/workflows/whisper-lib.yml::Package Windows runtime Design: new stringly-typed @ .github/workflows/whisper-lib.yml::matrix.platform Deferred: Verifying the windows-x86_64 row and the other rows unchanged on the first push run to master waits for the branch to land. Plan: vibe/2026-09-28-2-whisper-cuda-backend.md
A manual release run of the whisper runtime workflow now also compiles whisper with CUDA 12.8 on a hosted Ubuntu 22.04 runner and adds a Linux x86-64 CUDA archive to the tag's release. Pushes skip this build because the hosted CUDA compile is slow, and publishing waits for it, so a failed CUDA build publishes nothing from that run. The archive carries the CUDA runtime and cuBLAS libraries beside whisper and ggml, so a host needs only an NVIDIA driver. Before the upload, the run loads the library against the toolkit's driver stub and fails unless whisper reports its CUDA backend and every dependency outside a short list of system libraries resolves inside the archive.
- `build-linux-cuda`: The CUDA build is a job of its own beside the `build` matrix, and it repeats that job's checkout, staging cleanup, and Linux configure flags instead of sharing them. Only a `workflow_dispatch` run starts it, so a push to master never compiles or checks the CUDA packaging.
- `Package the CUDA runtime`: Each ggml library ships once, under the SONAME that `objdump -p` reports, beside `libwhisper.so` and the toolkit's `libcudart.so.12`, `libcublas.so.12`, and `libcublasLt.so.12`. Copying every file and symlink name, as the CPU Unix archives do, would store the large CUDA backend three times.
- `needs: [build, build-linux-cuda]`: The publish job now waits for the CUDA job and picks up its `whisper-linux-x86_64-cuda-${{ env.WHISPER_TAG }}` artifact through the existing `whisper-*-${{ env.WHISPER_TAG }}` download pattern, so the release gains `whisper-${WHISPER_TAG}-linux-x86_64-cuda.zip` and its `SHA256SUMS` line. A failed or timed-out CUDA build therefore publishes none of that dispatch's archives.
- `--parallel "$(nproc)"`: The CUDA build caps make at one compile job per CPU, because a bare `--parallel` starts every nvcc compile at once and can exhaust the runner's memory.
- `Smoke-load the CUDA runtime against the driver stub`: Because hosted runners have no NVIDIA driver, the step puts the toolkit's stub `libcuda.so` on the loader path as `libcuda.so.1` and fails unless `whisper_print_system_info()` lists a CUDA backend. It then fails when `ldd` shows any package library's dependency missing or resolved outside the package, other than glibc, `libstdc++`, `libgcc_s`, `libgomp`, and `libcuda.so.1`, and names each such dependency.
- `cuda-nvcc-12-8`: The toolkit packages name the 12.8 series but no patch release, and the job installs NVIDIA's apt keyring from a hard-coded URL without a checksum, so the CUDA runtime the archive bundles is whatever NVIDIA's repository serves at dispatch time.
Design: new surface-growth @ .github/workflows/whisper-lib.yml::build-linux-cuda boundary: persisted instead-of: parallel-abstraction: a separate workflow and release would split one whisper tag across two releases
Design: new hidden-dependency @ .github/workflows/whisper-lib.yml::build-linux-cuda
Design: new clone-block @ .github/workflows/whisper-lib.yml::build-linux-cuda
Deferred: Verifying the build-linux-cuda skip on the first push run to master waits for the branch to land.
Plan: vibe/2026-09-28-2-whisper-cuda-backend.md
Windows x86-64 and Linux x86-64 each have more than one whisper runtime build, so the gateway configuration gains a speech-to-text setting that names the one to download: automatic selection, the CPU build, or the CUDA build. Automatic selection is the default, and serialized configuration leaves the setting out while it holds that value. An unrecognized value fails to parse with an error that names the accepted values. This change defines and exposes the setting without adding code that acts on it. - `WhisperBackend`: A new `#[non_exhaustive]` enum, `Auto` (the default), `Cpu`, or `Cuda`, with kebab-case serde spellings, re-exported from the crate root beside the other configuration types. - `SttPipelineConfig`: Gains `whisper_backend` in `[stt]`, as its parse form `RawSttPipelineConfig` does, because the build choice is a speech concern and `[local]` configures llama-server and the artifact cache. - `WhisperBackend::is_auto`: Both structs skip the field when it returns true, so `Config::to_json` and the derived `SttPipelineConfig` serializer omit `whisper_backend` at `auto` and write the spelling otherwise. A file that sets the key fails on a build without this change, because `RawSttPipelineConfig` denies unknown fields. - `rejects_an_unknown_whisper_backend_naming_the_accepted_values`: An unknown value such as `vulkan` fails as a `ConfigErrorKind::Parse` error whose source chain names the rejected value and `auto`, `cpu`, and `cuda`. - `SttPipelineConfig::whisper_backend`: The change adds no caller outside tests and doc examples, although its doc comment says the gateway consults the setting on Windows x86-64 and Linux x86-64. - `TryFrom<RawSttPipelineConfig>`: Copies `whisper_backend` without a platform check, so `cuda` parses even on hosts outside the two the doc comments name. The doc comments say where the setting applies instead of warning when it is set elsewhere. Design: new surface-growth @ crates/gateway/config/src/config/stt.rs::WhisperBackend boundary: persisted Design: new surface-growth @ crates/gateway/config/src/config/stt.rs::SttPipelineConfig::whisper_backend boundary: pub Uncertain: A5 - the touched code does not show whether changing whisper_backend reports restart_required Plan: vibe/2026-09-28-2-whisper-cuda-backend.md
On Windows x86-64 and Linux x86-64 the gateway now picks between a CPU and a CUDA build of the pinned whisper speech runtime, so a Windows host without an NVIDIA GPU and a Linux host with one can each run a build that suits it, while every other platform keeps its single build. The speech backend setting drives the pick as the local inference server's backend setting drives its own. Automatic selection, the default, asks the NVIDIA driver tool for GPUs and takes the CUDA build when it reports any and the CPU build when it reports none or fails, and an explicit choice takes its build without probing. Speech preparation passes the configured choice to runtime provisioning and logs the provisioned library path, and the documentation and example configuration describe the setting and the per-build cache directories. The two added builds carry placeholder digests, so neither can be provisioned until its real digest is pinned. - `WhisperAsset`: Gains `backend: Option<WhisperBackend>`, set on the four Windows x86-64 and Linux x86-64 rows and `None` on the three single-build rows, so its fields now mirror `ServerAsset`'s beside it. - `whisper_asset`: Takes the setting and the probed GPU list as arguments and does no I/O. Its body repeats `server_asset`'s in the same file, with the whisper table, `auto_whisper_backend`, and `whisper_backend_applies` in place of the llama-server table, `auto_backend`, and the hard-coded Windows x86-64 check. - `whisper_backend_applies`: Derives the platforms with a choice from the table instead of naming them, so a platform has a choice exactly when one of its rows carries a backend. On such a platform, a row without a backend can never be selected. - `whisper_asset_with_probe`: Runs its probe once for `auto` on a platform with a choice and never otherwise, as the probe-count tests assert. `provision_whisper_library`, whose public signature now takes the `WhisperBackend`, passes `nvidia_compute_caps` as the probe. - `prepare_impl`: Holds the former body of `prepare`, which passes it `ArtifactStore::provision_whisper_library`, so a test can observe the backend handed to the library provision without probing or downloading. The body reads the backend from `[stt]`, falling back to `auto` when the section is absent, and logs `provisioned whisper library` with the library path at info level. - `auto_whisper_backend`: Any reported compute capability selects the CUDA build, and no GPU or a failed probe selects the CPU build. Unlike `auto_backend`, it does not tell GPU generations apart. - `WHISPER_ASSETS`: The new `windows-x86_64` and `linux-x86_64-cuda` rows pin an all-zero sha256, so provisioning either build cannot pass digest verification until real pins replace them. Under the default `auto`, those rows are the pick on Linux x86-64 hosts with an NVIDIA GPU and on Windows x86-64 hosts where the probe finds none, which the parent commit served with pinned builds. - `whisper_assets_cover_the_seven_release_builds`: Checks each digest only for 64 hex digits, so it passes the all-zero placeholders. - `provision_whisper_library_reuses_a_verified_install`: Runs with explicit `Cpu` and `Cuda` backends to keep the host's probe out, and no test in this change calls `provision_whisper_library` under `auto`, so its `nvidia_compute_caps` wiring has no test. Design: new parallel-abstraction @ crates/gateway/local/src/artifacts/assets.rs::WhisperAsset Design: new surface-growth @ crates/gateway/local/src/artifacts/assets.rs::WHISPER_ASSETS boundary: persisted Design: new pure-function @ crates/gateway/local/src/artifacts/assets.rs::auto_whisper_backend deps: Option<&[(u64, u64)]> Design: new pure-function @ crates/gateway/local/src/artifacts/assets.rs::whisper_backend_applies deps: &str,&str Design: new clone-block @ crates/gateway/local/src/artifacts/assets.rs::whisper_asset deps: &str,&str,Option<&[(u64, u64)]>,WhisperBackend Design: new shared-parameter-cluster @ crates/gateway/local/src/artifacts/assets.rs::whisper_asset deps: &str,&str,Option<&[(u64, u64)]>,WhisperBackend Design: new shared-parameter-cluster @ crates/gateway/local/src/artifacts/assets.rs::whisper_asset_with_probe deps: &str,&str,WhisperBackend,impl FnOnce() -> Option<Vec<(u64, u64)>> Design: new shared-parameter-cluster @ crates/gateway/local/src/artifacts/assets.rs::tests::pick_with_probe deps: &str,&str,Option<Vec<(u64, u64)>>,WhisperBackend Design: new surface-growth @ crates/gateway/local/src/artifacts.rs::ArtifactStore::provision_whisper_library boundary: pub Uncertain: A1 - the touched code does not show that prepare, which now reaches the nvidia-smi probe, runs only after the listener binds Plan: vibe/2026-09-28-2-whisper-cuda-backend.md
wpak-ai
marked this pull request as draft
September 29, 2026 15:10
The native speech test fixtures now name the whisper backend in their configs, so which build a suite loads no longer depends on whether the host has an NVIDIA GPU. The value comes from a test-only environment variable and defaults to the CPU build. The fixtures pass it through unchecked, so an unknown value fails the config parse with the error that names the accepted values. CI's native job sets the variable to CUDA, so its Windows GPU runner keeps testing the CUDA build. - `fixture_whisper_backend` is a new public function beside the `require_fixture` re-export, and the gateway crate's realtime test imports it from there. It reads `PROMPTFORGE_WHISPER_BACKEND` directly, returns `cpu` when the variable is unset, and returns the spelling as an unchecked `String`. - `native_speech_service`, `fixture_service_with_models_on_dedicated_thread`, and `verbose_round_trip_accepts_literal_timestamp_granularities_field` write that spelling as `[stt] whisper_backend`. The batch test's config previously had no `[stt]` section and now has one holding only that key. - `.github/workflows/stt-miri.yml` exports `PROMPTFORGE_WHISPER_BACKEND=cuda` beside the other fixture variables. Its `gateway-stt` unit and integration steps therefore keep testing the CUDA build on the self-hosted Windows CUDA runner. - `fixture_whisper_backend` has no test of its own. Its `cpu` default and its pass-through are exercised only through the native fixtures that call it. Design: extends surface-growth @ crates/gateway/stt/api/src/test_fixtures/native.rs::fixture_whisper_backend boundary: pub Design: new hidden-dependency @ crates/gateway/stt/api/src/test_fixtures/native.rs::fixture_whisper_backend boundary: pub Design: new stringly-typed @ crates/gateway/stt/api/src/test_fixtures/native.rs::fixture_whisper_backend boundary: pub Design: extends temporal-coupling @ .github/workflows/stt-miri.yml Deferred: Verifying the native-whisper job on the CUDA build waits for merge or a maintainer dispatch of .github/workflows/stt-miri.yml. Plan: vibe/2026-09-28-2-whisper-cuda-backend.md
Whisper runs inside the gateway process, so a runtime build the host cannot execute ends the gateway instead of failing only the speech load. Automatic selection on Linux x86-64 now takes the CUDA build only when the NVIDIA driver's major version is 570 or later, the CUDA 12.8 floor, and otherwise takes the CPU build, including when the version cannot be read; Windows sets no floor, and an explicit CUDA choice is still honored. On every x86-64 platform and under every setting, selection also requires the SSE4.2, AVX, AVX2, BMI2, FMA, and F16C extensions the pinned builds are compiled to use, and a CPU missing any of them fails the speech load with an error that names the build, the required extensions, and the missing ones. The GPU probe now reads the driver version beside each GPU's compute capability, and the host's GPU and CPU facts reach selection as arguments, so the selection tests run the same on any machine. The documentation and example configuration describe the floor, the fallback, and the CPU requirement. - `NvidiaProbe`: The probe's answer is now a type holding each GPU's compute capability and one driver major version, which the pure `parse_nvidia_probe` builds from `nvidia-smi` CSV lines. The llama-server pick runs the same wider query but receives only `compute_caps`, so `server_asset` keeps its signature. - `min_driver_major`: The driver floor is per-row data on `WhisperAsset`, `Some(570)` on `linux-x86_64-cuda` and `None` on every other row, and `auto_whisper_backend` reads it from the platform's CUDA row through the new `whisper_row` lookup. Only `auto` consults it. - `X86_BASELINE`: The six required extensions are bare strings in `is_x86_feature_detected!` spelling, and the same spellings recur in the match arms of `host_x86_extensions`, in both `x86_extensions: &[&str]` parameters, and in the `Vec<String>` fields of `UnsupportedCpu`. A baseline spelling without its own match arm reads as absent, so it fails selection on every x86-64 host rather than passing unchecked. - `x86_extensions`: `whisper_asset` takes the host's detected extensions as an argument, as it takes the probe's answer, and does no I/O; `provision_whisper_library` gathers them with `host_x86_extensions()` and passes `nvidia_probe` as the probe. `whisper_asset` and `whisper_asset_with_probe` both gain it as their fifth parameter. - `auto_whisper_backend`: On Linux x86-64, a GPU on a driver major below 570 or with an unreadable version now gets the CPU build, and `parse_nvidia_probe` keeps the lowest reading across GPUs, counting an unreadable one as lowest. An explicit `cuda` still selects `linux-x86_64-cuda` at any driver version, which the docs say can end the gateway. - `UnsupportedCpu`: When the CPU lacks a baseline extension, every x86-64 row, the macOS one included, fails selection under every setting before any download, with a message that names the build, every required extension, and the missing ones. Off x86-64, `host_x86_extensions` reports nothing and the aarch64 rows skip the check. - `nvidia_probe`: A probe that cannot start `nvidia-smi`, exits unsuccessfully, or finds no readable compute capability still returns `None` without a log line, and `auto` then falls back to the Vulkan llama-server build and the CPU whisper build. - `windows-x86_64-cuda`: Its row sets no driver floor, so Windows `auto` takes the CUDA build at any driver version, as `auto_takes_the_windows_cuda_whisper_build_at_any_driver_version` asserts. - `host_x86_extensions`: Neither it nor `nvidia_probe` has a direct test. `provision_whisper_library_reuses_a_verified_install` runs it on the real CPU, so that test now fails on an x86-64 host below the baseline, and no test runs `provision_whisper_library` under `auto`. Design: new value-object @ crates/gateway/local/src/artifacts/assets.rs::NvidiaProbe Design: new stringly-typed @ crates/gateway/local/src/artifacts/assets.rs::X86_BASELINE Design: new pure-function @ crates/gateway/local/src/artifacts/assets.rs::whisper_row deps: &str,&str,Option<WhisperBackend> Design: extends pure-function @ crates/gateway/local/src/artifacts/assets.rs::auto_whisper_backend deps: &str,&str,Option<&NvidiaProbe> Design: new shared-parameter-cluster @ crates/gateway/local/src/artifacts/assets.rs::auto_whisper_backend deps: &str,&str,Option<&NvidiaProbe> Design: new pure-function @ crates/gateway/local/src/artifacts/assets.rs::whisper_asset deps: &[&str],&str,&str,Option<&NvidiaProbe>,WhisperBackend Design: extends shared-parameter-cluster @ crates/gateway/local/src/artifacts/assets.rs::whisper_asset deps: &[&str],&str,&str,Option<&NvidiaProbe>,WhisperBackend Design: extends shared-parameter-cluster @ crates/gateway/local/src/artifacts/assets.rs::whisper_asset_with_probe deps: &[&str],&str,&str,WhisperBackend,impl FnOnce() -> Option<NvidiaProbe> Design: new pure-function @ crates/gateway/local/src/artifacts.rs::parse_nvidia_probe deps: &str Design: new swallowed-exception @ crates/gateway/local/src/artifacts.rs::nvidia_probe Design: extends surface-growth @ crates/gateway/local/src/artifacts.rs::ArtifactStore::provision_whisper_library boundary: pub Design: new surface-growth @ crates/gateway/local/src/error.rs::LocalError::UnsupportedCpu boundary: pub Uncertain: A1 - the touched code does not show that provision_whisper_library, which now also reads the host CPU's extensions, runs only after the listener binds Pending: N2 - compounds Plan: vibe/2026-09-28-2-whisper-cuda-backend.md
The Windows x86-64 CPU and Linux x86-64 CUDA whisper builds have no archive in the b4938 release yet, so their rows in the runtime download table keep all-zero digests. A comment above each of those rows now says that the digest is filled in once the release holds the archive, and that until then the row fails closed because no download can match it. The change lands with both rows in that state, and a single later commit pins both digests once the release publishes the archives. No code behavior changes. - `WHISPER_ASSETS` keeps the all-zero `sha256` values of its `windows-x86_64` CPU row and its `linux-x86_64-cuda` row. Each of the two rows gains a comment stating that the pin is filled in once the `whisper-lib-b4938` release holds the archive, and that the row is fail-closed until then. - `WHISPER_ASSETS` gains comments only: no digest, row, or field changes, and no test is added or changed. Deferred: The windows-x86_64 and linux-x86_64-cuda sha256 pins wait for the whisper-lib-b4938 release after merge. Plan: vibe/2026-09-28-2-whisper-cuda-backend.md
Packaging the Windows CUDA whisper runtime now fails when the bundle lacks the CUDA runtime, cuBLAS, or cuBLASLt library copied from the toolkit. Before, the toolkit search ignored its own errors and the packaging never checked what it found, so an archive could be built without one of those libraries. The check follows each library's copy and runs before the archive is compressed, and its error names the missing library's file pattern. The Windows CPU build's packaging is unchanged.
- `foreach ($runtime in @("cudart64_*.dll", "cublas64_*.dll", "cublasLt64_*.dll"))`: A `foreach` statement replaces the pipeline over the three toolkit patterns, and each pattern's check sits right after its own copy inside the CUDA-only branch. The added loop and nested check put the packaging script's cognitive complexity at 16, up from 11.
- `throw "$runtime missing from the runtime bundle"`: The script stops at the first pattern, in list order, that matches no file in the package directory, before it compresses the archive, so a toolkit missing several libraries reports only the first. The message has the form the `ggml*.dll` check later in the script already throws.
- `-ErrorAction SilentlyContinue`: The toolkit search still discards its own errors, such as an unreadable directory or a toolkit root that does not exist, so a failing row reports the missing pattern and not the search error behind it.
- `if (-not (Get-ChildItem package -Filter $runtime))`: The change adds no test for the check, and the workflow has no pull-request trigger, so the check first runs on the self-hosted CUDA runner when the workflow next runs on master or by manual dispatch.
Design: removes swallowed-exception @ .github/workflows/whisper-lib.yml::Package Windows runtime boundary: persisted
Design: new oversized-unit @ .github/workflows/whisper-lib.yml::Package Windows runtime boundary: persisted
Plan: vibe/2026-09-28-2-whisper-cuda-backend.md
The Windows x86-64 CPU whisper build now configures ggml without OpenMP, so ggml runs its own thread pool instead of the MSVC OpenMP runtime. With OpenMP on, freeing the whisper library after a CPU transcription unloaded that runtime under its live worker threads and ended the host process. The Windows CUDA build and the other platforms' builds keep their configuration. - `-DGGML_OPENMP=OFF`: Only the Windows CPU configure step gains the flag, and the DLLs it builds no longer import the MSVC OpenMP runtime, `VCOMP140.DLL`. The Windows CUDA configure step keeps OpenMP, so the two Windows builds, which otherwise share every flag except `-DGGML_CUDA=ON`, now differ in their threading runtime. - `.github/workflows/whisper-lib.yml`: No step checks that the CPU build is free of OpenMP: the Windows smoke-load prints the system information without asserting on it, and the CPU row's packaging checks only that `whisper.dll` and `ggml*.dll` files are present. The change adds no test that unloads the library, and with no pull-request trigger the changed row first builds on a push to master or a manual dispatch. Plan: vibe/2026-09-28-2-whisper-cuda-backend.md
Every native whisper test now shuts down each speech engine it builds and expects the shutdown to succeed, where it used to drop the engine. Shutdown joins the engine's decode workers, so each test releases its decoders and its hold on the speech library before it returns, and the process no longer exits while a worker is still freeing whisper state. Dropping an engine only signals its workers and detaches them; on the Linux CUDA build, the detached workers raised CUDA errors while freeing whisper state, and the suite then aborted during its last test. No engine code changes, and the existing assertions stay as they were. - `std::fs::remove_file(model)`: The transcription-contract test now shuts down `glossary_prompted` and then `unprompted` before it removes its copied model, and since `shutdown` takes the engine by reference, both engines are still in scope at the removal. The removal's expectation now credits shutting down, not dropping, with releasing the model. - `.shutdown()`: A worker that panicked now fails its test, because shutdown returns an error for it and every call expects success; dropping the engine detached that worker without joining it. - `engine.shutdown()`: Each shutdown follows the test's assertions, so a test that panics first, on an assertion or a failed decode, still drops its engines and leaves their workers detached. - `crates/gateway/stt/backend-whisper/tests/native_whisper.rs`: Every test that builds an engine is ignored by default and needs the packaged whisper and model fixtures, and none checks the process's exit status or error output. A worker still freeing whisper state at exit can therefore surface only when the suite runs against a packaged build. Design: new oversized-unit @ crates/gateway/stt/backend-whisper/tests/native_whisper.rs::packaged_runtime_preserves_native_transcription_contract Plan: vibe/2026-09-28-2-whisper-cuda-backend.md
A graceful stop now retires the speech service before the gateway finishes serving. Admission closes, which ends Realtime sessions, admitted speech work drains, and the engine's workers are joined, so native decoders free their contexts while the process is still whole. CUDA memory freed during process exit fails with a driver-shutting-down error, which this ordering avoids. The retirement shares the command worker's join deadline, so the overall shutdown bound is unchanged, and a retirement still draining at that deadline is abandoned with a warning, as a stuck command already is. A speech service handed to the gateway is retired with it, so one service serves one run of the server.
- `with_speech_service`: A graceful stop of `serve` retires the service it was given, clones included, so one service serves one `serve` call. The limit is documented, not typed.
- `deadline`: Read once before the worker join as now plus `WORKER_JOIN_TIMEOUT`, it bounds that join and then the speech retirement, both through `tokio::time::timeout_at`, so the two waits together take at most one bound.
- `tokio::task::spawn_blocking`: The retirement blocks until admitted speech work drains, so it runs on the blocking pool. Past the deadline, `serve` stops waiting and leaves the task to the runtime teardown bound.
- `serve_until_stopped`: The worker-join test's serving setup and parked command move into this helper and `park_a_command`, which the speech tests reuse.
- `speech did not retire within {WORKER_JOIN_TIMEOUT:?}; abandoning it`: `serve` logs this warning when the deadline passes first, then returns the server's result.
- `serve_retires_speech_before_it_returns`: By the time `serve` returns `Ok`, both scripted workers are dropped and the service reports not ready.
- `serve_abandons_a_speech_retirement_that_outlasts_the_join_bound`: A held worker job keeps the retirement waiting, so `serve` returns only after the whole join bound. The abandoned retirement joins both workers once the job drops.
- `serve_bounds_the_worker_join_and_the_speech_retirement_by_one_deadline`: With a stuck command and a held job, `serve` returns within one and a half join bounds, so the retirement gets no bound of its own. It still joins both workers once the job drops.
- `gateway_auth_origin_query_and_final_speech_surfaces_precede_upgrade`: The trusted server gets its own scripted service, because the strict server's stop retired the shared one.
- `timeout_at`: The retirement's check reads only the timeout, so a panic inside the blocking task is discarded with its join result, and `serve` neither logs it nor returns an error. The worker join discards its join result the same way.
- `ScriptedDecoder`: The new tests run on scripted decoders, so none loads a native decoder or reproduces the CUDA error itself.
Design: new temporal-coupling @ crates/gateway/app/src/runner.rs::Gateway::with_speech_service boundary: pub
Design: new surface-growth @ crates/gateway/app/src/runner.rs::Gateway::serve boundary: pub instead-of: global-state: skipping the native free at process exit needs process-global state to detect exit
Design: extends swallowed-exception @ crates/gateway/app/src/runner.rs::Gateway::serve boundary: pub
Design: new oversized-unit @ crates/gateway/app/src/runner.rs::Gateway::serve boundary: pub
Design: new clone-block @ crates/gateway/app/src/runner.rs::drain_tests::serve_bounds_the_worker_join_and_the_speech_retirement_by_one_deadline
Design: extends oversized-unit @ crates/gateway/app/tests/it/realtime_stt/authentication.rs::gateway_auth_origin_query_and_final_speech_surfaces_precede_upgrade
Repairs: a graceful stop retires speech before serve returns @ crates/gateway/app/src/runner.rs::Gateway::serve - serve returned before the speech engine's workers were joined
Plan: vibe/2026-09-28-2-whisper-cuda-backend.md
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #90
What changes
whisper-lib.ymlSHA256SUMS, and refuses a release whose archives and checksum lines disagree. The five pinnedb4938archives stay untouched.windows-x86_64CPU row on hostedwindows-2022, built on push like the others.build-linux-cudajob (CUDA 12.8, hostedubuntu-22.04), dispatch only. Its archive bundles cudart and cuBLAS, so hosts need only the NVIDIA driver.gateway-config:[stt] whisper_backend = auto | cpu | cuda, like[local] llama_backend.gateway-localselects the build asserver_assetdoes;autoreusesnvidia_compute_caps()on Windows and Linux x86-64.prepare()passes the setting and logsprovisioned whisper library. Docs, guide, and example config updated.Before merge
The two new rows fail closed (all-zero SHA-256) until pinned.
Testing
libcuda.so.1.build-xtask, books, and fmt pass. Clippy passes on the three changed crates; workspace clippy is left to CI.