Repository navigation
feat: extract phoneme timing from a recording with the TIFA aligner - #2484
Open
KakaruHayate wants to merge 7 commits into
Open
KakaruHayate wants to merge 7 commits into
KakaruHayate wants to merge 7 commits into
Conversation
Adds a phoneme-timing extraction mode to the audio transcription dialog. Where SOME and GAME turn a recording into notes, this aligns the phoneme timing of an existing track to a recording and writes the result back as per-phoneme offsets, one undoable edit. Notes and lyrics are untouched, so the feature fits a project whose MIDI and lyrics are already finished. The aligner is tifa.cpp's CLI (TIFA, a token-imputing forced aligner), installed as an .oudep package next to the other dependencies. It takes an explicit phone list, so no text G2P is involved: OpenUtau already knows how each note is pronounced. Phoneme mapping: * The phone list is built per note from the aligner's own syllable/mora tables (Mandarin pinyin, Cantonese jyutping, Japanese mora, English ARPAbet). A 3-segment bank and the aligner's 2-segment table therefore meet in the middle instead of relying on symbol-by-symbol guesses. * VC-style phonemes that repeat the previous note's tail are folded into a single occurrence, using a longest suffix/prefix match that only trims multi-phone units, so a repeated vowel across two notes survives. * Per note, the returned spans are collapsed into one window and the note's phonemes are remapped onto it proportionally, which preserves the bank's internal duration allocation (compound vowels, VC transitions). * Unmappable notes are skipped and reported; windows too short to trust are left untouched and reported separately. Classic banks are positioned by their audible onset (position - preutter) so the rendered consonant lands on the measured boundary while the bank's overlap behaviour stays intact; model singers have no oto, so their phoneme position is already the audible timing.
Adds the repack script that turns a tifa.cpp CLI bundle into an .oudep package (manifest, platform tag, flat archive) and rejects a package built for another platform before spawning the CLI.
The aligner is trained on short utterances and tifa.cpp caps the CLI at 6000 mel frames (60 s). A part longer than 40 s is now split at note boundaries, each chunk is aligned against its own audio slice (with half a second of margin), and the spans are rebased onto the project timeline before the moves are computed. Chunk boundaries land on pauses because a note is never split, and per-chunk cropping also keeps unrelated audio out of the alignment window.
* Zero-width placeholders keep the span list parallel to the phone list when a chunk has no audio or the CLI returns an unexpected span count, instead of aborting the whole alignment. * The package script keys its download cache by release tag, so packaging a different tag cannot silently reuse another release's bundle. * The result message labels the unresolved and uncertain note counts; both were formatted into the string without appearing in it.
tifa.cpp v0.1.5 fills the stretches no phone covers in every TextGrid tier with an interval labelled by --fill-gaps (default SP), so the phones tier no longer holds exactly the phone list and the span count check rejected every chunk. Drop the filler intervals when that makes the counts line up and fall back to the raw tier otherwise, which keeps older bundles working. The packaging script now drops what the align path never reads (the breath/AP detector weights of tifa.cpp's dataset workflow, the English LSTM G2P and the text G2P dictionaries): 96.8 -> 44.8 MB on Windows, 121.1 -> 69.1 MB on Linux, 82.0 -> 30.1 MB on macOS, verified end to end against v0.1.5 with the trimmed model only. --keep-all restores the verbatim bundle.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Adds a fourth algorithm to the audio transcription dialog: phoneme timing
alignment. SOME and GAME turn a recording into notes; this instead aligns
the phoneme timing of an existing track to a recording and writes the result
back as per-phoneme offsets, in one undoable edit. Notes and lyrics are not
touched, and the piano roll's existing phoneme layer shows the new
boundaries. It fits the workflow where MIDI and lyrics are already finished
and the recording tells us where the phonemes actually are.
The aligner is tifa.cpp (TIFA, a
token-imputing forced aligner) driven through its CLI, installed as an
.oudepdependency next to the other ones. The CLI takes an explicit phonelist, so no text G2P is involved: OpenUtau already knows how each note is
pronounced.
Phoneme mapping
This is the part that needed the most care, because a bank's dictionary
splits a syllable differently from the aligner's tables:
tables (Mandarin pinyin, Cantonese jyutping, Japanese mora, English
ARPAbet), not by guessing symbol by symbol. A 3-segment bank and the
aligner's 2-segment table therefore meet in the middle: the aligner always
receives its canonical sequence, which is what its accuracy depends on.
a k+k afora ka) are folded into one occurrence using a longest suffix/prefix matchthat only trims multi-phone units, so a deliberate repeated vowel across
two notes survives.
phonemes are remapped onto it proportionally to their original timing,
which preserves the bank's internal duration allocation (compound vowels,
VC transitions) - exactly the requirement for a 3-segment compound vowel
meeting a 2-segment table.
syllable resolved into one-frame phones), are left untouched and reported
in the result dialog.
set falls back to "unresolved" and is reported rather than mis-aligned.
Classic vs model singers
Model singers are the primary target: they have no oto, so a phoneme's
positionis already its audible timing and the offsets translate directlyinto rendered durations.
Classic (UTAU) banks overlap phonemes by design, which the aligner cannot
measure - but it does not need to. The aligner measures boundaries, and in
the classic renderer a phoneme's audible onset is
position - preutter,with the overlap coming from the oto. So the alignment moves
positionandleaves the bank's overlap behaviour intact, the same operation as dragging
the phoneme start handle in the piano roll. Consonant length itself still
comes from the sample and VEL, as before.
Long recordings
TIFA is trained on short utterances and tifa.cpp caps the CLI at 6000 mel
frames (60 s). Parts longer than 40 s are split at note boundaries, each
chunk is aligned against its own audio slice (0.5 s margin), and the spans
are rebased onto the project timeline before the moves are computed. The
progress dialog shows the chunk counter. Cropping also keeps unrelated audio
out of the alignment window, which measurably helps: on a test clip the
agreement metric went from 0.81 (single pass) to 0.94 (chunked).
Dependency
The package is a flat zip with
oudep.yaml+config.json+ the CLI, itsggml runtime libraries and the q4 GGUF model.
Misc/tifa-oudep/package.pyrebuilds it from the tifa.cpp release bundles (q4 only, three platforms:
windows-x64, linux-x64, macos-arm64). For testing I attached the three
packages to https://github.com/KakaruHayate/tifa.cpp/releases/tag/oudep;
maintainers may want to host their own copy. The code does not hardcode a
download URL - the user installs the
.oudeplike any other dependency.Testing
TextGrid parser, the VC folding, the 3-segment-to-2-segment case, and the
proportional remap (
OpenUtau.Test/Core/Analysis/TifaPhonemeAlignerTest.cs).dependency: 20/20 phones placed, agreement 0.81 single-pass / 0.94 chunked,
monotonic offsets, clamps reported.
Not in this PR: a settings entry for the CLI backend (currently
auto),preutter adjustment for classic banks (optional follow-up), and per-phoneme
UI preview before applying.
Update: tifa.cpp v0.1.5
The package now targets tifa.cpp v0.1.5. Two things changed there that this
PR handles:
interval labelled by
--fill-gaps(defaultSP), so the phones tier is nolonger exactly the phone list. The parser drops those fillers when that
makes the counts line up and falls back to the raw tier otherwise, so both
v0.1.3 and v0.1.5 bundles work. Model, vocabulary and dictionaries are
unchanged between the two versions (verified with
inspectand by diffingthe bundled dictionaries), so the mapping tables needed no update.
ggml libraries and
models/tifa.gguf. The breath/AP detector weights belongto tifa.cpp's dataset workflow (
align -> breathe --merge -> align) and theLSTM G2P plus dictionaries are only used for text input, which OpenUtau
never triggers because it always passes an explicit phone list. That is
96.8 -> 44.8 MB (Windows), 121.1 -> 69.1 MB (Linux) and 82.0 -> 30.1 MB
(macOS), verified end to end against v0.1.5 with only the aligner weights
present.
--keep-allstill builds the verbatim bundle.