Skip to content

feat: extract phoneme timing from a recording with the TIFA aligner - #2484

Open
KakaruHayate wants to merge 7 commits into
openutau:masterfrom
KakaruHayate:feature/tifa-phoneme-timing
Open

KakaruHayate wants to merge 7 commits into
openutau:masterfrom
KakaruHayate:feature/tifa-phoneme-timing

Conversation

@KakaruHayate

@KakaruHayate KakaruHayate commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

What

Adds a fourth algorithm to the audio transcription dialog: phoneme timing
alignment
. SOME and GAME turn a recording into notes; this instead aligns
the phoneme timing of an existing track to a recording and writes the result
back as per-phoneme offsets, in one undoable edit. Notes and lyrics are not
touched, and the piano roll's existing phoneme layer shows the new
boundaries. It fits the workflow where MIDI and lyrics are already finished
and the recording tells us where the phonemes actually are.

The aligner is tifa.cpp (TIFA, a
token-imputing forced aligner) driven through its CLI, installed as an
.oudep dependency next to the other ones. The CLI takes an explicit phone
list, so no text G2P is involved: OpenUtau already knows how each note is
pronounced.

Phoneme mapping

This is the part that needed the most care, because a bank's dictionary
splits a syllable differently from the aligner's tables:

  • The phone list is built per note from the aligner's own syllable/mora
    tables (Mandarin pinyin, Cantonese jyutping, Japanese mora, English
    ARPAbet), not by guessing symbol by symbol. A 3-segment bank and the
    aligner's 2-segment table therefore meet in the middle: the aligner always
    receives its canonical sequence, which is what its accuracy depends on.
  • VC-style phonemes that repeat the previous note's tail (a k + k a for
    a ka) are folded into one occurrence using a longest suffix/prefix match
    that only trims multi-phone units, so a deliberate repeated vowel across
    two notes survives.
  • Per note, the returned spans are collapsed into one window and the note's
    phonemes are remapped onto it proportionally to their original timing,
    which preserves the bank's internal duration allocation (compound vowels,
    VC transitions) - exactly the requirement for a 3-segment compound vowel
    meeting a 2-segment table.
  • Notes whose lyrics cannot be mapped, and windows too short to trust (a
    syllable resolved into one-frame phones), are left untouched and reported
    in the result dialog.
  • English currently covers the ARPAbet-based phonemizers; the VCCV symbol
    set falls back to "unresolved" and is reported rather than mis-aligned.

Classic vs model singers

Model singers are the primary target: they have no oto, so a phoneme's
position is already its audible timing and the offsets translate directly
into rendered durations.

Classic (UTAU) banks overlap phonemes by design, which the aligner cannot
measure - but it does not need to. The aligner measures boundaries, and in
the classic renderer a phoneme's audible onset is position - preutter,
with the overlap coming from the oto. So the alignment moves position and
leaves the bank's overlap behaviour intact, the same operation as dragging
the phoneme start handle in the piano roll. Consonant length itself still
comes from the sample and VEL, as before.

Long recordings

TIFA is trained on short utterances and tifa.cpp caps the CLI at 6000 mel
frames (60 s). Parts longer than 40 s are split at note boundaries, each
chunk is aligned against its own audio slice (0.5 s margin), and the spans
are rebased onto the project timeline before the moves are computed. The
progress dialog shows the chunk counter. Cropping also keeps unrelated audio
out of the alignment window, which measurably helps: on a test clip the
agreement metric went from 0.81 (single pass) to 0.94 (chunked).

Dependency

The package is a flat zip with oudep.yaml + config.json + the CLI, its
ggml runtime libraries and the q4 GGUF model. Misc/tifa-oudep/package.py
rebuilds it from the tifa.cpp release bundles (q4 only, three platforms:
windows-x64, linux-x64, macos-arm64). For testing I attached the three
packages to https://github.com/KakaruHayate/tifa.cpp/releases/tag/oudep;
maintainers may want to host their own copy. The code does not hardcode a
download URL - the user installs the .oudep like any other dependency.

Testing

  • Unit tests for the syllable tables, pinyin/jyutping normalization, the
    TextGrid parser, the VC folding, the 3-segment-to-2-segment case, and the
    proportional remap (OpenUtau.Test/Core/Analysis/TifaPhonemeAlignerTest.cs).
  • End-to-end runs against a real singing recording with the packaged
    dependency: 20/20 phones placed, agreement 0.81 single-pass / 0.94 chunked,
    monotonic offsets, clamps reported.
  • Full suite: 619 tests pass on Windows.

Not in this PR: a settings entry for the CLI backend (currently auto),
preutter adjustment for classic banks (optional follow-up), and per-phoneme
UI preview before applying.


Update: tifa.cpp v0.1.5

The package now targets tifa.cpp v0.1.5. Two things changed there that this
PR handles:

  • v0.1.5 fills the stretches no phone covers in every TextGrid tier with an
    interval labelled by --fill-gaps (default SP), so the phones tier is no
    longer exactly the phone list. The parser drops those fillers when that
    makes the counts line up and falls back to the raw tier otherwise, so both
    v0.1.3 and v0.1.5 bundles work. Model, vocabulary and dictionaries are
    unchanged between the two versions (verified with inspect and by diffing
    the bundled dictionaries), so the mapping tables needed no update.
  • The .oudep packages are trimmed to what the align path reads: the CLI, its
    ggml libraries and models/tifa.gguf. The breath/AP detector weights belong
    to tifa.cpp's dataset workflow (align -> breathe --merge -> align) and the
    LSTM G2P plus dictionaries are only used for text input, which OpenUtau
    never triggers because it always passes an explicit phone list. That is
    96.8 -> 44.8 MB (Windows), 121.1 -> 69.1 MB (Linux) and 82.0 -> 30.1 MB
    (macOS), verified end to end against v0.1.5 with only the aligner weights
    present. --keep-all still builds the verbatim bundle.

Adds a phoneme-timing extraction mode to the audio transcription dialog.
Where SOME and GAME turn a recording into notes, this aligns the phoneme
timing of an existing track to a recording and writes the result back as
per-phoneme offsets, one undoable edit. Notes and lyrics are untouched, so
the feature fits a project whose MIDI and lyrics are already finished.

The aligner is tifa.cpp's CLI (TIFA, a token-imputing forced aligner),
installed as an .oudep package next to the other dependencies. It takes an
explicit phone list, so no text G2P is involved: OpenUtau already knows how
each note is pronounced.

Phoneme mapping:

* The phone list is built per note from the aligner's own syllable/mora
  tables (Mandarin pinyin, Cantonese jyutping, Japanese mora, English
  ARPAbet). A 3-segment bank and the aligner's 2-segment table therefore
  meet in the middle instead of relying on symbol-by-symbol guesses.
* VC-style phonemes that repeat the previous note's tail are folded into a
  single occurrence, using a longest suffix/prefix match that only trims
  multi-phone units, so a repeated vowel across two notes survives.
* Per note, the returned spans are collapsed into one window and the note's
  phonemes are remapped onto it proportionally, which preserves the bank's
  internal duration allocation (compound vowels, VC transitions).
* Unmappable notes are skipped and reported; windows too short to trust are
  left untouched and reported separately.

Classic banks are positioned by their audible onset (position - preutter)
so the rendered consonant lands on the measured boundary while the bank's
overlap behaviour stays intact; model singers have no oto, so their
phoneme position is already the audible timing.
Adds the repack script that turns a tifa.cpp CLI bundle into an .oudep
package (manifest, platform tag, flat archive) and rejects a package built
for another platform before spawning the CLI.
The aligner is trained on short utterances and tifa.cpp caps the CLI at
6000 mel frames (60 s). A part longer than 40 s is now split at note
boundaries, each chunk is aligned against its own audio slice (with half a
second of margin), and the spans are rebased onto the project timeline
before the moves are computed. Chunk boundaries land on pauses because a
note is never split, and per-chunk cropping also keeps unrelated audio out
of the alignment window.
KakaruHayate and others added 3 commits October 2, 2026 23:47
* Zero-width placeholders keep the span list parallel to the phone list when
  a chunk has no audio or the CLI returns an unexpected span count, instead
  of aborting the whole alignment.
* The package script keys its download cache by release tag, so packaging a
  different tag cannot silently reuse another release's bundle.
* The result message labels the unresolved and uncertain note counts; both
  were formatted into the string without appearing in it.
tifa.cpp v0.1.5 fills the stretches no phone covers in every TextGrid tier
with an interval labelled by --fill-gaps (default SP), so the phones tier no
longer holds exactly the phone list and the span count check rejected every
chunk. Drop the filler intervals when that makes the counts line up and fall
back to the raw tier otherwise, which keeps older bundles working.

The packaging script now drops what the align path never reads (the breath/AP
detector weights of tifa.cpp's dataset workflow, the English LSTM G2P and the
text G2P dictionaries): 96.8 -> 44.8 MB on Windows, 121.1 -> 69.1 MB on
Linux, 82.0 -> 30.1 MB on macOS, verified end to end against v0.1.5 with the
trimmed model only. --keep-all restores the verbatim bundle.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant