Skip to content

Serialize DirectML session creation, inference and disposal to stop native crashes - #2494

Open
KakaruHayate wants to merge 6 commits into
openutau:masterfrom
KakaruHayate:fix/dml-serialization
Open

KakaruHayate wants to merge 6 commits into
openutau:masterfrom
KakaruHayate:fix/dml-serialization

Conversation

@KakaruHayate

@KakaruHayate KakaruHayate commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

What

The Windows DirectML build of ONNX Runtime kills the process natively when session creation, Run and disposal overlap on the same device: 0xC0000005, no managed exception, nothing in the log, only a Windows ".NET Runtime" 1026 event whose recorded address is the CLR's fail-fast site rather than the original fault. This serializes all three behind Onnx.DmlLock, engaged only for the DirectML runner, and falls back to CPU when the DML graph compiler rejects a model instead of failing the render.

Why this is real (and not just the 1.24.4 issue)

Measured on an RTX 2070 (driver 32.0.16.1088) with onnxruntime 1.23.0 + DirectML 1.15.4, binaries SHA-256 identical to the official Windows package, 4 threads on one DML session:

workload before after
concurrent Run on one session segfault 0xC0000005 in seconds survives
Run while another session is created segfault survives
Run while a session is disposed segfault survives
concurrent creation only (no inference) safe safe

Crash dumps from user reports land in this path both with 1.24.4 and with the pinned 1.23.0 (same signature, same DiffSinger session-creation stack), so the version pin alone does not cover machines where the overlap happens.

Related onnxruntime issues - a documented bug class that cannot reach the frozen package

This matches a class of DirectML EP bugs whose fixes land on main but can never ship in the Microsoft.ML.OnnxRuntime.DirectML package (frozen at 1.24.4; NuGet has 1.23.0 then 1.24.1-1.24.4 and nothing newer):

  • #27118 - transformer-based models fail with E_INVALIDARG (80070057) at init or Run on the DML EP (Int64 indices in Gather/Scatter/TopK ops); DiffSinger acoustic models are transformer-based and fail exactly this way on some machines (catchable on some, native crash on others).
  • #31665 - unvalidated constant tensor byte size in the DML OnnxTensorWrapper handed bad data to kernels instead of failing (merged 2026-08-08).
  • #32745 - unvalidated DML upload heap allocations; oversized allocations crashed instead of returning out-of-memory (merged 2026-10-02).
  • #22147 / #20713 / #28007 - the DML EP is not thread-safe; the per-instance mutex fix is still an unmerged PR.

Unguarded paths closed

  • DiffSingerSinger.FreeMemory disposed sessions while a render could be mid-Run - SessionLock only covers the getters, not inference. This is the singer-switch crash.
  • LoadRenderedPitch (Ctrl+R and the realtime pitch service), LoadRenderedRealCurves and the script/export path shared no lock with the render path.
  • The other DML users (Worldline vocoder, SOME, Hnsep, Hifisampler, game backend) shared no lock with DiffSinger at all.

Notes

  • Onnx.EnterDmlScope() is a no-op for the CPU runner; DirectML work is serialized (within DiffSinger renders it already was, via the static render lock).
  • Lock order is uniformly DmlLock -> SessionLock; all SessionLock acquisition sites were audited.
  • Once DirectML has been used, the scope stays engaged even if the preference later switches to CPU, because cached sessions keep running on DirectML; a process that never used DirectML pays nothing.
  • Session creation now falls back to CPU with a warning (Failed to create a DirectML inference session, falling back to CPU.) when the DML EP cannot compile a model, so affected machines keep rendering.
  • Because a native DirectML crash cannot be caught in-process, a marker file is written while DirectML work is in flight; if the next start finds it, the runner is switched to CPU automatically (persisted, logged, toasted) so affected machines self-heal instead of crash-looping.
  • Tests: 613 passed, 0 failed, 2 skipped on the merge result.

The Windows DirectML build of ONNX Runtime kills the process natively
(0xC0000005, no managed exception) whenever session creation, Run and
disposal overlap on the same device. Reproduced on an RTX 2070 with
onnxruntime 1.23.0 + DirectML 1.15.4: concurrent Run on one session,
Run while a session is created, and Run while a session is disposed all
die in seconds. Concurrent creation alone is safe, so the crash needs
an overlap, exactly the pattern the app reached.

Unguarded paths in the app, all now behind the scope:
- DiffSingerSinger.FreeMemory disposed sessions while a render could be
  mid-Run: SessionLock only covers the getters, not inference.
- LoadRenderedPitch, the script/export path and the realtime pitch
  service shared no lock with the render path.
- The other DML users (Worldline vocoder, SOME, Hnsep, Hifisampler,
  game backend) shared no lock with DiffSinger at all.

Onnx.DmlLock/EnterDmlScope only engages for the DirectML runner.
Session creation also falls back to CPU when the DML graph compiler
rejects a model instead of failing the render.
Review finding: EnterDmlScope() read the runner preference at call
time, so switching the preference to CPU while DirectML sessions were
still cached (or while DirectML work was in flight) turned the scope
into a no-op and let those sessions be created, run and disposed
concurrently again - the same native crash, just behind a preference
switch.

Track whether DirectML is or was in use and keep the lock engaged for
the rest of the process. A process that never used DirectML still gets
the no-op scope, so pure CPU users pay nothing.
@KakaruHayate
KakaruHayate requested a review from a team October 4, 2026 14:09
A native DirectML crash cannot be caught in-process, so neither the
serialization nor the catchable-failure CPU fallback can save machines
whose DML stack dies inside session creation itself. Reported on
0.1.572.9016: switching singers still crashed with createSession on the
stack - the global lock rules out an overlap, the DML call dies alone.

Write a marker file while DirectML work is in flight and clear it when
the scope ends. If the marker is still present at the next start, the
previous run died inside DirectML: switch the runner to CPU, persist
the preference, log it and toast the user. The switch can be reverted
in Preferences; a process that never touches DirectML never writes the
marker.
@KakaruHayate

KakaruHayate commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor Author

Reverted acbf605 (the crash-marker auto-fallback) for now, keeping this PR focused on the serialization fix.

Context: that commit was added after a single user still crashed on the serialized build. Their crash is inside InferenceSession.Init itself, with no concurrent DML work possible under the global lock - i.e. most likely a machine-specific DirectML failure (the #32745/#31665-class paths that cannot ship in the frozen package) rather than something this PR can address in-process. We would rather gather feedback with the serialization alone first, to rule out machine-specific issues, before proposing the self-healing behavior separately.

@stakira

The real-curve refresh and realtime pitch services introduced in openutau#2175 run on
thread-pool threads and used to take the predictor reference outside
SessionLock, then enter the lock to call Process. Between the getter returning
and the lock being taken, a singer switch's FreeMemory (which disposes the
predictors under SessionLock) could free the native session, leaving the
background thread to call Process on a disposed handle.

Take the getter result and use it entirely inside SessionLock in the three
affected paths - LoadRenderedPitch (plain and retake) and
LoadRenderedRealCurves - so the reference cannot be disposed while in use.
Lock order stays DmlLock -> SessionLock, matching every other site.
DiffSingerScript's variance export took the predictor reference before
entering both the DML scope and SessionLock, so a concurrent FreeMemory could
dispose the native session before Process ran - the same pattern fixed for the
other three call sites. Take the reference inside SessionLock.

EnterDmlScope decided whether to engage the lock before acquiring it, reading
the runner outside the lock. A CPU->DirectML switch landing while a thread
waited on the lock would hand it a no-op scope that cannot be upgraded,
leaving the following DirectML work outside DmlLock. Acquire the lock first,
then re-check the runner under it.

CPU fallback handling is unchanged; this PR only addresses races.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant