Serialize DirectML session creation, inference and disposal to stop native crashes - #2494
Open
KakaruHayate wants to merge 6 commits into
Open
KakaruHayate wants to merge 6 commits into
KakaruHayate wants to merge 6 commits into
Conversation
The Windows DirectML build of ONNX Runtime kills the process natively (0xC0000005, no managed exception) whenever session creation, Run and disposal overlap on the same device. Reproduced on an RTX 2070 with onnxruntime 1.23.0 + DirectML 1.15.4: concurrent Run on one session, Run while a session is created, and Run while a session is disposed all die in seconds. Concurrent creation alone is safe, so the crash needs an overlap, exactly the pattern the app reached. Unguarded paths in the app, all now behind the scope: - DiffSingerSinger.FreeMemory disposed sessions while a render could be mid-Run: SessionLock only covers the getters, not inference. - LoadRenderedPitch, the script/export path and the realtime pitch service shared no lock with the render path. - The other DML users (Worldline vocoder, SOME, Hnsep, Hifisampler, game backend) shared no lock with DiffSinger at all. Onnx.DmlLock/EnterDmlScope only engages for the DirectML runner. Session creation also falls back to CPU when the DML graph compiler rejects a model instead of failing the render.
Review finding: EnterDmlScope() read the runner preference at call time, so switching the preference to CPU while DirectML sessions were still cached (or while DirectML work was in flight) turned the scope into a no-op and let those sessions be created, run and disposed concurrently again - the same native crash, just behind a preference switch. Track whether DirectML is or was in use and keep the lock engaged for the rest of the process. A process that never used DirectML still gets the no-op scope, so pure CPU users pay nothing.
A native DirectML crash cannot be caught in-process, so neither the serialization nor the catchable-failure CPU fallback can save machines whose DML stack dies inside session creation itself. Reported on 0.1.572.9016: switching singers still crashed with createSession on the stack - the global lock rules out an overlap, the DML call dies alone. Write a marker file while DirectML work is in flight and clear it when the scope ends. If the marker is still present at the next start, the previous run died inside DirectML: switch the runner to CPU, persist the preference, log it and toast the user. The switch can be reverted in Preferences; a process that never touches DirectML never writes the marker.
This reverts commit acbf605.
Contributor
Author
|
Reverted acbf605 (the crash-marker auto-fallback) for now, keeping this PR focused on the serialization fix. Context: that commit was added after a single user still crashed on the serialized build. Their crash is inside |
The real-curve refresh and realtime pitch services introduced in openutau#2175 run on thread-pool threads and used to take the predictor reference outside SessionLock, then enter the lock to call Process. Between the getter returning and the lock being taken, a singer switch's FreeMemory (which disposes the predictors under SessionLock) could free the native session, leaving the background thread to call Process on a disposed handle. Take the getter result and use it entirely inside SessionLock in the three affected paths - LoadRenderedPitch (plain and retake) and LoadRenderedRealCurves - so the reference cannot be disposed while in use. Lock order stays DmlLock -> SessionLock, matching every other site.
DiffSingerScript's variance export took the predictor reference before entering both the DML scope and SessionLock, so a concurrent FreeMemory could dispose the native session before Process ran - the same pattern fixed for the other three call sites. Take the reference inside SessionLock. EnterDmlScope decided whether to engage the lock before acquiring it, reading the runner outside the lock. A CPU->DirectML switch landing while a thread waited on the lock would hand it a no-op scope that cannot be upgraded, leaving the following DirectML work outside DmlLock. Acquire the lock first, then re-check the runner under it. CPU fallback handling is unchanged; this PR only addresses races.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
The Windows DirectML build of ONNX Runtime kills the process natively when session creation,
Runand disposal overlap on the same device:0xC0000005, no managed exception, nothing in the log, only a Windows ".NET Runtime" 1026 event whose recorded address is the CLR's fail-fast site rather than the original fault. This serializes all three behindOnnx.DmlLock, engaged only for the DirectML runner, and falls back to CPU when the DML graph compiler rejects a model instead of failing the render.Why this is real (and not just the 1.24.4 issue)
Measured on an RTX 2070 (driver 32.0.16.1088) with onnxruntime 1.23.0 + DirectML 1.15.4, binaries SHA-256 identical to the official Windows package, 4 threads on one DML session:
Runon one session0xC0000005in secondsRunwhile another session is createdRunwhile a session is disposedCrash dumps from user reports land in this path both with 1.24.4 and with the pinned 1.23.0 (same signature, same DiffSinger session-creation stack), so the version pin alone does not cover machines where the overlap happens.
Related onnxruntime issues - a documented bug class that cannot reach the frozen package
This matches a class of DirectML EP bugs whose fixes land on main but can never ship in the
Microsoft.ML.OnnxRuntime.DirectMLpackage (frozen at 1.24.4; NuGet has 1.23.0 then 1.24.1-1.24.4 and nothing newer):E_INVALIDARG(80070057) at init orRunon the DML EP (Int64 indices in Gather/Scatter/TopK ops); DiffSinger acoustic models are transformer-based and fail exactly this way on some machines (catchable on some, native crash on others).OnnxTensorWrapperhanded bad data to kernels instead of failing (merged 2026-08-08).Unguarded paths closed
DiffSingerSinger.FreeMemorydisposed sessions while a render could be mid-Run-SessionLockonly covers the getters, not inference. This is the singer-switch crash.LoadRenderedPitch(Ctrl+R and the realtime pitch service),LoadRenderedRealCurvesand the script/export path shared no lock with the render path.Notes
Onnx.EnterDmlScope()is a no-op for the CPU runner; DirectML work is serialized (within DiffSinger renders it already was, via the static render lock).DmlLock->SessionLock; allSessionLockacquisition sites were audited.Failed to create a DirectML inference session, falling back to CPU.) when the DML EP cannot compile a model, so affected machines keep rendering.