Keep nmshd alive when a resize races a shell's exit (#199) - #205
Merged
Merged
Conversation
node-pty closes a PTY's descriptor when its socket closes, which can be before it reports the shell's exit, and a frontend can always send a resize just before it hears of that exit. pty.resize() then threw 'ioctl(2) failed, EBADF' synchronously. In the session service that happened inside the socket data handler (and the attach redraw timer), so one late resize could end nmshd and every live session it owns. ShellSession now tracks its own lifecycle: resizes after the shell has exited are no-ops, and only the closed-descriptor error is classified as that teardown race; any other failure still throws. The service routes every resize through one guarded path that reports an unexpected failure to the client that caused it instead of letting it end the service. Regression tests reproduce the EBADF deterministically on the old code: resize after exit and after kill, a burst of resizes racing the exit, a service that must survive late resizes while another session keeps working and resizing, and an unexpected resize error reported as a protocol error.
After service.close() the shells exit asynchronously and write their spool, recreating the runtime directory the test had just removed. The tests now wait for those shells before cleaning up, so no directories are left in TMPDIR.
This was referenced Sep 29, 2026
Owner
Author
|
CI at current head |
raiseCatError
added this pull request to stack #210
September 29, 2026 10:19
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #199
Root cause
UnixTerminal.resize()callspty.resize(fd, …)synchronously with no state check._close()), which can happen before it emitsexit. A frontend can also send a resize just before it hears of the exit.ioctl(2) failed, EBADF. This reproduces deterministically.Production paths that could hit it:
resizemessages, handled inside the socketdatahandler: the connection keeps itsownedsession after the shell exits, so a synchronous throw there is uncaught and ends nmshd with every live session.sessions.has()stays true between the fd closing andexit.Conclusion: resize after close is an expected teardown race, not a separate ordering bug.
Fix
ShellSession: it tracksexited(set onexit), and resizes after that are no-ops. OnlyEBADF(isClosedPtyError) is classified as the race; it marks the session closed and is ignored. Any other error still throws.SessionService: every resize (client messages and both redraw steps) goes through one guardedresize(). An unexpected failure is reported to the client that caused it as{type: 'error', code: 'resize'}instead of escaping and ending the service.Tests (
tests/ptyResizeRace.test.ts)kill.EBADFclassified; other errors still thrown.stty size→33x111).EBADF, deterministically rather than under load.git diff --checkare clean. The full suite passed 604/604 twice.CI
This repository's CI runs automatically only on PRs into
dev/master, so a stacked PR has no automatic checks. It was run manually on this branch withworkflow_dispatch: run 36518253918, Node 22 and Node 26 both passed.