Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 5 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,11 @@

All notable changes to `@devalade/shipnode` will be documented here.

## [Unreleased]

### Added
- **Opt-in `watt` runtime (wattpm).** `.runtime('watt', { main: 'dist/server.js' })` runs the web app as worker threads sharing one port via `SO_REUSEPORT` instead of PM2 processes — no supervisor in the request path and no per-process V8 duplication. `instances` becomes the thread count; workers run as systemd units (`shipnode-<app>[-<colour>]`). Blue-green, `rollback`, `restart`, `stop`, `logs`, `env`, `deploy --watch`, `doctor`, `status`, `metrics` and the monitor all support it. PM2 stays the default. Your app must depend on `wattpm` and `@platformatic/node`; scaling past one worker needs Linux. See [ADR-0009](docs/adr/0009-watt-runtime.md).

## [3.2.0-beta.1] - 2026-09-18

### Added
Expand Down
19 changes: 19 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -85,6 +85,25 @@ export default shipnode

Each app gets its own release directory, Caddy site, PM2 ecosystem, and health check. Deploy everything with `shipnode deploy`, or target one app with `shipnode deploy --app api`.

### Opt-in: wattpm runtime

Run the web app as worker threads sharing one port (`SO_REUSEPORT`) instead of PM2 processes — no supervisor in the request path and no per-process V8 duplication. PM2 remains the default.

```ts
export default shipnode
.backend()
.ssh({ host: '1.2.3.4', user: 'deploy' })
.deployTo('/var/www/myapp')
.pm2('myapp', { instances: 4, maxMemory: '512M' }) // instances = worker threads
.port(3000)
.domain('api.example.com')
.runtime('watt', { main: 'dist/server.js' }) // file each thread loads; must listen on process.env.PORT
.worker({ name: 'mailer', command: 'node dist/worker.js' }) // runs as its own systemd unit
.build();
```

Shipnode installs `wattpm` and `@platformatic/node` for you if your app doesn't list them; add them to your own dependencies to pin the versions. Zero-downtime blue-green, `rollback`, `logs`, `stop` and `env` work as with PM2; supervision is systemd (`shipnode-<app>[-<colour>]`). Scaling past one worker needs Linux. See `docs/adr/0009-watt-runtime.md` for the trade-offs.

### Web + workers

A backend can run additional long-running processes alongside the web server. PM2 supervises all of them under one deployment.
Expand Down
24 changes: 24 additions & 0 deletions docs/adr/0009-watt-runtime.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,24 @@
# Opt-in `watt` runtime (wattpm) alongside PM2

PM2 stays the default. `runtime: 'watt'` (builder: `.runtime('watt', { main })`) runs the web app under wattpm instead: N worker threads in one process, each accepting straight from the kernel via `SO_REUSEPORT`, so no supervisor process sits in the request path and V8's startup state is not duplicated per worker.

## Mechanism

- **Web app** → wattpm. `instances` becomes the worker-thread count. Shipnode renders two per-release files into the app root (ADR-0001): `shipnode.watt.json` (runtime config, `port: "{PORT}"`) and `shipnode.platformatic.json` (tells `@platformatic/node` which file to load — `watt.main`). The runtime file has its own name and is passed with `wattpm start -c`: an app at path `.` would otherwise pick up a same-directory `watt.json` as *its own* config and fail to start.
- **Workers** (declared processes without a port) → one plain systemd unit each. wattpm supervises HTTP threads, not arbitrary commands, so workers keep an OS process and get systemd's restart/logging.
- **Supervision** → systemd units `shipnode-<namespace>[-<process>][-<colour>]`, each running a per-release launcher script (`shipnode-run-*.sh`). A script file sidesteps systemd's `%`/`$` expansion and keeps dotenv handling identical to PM2 (ADR-0003): the env file is parsed as data, never sourced.
- **Dependencies** → if the release does not already have `wattpm` and `@platformatic/node`, shipnode installs both at the version its rendered configs target (`WATT_VERSION`, also the schema version) with the app's own package manager, after the normal install and relink so nothing prunes them. An app that lists them itself keeps its own versions and nothing is installed. Nothing is installed globally, so every release carries the runtime it was deployed with and a rollback restores the matching one. The auto-installed packages are not in the app's lockfile; pin them there for fully reproducible production installs. A custom `module` is never auto-installed.

## Zero-downtime

Blue-green (ADR-0005) works unchanged in shape: the idle colour is a separate unit (`shipnode-api-green`) on its own port, health-checked before Caddy flips. Workers are a single unit set restarted in `afterHealthy`. Rollback flips Caddy to the previous colour after checking its unit is `active`. Adopting watt on a host that ran PM2 retires the PM2 process after the first successful flip (`pm2 delete <namespace>`, a no-op when PM2 was never there).

## Trade-offs

- **`SO_REUSEPORT` is Linux-only.** On other OSes wattpm forces a single worker for the entrypoint (observed on macOS). Production VPSes are Linux; local runs are not representative of scaling.
- **Thread isolation is weaker than processes.** A native crash or OOM takes down every worker in the process; `health.maxHeapUsed` (from `maxMemory`) recycles a bloated worker before that.
- **`shipnode restart` is an in-place systemd restart**, not PM2's rolling `reload`. Use `deploy` (blue-green) for a zero-drop roll.
- **Load spread is kernel-hashed**, so clients that share few source ports (e.g. a local proxy over loopback) can land unevenly. Validate with a real traffic split before rolling out widely.
- **Observation reads systemd, not PM2.** `status`, the monitor and `--json` sample `systemctl show` for each candidate unit. CPU is a rate but systemd only exposes a cumulative counter, so the observe script samples it twice ~200ms apart (adds ~0.2s per poll on hosts running watt apps). `metrics` shows a refreshing `systemctl status` since there is no `pm2 monit` equivalent.
- **Deploy health is systemd-aware.** A unit must be `active` with `NRestarts=0`; a crash loop under `Restart=always` would otherwise look healthy between crashes. Failures include the last journal lines.
- `harden`'s PM2 steps do not apply to watt apps: units are enabled at install time and start at boot without a saved process list.
214 changes: 214 additions & 0 deletions docs/superpowers/specs/2026-08-30-monitor-redesign-design.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,214 @@
# Monitor redesign: a shared observation layer

## Problem

`shipnode monitor` is a single-server Ink TUI with its data layer welded to React.
Four things are wrong with it at once:

1. **Fleet-blind.** `resolveMonitorSession` narrows to one replica because "the
monitor holds one live connection". The one failure mode v3 introduced — a
partly-rolled fleet — is therefore invisible in the live view and visible only
in `status`, a one-shot text command.
2. **Overlapping surfaces.** `monitor`, `status`, and `logs` each reach the server
their own way, so the same facts are gathered by three code paths.
3. **No machine-readable output.** No `--once`, no `--json`. Nothing to pipe into
a script, a CI gate, or an alert.
4. **Layout hand-tuned to a 24-row terminal.** `releaseBoxExtra = min(6, rows-30)`,
a fixed 60/40 split, fixed panel heights, and a 389-line `App.tsx` that owns
keybindings, overlay routing, action dispatch, and layout at once.

## Scope

In scope: a shared observation layer, fleet visibility, `--once`/`--json`, and
rebuilding `status` and `logs` on the shared layer (they may change shape).

Out of scope for this spec: the TUI layout, the two views, and the keymap. Those
are designed separately once the foundation lands.

`shipnode metrics` stays as it is. It hands the terminal to PM2's own `pm2 monit`
over `ssh -t`, which is a deliberate escape hatch to the unfiltered truth when
shipnode's view and reality disagree — `monitor` filters PM2 to the app's
namespace and structurally cannot show that. It gains `--on` (so a fleet app is
targeted deliberately rather than by whichever host `getServerTargetResult`
returns) and honest help text; nothing else.

## Decision: collection is server-shaped, presentation pivots to app-shaped

Three models were considered.

- **App-shaped** — the subject is one app across its replicas. Fits v3's model
(`on`, roll, convergence are all app-scoped), but duplicates and mis-attributes
host-level facts: a box hosting three apps reports its load three times.
- **Server-shaped** — the subject is one server, N apps. Closest to today, and
system stats and accessories land naturally, but an app can only be compared
across replicas by switching servers and remembering.
- **Two-level (chosen)** — collect server-shaped, pivot to app-shaped for
presentation.

Collection stays server-shaped, so nothing is polled twice and host-level facts
have one home. The pivot is a pure function over snapshots — no I/O, trivially
testable — and it is what lets `status` be a renderer rather than a second data
path.

## Architecture

```
domain/observe/
snapshot.ts ServerSnapshot / AppSnapshot / FleetView types
parse.ts section parsers (moved from cli/monitor/state.ts)
script.ts buildObserveScript — composite shell, section-selectable
collector.ts MetricsCollector — one server, N apps -> ServerSnapshot
pivot.ts pivotByApp(ServerSnapshot[]) -> FleetView[]

services/observe/
session.ts ObserveSession — N collectors, scheduling, history, events
history.ts MetricsHistory

cli/monitor/ Ink TUI: subscribes to an ObserveSession
cli/commands/ status / logs / monitor — three presenters, one session
```

**Boundary:** nothing under `domain/observe/` or `services/observe/` imports
React, Ink, or chalk. Today `use-monitor-data.ts` mixes polling, alerting, and
`chalk.red(...)` event strings — presentation decisions baked into the data
layer. Events become typed values; colour is chosen at render.

This follows the existing `domain/deploy` + `services` split rather than
inventing a new shape.

## Types

Today's `MetricsSnapshot` conflates three scopes: host-level (`system`),
app-level (`processes`, `releases`, `health`, `caddy`), and server-level
(`accessories`, `deployLock`). The split follows those seams.

```ts
interface ServerSnapshot {
server: string; // the host string — identity everywhere, per ADR-0008
timestamp: string;
system: SystemInfo; // collected once, not once per app
accessories?: AccessoryInfo[];
deployLock: DeployLockInfo | null;
apps: AppSnapshot[];
error?: string; // whole-server failure: unreachable, timeout
}

interface AppSnapshot {
app: string;
appType: 'backend' | 'frontend';
processes: ProcessInfo[];
currentRelease: string | null;
releases: ReleaseRecord[];
health?: HealthInfo;
caddy?: CaddyInfo;
error?: string; // per-app failure: PM2 down for this namespace
}

interface FleetView {
app: string;
appType: 'backend' | 'frontend';
replicas: Array<{
server: string;
snapshot: AppSnapshot;
system: SystemInfo;
reachable: boolean;
}>;
convergence: FleetConvergence;
}

function pivotByApp(snapshots: ServerSnapshot[]): FleetView[];
```

`SystemInfo`, `ProcessInfo`, `HealthInfo`, `AccessoryInfo`, `CaddyInfo`, and
`ReleaseRecord` move over unchanged — they are already well-shaped.

`deployLock` sits on the server because that is where the lock lives
(`{remotePath}/.shipnode/deploy.lock`), which today's per-app snapshot quietly
misrepresents.

`assembleConvergence` is reused as-is: `assessConvergence` already takes
`ReplicaObservation[]` (`{ server, release }`), so the pivot feeds it directly.
No new convergence logic, and `status`'s fleet reporting keeps working through a
path now shared with the live view.

## Data flow

```
ObserveSession.tick()
-> for each server, bounded-parallel: MetricsCollector.collect({ apps, sections })
-> ServerSnapshot[] (partial: an unreachable server yields error, not a throw)
-> history.push per (server, app)
-> health streaks + typed events
-> pivotByApp() -> FleetView[]
-> notify subscribers with { servers, fleets, events, lastUpdate }
```

Both shapes are published every tick. The TUI's fleet view reads `fleets`, its
server view reads `servers`, `--json` emits `servers` (the collected truth) with
`fleets` derivable, and `status` renders the pivot. No mode re-polls.

## Collection decisions

Carried forward from today's poller:

- One SSH round trip per server per poll. The composite shell script with
`@@SHIPNODE:<name>@@` section markers is the good part of the existing poller
and survives intact — it now emits N apps per script instead of one, since
`getAppsForServer` already says what is on the box.
- Accessories stay sampled on a slower cadence (~10s) than the main poll, with
the previous value carried forward. `docker inspect` is the expensive section.
- The health probe stays bounded by `--max-time` derived from the interval, so a
hanging probe cannot stretch a tick.
- Each section carries its own fallback, so one failing probe cannot blank the
rest of the snapshot.

New:

- Parallel collection across servers is bounded at 4 concurrent connections; the
rest queue within the tick. A 12-host fleet must not open 12 SSH connections.
- A tick that cannot finish within the interval is skipped rather than
overlapped — today's `inFlightRef` guard, generalised per server.

## Error handling

Per AGENTS.md, `better-result` with typed `TaggedError` classes for expected
domain failures; exceptions only for programmer defects and adapter-boundary
infrastructure failures.

Failure is per-scope and never aborts a tick:

- A server that is unreachable or times out yields a `ServerSnapshot` with
`error` set and no app data. Other servers still report.
- An app whose PM2 section fails yields an `AppSnapshot` with `error` set. Other
apps on the same server still report.
- `assessConvergence` over a fleet with an unreachable replica reports the skew
it can see and names the replica it could not reach, rather than claiming
convergence from a partial observation. This matches the existing rule in
`reportFleetConvergence` that a narrowed run reports nothing rather than
reporting "converged" from one observation.

## Testing

`FakeRemoteExecutor` (`tests/testing/fake-executor.ts`) already supports
predicate-matched canned responses and command-history assertions, which covers
the collector without a network.

- `parse.ts` — the existing parser tests in `tests/unit/monitor.test.ts` move
across unchanged; they are good and already cover malformed input.
- `script.ts` — section selection: which sections appear for backend vs frontend,
with and without health checks and accessories, and N apps in one script.
- `collector.ts` — a canned multi-section stdout parses into a `ServerSnapshot`;
a failing section degrades that section only; a timeout yields a server-level
error.
- `pivot.ts` — pure function, so table-driven: converged fleet, skewed fleet,
unreachable replica, single-server app, app absent from one server.
- `session.ts` — with fake collectors and fake timers: concurrency cap honoured,
overlapping ticks skipped, accessory cadence, health-streak transitions in both
directions, subscriber notification.

## Migration

The existing `cli/monitor/` TUI keeps working throughout: `state.ts` and
`poller.ts` become thin re-exports over `domain/observe/` until the TUI is
rewritten in the follow-up spec. No user-visible change lands until the surfaces
are rebuilt.
14 changes: 12 additions & 2 deletions src/cli/commands/config.ts
Original file line number Diff line number Diff line change
Expand Up @@ -15,17 +15,27 @@ function showApp(app: ShipnodeApp, nodeVersion: string): void {
['keepReleases', String(app.keepReleases)],
]);

if (app.runtime === 'watt' && app.watt) {
ui.section('Runtime: watt (wattpm)', [
['main', app.watt.main],
['module', app.watt.module ?? '@platformatic/node'],
['maxHeapUsed', app.watt.maxHeapUsed ?? '(from maxMemory)'],
]);
}

if (app.pm2) {
const watt = app.runtime === 'watt';
for (const pm2App of app.pm2.apps) {
const rows: [string, string][] = [['name', pm2App.name]];
if (pm2App.command) rows.push(['command', pm2App.command]);
if (pm2App.port !== undefined) rows.push(['port', String(pm2App.port)]);
if (pm2App.instances !== undefined) rows.push(['instances', String(pm2App.instances)]);
if (pm2App.instances !== undefined) rows.push([watt && pm2App.port !== undefined ? 'threads' : 'instances', String(pm2App.instances)]);
if (pm2App.maxMemory !== undefined) rows.push(['maxMemory', pm2App.maxMemory]);
if (pm2App.env) {
for (const [k, v] of Object.entries(pm2App.env)) rows.push([`env.${k}`, v]);
}
ui.section(pm2App.port !== undefined ? `PM2 process: ${pm2App.name} (web)` : `PM2 process: ${pm2App.name}`, rows);
const label = watt ? 'Process' : 'PM2 process';
ui.section(pm2App.port !== undefined ? `${label}: ${pm2App.name} (web)` : `${label}: ${pm2App.name}`, rows);
}
}

Expand Down
8 changes: 7 additions & 1 deletion src/cli/commands/doctor.ts
Original file line number Diff line number Diff line change
Expand Up @@ -42,6 +42,10 @@ function checkLocal(config: ShipnodeConfig): void {
issues.push('PM2 apps are not configured for backend app');
}

for (const app of config.apps) {
if (app.runtime === 'watt' && !app.watt?.main) issues.push(`App '${app.name}' uses runtime 'watt' but has no watt.main`);
}

if (issues.length === 0) {
ui.success('Local configuration looks good');
} else {
Expand All @@ -59,7 +63,8 @@ export async function checkRemote(
ui.info('Checking remote server...');

const needsNode = config.apps.length > 0;
const needsPm2 = config.apps.some((app) => app.appType === 'backend' && app.pm2);
const needsPm2 = config.apps.some((app) => app.appType === 'backend' && app.pm2 && app.runtime !== 'watt');
const needsSystemd = config.apps.some((app) => app.appType === 'backend' && app.runtime === 'watt');
const needsCaddy = config.apps.some((app) => app.domain);
const needsDocker = Object.keys(config.accessories ?? {}).length > 0;

Expand All @@ -68,6 +73,7 @@ export async function checkRemote(
const checks = [
...(needsNode ? [{ name: 'Node', cmd: `${mise}; mise exec "node@${nodeVersion}" -- node --version` }] : []),
...(needsPm2 ? [{ name: 'PM2', cmd: `${mise}; mise exec "node@${nodeVersion}" -- pm2 --version` }] : []),
...(needsSystemd ? [{ name: 'systemd', cmd: 'systemctl --version' }] : []),
...(needsCaddy ? [{ name: 'Caddy', cmd: 'caddy version' }] : []),
...(needsDocker ? [{ name: 'Docker', cmd: 'docker --version' }] : []),
{ name: 'rsync', cmd: 'rsync --version' },
Expand Down
9 changes: 9 additions & 0 deletions src/cli/commands/env.ts
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,7 @@ import { resolve } from 'path';
import { runRemoteCommandForTargets } from '../runner.js';
import { ui } from '../ui.js';
import { getDeploymentName } from '../../domain/pm2/apps.js';
import { isWatt, resolveWattUnits, restartUnitCommand } from '../../domain/runtime/watt.js';
import type { RemoteExecutor } from '../../domain/remote/executor.js';

function shellSingleQuote(value: string): string {
Expand Down Expand Up @@ -82,6 +83,14 @@ export async function cmdEnv(
continue;
}

if (isWatt(app)) {
const units = await resolveWattUnits(executor, appPath, app, { colors: 'active' });
ui.info(`Restarting '${app.name}' (${units.join(', ')}) to pick up environment variables...`);
for (const unit of units) await executor.execOrThrow(restartUnitCommand(unit));
ui.success(`'${app.name}' restarted with new environment variables`);
continue;
}

const nodeVersion = config.nodeVersion === 'lts' ? '24' : config.nodeVersion;
const mise = `export PATH="$HOME/.local/bin:$HOME/.local/share/mise/shims:$PATH"`;
const checkResult = await executor.exec(
Expand Down
Loading
Loading