An open DMF/MXL production lab — clone it, run one script, and put your own phone on air through an EBU MXL shared-memory domain in minutes. The full build is a complete live broadcast production (two real cameras, open guest contribution, file playout, layouts, graphics keyer, program audio, a native multiview), controllable by anyone with a browser.
On a fresh Ubuntu x86-64 VM with AVX (any Azure D-v5, AWS m5/m6i, GCP n2):
git clone https://github.com/guycochran/mxl-switcher
cd mxl-switcher
sudo scripts/quickstart.shYou get a cuttable, keyed, browser-watchable MXL switcher:
- ✓ Test-pattern source + file playout
- ✓ HTML5 lower-third graphics keyer
- ✓ WebRTC program output (watch in any browser)
- ✓ Two SRT contribution slots — point Larix / OBS / vMix at
srt://YOUR-IP:8890 - ✓ A browser switcher — preview/program with TAKE + a live multiview of every source
Then check it's healthy and cut between sources:
sudo scripts/mxl-doctor # containers · flow presence · program path (add --deep for unique-frame liveness)
curl -X POST -d '{"slot":1}' http://127.0.0.1:9604/pipeline/active-input # cut to the clipOr drive it from a browser — quickstart.sh launches a self-contained switcher UI
(web/local.html, localhost-only) at http://127.0.0.1:3100/:
arm a source on preview, press TAKE to cut it to program, watch every input in
the multiview. No prodbots backend, no external hosts — just the open /api/mxl/*
routes. See the control-plane doc.
Cold-clone verified Oct 2026: GCP 2m33s · AWS 3m59s · Azure 4m17s, each → ON AIR → cut landed.
Next: What did I just build? · Understand MXL · All the production findings · Full quickstart
Built by Office Hours Global ahead of IBC 2026 to show that the Dynamic Media Facility vision isn't just for broadcasters with NVIDIA partnerships — one person can stand up cloud shared-memory production in a weekend with the open tooling.
The switcher ran live through IBC 2026 — visitors cut the program, built layouts, and panned the real studio camera from a browser, no login. It's now offline: the cloud facility is deallocated when idle (that's the whole point — see the cost breakdown below). Everything it did is documented with screenshots throughout this README, and the full stack rebuilds from this repo in minutes (cold-clone measured ~3 min to on-air across three clouds) — see Build one yourself.
What it actually costs (pay-as-you-go, deallocate when idle): the full three-VM facility below is ≈ $2.11/hr — production D32s_v5 $1.54 + contribution D8s_v5 $0.38 + TAMS island D4s_v5 $0.19. The production VM runs the entire switcher (both cams, guests, layouts, keyer, encoder, multiview wall, thumbnails) at a steady load ~10–11 of 32 cores; a D16 ran it before the multiview and second camera existed, at ~14/16 — tight. Sizing details in Build one yourself.
Prefer to own it? Nothing here needs the cloud — the whole facility is Linux
- shared memory + software encode (no GPU, no SDI, no capture cards). A capable single-box studio is ~$2,500 of commodity hardware (a Ryzen 9 7950X has the required AVX-512). Full COTS spec + cloud-vs-own break-even: docs/ON-PREM-COTS.md.
📚 Learn MXL: we wrote down everything we learned running this live — indexed and readable at mxlswitcher.com:
- What is MXL? — the plain-English explainer + MXL vs NDI/ST 2110/SRT; MXL architecture — grains/flows/domains/ring-buffer/TAI, quoted from the SDK docs.
- Field Findings — the reader-lifecycle bug taxonomy, cross-host fabric numbers, the whole switcher verified end-to-end on AWS (cross-cloud, real cameras), the flow-stabilizer fix, and why rate metrics lie. Measured, not guessed.
- EBU DMF context · who's building on MXL (source-cited adoption tracker).
- MXL → TAMS recipe — live program to a clippable time-addressable store.
- Learn MXL — a guided reading path through the canonical EBU/AMWA/CBC sources.
- Glossary · The journey · Why shared memory is the future · run it on AWS or GCP.
This is the open core: the MXL/TAMS switcher — contribution ingest, the shared-memory domain wiring, the selector/keyer/audio writers, the browser switcher UI, multiview, and the TAMS clipper. Everything the standards story depends on is here and buildable.
Camera control (PTZ pan/tilt/zoom, presets, voice) is not in this repo. In
the live demo that panel is ProdBots, a separate
commercial project of ours. The switcher treats camera control as a replaceable
media function — the PTZ tab simply embeds an operator app at /ptz_op.html and
any VISCA / RTSP / SRT camera works as the source. So:
- The video path is fully open — a real camera flows through MXL end to end with nothing proprietary in the chain.
- To drive PTZ, plug in your own operator app at that embed point, or use
ProdBots. The camera ingest (
tools/cam_ingest.py) is here; the operator console is the piece you bring.
We kept it this way on purpose: it mirrors the MXL thesis (compose a plant from replaceable functions) and lets us open the switcher without giving away ProdBots. If you build from this repo, expect a complete switcher with a documented seam where camera control plugs in — not the turnkey voice-driven PTZ shown in the demo video.
Live PTZ camera (US studio) → SRT → MXL domain in Azure → HTML5 graphics keyed in-cloud →
WebRTC to the browser. The visitor pans the real camera from the right-hand console.
One click later: real episode playout through the same shared-memory chain, audio following the cut.
- Remote camera as a first-class MXL source. RTSP/SRT contribution lands in the domain
with its grain index aligned to locally-generated flows, so a remote camera is an
instantly-cuttable selector input (~30 ms cuts, graphics stay up). See
tools/cam_ingest.pyand the timing discussion indocs/FINDINGS.md. - Audio-follow-video assembled inside the domain by a fourth writer
(
tools/audio_pgm.py): episode audio on playout, bars-and-tone on pattern, silence on camera. - Sub-second glass-to-cloud-graphics latency, proven by the camera's burnt-in clock and the cloud-keyed clock reading the same second in a single frame.
- The DMF white paper's use cases in miniature: "minimise small-site footprint", "shipping compute", "outsourcing peaks", and (deallocate when idle) "supporting sustainability" — on a general-purpose VM with zero specialised hardware.
Three machines, two sites, one shared-memory domain:
┌─ US STUDIO ────────────────────────┐ ┌─ AZURE VM (D8s_v5, Ubuntu 24.04) ──────────────┐
│ │ │ │
│ PTZ camera (VISCA + RTSP 1080p60) │ │ mediamtx ◄─SRT─┐ ┌────────────────────────┐ │
│ │ RTSP │ │ │ RTSP │ │ MXL domain (/dev/shm) │ │
│ ▼ │ │ ▼ │ │ │ │
│ Relay/kiosk box (Linux) │ │ cam_ingest.py ─┼──► │ CAM Live │ │
│ ├─ ffmpeg re-encode 1080p30 ─────┼─SRT──┼─────────────────┘ │ Clip Video/Audio ◄──── │ │ file-player (episode)
│ ├─ Express backend: │ │ audio_pgm.py ─────► │ PGM Audio │ │
│ │ /api/mxl/* control proxy │ │ test-generator ───► │ TG Video/Audio │ │
│ └─ serves mxl.html (kiosk page) │ │ │ │ │
│ │ │ input-selector ◄──► │ Selector PGM │ │
└────────────────────────────────────┘ │ html5-keyer ◄─────► │ Keyer PGM │ │
│ mxl2webrtc ◄──────── └───────────────────────┘ │
Viewer's browser │ │ │
───────────────── │ ▼ │
page + WHEP signaling ◄──named──────┼── mediamtx :8889 │
(https, Cloudflare tunnel │ │
tunnel, stable URL) │ │
media (RTP) ◄───direct UDP 8189─────┼── (NSG: media port open; control ports locked) │
└─────────────────────────────────────────────────┘
Note the split delivery path: the player page and WHEP signaling ride a named Cloudflare tunnel (stable HTTPS URL, survives reboots via systemd), while WebRTC media flows directly to the VM's IP — so only the media port is internet-open and every control surface stays IP-locked.
The Fabrics API experiment landed. The same live program (camera + keyed graphics, produced on VM1) now crosses hosts through the MXL Fabrics API (TCP provider) and is served out of a different VM's shared memory:
The fabric-delivered program was served live from VM2's shared memory during the demo (now offline — the cluster is deallocated when idle):
VM1 "production" (D8s_v5) VM2 "fabric peer" (D8s_v5)
┌─────────────────────────┐ ┌──────────────────────────┐
│ camera/playout/patterns │ fabric │ domain (/dev/shm) │
│ selector → keyer → PGM ─┼─(tcp)─►│ "VM1 Program via │
│ domain (/dev/shm) │ ~1.3 │ Fabric" → mxl2webrtc ──┼─► public viewer
│ ▲ │ Gbps │ │ (named tunnel)
│ └── TG return leg ◄┼────────┼── test-generator │
└─────────────────────────┘ └──────────────────────────┘
│ fabric (tcp, via PUBLIC IP — sockaddr patched)
▼
VM3 "network island" (D4s_v5, its own unpeered VNet)
┌─────────────────────────┐
│ domain ← same program │ ← proves the cross-region/cross-cloud
│ mxl2webrtc viewer │ recipe: only the address changes
└─────────────────────────┘
Measured: 30 grains/s sustained (1080p30 v210, ~1.3 Gbps), zero drops over 100k+ grains, initiator ~10% of one core, receiving host ~0% CPU — remote writes really do land without target CPU involvement. With both hosts on NTP, the remote flow's head index matched the locally-generated flows, so the stock input selector briefly cut the fabric-delivered flow on air. Whole cluster at the time: ~$0.95/hr (the production VM has since grown to a D32 as the facility did — current numbers at the top). Details and gotchas (libfabric ≥ 2.x required, TargetInfo sockaddr patching for non-routed networks) in docs/FINDINGS.md.
Field notes for engineers: the honest production report — measured numbers, the v1.1.0 reader-lifecycle failure taxonomy, every workaround labeled and mapped to the SDK issue it stands in for, and the NMOS (IS-04/05/07/08) control-plane roadmap — is in docs/FIELD-NOTES.md.
Cloud-neutral by design: everything here runs on any cloud (or metal) with AVX and RAM for the shared-memory domain. A complete replay plan for AWS as a production deployment — a two-VPC design (production + a contribution DMZ for stranger-facing guest ingest), instance mapping, the data-transfer cost rules, native-S3 TAMS, and a two-day phased build — is in docs/AWS-BUILD-PLAN.md.
Verified on AWS (Oct 2026): this is no longer just a plan. The entire
switcher — not only the fabric leg — was stood up end-to-end on 2×
c5n.9xlarge in us-west-2, built from this repo, fed by the real studio
cameras (PTZ + Haivision Makito X4) over cross-WAN SRT. Measured: both cameras
30.00 fps, cross-host program 31.3 grains/s at 3.99 ms p50 over the
fabric (TCP) — reproducing the Azure numbers almost exactly, confirming the
whole facility is cloud-portable. Two honest caveats: the program shown is the
clean selector output (the graphics keyer was bypassed this run — one wiring fix
away, not a GPU limit), and the fabric leg is TCP (EFA/RDMA is blocked by
upstream bug jonasohland/mxl-fabrics-proxy#2). Full write-up + screenshots:
Field Findings §7.5
(docs/FINDINGS.md).
The DMF white paper lists the MXL↔TAMS relationship as an open topic and notes
that "MXL Grains can be grouped as TAMS Flow Segments." We built it: the
fabric-delivered program on VM2 is cut into 1-second segments and registered —
with capture-derived TAI timeranges — into a TAMS
store (Eyevinn tams-gateway +
MinIO + CouchDB) running on the island VM. The store's HLS endpoint plays any
timerange of the show while it's still being recorded — we pulled a frame
from 11½ minutes in the past whose in-picture cloud-keyed clock matched its
TAMS timerange to the second. Capture timing preserved from camera → fabric →
store → playback: live production into time-addressable storage, on the same
then-$1/hr cluster. Bridge code: a ~90-line shipper (segment → presigned PUT →
POST /flows/{id}/segments) plus one ffmpeg segmenter. Full replication
recipe — architecture, grain→segment mapping, and eight earned gotchas — in
docs/TAMS.md.
The demo ran a TAMS "time machine" side-by-side with the live program (now offline), with live clipping: you could scrub the show while it was still being recorded, mark IN/OUT anywhere in the archive and the clip already exists — it's just a timerange URL against the store (zero media copied). One click more muxes it to a take-home MP4 by segment concat: a 15s clip of the live show exports in ~1.5s. This is TAMS's reason to exist — live-to-clip while the event runs — working against an MXL production. There's also a shared clip bin: saving a clip creates a new TAMS flow that re-registers the same media objects under the clip's timerange — zero bytes copied, and the store refcounts so deleting a clip never touches the archive. Every visitor sees the same bin (it lives in the store, not the browser).
Next on the bench: a grain-native segmenter — reading v210 grains straight from the domain and grouping them into TAMS Flow Segments without the H.264 detour, i.e. the white paper's sentence implemented literally.
Runtime apps are the stock cbcrc/mxl-hands-on containers (test generator, file player, input selector, HTML5 keyer, mxl2webrtc), orchestrated with CLOUDflex-broadcast/easy-mxl. This repo adds the glue that made it a usable remote production:
| Piece | What it does |
|---|---|
tools/cam_ingest.py |
Low-latency RTSP→MXL ingest. 150 ms jitterbuffer, decode to v210, cadence-preserving PTS re-stamp so the remote camera's grain index aligns with local flows. This is what makes cross-flow cutting of a remote source work. |
tools/audio_pgm.py |
Program audio mixer (v2): GStreamer audiomixer over a live silence anchor + episode audio + per-guest audio, with per-input volume/mute driven from the kiosk's fader strip. Auto-adopts guest audio flows as contributors join. |
tools/layout_pgm.py |
2-up / PiP compositor as a switcher input: all six sources behind two selectors feeding a compositor; every layout change is a live pad-property flip, so the output flow is never recreated (wedge-proof). Five timestamp iterations documented in-file. |
tools/guest_ingest.py |
Open contribution: anyone's SRT (phone/OBS/vMix, any res/fps) conformed to 1080p30 v210, self-announcing so the slot goes live hands-off. |
tools/guest_audio.py |
Contributor audio companion — pulls the guest's audio across the VNet into its own MXL flow for the mixer. |
tools/mxl_thumbs.py |
Per-input JPEG thumbnails rendered from raw grains (no decode), with content-hash detection of repeat-wedged readers. |
tools/mxl_multiview.py |
The multiview wall as an MXL flow: 8 domain sources → CPU compositor (3×3, v210) → new "Multiview PGM" flow. Hardware-multiviewer architecture in software. |
tools/mv_encode.py |
The wall's single browser encoder: one x264 leg serves every viewer 8 sources. Pipeline is proven-verbatim — the header explains which parts are load-bearing. |
tools/grain_probe.py |
Health board: persistent readers on every flow reporting bps + unique-fps — the probe that catches repeat-last-grain wedges. |
tools/pgm_lite.py |
960×540 program copy (~0.33 Gbps) for fabric receivers behind GigE. |
tools/patch-target-ip.py |
The dmf-mxl#714 NAT workaround as a tool: rewrites the sockaddr inside a fabric TargetInfo to a public IP. |
scripts/mxl-doctor |
One front door for lab health. Read-only report by default; --deep adds unique-fps liveness (transient probe); heal selector|program|guest runs the live-facility auto-healers below. |
scripts/doctor.sh |
The read-only inspector mxl-doctor wraps: containers · flow presence · program path, backend-free. Still runnable directly. |
tools/guest-leg-doctor.sh |
15s two-end healer for the guest fabric legs (initiator wedges and the target frozen-slices state). mxl-doctor heal guest. |
tools/cam_relay.py |
The first-generation fixed-offset latency normalizer (superseded by cam_ingest.py, kept for the record — see FINDINGS). |
backend/mxl-routes.js |
Express routes proxying browser clicks to the pipeline APIs (cut / key / pattern / one-call cascade repair). |
web/mxl.html |
The kiosk page: WebRTC program feed + camera-control console + switcher bar. |
web/lower-third.html |
Transparent OGraf-style lower-third + live clock + bug, rendered by the CBC HTML5 keyer's CEF. |
scripts/bring-up-mxl.sh |
One command from cold VM to running demo: containers → writers → ingest → selector → keyer → encoder → tunnel → page. |
The kiosk switcher now has a Studio Cam 2 button: a static SDI camera feeding a Haivision Makito X4, which calls into the cloud VM directly over SRT (HEVC Main10 1080p30, 20 Mbps, deinterlaced on the encoder) and lands in the MXL domain as its own v210 flow on selector slot 3 — the classic broadcast-contribution pattern, terminated in shared memory instead of a hardware decoder. Cuts to and from it are the same ~25 ms selector cuts. See FINDINGS §8 for the decoder-threading and multi-slice gotchas this surfaced, and §9 for why the feed runs video-only and the cam is H.264 (both load-shedding decisions taken live while the demo was being shown).
Everything above still runs — and grew into this. A day of live debugging with real phone contributors, real visitors, and one full machine crash produced a production facility that self-heals around contributor churn:
CONTRIBUTION (anyone) PRODUCTION (VM1, D32s_v5, 32 cores)
┌────────────────────────────────┐ ┌───────────────────────────────────────────────┐
│ Studio PTZ cam 1080p60 H.264 ─┼─SRT──┼─► cam_ingest (frame-threaded decode, 60→30) │
│ (camera-native, ZERO local │ copy │ │
│ transcode, 20 ms latency) │ │ MXL domain (/dev/shm) — one memory, 14 flows│
│ SDI cam2 → Makito X4 ──────────┼─SRT──┼─► cam2_ingest │
└────────────────────────────────┘ │ │
┌────────────────────────────────┐ │ 7-input selector ◄─ CAM·CAM2·Clip·TG· │
│ Your phone: scan the Larix QR │ │ │ Guest1·Guest2·Layout │
│ on the kiosk page → app opens │ │ ▼ │
│ pre-configured → tap = on air ─┼─SRT─┐│ layout_pgm: 2-up / PiP compositor │
│ (OBS / vMix / ffmpeg too) │ ││ audiomixer: episode + guest audio, │
└────────────────────────────────┘ ││ per-input faders/mute on the kiosk │
VM2 "contribution host" (D8s_v5)││ │ │
┌────────────────────────────────┘│ ▼ │
│ mediamtx :8890 (world-open SRT) │ HTML5 keyer (lower-third + clock) │
│ └► guest_ingest ×2 → v210 flows│ │ │
│ │ ↑ guest-leg-doctor │ ▼ │
│ ▼ │ (15 s, heals both │ encoder → mediamtx ─┬► WebRTC viewers │
│ MXL FABRIC legs (tcp) ─────────┼─►(forced IDR per cut) └► SRT egress: │
│ + audio pulled over VNet │ │ `streamid=read:mxl2webrtc`│
└─────────────────────────────────┘ ▼ PGM fabric leg (1.3 Gbps v210) │
└────────┼──────────────────────────────────────┘
VM3 "island" (D4s_v5) ▼
┌──────────────────────────┐ VM2: segmenter+shipper ─► TAMS store on VM3
│ TAMS: gateway+MinIO+ │ (12 h scrub/clip archive, storyboard sprites,
│ CouchDB — 12 h archive │ instant zero-copy clips + MP4 export)
└──────────────────────────┘
What changed since the diagrams above:
- The camera path lost its last transcode. The PTZ's native 1080p60 H.264 is copy-remuxed to SRT (20 ms latency; measured 6 ms site→Azure RTT) and decoded once, in the domain — pans are now butter, and the "relay box" re-encode that quantized network jitter into judder is gone.
- Open contribution with scan-to-join. The kiosk's "Send us your feed" panel carries per-slot Larix QR codes: a phone scans, the app opens with our SRT pre-configured, and the slot self-attaches — contributor audio joins the program mixer automatically. Contributions land on a separate VM and cross to production over the MXL Fabrics API, so the switcher host never decodes a stranger's stream.
- Layouts as an input. A 2-up/PiP compositor writes a "Layout" flow that the selector cuts like any camera, with a live preview panel while adjusting.
- PVW/PGM buses with a live multiview (grain-rendered thumbnails on every
preview button) and clean cuts (a
/pipeline/keyframeendpoint patched into mxl2webrtc forces an IDR at every cut — no mid-GOP smear). - Crowd-proof delivery. One encode fans out to any audience; polling endpoints micro-cache (60 simultaneous thumbnail requests → 1 origin fetch); the whole kiosk — page, WHEP, thumbnails, TAMS playlists, presigned segments — is served same-origin, because venue networks that block unfamiliar domains are real.
- Self-healing everywhere, earned the hard way: watchdogs for wedged fabric readers (both failure species — silent and repeat-last-grain), frozen-slice targets, zombie relays, stale thumbnails, contributor churn. A full production-VM resize (8→16 cores) was recovered to on-air in ~7 minutes, mostly by systemd.
- Take our program with you: any receiver can pull the finished program
today via
srt://<vm>:8890?streamid=read:mxl2webrtc— or go MXL-native and receive raw grains over the fabric: docs/JONAS-FABRIC-HANDOFF.md.
The demo grew a front door (mxlswitcher.com, a
fresh control-room UI) and the piece every real switcher has: a multiview
wall built the way hardware does it. tools/mxl_multiview.py
reads all seven inputs plus the keyed program straight out of the domain,
composites a 3×3 wall on the CPU (v210 end-to-end, no source ever decoded),
and writes it back as a new flow — which
tools/mv_encode.py encodes once for every browser
viewer. Compose in the domain, encode once: eight sources cost one encoder.
The wall is itself a peer flow — probeable, recordable, even cuttable.
Cost: ~2 cores at 15 fps (~3.5 at 30, but see FINDINGS §10 for why the
production setting is 15: the multiview is the first thing to de-rate; the
program is the product). §10 also documents the three integration
landmines: aggregator EOS on ignore-inactive-pads, tsmux latency handling
outside gst-launch, and MediaMTX silently discarding non-1316-byte SRT
payloads (mpegtsmux alignment=7 is mandatory).
The facility now speaks the industry's control-plane standards, not just a private REST dialect.
NMOS (AMWA BCP-007-03). tools/nmos_node.py is a
stdlib-only IS-04 v1.3 Node that presents every MXL flow as a standard
NMOS Sender/Receiver — transport urn:x-nmos:transport:mxl, tagged with
mxl_domain_id/mxl_flow_id exactly as AMWA's published
BCP-007-03 "NMOS Support for MXL"
prescribes, with live PGM/PVW tally carried as grouphint tags. Registered
into a standard registry, any broadcast controller discovers the cloud
switcher — verified live against Bitfocus Buttons (its NMOS registry
client) and shown here straight from the registry's IS-04 Query API:
To our knowledge this is the first time a commercial broadcast controller has discovered an MXL production facility through a standard NMOS registry — three days after IBC 2026, on the published spec.
A hardware panel. companion-module-mxl-switcher/
is a Bitfocus Companion module: an Elgato Stream Deck XL cutting the
cloud switcher with real broadcast tally — the on-air source burns red, the
armed preview source glows green, feedless inputs dim — over a 1 s status
poll. Actions (cut, TAKE, keyer, warm-up, SuperSource layouts), feedbacks
(PGM/PVW tally, no-feed, keyer state), and drop-in presets. Confirmed
running on hardware (Companion 5.0.5). Setup — module and a zero-install
one-click page import — in
companion-module-mxl-switcher/TURNKEY-COMPANION.md.
The arc: private REST → physical panel with tally → NMOS-discoverable per published spec. That's the difference between a demo and a facility. Full control-plane roadmap (IS-05 routing, IS-07 tally, IS-08 audio) in docs/FIELD-NOTES.md.
Coming back after teardown, or rebuilding elsewhere? Start with docs/RESUME.md (warm-start map) and docs/FINDINGS.md §12–17 (the week's operational lessons: the reader-wedge taxonomy, the flow stabilizer, measurement doctrine, self-heal coordination, the runbook, and standards/NMOS).
Fastest path: docs/QUICKSTART.md — one script,
fresh Ubuntu VM → cuttable, keyed, browser-watchable MXL switcher in ~10
minutes (sudo scripts/quickstart.sh). Then come back here for the map.
Could you stand up the full facility from this repo? Yes — here's the honest map of what's here, what's external, and what you'd bring.
The path: (1) one Ubuntu VM (AVX required — see FINDINGS §11), Docker,
the stock cbcrc/mxl-hands-on
containers and easy-mxl
give you a working domain with test generator, file player, selector, keyer,
and WebRTC out — that alone is a cuttable "switcher" with zero code from us.
(2) Add our tools in this order as you need them: cam_ingest.py (a real
camera as an instantly-cuttable input — the cadence re-stamp in it is the
single most load-bearing idea in the repo), mxl_thumbs.py (preview
thumbnails), audio_pgm.py, guest_ingest.py + guest_audio.py (open
contribution), layout_pgm.py (2-up/PiP as an input), mxl_multiview.py +
mv_encode.py (the wall). (3) scripts/bring-up-mxl.sh
shows the exact assembly order, every pipeline body, and the supervisor
pattern that keeps it alive; backend/mxl-routes.js
is the browser→pipeline control layer; web/ is the UI.
Read docs/FINDINGS.md before debugging anything — every
multi-hour hole we fell into is labeled.
VM sizing, from measured load (1080p30 chain):
| Setup | VM | Steady load |
|---|---|---|
| Core switcher (1 cam, playout, TG, keyer, encoder) | D8s_v5 (8c) | ~4–6/8 |
| + 2nd cam, guests, layouts, audio mixer | D16s_v5 (16c) | ~14/16 — works, no headroom |
| Everything incl. multiview wall + visitors | D32s_v5 (32c) | ~10–11/32 |
Rules of thumb from FINDINGS §9: every mxlsrc reader busy-spins toward a
core even on silence; software HEVC decode is 2–3× H.264 (codec choice is a
scheduling decision); the CEF keyer needs protected headroom or the program
freezes first.
What is NOT in this repo (you bring your own): the general web backend
that hosts the routes (any Express app works — mxl-routes.js is the MXL
part); MediaMTX config (near-stock: SRT + WebRTC + RTSP enabled,
MTX_WEBRTCADDITIONALHOSTS=<public-ip>); Cloudflare tunnel / TLS fronting
(any reverse proxy works — one gotcha: CNAMEs to cfargotunnel.com must be
proxied or CNAME-flattening returns NODATA); API tokens and per-site IPs
(grep the scripts for obvious placeholders); and the cameras themselves —
any RTSP or SRT source works, the Makito/PTZ specifics are just our studio.
Multi-VM (fabric contribution, TAMS recording) is optional and documented in
docs/JONAS-FABRIC-HANDOFF.md and
docs/TAMS.md — start with one VM.
The interesting engineering is in docs/FINDINGS.md — including:
- Why a remote source stutters when cut in a grain-indexed shared-memory domain, and the cadence-preserving re-stamp that fixes it.
- A measured case of the white paper's open topic "robustness when Grains are missing" (single-frame content flashes from stale ring slots) and the write-ahead margin that prevents it.
- The GStreamer
rtspsrcdefault 2000 ms jitterbuffer hiding insideuridecodebin— where almost all our latency was. mxlsinkrequires an explicitflow-id; an empty one fails as a misleadingnot-negotiated (-4).- Reader wedge on flow recreation, and the cascade-repair pattern that recovers.
- Grain timestamps are ring addresses, not metadata: every writer needs an offset-locked, drift-bounded restamp — free-running counters and raw passthrough each fail in a distinct, delayed way.
- The compositor's pad scaling runs in its single aggregation thread and caps a 1080p30 chain below realtime — scale per-branch instead.
- Wedged readers have TWO species: silent, and repeating the last grain at full rate — the second passes every liveness check except content hashing.
- SRT latency is per-path physics: 20 ms is right for a 6 ms wired RTT and catastrophically wrong for cellular contributors (use ~1000 ms).
- EBU DMF / MXL community — the MXL SDK and the Dynamic Media Facility white paper.
- CBC/Radio-Canada — the mxl-hands-on media function containers.
- CLOUDflex-broadcast/easy-mxl — the control panel that makes the domain approachable.
MIT licensed. Not affiliated with EBU or CBC; all findings offered upstream with love.

