Repository navigation
DDP Compression Extension — Gauging Upstream Interest #5810
Description
Activity
- added a commit that references this issue
on Aug 18, 2026 - addedAIPartly generated by an AI. Make sure that the contributor fully understands the code!Partly generated by an AI. Make sure that the contributor fully understands the code!
on Aug 19, 2026 @aenertia thanks, I think we'd be generally interested in a "compressed DDP" feature.
We would need to look at some details, once we see your proposed code
- "abusing" an unused DDP header flag has a risk of becoming incompatible with future DDP spec updates. It might be better if we define our own "DDP compressed" protocol ID, and leave "DDP standard" as it is.
- details of the RLE algo - i don't understand yet how you can do all compression in-place. The worst-case for RLE is when the encoder needs to duplicate the RLE "magic marker" too often. In this case, the compressed stream can become bigger than the original input. Likewise, if the stream that gets compressed temporarily requires more bytes than the uncompressed equivalent, you cannot do this in-place because the compressed data will overwrite some not-yet-compressed input data.
- It might be better to perform RGB-based (3 color streams) RLE, instead of byte-based RLE
- Possibility of adding the compression feature to DDP-over-ws, including our JS frontend code
Line 216 in 42b2399
function sendDDP(ws, start, len, colors, isESP8266=false) {
interesting references:
- ITU recommendation for color run-length encoding: https://www.itu.int/rec/T-REC-T.45-200002-I/en
- Wikipedia: RLE https://en.wikipedia.org/wiki/Run-length_encoding#Variants
@softhack007 thanks for looking at this -- good to know there's interest. These are all solid points and I think they should be worked into the spec doc rather than decided upfront. Working through each one:
1. Reserved-bit vs own protocol ID
Fair point about borrowing reserved bits. I've been looking at this more and there are actually a few options worth considering:
- Current approach: flags byte bit 5 (0x20). Simple, but as you say, borrowing against future spec changes.
- Customer-defined data type: byte 2 has a
Cbit (bit 7) explicitly defined in the spec as "1 for Customer defined". SodataType = 0x80 | original_typewould signal "this is a vendor extension" using a mechanism the spec already provides for exactly this purpose. - Separate protocol discriminator: a new port or a different identification scheme entirely.
The
Cbit approach is interesting because it's the one place the DDP spec explicitly invites extensions. The tradeoff is that compression is really a transport concern rather than a data type concern, but the spec doesn't have a dedicated extension mechanism so you use what's available.I'll add a section to the spec doc covering all three options with pros/cons so we can pick the right one before any code lands upstream. No point locking in a wire format prematurely.
2. In-place RLE / worst-case expansion
You're right about the general principle -- RLE can expand (worst case ~0.8% for PackBits). The current implementation uses separate buffers on the encoder side and a streaming decoder on the receiver side, so there's no in-place encoding anywhere. But your question raises something worth documenting properly in the spec: what guarantees does a receiver need about maximum expansion, and what MUST a sender do when compression doesn't help?
Current behaviour:
rle_encode_adaptive()compares compressed vs raw size and returnsDDP_COMP_TYPE_NONEif compressed >= raw, so expansion never hits the wire. The decoder (RLEDecoder) is a streaming state machine that reads one byte at a time from the packet buffer -- no intermediate allocation, no in-place decode. But these are implementation details that should be spec-level requirements ("a sender MUST fall back to uncompressed if the encoded output exceeds the raw payload size").I'll tighten up section 4 of the spec doc with explicit worst-case bounds and sender/receiver obligations. The code is in
ddp_compress.h(~150 lines) if you want to look at the actual encoder/decoder.3. RGB-tuple vs byte-level RLE
This is worth exploring properly. The current codec does byte-level RLE on XOR deltas, which works well for the sparse-change case (unchanged pixels become zero-byte runs regardless of channel boundaries). But there are at least three approaches worth benchmarking:
- Byte-level RLE (current): simple, works well with XOR delta, 300 unchanged RGB pixels = 900 zero bytes = one 2-byte RLE token. Partially-changed pixels still compress well per-channel.
- RGB-tuple RLE: each run unit is 3 bytes (or 4 for RGBW). Identical compression for unchanged pixels, but loses per-channel run-length on partial changes.
- Separate colour planes (closer to the ITU T.45 reference you linked): split into R/G/B planes, RLE each independently. Better for smooth gradients. But needs three decode passes plus reassembly on the receiver, which breaks the streaming decoder model and requires buffering.
The codec is already structured for multiple compression types in the upper nibble (delta+RLE, RLE-only, transform). Adding an RGB-tuple or plane-split variant would be a new type code -- the framework supports it, it's just a question of what performs best on real LED animation data.
I'll add a section to the spec covering these variants with the tradeoff analysis, and I can run benchmarks on real animation patterns (rainbow cycle, sparse twinkle, gradient fade, chase) to get actual numbers rather than theoretical arguments. Would be good to know what kind of patterns matter most for the upstream use case.
4. DDP-over-WebSocket + JS frontend
Agreed this would be valuable. The codec is pure byte manipulation, so a JS PackBits encoder/decoder is straightforward -- maybe 30 lines each. The receiver side already handles compressed DDP regardless of transport (same
handleDDPPacket()path for UDP and WS).One thing worth working through in the spec: WebSocket already has per-message compression (permessage-deflate, RFC 7692). If that's enabled, layering PackBits on top is mostly redundant. The delta framing (only send changed pixels) might be more useful over WS than the RLE encoding itself. Or it might make sense to define a WS-specific mode that does delta-only without RLE, relying on permessage-deflate for the entropy removal. Worth thinking through before writing code.
I'll add a transport-specific section to the spec covering the WS path, including interaction with permessage-deflate and what the JS implementation surface looks like. Probably makes sense as a follow-on PR after the core codec, since it's a different review surface (JS in
common.js).
I'll update the spec doc with sections covering all of these open questions -- wire format options, sender/receiver obligations, compression variant benchmarks, and WS transport considerations. Better to get the spec right before submitting code PRs.
- added a commit that references this issue
on Aug 20, 2026 WebSocket already has per-message compression
I dont think this is implemented.
regarding compression: another approach would be to use what GIF does: use a palette look-up, this would only be helpful on larger setups though and can be quite lossy (and maybe too computationally expensive).
Another addition which would be especially helpful when using RGB split planes: use a very simple adaptive color depth reduction. While this is also lossy it would be quite effective in video streams: since WLED uses gamma correction and colors are compressed especially at the higher end, the loss would only be visible when using it side-by-side with an uncompressed section (i.e. a matrix consisting of multiple ESP's) - and maybe not even then. What I am thinking of is something like this:
if (color > 196) color = color &0xF8; else if (color > 128) color = color &0xFC; else if (color > 64) color = color &0xFE;
so stripping out the LSB's as those small variations get compressed by gamma and those variations are also almost indistinguishable even if not using gamma. On the lower end, no compression is used to preserve color gradients and "fading trails".
I also did run experiments at one point using RGB565 encoding and found that the color loss does not cause significant banding but those tests were brief based on rainbow patterns, I would expect fading trails to be affected.Yeah, permessage-deflate isn't in the ESP32 WS implementation -- found that out pretty quickly.
Been running tests over the weekend with planar and channel variants alongside the byte-RLE already in the public repo. A bit more overhead but worth it for several input patterns without blowing heap.
The bigger issue is WS over PPP -- layered framing adds real overhead. It works but sustainable FPS is noticeably worse than the UDP path. Same on WiFi but less visible. On beefier variants with PSRAM, ROM
tinflfor decompress-only might be worth exploring.Numbers from extended testing on M5StickC (ESP32-PICO-D4, 40x80 TFT, PPP @ 1.5Mbaud):
Transport Encoding FPS Notes UDP / WiFi raw (rainbow) 661 link not the bottleneck UDP / WiFi raw (pulse) 997 UDP / WiFi effective (TFT active) 83 stable 30s UDP / PPP raw (rainbow) 41.8 link-limited UDP / PPP compressed (twinkle) 435 compression hides the link cost UDP / PPP effective (TFT active) 76 stable 30s WS / WiFi ceiling (256px seg) 150 WS / WiFi effective (TFT active) 53 stable 30s WS / PPP compressed (256px seg) 150 at ceiling WS / PPP effective (TFT active) 51 died at 5.4s -- WS+PPP backpressure WS+PPP is the bad combination -- TCP framing overhead compounds with serial link serialisation and backpressure kills it under sustained load. UDP+PPP+compression is surprisingly good; on sparse patterns the compression ratio offsets the link cost almost entirely.
The TFT active bits are to do with the eSPI as Matrix bus which I am co-testing - effectively new bus-type use your SPI screen as a pixel segment; has to careful manage dma ping-pong . But as part of that I have implemented a skip segment/DDP eligible segment schema which is actually super useful generally. My primary use case is to use eSPI as a 'preview pane' and then Copy to Effect to actual real matrix segments.
Also for the new esps3.1 and the P4 - they have jpeg decoding on chip. In theory using DDP transport with compression extension alows you to do realtime video elegantly to i.e hub75 panels walls etc.
@coderabbitai the proposed (experimental, see https://github.com/aenertia/WLED/blob/dev/ppp-wifi/wled00/e131.cpp) code seems to compress only single DDP packets, but not a full frame. I'm wondering if this would create any benefit in a normal WiFi or ethernet setup? A single DDP message is already size-optimised to fit to the network MTU size. If the message get smaller than one MTU, this could cause additional message delays. What do you think?
coderabbitai commented
on Aug 23, 2026 coderabbitaiboton Aug 23, 2026 – with coderabbitaiContributorMore actions🧩 Analysis chain
🏁 Script executed:
#!/bin/bash set -euo pipefail # Locate the local source file, if the checked-out repository contains it. if fd -a '^e131\.cpp$' . | grep -q .; then file=$(fd -a '^e131\.cpp$' . | head -n1) echo "LOCAL_FILE=$file" ast-grep outline "$file" --items all --view expanded || true rg -n -C 4 -i 'ddp|compress|rle|mtu|packet|payload' "$file" || true fi # Inspect the exact public branch referenced in the issue discussion. curl -fsSL \ 'https://raw.githubusercontent.com/aenertia/WLED/dev/ppp-wifi/wled00/e131.cpp' \ -o /tmp/aenertia-e131.cpp echo 'REMOTE_SYMBOLS:' rg -n -C 5 -i 'ddp|compress|rle|mtu|packet|payload' /tmp/aenertia-e131.cpp echo 'REMOTE_RELEVANT_RANGES:' nl -ba /tmp/aenertia-e131.cpp | sed -n '1,260p'
Length of output: 38848
🏁 Script executed:
#!/bin/bash set -euo pipefail tmpdir=$(mktemp -d) trap 'rm -rf "$tmpdir"' EXIT git clone --quiet --depth 1 --branch dev/ppp-wifi \ https://github.com/aenertia/WLED.git "$tmpdir/WLED" cd "$tmpdir/WLED" echo 'COMPRESSION_FILES:' fd -a -t f 'ddp|compress|e131' . | sort echo 'ENCODER_AND_DDP_SEND_REFERENCES:' rg -n -i -C 5 \ 'rle_encode_adaptive|DDP_COMP_TYPE|DDP_FLAGS_COMPRESSED|ddp_compress|beginPacket|channelOffset|DDP_HEADER_LEN' \ -g '*.{c,cc,cpp,cxx,h,hpp,ino}' . echo 'DDP_COMPRESS_HEADER:' if test -f wled00/ddp_compress.h; then sed -n '1,280p' wled00/ddp_compress.h fi
Length of output: 52510
@softhack007, yes. The current experimental bridge code compresses each DDP packet independently.It accepts at most
DDP_CHANNELS_PER_PACKETbytes, which is 1440 bytes. It replaces that packet payload with its compressed form. It updatesdataLento the compressed length. It does not combine several DDP packets into one compressed frame.For a normal Ethernet or good WiFi link, this usually gives little benefit:
- The sender still sends the same number of UDP datagrams.
- Each datagram still has IP, UDP, DDP, MAC, and WiFi MAC overhead.
- The receiver still processes the same number of packet callbacks.
- The
PUSHpacket still arrives after the same sequence of frame packets. - The frame completion latency therefore does not decrease much.
A smaller UDP datagram does not wait until it reaches the MTU. UDP has no equivalent of TCP Nagle coalescing. The sender sends the smaller datagram when it calls
endPacket(). Therefore, reduced payload size does not itself add a message delay.The benefit is the reduced serialization time and, on WiFi, reduced airtime. This can matter on a slow or congested link. It matters much more for PPP and similar serial links. It can also help low-rate WiFi, weak-signal WiFi, or a busy shared RF channel. It is almost irrelevant on 100 Mbps or 1 Gbps Ethernet unless the sender or receiver is already bandwidth-limited.
There is also a cost. Per-packet RLE has less opportunity to find runs across the 1440-byte packet boundary. It adds encoder CPU time for every packet. For incompressible animation data, the code correctly falls back to raw data, but it still paid the attempted encoding cost.
I think the primary value of frame-level compression would be different:
- Compress the complete frame, or a larger frame region.
- Split the compressed byte stream into MTU-safe DDP packets.
- Reassemble the stream before decoding, or define a streaming decoder with explicit frame state.
- Render only after the complete compressed frame is available.
That design can reduce both byte count and, for sufficiently compressible frames, packet count. It has higher buffering requirements. It also makes packet loss and frame resynchronization more complex.
One protocol detail needs attention in the current packet-local format. The compressed DDP
dataLenis the compressed length. The receiver decodes until the compressed stream ends. The format does not carry the expected decoded length for that packet. A final format should include an explicit uncompressed byte count, or another unambiguous range length. The receiver should enforce that count. This prevents a malformed or truncated stream from writing past the intended packet range.So, I agree with the concern: packet-local compression is mainly useful for constrained transports. For ordinary WiFi and Ethernet, it is not likely to improve FPS or frame latency unless measured data shows the network link is the bottleneck.
You are interacting with an AI system.
Reacted by Frank Möhle@aenertia @DedeHai considering our previous ideas, I think we definitely must have a new protocol ID for "compressed DDP". It could be integrated into the normal DDP encoder/decoder source code, similar to the way that E1.31 and Art-Net share a lot of code.
Also we need to spend more thoughts on the compression methods (even jpeg and gzip could be an option) and stream format. I think that RLE compression has a benefit on encoding speed, because it's really a simple and straightforward method. "2D" delta-compression looks appealing, however it requires to keep a copy of the last pixels frame. If we have a delta-compression scheme, the next question is how to implement stream restart points, so the receiver gets a chance to recover from lost or re-ordered message chunks.
We don't need to stuff everything into to existing DDP format, but we can define our own format that might still be very close to DDP. And it's clear whatever compression we end with, it should apply to a full frame. Lossy compression could be an option, however i'm not sure it's going to be JPEG, because JPEG produces block artefacts that might be very visible on a relatively small set of pixxels.
Personally I think the best use case for "compressed DDP" is large matrix displays (64x64 and beyond). Maybe in combination with the WLED "Video Lab" tool by @DedeHai.
Update: I've pushed a device-validated implementation with spec and benchmarks to
aenertia/WLED:staging/ddp-spec.Honest context: this was developed over a weekend using a custom agentic coding harness (OpenCode with OhMyOpenCode) driving multiple LLM providers, with me directing, testing on hardware, and reviewing output. The code works and the numbers are real measurements from my M5StickC, but the volume of code and documentation reflects AI-assisted velocity, not months of hand-crafting. I've done a slop sweep and tried to match upstream style but there will be rough edges — I'd rather be upfront about that than pretend otherwise. Happy to address anything that reads wrong.
What's implemented and tested on device (M5StickC, ESP32-PICO-D4, WiFi + PPP):
- 6 codec types decoded in
handleDDPPacket(): Delta+RLE (0x10), RLE keyframe (0x20), Transform (0x30), Delta-only (0x40), Tuple-RLE (0x50), Planar-RLE (0x60) - Uses the DDP C bit (
dataType & 0x80) as the compression signal — the spec's own customer extension mechanism. Codec type insequenceNumupper nibble. Standard senders/receivers unaffected. - Per-segment DDP routing via destination byte (1-32 → segment) or eligibility bitmask for flat-stream distribution
- Mixed frozen/non-frozen segment rendering — internal effects and DDP coexist on separate segments without tearing
Measured compression ratios (PPP 1.5Mbaud, 3200px, worst-case hardware — no PSRAM, dual-stack, SPI TFT DMA active):
Content Codec Ratio Achieved FPS Ghost rider (sparse) Delta+RLE 33:1 120fps Solid pulse Planar-RLE 62:1 120fps Sparse twinkle Delta+RLE 8:1 20fps Rainbow (worst case) Raw 1:1 10fps (link-limited) Typical setups (WiFi-only, no SPI bus overhead) will do better — these are pessimistic baselines.
Addressing previous feedback:
- @softhack007 "new protocol ID" → C bit is the DDP spec's own extension point. Limitation acknowledged: if the DDP spec ever assigns meaning to the
sequenceNumupper nibble, we have a conflict. Discussed in spec §10.1. - @softhack007 "full frame compression" → implemented for Delta+RLE/RLE (multi-packet streaming). Tuple-RLE and Planar-RLE are single-packet only — the decoder has a known limitation with multi-packet continuation that I haven't fixed yet.
- @softhack007 "stream restart" → 0x20 keyframe resets receiver state. Configurable interval.
- @softhack007 "gzip/deflate" → analysed but not implemented. Heap budget on ESP32-PICO (no PSRAM) is too tight (~42KB needed during inflate vs ~67KB free). Viable on S3/P4 with PSRAM. See spec §10.6.
- @DedeHai "adaptive colour depth reduction" → implemented sender-side in
ddp_bench.py --lossy-depth. Firmware stays lossless. Your exact algorithm (LSB stripping by brightness band) is what I used. - @DedeHai "palette LUT" → I don't think this works well over lossy UDP — a lost palette packet corrupts everything until resync. Discussed in spec §10.5.
- @softhack007 "JPEG" → reserved codec slot 0x70 for hardware JPEG on S3/P4. Block artefacts are fine at hub75 scale (128x128+), not at small matrices. See spec §10.7.
Known gaps I haven't addressed yet:
- Tuple-RLE (0x50) glitches on multi-packet frames (decoder design issue, confirmed on device)
- Transform codec (0x30) is implemented but not benchmarked
- Mode A per-segment routing (explicit destination byte) hasn't been visually validated end-to-end, only Mode B (eligibility mask)
- WebSocket DDP compression path is implemented but not tested
- The prevFrame buffer uses RGB565 (2B/pixel) which introduces ~1% quantisation error per delta frame — invisible on LEDs, measurable in test fixtures
Full spec:
docs/ddp-readme.md
Branch README:README-ddp-spec.md
Python sender/benchmark:tools/ddp_bench.py- 6 codec types decoded in
@aenertia no problem, we are not in a hurry so take your time. Good quality code is more valuable for us than a quick-shot prototype. Just start a PR once you have something, and we'll discuss from that point on.
Just one comment that might be important for now:
softhack007 "new protocol ID" → C bit is the DDP spec's own extension point. Limitation acknowledged: if the DDP spec ever assigns meaning to the
sequenceNumupper nibble, we have a conflict. Discussed in [spec §10.1 (https://github.com/aenertia/WLED/blob/staging/ddp-spec/docs/ddp-readme.md#101-wire-format-c-bit-vs-new-protocol-id).This was not just a recommendation - the new "compressed DDP" header must be sufficiently distinct from the standard DDP protocol header. For example, it could carry a "protocol version" of 0b11 in byte0, plus the "custom" bit in byte2; or maybe prefix "DDC" to the header so it violated the DDP header specs. The objective is to ensure that other DDP clients (who can't handle compression) do not try to read this stream as DDP.
This is a hard constraint; I won't accept any PR that does not meet it.
@coderabbitai I think that some form of compression support (gzip, zip, xz, or other) is already included in esp-idf? Please look up the official documentation, both for an API overview and for information about memory (RAM) needed for compression and decompression.
coderabbitai commented
on Aug 24, 2026 coderabbitaiboton Aug 24, 2026 – with coderabbitaiContributorMore actions🧩 Analysis chain
🤖 web_search executed:
site:docs.espressif.com/projects/esp-idf/en/latest/esp32 "gzip" OR "deflate" OR "tinfl" compression decompression RAMLength of output: 18951
🤖 web_search executed:
site:docs.espressif.com/projects/esp-idf/en/latest/esp32 "miniz" "heap" OR "memory"Length of output: 19950
🤖 web_search executed:
site:github.com/espressif/esp-idf gzip deflate tinfl miniz component memory RAM officialLength of output: 13938
🏁 Script executed:
#!/bin/bash set -euo pipefail url='https://raw.githubusercontent.com/espressif/esp-idf/master/components/esp_rom/include/miniz.h' curl -fsSL "$url" -o /tmp/esp-idf-miniz.h echo 'KEY_CONSTANTS_AND_CONFIGURATION:' rg -n -C 3 \ 'MINIZ_NO_MALLOC|TINFL_LZ_DICT_SIZE|TDEFL_LESS_MEMORY|TDEFL_LZ_DICT_SIZE|TINFL_DECOMPRESS_MEM_TO_MEM_FAILED' \ /tmp/esp-idf-miniz.h echo echo 'PUBLIC_CODEC_API_COMMENTS:' rg -n -C 6 \ 'mz_(deflate|inflate)(Init|Init2|End|Reset|[[:space:]]*\()|tinfl_decompress(_mem_to_mem|_mem_to_callback)?[[:space:]]*\(' \ /tmp/esp-idf-miniz.h echo echo 'STATE_STRUCTURE_SIZES:' python3 - <<'PY' import re p = '/tmp/esp-idf-miniz.h' text = open(p, encoding='utf-8').read() for name in ('tinfl_decompressor', 'tdefl_compressor', 'mz_stream'): m = re.search(r'(typedef\s+struct.*?\}\s*' + name + r'\s*;)', text, re.S) if m: print(f'\n--- {name} ---') print(m.group(1)[:8000]) PY
Length of output: 25455
@softhack007, ESP-IDF does includeminizin the ROM component. It provides a subset of the zlib API and low-level DEFLATE APIs.Official API header:
https://github.com/espressif/esp-idf/blob/master/components/esp_rom/include/miniz.hAvailable formats and APIs
minizsupports:- Raw DEFLATE through
mz_deflateInit2()andmz_inflateInit2()with a negativewindow_bits. - zlib-wrapped DEFLATE through
mz_deflateInit(),mz_inflateInit(),mz_compress*(), andmz_uncompress(). - Low-level streaming decode through
tinfl_decompress().
A gzip stream is a DEFLATE stream with a gzip-specific wrapper. The documented
minizAPI explicitly identifies raw DEFLATE and zlib wrappers. It does not document gzip-wrapper parsing. A compressed-DDP implementation could use raw DEFLATE or zlib-wrapped DEFLATE. It would need an additional gzip parser if it accepts gzip data.minizalso contains ZIP archive APIs in its upstream interface. ZIP is an archive container. It is not a suitable frame-stream format. It adds archive metadata and does not solve packet loss or frame synchronization.I found no ESP-IDF-provided XZ/LZMA API in the official ESP-IDF documentation or ROM
minizheader.RAM requirements
The critical point is that ESP-IDF configures this ROM
minizbuild with:`#define` MINIZ_NO_MALLOC
This disables the convenience functions that allocate their own result buffer. In particular,
tinfl_decompress_mem_to_heap()cannot be used. The header also states that functions without caller-provided allocation callbacks do not work in this configuration.For decoding,
TINFL_LZ_DICT_SIZEis:`#define` TINFL_LZ_DICT_SIZE 32768
Therefore, a streaming DEFLATE decoder normally needs a 32 KiB LZ dictionary/output window. The documented
tinfl_decompress_mem_to_callback()helper also uses an internal 32 KiB buffer. However, on ESP32 ROM builds, that helper is not usable becauseMINIZ_NO_MALLOCdisables its internal allocation path. The implementation should instead use a caller-owned buffer and a lower-level streaming interface or zlib-compatible API with explicit allocation callbacks.The caller also needs:
- The DEFLATE decoder state.
- Input buffering for incoming DDC packets.
- Output handling or a full-frame destination buffer.
- A strict decoded-size limit from the DDC frame header.
The exact total runtime RAM is therefore not a fixed 32 KiB value. It depends on the selected API and the frame buffering design. The 32 KiB dictionary is the minimum major working buffer for normal DEFLATE decoding.
Compression is much more expensive than decompression. The shipped header enables:
`#define` TDEFL_LESS_MEMORY 1
Even in this reduced-memory configuration,
tdefl_compressorcontains:- A 32,768-byte LZ dictionary.
- A 24 KiB LZ code buffer.
- A 31,948-byte output buffer.
- A 65,536-byte
m_nexttable. - An 8,192-byte hash table.
- Huffman tables and state.
This is approximately 164 KiB plus structure overhead for the compressor state alone. This estimate excludes source data, destination data, packet buffers, and task stack. It makes on-device DEFLATE encoding unsuitable for memory-constrained ESP32 variants such as ESP32-PICO-D4. A desktop sender can encode the stream. WLED only needs to decode it.
Recommendation for compressed DDC
For a first version:
- Use a distinct DDC protocol header, as required above.
- Define one codec as raw DEFLATE or zlib-wrapped DEFLATE.
- Implement decode only in WLED.
- Require a keyframe and explicit uncompressed frame length in each independently decodable frame.
- Reject a stream when decoded output exceeds that declared length.
- Use a fixed caller-owned 32 KiB decode window.
- Enable the DEFLATE codec only on targets with sufficient free internal RAM or PSRAM.
This does not remove the need for a simpler codec. RLE remains useful because it has predictable and very small RAM use. DEFLATE is a reasonable optional codec for large matrix frames and ESP32-S3/P4 devices, especially when the sender performs compression.
The ESP-IDF memory and heap APIs can measure the available budget on the real target:
heap_caps_get_free_size(MALLOC_CAP_8BIT)heap_caps_get_largest_free_block(MALLOC_CAP_8BIT)heap_caps_get_minimum_free_size(...)
Official references:
You are interacting with an AI system.
Reacted by Frank Möhle- Raw DEFLATE through
@aenertia PS: saw your idea about "Per-segment routing". In fact I'd like to stay in scope of "DDP compressed" for whatever PR you may come up with. Destination routing is not implemented in any protocol that WLED speaks, and its not really needed IMHO. All other protocols we support respect the "use main segment only" settings option. Without "main seg only", we draw directly to the strip (faster), and some background tasks like audioreactive are suspended as the whole fixture is under external control (faster). With "main seg only", the main segment is the only target for streaming LEDs, all other segments keep showing WLED effects.
If you want to extend that, it'll be a separate PR.
@coderabbitai is there a JSON API command to define the "main segment" arbitrarily?
coderabbitai commented
on Aug 24, 2026 coderabbitaiboton Aug 24, 2026 – with coderabbitaiContributorMore actions🧩 Analysis chain
🏁 Script executed:
#!/bin/bash set -euo pipefail echo 'CANDIDATE_FILES:' fd -a -t f . | rg -i 'json|api|segment|state' echo echo 'MAIN-SEGMENT REFERENCES:' rg -n -i -C 5 \ 'mainSegment|main segment|"mainseg"|mainseg|mainSegmentId|mainSegmentId' \ -g '*.{c,cc,cpp,cxx,h,hpp,ino,json,md}' . echo echo 'JSON SEGMENT-STATE REFERENCES:' rg -n -i -C 4 \ 'root\["seg"\]|root\["main|doc\["seg"\]|doc\["main|JSON.*segment|segment.*JSON' \ -g '*.{c,cc,cpp,cxx,h,hpp,ino,md}' .
Length of output: 122777
@softhack007, yes. Send the top-levelmainsegproperty to the JSON state endpoint.{ "mainseg": 2 }This sets segment ID
2as the main segment. Segment IDs are zero-based.The target segment must exist and be active. If the ID is invalid or inactive, WLED falls back to the last active segment.
WLED does not change
mainsegwhile realtime mode is active. The current value appears in the state response asmainseg.For the HTTP API, the equivalent command is
SM=<segment-id>.You are interacting with an AI system.
Yeah that's a planned separate set of things. And quite critical for my use case ; but am still evolving it a bit so wasn't in first set of submissions. I've ended up in a mode A/B approach to support legacy behavior (main segment only DDP) and flexible eligibility marking per segment. This provides total coverage without breaking anything.
Compression is certainly an interesting idea, however my first thought is that any form of lossless compression is just going to produce something that is just going to be unreliable.
What is the point of having a system that only works if you keep the complexity of what is being displayed under a threshold that is dependent on the transport? There is a reason why most (all) realtime lighting protocols do not use compression and things like hub75 cards uses raw ethernet frames, no TCP/UDP
Nobody wants a lightning system that is really glitchy, that works during setup with basic tests, but then drops frames, inconsistent playback or data corruption.
This isn't too say that any form of compression is a total non-starter, but really only should be seen as adding extra safety margin, not delivering something that isn't possible without compression, but how you ensure every user knows that unless their setup works even when running at 1:1 definitely needs consideration.
Alternatively you need to say this is lossy compression with max bitrate set by the user and they accept that they are going to get image artifacts or dropped frames when the source exceeds that
I concur - my original need was very low bandwidth links. In a normal resourced system I completely agree this breaks the general layered systems architecture. This is something best dealt with in a transport L2/L3 compression layer IMNSHO. But MCU's are a different beast entirely - so doing application layer compression within the constraints of what is there is really the only option.
It certainly IS useful ; and I think you're spot on this is very much something that you would only turn on in need, and it absolutely does introduce failure modes for i.e certain patterns etc. BUT it likewise makes a significant portion of applications that are completely out of reach (live streaming text to large pixel count controllers; or distant low count ones being the two immediately unlocked by this for me) - which are otherwise unavailable.
DDP Compression Extension — Gauging Upstream Interest
I've been running WLED on an M5StickC over a PPP serial link (USB at 1.5Mbaud,
~172 KB/s effective) and hit the obvious bandwidth wall. To make it work I
implemented a compression extension for DDP — it uses the reserved bits in the
existing header so it's backwards compatible with standard senders and receivers.
Posting here to share what I've built and see if there's any appetite for
carrying something like this upstream. Happy to submit PRs if it's useful to
others, or keep it as a fork-only thing if not.
The problem
DDP is great for LED streaming. The bandwidth requirements are not:
WiFi is fine. Anything slower gets painful quickly.
What I built
A delta+RLE compression extension that uses the reserved bits in the DDP header:
COMPRESSEDflag (0x20)0x10= delta+RLE (XOR against previous frame, then PackBits RLE)0x20= RLE only (keyframe — no delta, for resync after packet loss)0x30= transform (global fade/scale + sparse explicit writes — decoder only so far)Standard DDP senders never set bit 5, so existing receivers are unaffected.
A receiver that doesn't understand compression will display noise on compressed
packets — which is the right behaviour for an opt-in extension.
Measured compression ratios (real hardware, M5StickC, 800 LEDs, 40fps)
The worst case (rainbow, every pixel different every frame) gets nothing. Most
real-world animations are somewhere between "solid pulse" and "sparse twinkle"
and compress well.
What this means at 95% compression (typical for sparse animations):
Implementation
The codec is a ~150-line header-only C file (
ddp_compress.h) with:RLEDecoderstruct: streaming decoder, operates directly on the receivedpacket buffer, no heap allocation
rle_encode(): PackBits encoder with caller-provided output bufferrle_encode_adaptive(): tries delta+RLE and RLE-only, returns the smallerThe receiver integration is in
handleDDPPacket()— guarded by#ifdef WLED_ENABLE_DDP_COMPRESSIONso it's zero-cost when disabled.I also wrote Python tools for testing:
ddp_bench.py: benchmark tool with raw vs compressed comparison, IFS fractalrenderer, per-segment targeting, RGBW support
ddp_codec.py: standalone codec library for sendersifs_ddp.py: IFS fractal → DDP streamer with auto MTU detectionAll of this is in my fork: https://github.com/aenertia/WLED/tree/dev/ppp-wifi
The relevant branches:
pr/ddp-rle-codec— codec header only, no WLED core changespr/ddp-compressed-receiver— receiver integration, depends on abovepr/ddp-compressed— full bundle: codec + receiver + tools + specUse cases beyond serial
Once I had it working I realised the same approach applies to a few other
scenarios that might be of broader interest:
Low-baud serial — RS-485 runs, long cable runs at 115200 baud. Raw DDP
gives 4.7fps for 800 pixels. Compressed gives ~70fps. That's the difference
between "technically works" and "actually useful" for wired installations where
WiFi isn't an option.
LoRa and other RF — I haven't tested this personally, but the maths works
out. At 250kbps LoRa, raw DDP gives 10fps for 800 pixels; compressed gives
~150fps. Outdoor installations, agricultural lighting, distributed art
installations — anywhere WiFi doesn't reach. DDP is already UDP-based so
running it over a LoRa UDP bridge needs no firmware changes beyond enabling
compression.
WLED-to-WLED streaming — WLED already has
realtimeBroadcast()forWLED-to-WLED sync. The receiver side already handles compressed DDP. Adding
compression to the send path would reduce WiFi airtime for sync setups and
make large-pixel-count sync feasible on constrained devices. The encoder is
already there; it just needs wiring into the send path.
Capture and replay — this one's speculative but I think it's interesting.
The DDP header has a 4-byte timecode field (bytes 10–13, T flag in byte 0)
that WLED currently ignores. Combined with compression, it makes recording
DDP sequences to flash/SPIFFS feasible:
The timecode field enables frame-accurate replay: read the next frame's
timestamp, wait until
millis()matches, render. No host required. Thiswould turn WLED into a standalone animation player for pre-recorded sequences —
useful for installations where a host PC isn't practical.
I haven't implemented the recorder/player yet — that would need the codec to
land first. But it's a natural next step if there's interest.
Relation to existing issues
A few upstream issues that this might be relevant to:
Also worth coordinating with PR #5774
(split udp.cpp into per-protocol files) since the receiver integration touches
the same file.
The ask
Is there appetite for something like this in WLED? I'm happy to:
The codec is small and the receiver integration is guarded by a build flag, so
the cost of carrying it is low. But I understand if serial/LoRa use cases
aren't a priority for the project.
References
docs/ddp-readme.md(§3–§6, §12.2, §15)wled00/ddp_compress.htools/ddp_bench.py,tools/ddp_codec.py