Skip to content

DDP Compression Extension — Gauging Upstream Interest #5810

Description

@aenertia

DDP Compression Extension — Gauging Upstream Interest

I've been running WLED on an M5StickC over a PPP serial link (USB at 1.5Mbaud,
~172 KB/s effective) and hit the obvious bandwidth wall. To make it work I
implemented a compression extension for DDP — it uses the reserved bits in the
existing header so it's backwards compatible with standard senders and receivers.

Posting here to share what I've built and see if there's any appetite for
carrying something like this upstream. Happy to submit PRs if it's useful to
others, or keep it as a fork-only thing if not.


The problem

DDP is great for LED streaming. The bandwidth requirements are not:

Transport Effective BW 800px RGB @ 30fps 12,800px RGB @ 30fps
WiFi (20 Mbps) 2.3 MB/s fine fine
PPP/USB serial (1.5 Mbaud) 172 KB/s 71 fps — tight 4.5 fps — unusable
Serial (115200 baud) 11.5 KB/s 4.7 fps impossible
LoRa (250 kbps) ~25 KB/s 10 fps impossible

WiFi is fine. Anything slower gets painful quickly.


What I built

A delta+RLE compression extension that uses the reserved bits in the DDP header:

  • Flags byte (byte 0), bit 5: COMPRESSED flag (0x20)
  • Sequence byte (byte 1), upper nibble: compression type
    • 0x10 = delta+RLE (XOR against previous frame, then PackBits RLE)
    • 0x20 = RLE only (keyframe — no delta, for resync after packet loss)
    • 0x30 = transform (global fade/scale + sparse explicit writes — decoder only so far)

Standard DDP senders never set bit 5, so existing receivers are unaffected.
A receiver that doesn't understand compression will display noise on compressed
packets — which is the right behaviour for an opt-in extension.

Measured compression ratios (real hardware, M5StickC, 800 LEDs, 40fps)

Pattern Raw Compressed Ratio Wire bandwidth
Rainbow (worst case) 2,400 B 2,400 B 1:1 93.7 KB/s
Solid colour pulse 2,400 B ~1,900 B 1.2:1 77.8 KB/s
Sparse twinkle (2% change/frame) 2,400 B 120 B 20:1 4.6 KB/s
Static (no change) 2,400 B ~16 B ~150:1 <1 KB/s

The worst case (rainbow, every pixel different every frame) gets nothing. Most
real-world animations are somewhere between "solid pulse" and "sparse twinkle"
and compress well.

What this means at 95% compression (typical for sparse animations):

Transport 800px before 800px after 12,800px before 12,800px after
PPP 1.5Mbaud 71 fps >100 fps 4.5 fps 89 fps
Serial 115200 4.7 fps ~70 fps impossible 4.4 fps
LoRa 250kbps 10 fps ~150 fps impossible 9 fps

Implementation

The codec is a ~150-line header-only C file (ddp_compress.h) with:

  • RLEDecoder struct: streaming decoder, operates directly on the received
    packet buffer, no heap allocation
  • rle_encode(): PackBits encoder with caller-provided output buffer
  • rle_encode_adaptive(): tries delta+RLE and RLE-only, returns the smaller

The receiver integration is in handleDDPPacket() — guarded by
#ifdef WLED_ENABLE_DDP_COMPRESSION so it's zero-cost when disabled.

I also wrote Python tools for testing:

  • ddp_bench.py: benchmark tool with raw vs compressed comparison, IFS fractal
    renderer, per-segment targeting, RGBW support
  • ddp_codec.py: standalone codec library for senders
  • ifs_ddp.py: IFS fractal → DDP streamer with auto MTU detection

All of this is in my fork: https://github.com/aenertia/WLED/tree/dev/ppp-wifi

The relevant branches:


Use cases beyond serial

Once I had it working I realised the same approach applies to a few other
scenarios that might be of broader interest:

Low-baud serial — RS-485 runs, long cable runs at 115200 baud. Raw DDP
gives 4.7fps for 800 pixels. Compressed gives ~70fps. That's the difference
between "technically works" and "actually useful" for wired installations where
WiFi isn't an option.

LoRa and other RF — I haven't tested this personally, but the maths works
out. At 250kbps LoRa, raw DDP gives 10fps for 800 pixels; compressed gives
~150fps. Outdoor installations, agricultural lighting, distributed art
installations — anywhere WiFi doesn't reach. DDP is already UDP-based so
running it over a LoRa UDP bridge needs no firmware changes beyond enabling
compression.

WLED-to-WLED streaming — WLED already has realtimeBroadcast() for
WLED-to-WLED sync. The receiver side already handles compressed DDP. Adding
compression to the send path would reduce WiFi airtime for sync setups and
make large-pixel-count sync feasible on constrained devices. The encoder is
already there; it just needs wiring into the send path.

Capture and replay — this one's speculative but I think it's interesting.
The DDP header has a 4-byte timecode field (bytes 10–13, T flag in byte 0)
that WLED currently ignores. Combined with compression, it makes recording
DDP sequences to flash/SPIFFS feasible:

  • 10 seconds at 30fps, 800 pixels, raw: 720 KB — doesn't fit in 256KB SPIFFS
  • Same, compressed at 95%: 36 KB — fits easily

The timecode field enables frame-accurate replay: read the next frame's
timestamp, wait until millis() matches, render. No host required. This
would turn WLED into a standalone animation player for pre-recorded sequences —
useful for installations where a host PC isn't practical.

I haven't implemented the recorder/player yet — that would need the codec to
land first. But it's a natural next step if there's interest.


Relation to existing issues

A few upstream issues that this might be relevant to:

  • #5755 — DDP-over-WebSocket fragmented packets (bandwidth pressure)
  • #5412 — DDP out-of-sequence packets on WiFi (compression reduces packet count)
  • #4320 — UI unresponsive during DDP realtime (compression reduces CPU load from packet processing)

Also worth coordinating with PR #5774
(split udp.cpp into per-protocol files) since the receiver integration touches
the same file.


The ask

Is there appetite for something like this in WLED? I'm happy to:

  • Submit the codec header as a standalone PR (no WLED core changes, easy to review)
  • Follow up with the receiver integration if the codec lands
  • Adjust the wire format if there are concerns about the reserved-bit usage
  • Keep it as a fork-only thing if it's too niche

The codec is small and the receiver integration is guarded by a build flag, so
the cost of carrying it is low. But I understand if serial/LoRa use cases
aren't a priority for the project.


References

Activity

  1. added a commit that references this issue on Aug 18, 2026
  2. added
    AIPartly generated by an AI. Make sure that the contributor fully understands the code!
    on Aug 19, 2026
  3. softhack007 commented on Aug 19, 2026

    @softhack007
    Member

    @aenertia thanks, I think we'd be generally interested in a "compressed DDP" feature.

    We would need to look at some details, once we see your proposed code

    • "abusing" an unused DDP header flag has a risk of becoming incompatible with future DDP spec updates. It might be better if we define our own "DDP compressed" protocol ID, and leave "DDP standard" as it is.
    • details of the RLE algo - i don't understand yet how you can do all compression in-place. The worst-case for RLE is when the encoder needs to duplicate the RLE "magic marker" too often. In this case, the compressed stream can become bigger than the original input. Likewise, if the stream that gets compressed temporarily requires more bytes than the uncompressed equivalent, you cannot do this in-place because the compressed data will overwrite some not-yet-compressed input data.
    • It might be better to perform RGB-based (3 color streams) RLE, instead of byte-based RLE
    • Possibility of adding the compression feature to DDP-over-ws, including our JS frontend code
      function sendDDP(ws, start, len, colors, isESP8266=false) {

    interesting references:

  4. aenertia commented on Aug 20, 2026

    @aenertia
    ContributorAuthor

    @softhack007 thanks for looking at this -- good to know there's interest. These are all solid points and I think they should be worked into the spec doc rather than decided upfront. Working through each one:

    1. Reserved-bit vs own protocol ID

    Fair point about borrowing reserved bits. I've been looking at this more and there are actually a few options worth considering:

    • Current approach: flags byte bit 5 (0x20). Simple, but as you say, borrowing against future spec changes.
    • Customer-defined data type: byte 2 has a C bit (bit 7) explicitly defined in the spec as "1 for Customer defined". So dataType = 0x80 | original_type would signal "this is a vendor extension" using a mechanism the spec already provides for exactly this purpose.
    • Separate protocol discriminator: a new port or a different identification scheme entirely.

    The C bit approach is interesting because it's the one place the DDP spec explicitly invites extensions. The tradeoff is that compression is really a transport concern rather than a data type concern, but the spec doesn't have a dedicated extension mechanism so you use what's available.

    I'll add a section to the spec doc covering all three options with pros/cons so we can pick the right one before any code lands upstream. No point locking in a wire format prematurely.

    2. In-place RLE / worst-case expansion

    You're right about the general principle -- RLE can expand (worst case ~0.8% for PackBits). The current implementation uses separate buffers on the encoder side and a streaming decoder on the receiver side, so there's no in-place encoding anywhere. But your question raises something worth documenting properly in the spec: what guarantees does a receiver need about maximum expansion, and what MUST a sender do when compression doesn't help?

    Current behaviour: rle_encode_adaptive() compares compressed vs raw size and returns DDP_COMP_TYPE_NONE if compressed >= raw, so expansion never hits the wire. The decoder (RLEDecoder) is a streaming state machine that reads one byte at a time from the packet buffer -- no intermediate allocation, no in-place decode. But these are implementation details that should be spec-level requirements ("a sender MUST fall back to uncompressed if the encoded output exceeds the raw payload size").

    I'll tighten up section 4 of the spec doc with explicit worst-case bounds and sender/receiver obligations. The code is in ddp_compress.h (~150 lines) if you want to look at the actual encoder/decoder.

    3. RGB-tuple vs byte-level RLE

    This is worth exploring properly. The current codec does byte-level RLE on XOR deltas, which works well for the sparse-change case (unchanged pixels become zero-byte runs regardless of channel boundaries). But there are at least three approaches worth benchmarking:

    • Byte-level RLE (current): simple, works well with XOR delta, 300 unchanged RGB pixels = 900 zero bytes = one 2-byte RLE token. Partially-changed pixels still compress well per-channel.
    • RGB-tuple RLE: each run unit is 3 bytes (or 4 for RGBW). Identical compression for unchanged pixels, but loses per-channel run-length on partial changes.
    • Separate colour planes (closer to the ITU T.45 reference you linked): split into R/G/B planes, RLE each independently. Better for smooth gradients. But needs three decode passes plus reassembly on the receiver, which breaks the streaming decoder model and requires buffering.

    The codec is already structured for multiple compression types in the upper nibble (delta+RLE, RLE-only, transform). Adding an RGB-tuple or plane-split variant would be a new type code -- the framework supports it, it's just a question of what performs best on real LED animation data.

    I'll add a section to the spec covering these variants with the tradeoff analysis, and I can run benchmarks on real animation patterns (rainbow cycle, sparse twinkle, gradient fade, chase) to get actual numbers rather than theoretical arguments. Would be good to know what kind of patterns matter most for the upstream use case.

    4. DDP-over-WebSocket + JS frontend

    Agreed this would be valuable. The codec is pure byte manipulation, so a JS PackBits encoder/decoder is straightforward -- maybe 30 lines each. The receiver side already handles compressed DDP regardless of transport (same handleDDPPacket() path for UDP and WS).

    One thing worth working through in the spec: WebSocket already has per-message compression (permessage-deflate, RFC 7692). If that's enabled, layering PackBits on top is mostly redundant. The delta framing (only send changed pixels) might be more useful over WS than the RLE encoding itself. Or it might make sense to define a WS-specific mode that does delta-only without RLE, relying on permessage-deflate for the entropy removal. Worth thinking through before writing code.

    I'll add a transport-specific section to the spec covering the WS path, including interaction with permessage-deflate and what the JS implementation surface looks like. Probably makes sense as a follow-on PR after the core codec, since it's a different review surface (JS in common.js).


    I'll update the spec doc with sections covering all of these open questions -- wire format options, sender/receiver obligations, compression variant benchmarks, and WS transport considerations. Better to get the spec right before submitting code PRs.

  5. DedeHai commented on Aug 23, 2026

    @DedeHai
    Collaborator

    WebSocket already has per-message compression

    I dont think this is implemented.

    regarding compression: another approach would be to use what GIF does: use a palette look-up, this would only be helpful on larger setups though and can be quite lossy (and maybe too computationally expensive).

    Another addition which would be especially helpful when using RGB split planes: use a very simple adaptive color depth reduction. While this is also lossy it would be quite effective in video streams: since WLED uses gamma correction and colors are compressed especially at the higher end, the loss would only be visible when using it side-by-side with an uncompressed section (i.e. a matrix consisting of multiple ESP's) - and maybe not even then. What I am thinking of is something like this:

    if (color > 196) color = color &0xF8;
    else if (color > 128) color = color &0xFC;
    else if (color > 64)  color = color &0xFE;

    so stripping out the LSB's as those small variations get compressed by gamma and those variations are also almost indistinguishable even if not using gamma. On the lower end, no compression is used to preserve color gradients and "fading trails".
    I also did run experiments at one point using RGB565 encoding and found that the color loss does not cause significant banding but those tests were brief based on rainbow patterns, I would expect fading trails to be affected.

  6. aenertia commented on Aug 23, 2026

    @aenertia
    ContributorAuthor

    Yeah, permessage-deflate isn't in the ESP32 WS implementation -- found that out pretty quickly.

    Been running tests over the weekend with planar and channel variants alongside the byte-RLE already in the public repo. A bit more overhead but worth it for several input patterns without blowing heap.

    The bigger issue is WS over PPP -- layered framing adds real overhead. It works but sustainable FPS is noticeably worse than the UDP path. Same on WiFi but less visible. On beefier variants with PSRAM, ROM tinfl for decompress-only might be worth exploring.

    Numbers from extended testing on M5StickC (ESP32-PICO-D4, 40x80 TFT, PPP @ 1.5Mbaud):

    Transport Encoding FPS Notes
    UDP / WiFi raw (rainbow) 661 link not the bottleneck
    UDP / WiFi raw (pulse) 997
    UDP / WiFi effective (TFT active) 83 stable 30s
    UDP / PPP raw (rainbow) 41.8 link-limited
    UDP / PPP compressed (twinkle) 435 compression hides the link cost
    UDP / PPP effective (TFT active) 76 stable 30s
    WS / WiFi ceiling (256px seg) 150
    WS / WiFi effective (TFT active) 53 stable 30s
    WS / PPP compressed (256px seg) 150 at ceiling
    WS / PPP effective (TFT active) 51 died at 5.4s -- WS+PPP backpressure

    WS+PPP is the bad combination -- TCP framing overhead compounds with serial link serialisation and backpressure kills it under sustained load. UDP+PPP+compression is surprisingly good; on sparse patterns the compression ratio offsets the link cost almost entirely.

  7. aenertia commented on Aug 23, 2026

    @aenertia
    ContributorAuthor

    The TFT active bits are to do with the eSPI as Matrix bus which I am co-testing - effectively new bus-type use your SPI screen as a pixel segment; has to careful manage dma ping-pong . But as part of that I have implemented a skip segment/DDP eligible segment schema which is actually super useful generally. My primary use case is to use eSPI as a 'preview pane' and then Copy to Effect to actual real matrix segments.

  8. aenertia commented on Aug 23, 2026

    @aenertia
    ContributorAuthor

    Also for the new esps3.1 and the P4 - they have jpeg decoding on chip. In theory using DDP transport with compression extension alows you to do realtime video elegantly to i.e hub75 panels walls etc.

  9. softhack007 commented on Aug 23, 2026

    @softhack007
    Member

    @coderabbitai the proposed (experimental, see https://github.com/aenertia/WLED/blob/dev/ppp-wifi/wled00/e131.cpp) code seems to compress only single DDP packets, but not a full frame. I'm wondering if this would create any benefit in a normal WiFi or ethernet setup? A single DDP message is already size-optimised to fit to the network MTU size. If the message get smaller than one MTU, this could cause additional message delays. What do you think?

  10. coderabbitai commented on Aug 23, 2026

    @coderabbitai
    Contributor
    🧩 Analysis chain

    🏁 Script executed:

    #!/bin/bash
    set -euo pipefail
    
    # Locate the local source file, if the checked-out repository contains it.
    if fd -a '^e131\.cpp$' . | grep -q .; then
      file=$(fd -a '^e131\.cpp$' . | head -n1)
      echo "LOCAL_FILE=$file"
      ast-grep outline "$file" --items all --view expanded || true
      rg -n -C 4 -i 'ddp|compress|rle|mtu|packet|payload' "$file" || true
    fi
    
    # Inspect the exact public branch referenced in the issue discussion.
    curl -fsSL \
      'https://raw.githubusercontent.com/aenertia/WLED/dev/ppp-wifi/wled00/e131.cpp' \
      -o /tmp/aenertia-e131.cpp
    
    echo 'REMOTE_SYMBOLS:'
    rg -n -C 5 -i 'ddp|compress|rle|mtu|packet|payload' /tmp/aenertia-e131.cpp
    
    echo 'REMOTE_RELEVANT_RANGES:'
    nl -ba /tmp/aenertia-e131.cpp | sed -n '1,260p'

    Length of output: 38848


    🏁 Script executed:

    #!/bin/bash
    set -euo pipefail
    
    tmpdir=$(mktemp -d)
    trap 'rm -rf "$tmpdir"' EXIT
    
    git clone --quiet --depth 1 --branch dev/ppp-wifi \
      https://github.com/aenertia/WLED.git "$tmpdir/WLED"
    
    cd "$tmpdir/WLED"
    
    echo 'COMPRESSION_FILES:'
    fd -a -t f 'ddp|compress|e131' . | sort
    
    echo 'ENCODER_AND_DDP_SEND_REFERENCES:'
    rg -n -i -C 5 \
      'rle_encode_adaptive|DDP_COMP_TYPE|DDP_FLAGS_COMPRESSED|ddp_compress|beginPacket|channelOffset|DDP_HEADER_LEN' \
      -g '*.{c,cc,cpp,cxx,h,hpp,ino}' .
    
    echo 'DDP_COMPRESS_HEADER:'
    if test -f wled00/ddp_compress.h; then
      sed -n '1,280p' wled00/ddp_compress.h
    fi

    Length of output: 52510


    @softhack007, yes. The current experimental bridge code compresses each DDP packet independently.

    It accepts at most DDP_CHANNELS_PER_PACKET bytes, which is 1440 bytes. It replaces that packet payload with its compressed form. It updates dataLen to the compressed length. It does not combine several DDP packets into one compressed frame.

    For a normal Ethernet or good WiFi link, this usually gives little benefit:

    • The sender still sends the same number of UDP datagrams.
    • Each datagram still has IP, UDP, DDP, MAC, and WiFi MAC overhead.
    • The receiver still processes the same number of packet callbacks.
    • The PUSH packet still arrives after the same sequence of frame packets.
    • The frame completion latency therefore does not decrease much.

    A smaller UDP datagram does not wait until it reaches the MTU. UDP has no equivalent of TCP Nagle coalescing. The sender sends the smaller datagram when it calls endPacket(). Therefore, reduced payload size does not itself add a message delay.

    The benefit is the reduced serialization time and, on WiFi, reduced airtime. This can matter on a slow or congested link. It matters much more for PPP and similar serial links. It can also help low-rate WiFi, weak-signal WiFi, or a busy shared RF channel. It is almost irrelevant on 100 Mbps or 1 Gbps Ethernet unless the sender or receiver is already bandwidth-limited.

    There is also a cost. Per-packet RLE has less opportunity to find runs across the 1440-byte packet boundary. It adds encoder CPU time for every packet. For incompressible animation data, the code correctly falls back to raw data, but it still paid the attempted encoding cost.

    I think the primary value of frame-level compression would be different:

    1. Compress the complete frame, or a larger frame region.
    2. Split the compressed byte stream into MTU-safe DDP packets.
    3. Reassemble the stream before decoding, or define a streaming decoder with explicit frame state.
    4. Render only after the complete compressed frame is available.

    That design can reduce both byte count and, for sufficiently compressible frames, packet count. It has higher buffering requirements. It also makes packet loss and frame resynchronization more complex.

    One protocol detail needs attention in the current packet-local format. The compressed DDP dataLen is the compressed length. The receiver decodes until the compressed stream ends. The format does not carry the expected decoded length for that packet. A final format should include an explicit uncompressed byte count, or another unambiguous range length. The receiver should enforce that count. This prevents a malformed or truncated stream from writing past the intended packet range.

    So, I agree with the concern: packet-local compression is mainly useful for constrained transports. For ordinary WiFi and Ethernet, it is not likely to improve FPS or frame latency unless measured data shows the network link is the bottleneck.

    You are interacting with an AI system.

  11. softhack007 commented on Aug 23, 2026

    @softhack007
    Member

    @aenertia @DedeHai considering our previous ideas, I think we definitely must have a new protocol ID for "compressed DDP". It could be integrated into the normal DDP encoder/decoder source code, similar to the way that E1.31 and Art-Net share a lot of code.

    Also we need to spend more thoughts on the compression methods (even jpeg and gzip could be an option) and stream format. I think that RLE compression has a benefit on encoding speed, because it's really a simple and straightforward method. "2D" delta-compression looks appealing, however it requires to keep a copy of the last pixels frame. If we have a delta-compression scheme, the next question is how to implement stream restart points, so the receiver gets a chance to recover from lost or re-ordered message chunks.

    We don't need to stuff everything into to existing DDP format, but we can define our own format that might still be very close to DDP. And it's clear whatever compression we end with, it should apply to a full frame. Lossy compression could be an option, however i'm not sure it's going to be JPEG, because JPEG produces block artefacts that might be very visible on a relatively small set of pixxels.

    Personally I think the best use case for "compressed DDP" is large matrix displays (64x64 and beyond). Maybe in combination with the WLED "Video Lab" tool by @DedeHai.

  12. aenertia commented on Aug 24, 2026

    @aenertia
    ContributorAuthor

    Update: I've pushed a device-validated implementation with spec and benchmarks to aenertia/WLED:staging/ddp-spec.

    Honest context: this was developed over a weekend using a custom agentic coding harness (OpenCode with OhMyOpenCode) driving multiple LLM providers, with me directing, testing on hardware, and reviewing output. The code works and the numbers are real measurements from my M5StickC, but the volume of code and documentation reflects AI-assisted velocity, not months of hand-crafting. I've done a slop sweep and tried to match upstream style but there will be rough edges — I'd rather be upfront about that than pretend otherwise. Happy to address anything that reads wrong.

    What's implemented and tested on device (M5StickC, ESP32-PICO-D4, WiFi + PPP):

    • 6 codec types decoded in handleDDPPacket(): Delta+RLE (0x10), RLE keyframe (0x20), Transform (0x30), Delta-only (0x40), Tuple-RLE (0x50), Planar-RLE (0x60)
    • Uses the DDP C bit (dataType & 0x80) as the compression signal — the spec's own customer extension mechanism. Codec type in sequenceNum upper nibble. Standard senders/receivers unaffected.
    • Per-segment DDP routing via destination byte (1-32 → segment) or eligibility bitmask for flat-stream distribution
    • Mixed frozen/non-frozen segment rendering — internal effects and DDP coexist on separate segments without tearing

    Measured compression ratios (PPP 1.5Mbaud, 3200px, worst-case hardware — no PSRAM, dual-stack, SPI TFT DMA active):

    Content Codec Ratio Achieved FPS
    Ghost rider (sparse) Delta+RLE 33:1 120fps
    Solid pulse Planar-RLE 62:1 120fps
    Sparse twinkle Delta+RLE 8:1 20fps
    Rainbow (worst case) Raw 1:1 10fps (link-limited)

    Typical setups (WiFi-only, no SPI bus overhead) will do better — these are pessimistic baselines.

    Addressing previous feedback:

    • @softhack007 "new protocol ID" → C bit is the DDP spec's own extension point. Limitation acknowledged: if the DDP spec ever assigns meaning to the sequenceNum upper nibble, we have a conflict. Discussed in spec §10.1.
    • @softhack007 "full frame compression" → implemented for Delta+RLE/RLE (multi-packet streaming). Tuple-RLE and Planar-RLE are single-packet only — the decoder has a known limitation with multi-packet continuation that I haven't fixed yet.
    • @softhack007 "stream restart" → 0x20 keyframe resets receiver state. Configurable interval.
    • @softhack007 "gzip/deflate" → analysed but not implemented. Heap budget on ESP32-PICO (no PSRAM) is too tight (~42KB needed during inflate vs ~67KB free). Viable on S3/P4 with PSRAM. See spec §10.6.
    • @DedeHai "adaptive colour depth reduction" → implemented sender-side in ddp_bench.py --lossy-depth. Firmware stays lossless. Your exact algorithm (LSB stripping by brightness band) is what I used.
    • @DedeHai "palette LUT" → I don't think this works well over lossy UDP — a lost palette packet corrupts everything until resync. Discussed in spec §10.5.
    • @softhack007 "JPEG" → reserved codec slot 0x70 for hardware JPEG on S3/P4. Block artefacts are fine at hub75 scale (128x128+), not at small matrices. See spec §10.7.

    Known gaps I haven't addressed yet:

    • Tuple-RLE (0x50) glitches on multi-packet frames (decoder design issue, confirmed on device)
    • Transform codec (0x30) is implemented but not benchmarked
    • Mode A per-segment routing (explicit destination byte) hasn't been visually validated end-to-end, only Mode B (eligibility mask)
    • WebSocket DDP compression path is implemented but not tested
    • The prevFrame buffer uses RGB565 (2B/pixel) which introduces ~1% quantisation error per delta frame — invisible on LEDs, measurable in test fixtures

    Full spec: docs/ddp-readme.md
    Branch README: README-ddp-spec.md
    Python sender/benchmark: tools/ddp_bench.py

  13. softhack007 commented on Aug 24, 2026

    @softhack007
    Member

    @aenertia no problem, we are not in a hurry so take your time. Good quality code is more valuable for us than a quick-shot prototype. Just start a PR once you have something, and we'll discuss from that point on.

    Just one comment that might be important for now:

    softhack007 "new protocol ID" → C bit is the DDP spec's own extension point. Limitation acknowledged: if the DDP spec ever assigns meaning to the sequenceNum upper nibble, we have a conflict. Discussed in [spec §10.1 (https://github.com/aenertia/WLED/blob/staging/ddp-spec/docs/ddp-readme.md#101-wire-format-c-bit-vs-new-protocol-id).

    This was not just a recommendation - the new "compressed DDP" header must be sufficiently distinct from the standard DDP protocol header. For example, it could carry a "protocol version" of 0b11 in byte0, plus the "custom" bit in byte2; or maybe prefix "DDC" to the header so it violated the DDP header specs. The objective is to ensure that other DDP clients (who can't handle compression) do not try to read this stream as DDP.

    This is a hard constraint; I won't accept any PR that does not meet it.

  14. softhack007 commented on Aug 24, 2026

    @softhack007
    Member

    @coderabbitai I think that some form of compression support (gzip, zip, xz, or other) is already included in esp-idf? Please look up the official documentation, both for an API overview and for information about memory (RAM) needed for compression and decompression.

  15. coderabbitai commented on Aug 24, 2026

    @coderabbitai
    Contributor
    🧩 Analysis chain

    🤖 web_search executed:

    site:docs.espressif.com/projects/esp-idf/en/latest/esp32 "gzip" OR "deflate" OR "tinfl" compression decompression RAM

    Length of output: 18951


    🤖 web_search executed:

    site:docs.espressif.com/projects/esp-idf/en/latest/esp32 "miniz" "heap" OR "memory"

    Length of output: 19950


    🤖 web_search executed:

    site:github.com/espressif/esp-idf gzip deflate tinfl miniz component memory RAM official

    Length of output: 13938


    🏁 Script executed:

    #!/bin/bash
    set -euo pipefail
    
    url='https://raw.githubusercontent.com/espressif/esp-idf/master/components/esp_rom/include/miniz.h'
    curl -fsSL "$url" -o /tmp/esp-idf-miniz.h
    
    echo 'KEY_CONSTANTS_AND_CONFIGURATION:'
    rg -n -C 3 \
      'MINIZ_NO_MALLOC|TINFL_LZ_DICT_SIZE|TDEFL_LESS_MEMORY|TDEFL_LZ_DICT_SIZE|TINFL_DECOMPRESS_MEM_TO_MEM_FAILED' \
      /tmp/esp-idf-miniz.h
    
    echo
    echo 'PUBLIC_CODEC_API_COMMENTS:'
    rg -n -C 6 \
      'mz_(deflate|inflate)(Init|Init2|End|Reset|[[:space:]]*\()|tinfl_decompress(_mem_to_mem|_mem_to_callback)?[[:space:]]*\(' \
      /tmp/esp-idf-miniz.h
    
    echo
    echo 'STATE_STRUCTURE_SIZES:'
    python3 - <<'PY'
    import re
    p = '/tmp/esp-idf-miniz.h'
    text = open(p, encoding='utf-8').read()
    for name in ('tinfl_decompressor', 'tdefl_compressor', 'mz_stream'):
        m = re.search(r'(typedef\s+struct.*?\}\s*' + name + r'\s*;)', text, re.S)
        if m:
            print(f'\n--- {name} ---')
            print(m.group(1)[:8000])
    PY

    Length of output: 25455


    @softhack007, ESP-IDF does include miniz in the ROM component. It provides a subset of the zlib API and low-level DEFLATE APIs.

    Official API header:
    https://github.com/espressif/esp-idf/blob/master/components/esp_rom/include/miniz.h

    Available formats and APIs

    miniz supports:

    • Raw DEFLATE through mz_deflateInit2() and mz_inflateInit2() with a negative window_bits.
    • zlib-wrapped DEFLATE through mz_deflateInit(), mz_inflateInit(), mz_compress*(), and mz_uncompress().
    • Low-level streaming decode through tinfl_decompress().

    A gzip stream is a DEFLATE stream with a gzip-specific wrapper. The documented miniz API explicitly identifies raw DEFLATE and zlib wrappers. It does not document gzip-wrapper parsing. A compressed-DDP implementation could use raw DEFLATE or zlib-wrapped DEFLATE. It would need an additional gzip parser if it accepts gzip data.

    miniz also contains ZIP archive APIs in its upstream interface. ZIP is an archive container. It is not a suitable frame-stream format. It adds archive metadata and does not solve packet loss or frame synchronization.

    I found no ESP-IDF-provided XZ/LZMA API in the official ESP-IDF documentation or ROM miniz header.

    RAM requirements

    The critical point is that ESP-IDF configures this ROM miniz build with:

    `#define` MINIZ_NO_MALLOC

    This disables the convenience functions that allocate their own result buffer. In particular, tinfl_decompress_mem_to_heap() cannot be used. The header also states that functions without caller-provided allocation callbacks do not work in this configuration.

    For decoding, TINFL_LZ_DICT_SIZE is:

    `#define` TINFL_LZ_DICT_SIZE 32768

    Therefore, a streaming DEFLATE decoder normally needs a 32 KiB LZ dictionary/output window. The documented tinfl_decompress_mem_to_callback() helper also uses an internal 32 KiB buffer. However, on ESP32 ROM builds, that helper is not usable because MINIZ_NO_MALLOC disables its internal allocation path. The implementation should instead use a caller-owned buffer and a lower-level streaming interface or zlib-compatible API with explicit allocation callbacks.

    The caller also needs:

    • The DEFLATE decoder state.
    • Input buffering for incoming DDC packets.
    • Output handling or a full-frame destination buffer.
    • A strict decoded-size limit from the DDC frame header.

    The exact total runtime RAM is therefore not a fixed 32 KiB value. It depends on the selected API and the frame buffering design. The 32 KiB dictionary is the minimum major working buffer for normal DEFLATE decoding.

    Compression is much more expensive than decompression. The shipped header enables:

    `#define` TDEFL_LESS_MEMORY 1

    Even in this reduced-memory configuration, tdefl_compressor contains:

    • A 32,768-byte LZ dictionary.
    • A 24 KiB LZ code buffer.
    • A 31,948-byte output buffer.
    • A 65,536-byte m_next table.
    • An 8,192-byte hash table.
    • Huffman tables and state.

    This is approximately 164 KiB plus structure overhead for the compressor state alone. This estimate excludes source data, destination data, packet buffers, and task stack. It makes on-device DEFLATE encoding unsuitable for memory-constrained ESP32 variants such as ESP32-PICO-D4. A desktop sender can encode the stream. WLED only needs to decode it.

    Recommendation for compressed DDC

    For a first version:

    1. Use a distinct DDC protocol header, as required above.
    2. Define one codec as raw DEFLATE or zlib-wrapped DEFLATE.
    3. Implement decode only in WLED.
    4. Require a keyframe and explicit uncompressed frame length in each independently decodable frame.
    5. Reject a stream when decoded output exceeds that declared length.
    6. Use a fixed caller-owned 32 KiB decode window.
    7. Enable the DEFLATE codec only on targets with sufficient free internal RAM or PSRAM.

    This does not remove the need for a simpler codec. RLE remains useful because it has predictable and very small RAM use. DEFLATE is a reasonable optional codec for large matrix frames and ESP32-S3/P4 devices, especially when the sender performs compression.

    The ESP-IDF memory and heap APIs can measure the available budget on the real target:

    • heap_caps_get_free_size(MALLOC_CAP_8BIT)
    • heap_caps_get_largest_free_block(MALLOC_CAP_8BIT)
    • heap_caps_get_minimum_free_size(...)

    Official references:

    You are interacting with an AI system.

  16. softhack007 commented on Aug 24, 2026

    @softhack007
    Member

    @aenertia PS: saw your idea about "Per-segment routing". In fact I'd like to stay in scope of "DDP compressed" for whatever PR you may come up with. Destination routing is not implemented in any protocol that WLED speaks, and its not really needed IMHO. All other protocols we support respect the "use main segment only" settings option. Without "main seg only", we draw directly to the strip (faster), and some background tasks like audioreactive are suspended as the whole fixture is under external control (faster). With "main seg only", the main segment is the only target for streaming LEDs, all other segments keep showing WLED effects.

    If you want to extend that, it'll be a separate PR.

  17. softhack007 commented on Aug 24, 2026

    @softhack007
    Member

    @coderabbitai is there a JSON API command to define the "main segment" arbitrarily?

  18. coderabbitai commented on Aug 24, 2026

    @coderabbitai
    Contributor
    🧩 Analysis chain

    🏁 Script executed:

    #!/bin/bash
    set -euo pipefail
    
    echo 'CANDIDATE_FILES:'
    fd -a -t f . | rg -i 'json|api|segment|state'
    
    echo
    echo 'MAIN-SEGMENT REFERENCES:'
    rg -n -i -C 5 \
      'mainSegment|main segment|"mainseg"|mainseg|mainSegmentId|mainSegmentId' \
      -g '*.{c,cc,cpp,cxx,h,hpp,ino,json,md}' .
    
    echo
    echo 'JSON SEGMENT-STATE REFERENCES:'
    rg -n -i -C 4 \
      'root\["seg"\]|root\["main|doc\["seg"\]|doc\["main|JSON.*segment|segment.*JSON' \
      -g '*.{c,cc,cpp,cxx,h,hpp,ino,md}' .

    Length of output: 122777


    @softhack007, yes. Send the top-level mainseg property to the JSON state endpoint.

    {
      "mainseg": 2
    }

    This sets segment ID 2 as the main segment. Segment IDs are zero-based.

    The target segment must exist and be active. If the ID is invalid or inactive, WLED falls back to the last active segment.

    WLED does not change mainseg while realtime mode is active. The current value appears in the state response as mainseg.

    For the HTTP API, the equivalent command is SM=<segment-id>.

    You are interacting with an AI system.

  19. aenertia commented on Aug 24, 2026

    @aenertia
    ContributorAuthor

    Yeah that's a planned separate set of things. And quite critical for my use case ; but am still evolving it a bit so wasn't in first set of submissions. I've ended up in a mode A/B approach to support legacy behavior (main segment only DDP) and flexible eligibility marking per segment. This provides total coverage without breaking anything.

  20. netmindz commented on Aug 25, 2026

    @netmindz
    Member

    Compression is certainly an interesting idea, however my first thought is that any form of lossless compression is just going to produce something that is just going to be unreliable.

    What is the point of having a system that only works if you keep the complexity of what is being displayed under a threshold that is dependent on the transport? There is a reason why most (all) realtime lighting protocols do not use compression and things like hub75 cards uses raw ethernet frames, no TCP/UDP

    Nobody wants a lightning system that is really glitchy, that works during setup with basic tests, but then drops frames, inconsistent playback or data corruption.

    This isn't too say that any form of compression is a total non-starter, but really only should be seen as adding extra safety margin, not delivering something that isn't possible without compression, but how you ensure every user knows that unless their setup works even when running at 1:1 definitely needs consideration.

    Alternatively you need to say this is lossy compression with max bitrate set by the user and they accept that they are going to get image artifacts or dropped frames when the source exceeds that

  21. aenertia commented on Aug 26, 2026

    @aenertia
    ContributorAuthor

    I concur - my original need was very low bandwidth links. In a normal resourced system I completely agree this breaks the general layered systems architecture. This is something best dealt with in a transport L2/L3 compression layer IMNSHO. But MCU's are a different beast entirely - so doing application layer compression within the constraints of what is there is really the only option.

    It certainly IS useful ; and I think you're spot on this is very much something that you would only turn on in need, and it absolutely does introduce failure modes for i.e certain patterns etc. BUT it likewise makes a significant portion of applications that are completely out of reach (live streaming text to large pixel count controllers; or distant low count ones being the two immediately unlocked by this for me) - which are otherwise unavailable.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    AIPartly generated by an AI. Make sure that the contributor fully understands the code!enhancement

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions