Skip to content

perf: preserve routing snapshots without copying unmatched rules - #3943

Open
djwhitt wants to merge 6 commits into
mainfrom
perf/routing-snapshot-cache
Open

djwhitt wants to merge 6 commits into
mainfrom
perf/routing-snapshot-cache

Conversation

@djwhitt

@djwhitt djwhitt commented Sep 7, 2026 •

Copy link
Copy Markdown
Contributor

Review scope

This PR owns the generation-safe ID-keyed routing snapshot cache and its production hardening: compact targets, batch reuse, missing-generation repair, and byte-aware capacity/telemetry.

It intentionally does not include zero-based positional addressing; that isolated change is stacked as #3946.

Boundary Revision Purpose
Main used for the original routing comparison 130bdeb6 (#3942) Per-rule Cachex lookup baseline
Initial generation-safe snapshot a943dc6a Original snapshot-cache implementation
This PR head 6820c6ee Hardened ID-keyed snapshot
Positional child 4754419c (#3946) Separate addressing optimization

What changes

  • Build the routing tree and compact {rule_id, backend_id, sink} targets from one database query.
  • Publish a small Cachex header and one generation-qualified ETS target tuple per source.
  • Binary-search a compact rule-ID index for sparse matches; fetch the tuple once for dense matches.
  • Prepare one immutable tree/snapshot per ingest batch rather than fetching it for every event.
  • Keep a compressed reader-owned target map as the exact fallback for replacement, eviction, or store restart.
  • After fallback, rehydrate only the still-current Cachex generation; stale readers continue with their own decoded snapshot.
  • Bound the store by 100,000 sources, a configurable 512 MiB estimated-byte limit, one-hour TTL, and five-minute sweep.
  • Emit [:logflare, :rules, :routing_snapshot_store] source/byte gauges on lifecycle changes.
  • Preserve nil rejection, routing-target validation, and spool per-source error isolation.

Architecture

flowchart LR
  Q["One rules query"] --> B["Routing tree + compact targets"]
  B --> C["Cachex header<br/>tree + generation + fallback"]
  B --> E["ETS generation<br/>target tuple"]

  P["Prepare once per batch"] --> C
  C --> M["Match rule IDs"]
  M --> H{"ETS generation available?"}
  H -->|Yes| R["Read matched targets"]
  H -->|No| F["Decode immutable fallback once"]
  F --> K{"Header still current?"}
  K -->|Yes| RH["Rehydrate ETS + Cachex header"]
  K -->|No| L["Carry decoded fallback through batch"]
Loading

Consistency guarantees

  • Tree and target snapshot come from the same query and are published together.
  • Generation-qualified keys prevent a reader from combining old and new targets.
  • Replacement retires old ETS data immediately without requiring reader acquisition/release bookkeeping.
  • Suspended, crashed, or stale readers remain correct through their immutable fallback.
  • Count, byte, TTL, and explicit invalidation retirement all preserve the same generation rules.

Performance

These are local microbenchmarks, not production-throughput claims. The corrected fixtures assert the intended 1/8/all match counts before measurement. Environment: Linux, 6 available cores, Elixir 1.19.5, OTP 27.3.4.6, JIT, Benchee 1.5.0, parallel 1, warmup 2s, measurement 5s, memory 2s.

1. Initial snapshot vs then-current main

Two alternating runs per revision with fully warmed caches:

Rules / matches Main 130bdeb6 Initial snapshot a943dc6a Mean change
100 / 8 13.35 / 13.58 us 7.63 / 7.20 us 1.82x faster
100 / 1 11.86 / 12.08 us 11.28 / 10.34 us 1.11x faster
100 / all 152.26 / 161.65 us 74.60 / 72.57 us 2.13x faster
1,000 / 8 47.52 / 48.30 us 39.04 / 38.97 us 1.23x faster
1,000 / 1 88.71 / 89.48 us 86.07 / 86.89 us comparable
1,000 / all 1.52 / 1.54 ms 0.88 / 0.90 ms 1.72x faster

2. Hardening commit vs initial snapshot

One adjacent run after adding compact ID-keyed targets, batch state, recovery, and byte-aware lifecycle handling:

Rules / matches Initial snapshot Hardened ID-keyed head Approximate change
100 / 8 7.63 / 7.20 us 7.02 us 1.06x faster
100 / 1 11.28 / 10.34 us 11.65 us comparable
100 / all 74.60 / 72.57 us 24.78 us 2.97x faster
1,000 / 8 39.04 / 38.97 us 35.97 us 1.08x faster
1,000 / 1 86.07 / 86.89 us 129.17 us 1.49x slower
1,000 / all 0.88 / 0.90 ms 0.348 ms 2.56x faster

The 1,000-rule/one-match regression is explicit. A matching-only probe found ID and positional keys within 1%, so this shape needs production-informed profiling rather than another speculative tree change.

3. Batch allocation and retained representation

For 1,000 rules / 8 matches, preparing the ID-keyed snapshot once per batch was:

Batch Prepare once vs fetch per event Allocation reduction
10 events 4.34x faster 8.85x lower
100 events 5.85x faster 11.33x lower

Replacing full %Rule{} values with compact targets reduced the modeled tree/tuple/index/fallback representation by 94.5-95.7%. The 1,000-rule fixtures fell from 1.20-1.42 MB to 58-71 KB. This component model excludes Cachex/ETS table overhead and allocator fragmentation; it is not process RSS.

Full hardening details: source_routing_cache_hardening_results.md.

The cumulative final stack vs current main retained-memory and batch-allocation comparison is in #3946. Those positional-head results are not attributed to this ID-keyed boundary.

Earlier #3937 comparison, fixture correction, and commands

Before the rebase, two runs against 1bff0f28 showed 6.1-11.0x faster sparse routing and 1.13-1.37x faster dense routing. Those gains are not claimed against current main. Stage isolation for 1,000 rules / 8 matches reduced cache fetch from 433.34 us to 30.40 us while resident-tree matching stayed around 6 us, identifying full-map retrieval as the dominant #3937 regression cost.

The previous unquoted metadata.rule_id:rule-100 parsed as equality to "rule" plus a negated message term, so the case labeled “one matching” actually matched zero rules. The fixture is now quoted and every case asserts its intended count. Earlier mislabeled v1.50.9/main numbers are not comparable.

Initial snapshot details: source_routing_snapshot_results.md.

MIX_ENV=test ../bin/x mix run test/profiling/source_routing_bench.exs
MIX_ENV=test ../bin/x mix run test/profiling/source_routing_batch_bench.exs
MIX_ENV=test ROUTING_BENCH_STAGES=1 ../bin/x mix run test/profiling/source_routing_bench.exs

A six-reader Benchee attempt was killed with exit -9 before results; no concurrent-throughput claim is made.

Validation

  • 90 focused snapshot/cache/tree/router tests, 0 failures, 1 existing excluded benchmark test.
  • 25 additional rules/context-cache/gossip/cache-buster/BigQuery consumer tests, 0 failures.
  • Concurrent replacement, suspended/crashed readers, restart, conditional repair, batch-local fallback reuse, count/byte capacity, expiry, telemetry, late retirement, sparse/dense/fallback equivalence, and threshold boundaries.
  • Formatter and MIX_ENV=test mix compile through project wrappers.
  • MIX_ENV=test mix lint.all; only 32 existing design suggestions.
  • MIX_ENV=test mix test.typings; 158 configured errors skipped, 10 existing unnecessary skips, task passed.
  • Hosted Elixir CI and Playwright pass at 6820c6ee.
  • Production-shaped concurrent throughput and rollout were not performed locally.

Remaining review points

  • The 50% dense threshold is conservative rather than demonstrated optimal across production rule shapes.
  • Estimated-byte accounting is approximate, and one oversized source is retained rather than made unroutable.
  • Cold publication is serialized; hot reads are not.
  • The 1,000-rule/one-match microbenchmark needs production-informed profiling.

Related: #3935, #3937, #3942. Positional follow-up: #3946.

Comment thread lib/logflare/rules.ex Outdated
def rules_tree_by_source_id(id) do
rules = list_by_source_id(id)
RulesTree.build(rules)
{RulesTree.build(rules), Map.new(rules, &{&1.id, &1})}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is this not still problematic if there are a lot of rules on the source?

@djwhitt djwhitt Sep 8, 2026 •

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Addressed in c984bfd

@djwhitt
djwhitt marked this pull request as ready for review September 8, 2026 21:21
@djwhitt
djwhitt requested review from Baishan and chasers September 8, 2026 21:25

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants