Skip to content

feat(microservices): notify users when processing falls behind - #3952

Open
ryanmelt wants to merge 1 commit into
mainfrom
feat/processing-lag-alerts
Open

ryanmelt wants to merge 1 commit into
mainfrom
feat/processing-lag-alerts

Conversation

@ryanmelt

Copy link
Copy Markdown
Member

What

Adds a TopicLagMonitor (Ruby openc3/utilities/topic_lag_monitor.rb and Python openc3/utilities/topic_lag_monitor.py). It is wired into every microservice that already reported a *_topic_delta_seconds gauge: decom, log, text_log, interface cmd, router tlm, and tsdb ingest.

The monitor keeps setting the existing gauge. On top of that it:

  • Sends lag notifications. It publishes a WARN notification when lag reaches the warning threshold and an ERROR notification when it reaches the critical threshold. While the microservice stays behind, it re-notifies at a limited rate.
  • Sends a recovery notification once the lag has stayed below half of the threshold for the hold time. This hysteresis stops the notifications from flapping. Stepping down from critical to warning makes no noise.
  • Detects lost data. Streams are trimmed by wall-clock time, so a consumer that falls far enough behind silently skips data. Only while lagging, when two consecutive message ids on a topic are more than 1 s apart, it runs XINFO STREAM (at most once per second per topic). It compares the stream's max-deleted-entry-id to the last id this consumer processed. If unread entries were deleted, it publishes one ALERT covering all affected topics, listing each topic and roughly how many seconds were skipped. This alert is rate limited.
  • Finds the right Redis instance for XINFO from the target hashtag in the topic name ({TARGET}), falling back to the microservice's db_shard.

Notifications go through the existing Logger notification / alert types, so they show up in the Notifications menu with no frontend changes.

Why

The lag gauges were only visible as metrics (charted in Enterprise System Health, invisible in Core). A microservice that couldn't keep up, or that lost data to stream trimming, gave users no indication.

Configuration (env vars)

Variable Default Meaning
OPENC3_LAG_WARN_SECONDS 5 Warning threshold. 0 disables notifications and alerts; the gauge is still set.
OPENC3_LAG_CRITICAL_SECONDS 30 Critical threshold
OPENC3_LAG_CLEAR_SECONDS 10 How long lag must stay below half the threshold before recovering
OPENC3_LAG_RENOTIFY_SECONDS 300 Minimum time between repeated lag notifications and between lost-data alerts

Notes

  • Recovery is evaluated as messages arrive. If a stream stops completely right after the consumer catches up, the recovery notification appears with the next message.
  • Lost-data detection needs Valkey or Redis 7+ for max-deleted-entry-id. On servers without that field it skips the check.
  • This does not add a lag field to the microservice status model.

Testing

  • New spec/utilities/topic_lag_monitor_spec.rb (16 examples) and test/utilities/test_topic_lag_monitor.py (16 tests). They cover the thresholds, hysteresis and recovery, stepping down from critical, rate limiting, env config, lost-data detection, alert aggregation, XINFO rate limiting, db_shard lookup, and error handling.
  • bundle exec rspec spec/microservices spec/utilities spec/models spec/topics: 1493 examples, 0 failures.
  • uv run pytest ./test: 3045 passed. uv run ruff check openc3: clean.

🤖 Generated with Claude Code

The *_topic_delta_seconds lag gauges were only visible as metrics, so a
microservice falling behind (and silently skipping data trimmed from its
streams) went unnoticed. Add a TopicLagMonitor (Ruby and Python) that
publishes warn/critical/recovery notifications with hysteresis and alerts
when data was trimmed before it could be processed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@sonarqubecloud

Copy link
Copy Markdown

Quality Gate Failed Quality Gate failed

Failed conditions
Bugs with severity Major found (required < Minor)
Code smells with severity Major found (required < Major)

See analysis details on SonarQube Cloud

@codecov

codecov Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 99.20635% with 1 line in your changes missing coverage. Please review.
✅ Project coverage is 80.21%. Comparing base (688f928) to head (a7eab7c).

Files with missing lines Patch % Lines
openc3/lib/openc3/utilities/topic_lag_monitor.rb 99.09% 1 Missing ⚠️
Additional details and impacted files
@@            Coverage Diff             @@
##             main    #3952      +/-   ##
==========================================
+ Coverage   80.08%   80.21%   +0.12%     
==========================================
  Files         901      903       +2     
  Lines       68356    68618     +262     
  Branches     2645     2698      +53     
==========================================
+ Hits        54743    55042     +299     
+ Misses      12946    12909      -37     
  Partials      667      667              
Flag Coverage Δ
frontend 67.03% <ø> (+0.08%) ⬆️
python 80.24% <ø> (+0.10%) ⬆️
ruby-api 82.65% <ø> (+0.65%) ⬆️
ruby-backend 85.71% <99.20%> (+0.05%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.


def db_shard_for(topic)
# Target topics carry the target name as a Redis hashtag, e.g. SCOPE__DECOM__{TGT}__PKT
match = topic.match(/\{([^}]+)\}/)

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants