Repository navigation
LiDAR state estimation: reproducible evidence and measurement baseline - #559
Conversation
There was a problem hiding this comment.
Copilot wasn't able to review any files in this pull request.
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
9af7734 to
6d8c799
Compare
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
There was a problem hiding this comment.
🔵 Needs a closer look
It makes safety-relevant state-estimation/tracker behaviour changes across many packages while being labelled documentation-only, so it needs human verification of correctness, tuning safety, and the mislabeled scope.
Review details
- Files reviewed: 52/55 changed files
- Comments generated: 1
- Review effort level: Balanced
c863b09 to
a4ebf9c
Compare
There was a problem hiding this comment.
🔵 Needs a closer look
The change is large and cross-component (config schema, tracking, clustering, sweep objective, perf harness), the shown diff is only a subset of the actual runtime changes, and the description understates the scope, so it needs human review.
Review details
- Files reviewed: 122/127 changed files
- Comments generated: 1
- Review effort level: Balanced
913f7aa to
cfe54be
Compare
🤖 Version Bump AdvisoryWarnings
Version Updates✅ Radar: 0.5.1-pre35 → 0.5.1-pre36 📖 See CHANGELOG.md for detailed guidelines. This is an automated advisory. Review the detected changes and update versions accordingly. |
f927ae5 to
ee2a449
Compare
5ab136d to
9e69f1f
Compare
Investigation into the reported defect where a tracked vehicle's bounding box steps roughly one metre laterally for a few frames and returns. The defect was reproduced using production l4perception.DBSCAN and EstimateOBBFromCluster against a synthetic Pandar40P with zero measurement noise, a rigid box vehicle, and a straight path at constant speed: maximum frame-to-frame excursion 1.119 m, mean lateral bias 0.676 m. Root cause is the measurement definition, not the filter. computeClusterMetrics sets WorldCluster.CentroidX/Y to the medoid: the real LiDAR return nearest the arithmetic mean. A medoid can only sit on a visible face, so it carries a viewpoint-dependent bias of up to W/2 and hops between faces as visibility changes. Because that error is a smooth function of viewing geometry and stays correlated over tens of frames, it violates the zero-mean white-noise assumption, and CV, CA, CTRA, UKF and IMM all inherit it unchanged. The first increment must therefore be an observation model, which reorders the sequencing recommended in pipeline-review-open-questions Q5. Measured comparison of candidate measurements over the same 40-frame pass: medoid centroid (today) 0.676 m mean bias / 0.680 m max hop OBB centre 0.279 m / 0.565 m nearest OBB corner 0.202 m / 1.398 m (corner identity flips) corner plus temporal identity 0.337 m / 0.675 m near edge plus dimension prior 0.035 m / 0.370 m Further findings measured against sensor_data.db (55,315 tracks and 3,526,860 observations): - moving tracks are associated on only 43.6 % of sensor frames, so the effective observation rate is about 5 Hz rather than the sensor's 10 Hz - 11.3 % of moving tracks show a lateral excursion above 0.5 m in the already filtered output - lidar_clusters holds 0 rows and InsertCluster has no caller, so raw observations are never persisted - lidar_track_observations stores the Kalman estimate, not the observation - six quality columns in lidar_tracks are never written by InsertTrack - the visualiser drops the observation for every tracked object, and the debug collector that captures innovations is implemented but never wired The plan covers current architecture with source references, the five requested comparison matrices with per-criterion scoring and decision gates, proposed Go types and interfaces, observation uncertainty and residual designs, persistence schema with retention policy, evaluation corpus and partitioning, diagnostic harness requirements, and a nine-phase roadmap with per-phase files, tests, migrations, costs, risks and acceptance criteria. Documentation only. No runtime code changes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…iment Three changes, in response to review. 1. Split the plan. The single document held two workstreams with a strict one-way dependency: producing a trustworthy trajectory (L4b/L5) and measuring road-user behaviour from it (L7/L8). They have different decision types, different gates and different audiences, and combining them pushed the document past 3,000 lines. lidar-state-estimation-plan.md (renamed) owns Phases 0-5 and 8 lidar-behaviour-analytics-plan.md (new) owns Phases 6 and 7 Phase numbering is shared across both so cross-references survive. The estimation plan keeps the jerk-observability and empirical-path conclusions, because both are grounded in sampling-rate evidence that belongs with the estimator, and gains a Section 13 stating the contract behaviour analytics depends on. 2. Behaviour analytics specification. Incorporates the attached specification with literature grounding. Structure: five-way distinction between observable, surrogate, legal, population-relative and longitudinal statements; benchmark taxonomy; opportunity normalisation; a thirteen-group feature matrix with definitions, units, scope, map dependency, roadside suitability and benchmark kind; equations for THW, TTC, PET and DRAC; uncertainty propagation with derived suppression rules; L8 data model; dataset strategy; phased roadmap 6A/6B/6C/7. Measured constraint that shapes it: a typical vehicle passage is 9.6 s median, 42 observations, 51 m of observed path at 6-10 m/s, over 1,710 confirmed tracks. Four conclusions that diverge from the brief's starting assumptions: - SDLP cannot be measured here. It is the outcome of a ~1 hour standardised on-the-road test (Verster and Roth 2011). A ten-second passage is three orders of magnitude short. The roadside quantity gets a different name and must never be compared to SDLP norms. - The two-second following rule is driver education, not a research threshold. The 100-Car study records headway but defines near-crashes by evasive manoeuvre, not a headway cutoff. THW bands are no_established_threshold. - Point-count-style confidence is wrong for passing clearance: its sigma is dominated by extent, not position, so cyclist width must come from a class prior rather than a per-track estimate. - PET needs no map. A conflict point can be derived from observed path intersections, which moves PET into Phase 6B. Inspection finding: there is no posted speed limit anywhere in the schema. The only one in the codebase is a per-request report parameter. Legal speed benchmarks need it in site_config_periods, which already carries the effective-date pattern. DOIs are given only where directly observed; publisher URLs are used elsewhere rather than constructing plausible-looking DOIs. The SSAM PET default of 5 s is flagged as secondary, since the FHWA techbrief and validation report state the TTC default and not the PET default. 3. Experiment E1 on the soma static captures. Verified the four recordings: 38 m 01 s total, ~22,810 frames, 5,062 MB, all Hesai Pandar40P on port 2369, same sensor, 2025-12-06 across four placements. "static" means the sensor was stationary, not that the scene was empty. They give viewpoint diversity and nothing else, which happens to be exactly the axis the medoid-bias hypothesis depends on, and the axis kirk0 cannot probe. There is no ground truth, so all four tests are designed to need none: the decisive one bins signed lateral offset by aspect angle, where random error has a zero conditional mean and a geometric bias does not. Includes partition assignment across the plan's three partitions, the settling budget constraint (soma2 is 69 s against a 60 s settling duration), and an explicit statement that a negative result fires Section 15's first invalidating condition rather than counting as experimental failure. Documentation only. No runtime code changes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…tion Architectural corrections from review. Preserves the investigation, matrices, gates, experiments and measured findings; changes the architecture around them. Cross-cutting: added an explicit principles section. No black boxes, and estimate the most probable physical path rather than cosmetically smoothing. Product language moves from "smoothed trajectory" to "final trajectory estimate" and "retrospectively refined estimate"; algorithm names stay technical. Both principles were being violated, which is why they are now stated first. State-estimation plan: - Split Observation into immutable DetectionObservation and derived, versioned MeasurementInterpretation. The previous type mixed sensor evidence with fields requiring a track prediction, so two estimator versions could not consume the same evidence reproducibly. Recorded as defect P12. - Resolved the state dimensionality contradiction. The draft proposed a six-element [x,y,psi,v,a,omega] state AND a linear filter; those are incompatible. Option A adopted: [x,y,vx,vy] with a 4x4 covariance, with orientation, dimensions and vertical position as separate beliefs. This keeps the filter honestly linear and isolates the observation-model change that G-GEO-1 exists to test. Gate G-EST-4 added for migrating to Option B, with low-speed conditioning as the deciding criterion. - Added road-user motion models: rigid_vehicle, two_wheeler, pedestrian, unknown, with prior strength scaled by class posterior so classification uncertainty never becomes motion certainty. - Added an estimation lifecycle distinct from the track lifecycle, with explicit initialisation rules. The position seed keeps the medoid but sizes its covariance to the known bias rather than the medoid's precision. - Removed the medoid fallback during model degradation. It silently redefined X and Y mid-track from estimated physical pose to raw cluster point, which was the most dangerous line in the previous draft. - Made abnormal-motion thresholds class-conditioned. A 90-degree heading change is a spin for a car and an ordinary turn for a pedestrian. - Replaced the dimension sketch with a frame-admissibility taxonomy, a bounded class-anchored quantile estimator, and three defences against the merged-cluster ratchet. Elevation, per the answer that deployment sites are graded: - Added defect P11: ground removal is a flat height band on absolute sensor-frame Z and is documented as not slope-aware. On a grade it clips vehicles at one end of the scene and admits ground at the other, corrupting the cluster extents the near-edge measurement depends on. Remedy pulled into Phase 1. - P11 also threatens Experiment E1: on a straight graded approach, range correlates with both filter error and aspect angle, so a ground artefact could fake or mask the aspect-conditioned bias E1 exists to detect. Added a three-step mitigation: measure grade first, record GroundClipped per observation, stratify E1.1 by range as well as aspect. - Added the road-surface frame, the separation of road-user dynamics from road-surface geometry, discontinuous intersection surfaces, and a capability split. Substance stays in the L7 scene plan; this plan owns the interface and the honest planar fallback. - Added class-general and grade-aware synthetic test coverage. Behaviour plan: - Development is unblocked from G-SMO-1; production emission is not. The gate now covers the genuinely dangerous case, computing metrics from today's biased tracks, while allowing work against fixture trajectory streams. - Added a metric framework ahead of the feature matrix: three-tier road-user scope, declared per-metric applicability, a closed SuppressionReason vocabulary, an observation-support taxonomy separating occluded_inferred from missed_unknown, and interaction classification as a posterior that gates which formulae may run. - Split results into BehaviourMeasurement and BehaviourOutcome so categorical outcomes keep references to the evidence they were derived from. - Generalised passing clearance to minimum synchronised surface-to-surface separation, with lateral clearance as a projection for confident overtakes. - Replaced scalar sigma with an Uncertainty struct supporting intervals, bounds and a propagation method, and carried three uncertainty scopes so pairwise common-mode cancellation is a later refinement rather than a migration. - Marked jerk experimental with a published-boundary acceptance test. - Separated engineering dependency order from product priority, with vehicle-cyclist passing clearance leading the product list. The two orderings are close to inverted at the top. Both documents end with a "Changes introduced by this revision" section and answer the questions the revision raised. Documentation only. No runtime code changes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Decisions taken, recorded in state-estimation Section 21.1: - D1: P11 (slope-unaware ground removal) is a current defect, not future work. Deployment sites are graded. Severity still to be confirmed by measuring the grade per capture. - D2: ship the OBB centre as an immediate measurement stopgap. 0.279 m mean lateral bias against the medoid's 0.676 m, a 2.4x improvement for a change of measurement source. Fenced with four conditions: it is still a visible-surface artefact, validate via E1.3 first, persist which source produced each estimate, and re-baseline G-GEO-1's regression numbers afterwards. Without the source field the stopgap silently splits the historical record into two incomparable regimes, which is the same class of error as the medoid fallback removed from Section 12. - D3: gate set confirmed as G-PER-1, G-GEO-1, G-UNC-1, G-EST-1, G-SMO-1. - D4: product priority leads with vulnerable-road-user interactions. Experiment E1 corpus corrected against the real capture set: - soma2 dropped: at 69 s against a 60 s settling duration it cannot yield a meaningful partition. - soma1 and soma3 re-split, now carrying a -0-1 segment suffix. - clar0-1 added, and it matters more than the rest because it is a different site rather than a fourth placement at the same one. It becomes the held-out regression partition for that reason. - Row values measured on the superseded splits are marked for re-measurement. /Volumes/lidar was not readable from this session, so only kirk0 was re-verified on disk. Added the two known-defect VRLOGs (0fb02f22 from clar0-1, 60a4774c from kirk1) as the labelling source for the regression set, with an explicit note that a VRLOG replays decisions already made and so cannot run a candidate measurement. They supply labelled cases; E1 runs from the source captures. The two were produced by different builds under different tuning hashes and are not comparable to each other. Simplifications: - Estimator matrix cut from six rows to three. CTRV, CTRA and a stationary mode all depend on evidence that does not exist, and carrying them implied a choice that is not open. - G-EST-2/3/4 folded into a deferred-conditions table. Writing thresholds for them now would be guessing. - Behaviour Section 8 gains a scope table: eight groups in, five deferred. Note this resolves toward the confirmed VRU-first product priority rather than the literal groups 1-3 plus 9 originally proposed, because that set would have deferred PET and yielding, which are product priorities 2 and 3. - Shared production-database figures now cite state-estimation Section 1.5 as canonical instead of being restated in both plans. Backlog: nine entries added, five under v0.5.5 tracker correctness and four under v1.0, including the missing posted speed limit in site config. Documentation only. No runtime code changes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…105d8) Run f84105d8-b3be-416f-8809-551ef6bfce10 supplies real-data confirmation of the mechanism that Section 3 previously demonstrated only synthetically, through a second symptom the synthetic model predicts but that had not been checked. 6,846 frames, 11 m 03 s, 2,038 tracks, replayed from soma1-static-0.pcap at playback rate 0.1. The slow playback is a control: the frame-rate throttle was not engaged, so nothing below is a throughput artefact. Estimated width collapses on moving tracks. Of 288 tracks with max speed at or above 3 m/s, 172 (60 %) carry an estimated width below 0.5 m, narrower than a pedestrian, and 210 (73 %) are below 1.0 m, narrower than any car. The mode is 0.1 to 0.3 m, which is the thickness of a single observed face rather than a plausible object width. That is the same cause as the position bias. With only the near face visible, the medoid sits at W/2 from the true centre, and the extent across the object collapses to the face's own thickness. One mechanism, two symptoms, and the second is now measured on real traffic. Corroborating, same run: 1,648 of 2,038 tracks (81 %) carry no class; the 15 classified car have a mean width of 0.85 m against a real 1.8 m; mean observations per moving track is 6. The track in the inspector screenshots, trk_4a27c73c, is classified car at 8.0 m/s with L x W x H of 0.8 x 0.2 x 0.2 m and Hits 0 while Confirmed. Consequences recorded: - P7 restated. It is not that a running mean is biased low by partial views; the per-frame extents are themselves the wrong quantity under single-face visibility. Severity raised to High, since it starves the classifier and breaks every downstream clearance measurement. Section 9.2's admissibility rules are load-bearing rather than fastidious. - The 81 % unclassified rate is arithmetic, not an independent defect. The classifier reads dimensions and speed; sub-metre widths starve it. - Mean 6 observations per moving track is well below the 42-observation median in Section 1.5. The tuning hash differs from production, so it is recorded as run-specific rather than as a new baseline. f84105d8 promoted to primary labelling source in 16.5, ahead of 0fb02f22 and 60a4774c. Two properties noted so they are not rediscovered: playback 0.1 as above, and its source being the superseded soma1-static-0.pcap split rather than the current soma1-static-0-1.pcap, so a re-run will not reproduce it frame for frame. Cases are to be labelled by what they show, not by frame index. Documentation only. No runtime code changes. All queries were read-only against a WAL database and no live server was touched. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…haviour analytics
…1.1) Nothing in the pipeline measured whether a bounding box points where its vehicle is going. AlignmentMeanRad compares Kalman velocity against displacement: both describe motion, neither describes the box. HeadingJitterDeg measures how much the box moves, so a box locked at the wrong angle scores perfectly on it. Add course alignment: |OBB heading - direction of travel|, folded to [0, 90]. An OBB is symmetric, so a box pointing backwards along the course is correctly oriented and folds to 0; a 90 degree length/width swap is the worst case. Sampled only on live frames at or above 2 m/s, where course is meaningful. The tracker holds it as a 20-bin histogram so per-track cost is constant on the Pi; the offline analysis path computes percentiles directly. Measured on run baf20f02 (600 frames, 346 tracks), across the 55 tracks with samples: median per-track course error 50.3 deg, p85 67.2 deg, worst 85.0 deg. Track 18952226, one half of the split car in the report, sits at 66.4 deg while its box moves 1.95 deg per frame. Jitter called that track healthy. Two defects surfaced while wiring it up, fixed here: - The offline and live HeadingJitterDeg measured different quantities. The analysis path used Track.HeadingRad, the Kalman course; the tracker used OBB heading deltas. These are the two paths an A/B comparison puts side by side. The box quantity is now reported separately as OBBHeadingJitterDeg, and both fields state which is which. - Run-level aggregates key on final track state, and 211 of 346 tracks end DELETED, so the confirmed-only rollups covered 13 tracks. Course alignment rolls up over every track that produced samples instead. The existing jitter and alignment aggregates still have this flaw and are understated. Also adds the two-day sprint plan this is the first task of, with the measured evidence for the heading lock ratchet and the objective function that rewards it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Guard 3 in tracking_update.go rejects any OBB heading delta between 60 and 120 degrees, measured against the smoothed heading. Once that smoothed value has itself drifted more than 60 degrees from the truth, every correct measurement is rejected and the lock cannot release. Nothing recorded whether that had happened to a track. Add per-track HeadingLockedFrames, LongestLockRun, EnteredSustainedLock and ReleasedAfterLock, with a derived LockTrapped for the case that matters: a sustained lock that never released. A release needs five consecutive unlocked frames, because Guard 3 rejects per frame and a single frame slipping through between rejections is not the lock letting go. Both the live tracker and the offline analysis path compute this, and a test asserts the two agree, since they are the two sides of every A/B comparison. Measured on run baf20f02, over the 239 tracks living at least five live frames: 83 tracks (35 per cent) enter a sustained lock and 54 of those (65 per cent) never release. The longest single run is 152 frames, 15.2 seconds. Splitting tracks by that outcome, trapped tracks sit at a median 60.1 degrees of course error against 38.5 for the rest, so the lock costs about 22 degrees. That is the measured case for the D1.3 ratchet fix. Fixes a third defect found while wiring it up. The offline reconstruction counted DELETED ghost frames, which carry heading source pca rather than the lock the track died in, so fifty frames of ghost turned a track trapped for its whole life into a clean one: trapped read 11 per cent of locked tracks instead of 65. Lock stats now skip non-live frames. Excluding them dropped the pca frame count from 15,069 to 4,459, a difference of 10,610, matching the 10,610 DELETED track-frames counted independently. Run-level ratios are also taken only over tracks that lived long enough for a sustained lock to be detectable. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… (D1.3) Guard 3 rejects any OBB heading delta between 60 and 120 degrees, measured against the smoothed heading. That comparison is what makes it self-sustaining: once the smoothed value has itself drifted more than 60 degrees from the truth, every correct measurement lands inside the rejection band and is thrown away, so the lock can never release. On run baf20f02, 65 per cent of locked tracks never recovered, sitting a median 60 degrees from their direction of travel against 38 for the rest. Add a rejection counter. After OBBHeadingLockMaxRejections consecutive Guard 3 rejections the tracker stops believing its own smoothed heading and accepts the measurement. Default 5; zero restores the previous behaviour. Two details the fix depends on: - The release snaps rather than eases. The EMA moves 8 per cent of the gap per update, so easing across a delta wide enough to be rejected would re-trigger the guard on the next frame and never converge. - Only Guard 3 drives the counter. Guards 1 and 2 fire when the cluster is genuinely unusable, too few points or too near square, and snapping to such a measurement would replace a wrong answer with a random one. Releasing also restores dimension updates, which are only written when the heading update is accepted, so a released track stops carrying the frozen length and width it locked with. Proven on synthetic sequences reproducing the trap: a confirmed track whose smoothed heading sits 90 degrees from its course, fed correct measurements, stays above 80 degrees of course error with the release disabled and converges below 5 with it armed. The two pre-existing Guard 3 tests still pass. The end-to-end figure on baf20f02 is not measured here. A VRLOG replays decisions already made and cannot re-run the estimator, so confirming the fix needs a pipeline re-run over the source PCAP. Config is strict on missing keys by design, so the new parameter is added to the three shipped tuning files, to config-migrate (writing the shipped default rather than the zero value, which would quietly migrate a config back onto the ratchet), and to the runtime tuning endpoint's accepted keys. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
….5, D1.6) D1.5, fragment guard. The association gate is a 6 m radius on position alone, with nothing to stop a scrap of a cluster capturing a vehicle-sized track. On run baf20f02 a 0.11 by 0.08 metre cluster was associated to a track carrying a 4.33 metre car and became its dimensions for the next 28 frames. Forbid the pairing in the cost matrix rather than refusing it at update time, so the fragment stays unassociated and can seed its own track instead of being consumed. The guard fires only when a track has at least three observations, believes it is at least two metres, and the cluster's longest extent is below min_associable_extent_metres (default 0.5). It reads the running average extent, not the latest frame, because the latest frame is exactly the value a fragment would already have corrupted. The guard is deliberately narrow. Pedestrian and cyclist tracks legitimately carry sub-metre extents whose clusters vary by a similar amount frame to frame, so a small-cluster rule applied there would reject ordinary observations. Dimension consistency in the assignment cost proper is a separate change. D1.6, ghost fade. Deleted tracks are published with a fade-out alpha, which is a deliberate rendering feature. The defect was that its duration was deleted_track_grace_period: one number doing two unrelated jobs. That period exists so re-association can still find a track seconds later, and borrowing it as a render timer held a frozen box on screen for five seconds while the real object drove out from under it. It also corrupted every metric computed over the recorded stream, as D1.2 found. Split out deleted_track_render_fade, default 500ms. Re-association is untouched: the track stays in the map for the full grace period, it simply stops being drawn. Measured exactly on run baf20f02, since the fade is a pure filter on publication and needs no pipeline re-run: published DELETED track-frames fall from 10,610 (45.8 per cent of all track-frames) to 1,170 (5.0 per cent), an 89 per cent reduction in ghost frames. Both parameters follow the D1.3 pattern through the strict config schema: the three shipped tuning files, config-migrate defaults, and the runtime tuning endpoint's accepted keys. GetRecentlyDeletedTracks now windows on the render fade, so its tests are renamed to say so rather than asserting the old semantics. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
9e69f1f to
3096909
Compare
There was a problem hiding this comment.
Copilot review overview
🔵 Needs a closer look
The change is very large and touches sensor-data pipeline internals, DB migrations, and concurrency-sensitive code, which per project guidance require final human review beyond the concrete web bug flagged here.
Review effort: Balanced
Findings: 1
Open (1)
Resolved since last review (1)
4c37371 registered all 24 S2 archive sites in the corpus without updating this test, which still expected the original three cases and 16 captures. Pin the corpus as registered (24 cases, 110 captures), require it to lead with the Phase 0/1 cases, and keep those three multi-file, since they exist to exercise PCAP roll-over joins. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ss 7 The campaign put its supervisor, its runtime state, 1.1 MB result tables and one-off launchers in data/experiments/try/, which on main holds experiment write-ups only, and its manifests carry one workstation's paths. Move them to backup/state-est-experiment-data-20260924 and the LiDAR volume (velocity-campaign/state-est-experiments-20260924/, with the untracked raw output that was never in git). Keep the write-ups, OBJECTIVES.md and the pass-7 campaign definitions the analysis worker documents as its format. References to archived files become plain archive names, which also clears the five dead links in the campaign plan and the CSV link the offline docs site could not resolve (it does not copy .csv), breaking Build, Static build and Offline docs CI. Pass 7's findings lived only in commit messages and the moved tables, so the campaign plan now records them. Every ground-truth score in the campaign was matched over lidar_run_tracks rows written before #584 stored tracks as they finished, so the plan, the L4 write-up and the gap analysis now say those counts are not evidence, and the gap analysis no longer calls K10's gate reading corroborated. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…s volume lidar-state-estimation-baseline defaulted -pcap-root to /Volumes/lidar/lidar, the path main moved out of the Makefile into local.mk. The Makefile's evidence targets and the analysis worker both pass it explicitly; an empty root would otherwise resolve captures against the working directory. Add it to the E1 reproduction command. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Restore main's requirements (ruff unpinned in requirements.in, 0.16.8 in requirements.txt) instead of pinning back to 0.15.17. Ruff 0.16 widened its default rules: against this tree `ruff check .` reports 229 findings and `make format-python`, which the pre-commit hook runs over the whole repository, would rewrite 31 files, including removing noqa directives. Select the rules ruff applied before (E4, E7, E9, F) in pyproject.toml, so both versions pass `ruff check .` and 0.16.8 proposes no fixes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
scripts/test_lidar_jump_candidates.py was added without being listed in PYTHON_TEST_PATHS, so neither make test-python nor CI ever ran it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
runScan(true) set scanning and relied on the scan poll to clear it, but a failed scanCaptureRoots request never starts that poll, and the finally block only cleared it for quick scans. The spinner stayed on and both scan buttons stayed disabled until the page was reloaded. Route pages are not mountable under this Jest setup (.svelte imports map to a stub), so this was checked in the browser instead: with the scan endpoint returning 500, the buttons stayed disabled without the fix and re-enable with it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The root package.json pins packageManager to pnpm@12.6.0 (#592). pnpm 11+ records that pin, and the platform binaries it would install, as a leading document in pnpm-lock.yaml. Without it, `pnpm install --frozen-lockfile` at the repository root fails ("Cannot update packageManagerDependencies with frozen-lockfile"), and any non-frozen root install (the render target) rewrites the file and leaves the tree dirty. Checked before committing: all 15 entries (pnpm and the 14 @pnpm/exe platform binaries, musl variants included) match the npm registry's integrity hashes, the packageManager sha512 matches the published tarball, and the provenance attestations point at pnpm/pnpm release.yml at v12.6.0. The platform packages are optional: each machine installs only the one matching its os, cpu and libc. CI installs only in web/, public_html/ and docs_html/, each with its own lockfile, so this file does not reach CI. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A mid-September renumbering on this branch moved six plans from v0.6.0 to v0.6.2; the later backlog rewrite put Deployment & packaging back at v0.6.0 and made v0.6.2 backpack tailgating, so those plans pointed at the wrong release. Four differed from main only in that number and are restored to main's text; platform simplification's row reference and the multi-file plan are corrected in place. - TDL: steps 1-2 (the LiDAR-only transit record and behaviour labels) stay at v0.5.3 with behaviour analytics; the query language, radar union and description interface are unscheduled, as the backlog has them and D-01 defers the fused transit schema. Metrics registry follows. - Multi-file cases: complete in #569; only the pcap_file projection retirement remains (v0.6.6). - Milestones matched to the backlog: cmd extraction v0.5.8, stream robustness v0.5.5, replay-case terminology v0.5.8, structured logging unscheduled, heading sprint D2.4 at v0.5.7, state plan's Pi cost v0.6.7. - Motion-capture architecture now points at the route plan, which owns the near-term backpack and bike path. - Anchors: the sprint 0.5.2.0 backlog link, the behaviour plan's target-platform link, and the imager-fork section 8 link. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ons record Removed, as agreed: - data-sqlite-schema-v2-freeze-plan: froze the SQLite schema for a v2 rebuild, which contradicts the shared VRLOG storage design (VRLOG as the observation authority, SQLite as rebuildable catalogue) and this branch's own new tables. Unscheduled; nothing linked to it. - lidar-state-estimation-branch-audit: a dated review snapshot of PR #559 that claimed to own the cross-plan sequence, which the backlog owns. Its durable rule (Phases 0-2 are the core; a lower course-error statistic, an annotation exporter or a new motion filter is no substitute for the position correction) is now in the state plan's delivery declaration. - lidar-heading-d2-readiness-review: the pre-implementation handoff, superseded by the D2 implementation report. The 2026-09 campaign is a finished record rather than a plan, so it moves to docs/lidar/operations/parameter-experiment-campaign-2026-09.md with its links rebased. Every inbound link is repointed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…sured
The label-free scorecard classifies how a track ended by comparing its
prediction with the clusters that follow, and it read every cluster at its
OBB centre whatever the run measured. For a medoid run that compared
centroid-based predictions with points the tracker never followed, which
skewed the contested / unassigned-nearby / vanished classes, including in
the D2 A/B, where that comparison is now discounted.
ListClusterSummariesBySource takes a ClusterPosition, and the scorecard
picks it from the run's own observation_model_id. A test pins it: a medoid
run whose OBBs sit 3 m off the path still finds its object beside the
prediction ("vanished" before the fix).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…tre opt-in D2 shipped the OBB centre as the live Kalman measurement on the condition that it be validated on real data first. E1 could not do that (its reference path was fitted to OBB-centre estimates), so it was A/B tested against kirk0's annotated road users: 30 objects, 771 frames, a footprint-scaled gate both positions sit inside. The OBB centre matched recall (0.534 vs 0.527) but switched identity 146 times against 92, fragmented more (64 vs 55), and gave noisier, slightly less accurate vehicle speeds. The direction held at every gate tried. Headway pairing needs stable identity, so the OBB centre does not ship. Empty measurement mode now means the medoid, as every release before D2; obb_centre_v1 remains selectable in the tracker, replay harness and baseline tool, and InterpretMeasurement still records the OBB-centre candidate beside the near face for E1, so Phase 2's comparisons are unaffected. On kirk0 the default now confirms 25 tracks, main's figure. The throttle test counted association samples, which pool live tracks only, so it measured whether one synthetic track survived a 200 fps burst rather than whether frames reached tracking. It now counts tracker updates. The state plan records the outcome as D5 with the evidence; D2's condition-4 re-baseline no longer applies. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The whole capture replayed through the live server once with this branch (merged with main's #593) and once with main. Heading acceptance rose from 0.605 to 0.789, median course alignment fell from 51.9 to 24.0 degrees and tracks ending in an unrecovered lock fell from 0.449 to 0.099, while track counts, lifetimes and speed percentiles agree within noise. Box-overlap frames rose by about 8%, recorded as the figure to check against annotated truth before headway pairing relies on it. The record also notes the start-up panic outside the repository tree that #593 fixes on merge, and that replay filters on the server's own UDP port, which is why each arm replayed a port-rewritten copy verified packet by packet. Linked from the heading sprint plan (6.3) and the state plan. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The OBB-centre A/B and the medoid staying the production measurement, the scorecard position fix, the Columbus Broadway replay against main, the start-up panic #593 resolves on merge, the campaign archive, and the plan realignment and CI repair. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e new plans Adds September 22-25 entries: the annotation branch's plans PR (#585), a heavy planning-and-engineering day covering the job runner, two annotation fix rounds, and five new design docs (#580, #582-592), the job runner's first real-VM validation and its three bug fixes (#593), and today's state-estimation landing, plan-hygiene graduation, and the TrueNAS worker VM storage fix (#559, #594, this branch). Also cleans up the {dd/lidar/annotation} branch tag from all September 7, 8, 16, 17, 19, 20, and 21 entries now that PR #579 has merged to main. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
* [ai][docs] add missing blank line after lidar-market-watch shortlist table The table's last row ran straight into the following prose with no blank line, so Markdown table parsers (including the format-docs prettier hook) swallowed the prose into the table as malformed extra rows. Found while committing an unrelated doc from the same worktree: format-docs reformats the whole docs tree, not just staged files, so this was blocking every docs commit until fixed. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * [ai][docs] TrueNAS arrow to VM bansheeworker local shared storage runbook Records the investigation into giving VM bansheeworker a local, Tailscale- independent path to mount NFS exports from TrueNAS host arrow's pool: the macvtap isolation problem, six failed bridge attempts and their common cause, a CLI shortcut that caused a real outage, the confirmed root cause (eno1 never had a database row, plus a kernel-level bridge/macvtap rx_handler conflict), and the working solution (register eno1 alone first, then the isolated bridge, then attach the VM's second NIC and repoint fstab). Verified end to end, including across a guest reboot. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * [ai][docs] update devlog: state estimation, job runner hardening, five new plans Adds September 22-25 entries: the annotation branch's plans PR (#585), a heavy planning-and-engineering day covering the job runner, two annotation fix rounds, and five new design docs (#580, #582-592), the job runner's first real-VM validation and its three bug fixes (#593), and today's state-estimation landing, plan-hygiene graduation, and the TrueNAS worker VM storage fix (#559, #594, this branch). Also cleans up the {dd/lidar/annotation} branch tag from all September 7, 8, 16, 17, 19, 20, and 21 entries now that PR #579 has merged to main. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * [ai][docs] fix overclaiming status line in TrueNAS runbook Addressed PR review comment: the opening status said "solved and verified" while section 6 explicitly lists a host reboot as untested. The status line now names exactly what has and has not been observed (guest reboot verified, host reboot not yet verified). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
…1 hang and hardening (#620) * [ai][docs] add missing blank line after lidar-market-watch shortlist table The table's last row ran straight into the following prose with no blank line, so Markdown table parsers (including the format-docs prettier hook) swallowed the prose into the table as malformed extra rows. Found while committing an unrelated doc from the same worktree: format-docs reformats the whole docs tree, not just staged files, so this was blocking every docs commit until fixed. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * [ai][docs] TrueNAS arrow to VM bansheeworker local shared storage runbook Records the investigation into giving VM bansheeworker a local, Tailscale- independent path to mount NFS exports from TrueNAS host arrow's pool: the macvtap isolation problem, six failed bridge attempts and their common cause, a CLI shortcut that caused a real outage, the confirmed root cause (eno1 never had a database row, plus a kernel-level bridge/macvtap rx_handler conflict), and the working solution (register eno1 alone first, then the isolated bridge, then attach the VM's second NIC and repoint fstab). Verified end to end, including across a guest reboot. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * [ai][docs] update devlog: state estimation, job runner hardening, five new plans Adds September 22-25 entries: the annotation branch's plans PR (#585), a heavy planning-and-engineering day covering the job runner, two annotation fix rounds, and five new design docs (#580, #582-592), the job runner's first real-VM validation and its three bug fixes (#593), and today's state-estimation landing, plan-hygiene graduation, and the TrueNAS worker VM storage fix (#559, #594, this branch). Also cleans up the {dd/lidar/annotation} branch tag from all September 7, 8, 16, 17, 19, 20, and 21 entries now that PR #579 has merged to main. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * [ai][docs] fix overclaiming status line in TrueNAS runbook Addressed PR review comment: the opening status said "solved and verified" while section 6 explicitly lists a host reboot as untested. The status line now names exactly what has and has not been observed (guest reboot verified, host reboot not yet verified). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * [ai][docs] runbook: read-only captures confirmed, media share added read-only Records that captures was already ro: true server-side, so no change was needed there. Adds the new read-only media NFS share for the VM: the NFS ro flag alone does not grant filesystem read access to the mapped user, so an NFSv4 ACL grant (recursive) was needed on top of it, plus a note on the one directory the recursive apply missed and how it was found. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * [ai][docs] runbook: disk relocation, CPU passthrough fix, eno1 hardware hang Adds three new incident sections: relocating the guest's /home onto the already-provisioned /srv partition to relieve root disk pressure, the VM CPU Mode fix (Custom -> Host Passthrough) for a Bun native-binary spin bug on a CPU model missing SSE4.2/POPCNT, and a hardware e1000e TX hang on the host's eno1 that also orphaned the VM's macvtap NIC and needed a VM restart to recover. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * [ai][docs] runbook: eno1 hardening (EEE/TSO, watchdog), bridge attempt failed Adds section 12 documenting three mitigations attempted after the hardware hang in section 11: disabling EEE/TSO on eno1 (both persisted via Init Script), a systemd watchdog that auto-recovers from a recurrence (driver reload + VM restart, 120s cooldown), and an attempted br0-with-eno1-member bridge that failed after a full 5-minute test window with no DHCP lease and no error, most likely a switch-side BPDU Guard reaction. The bridge attempt is documented but not pursued further; the other two mitigations are live. Updates the stale section 7 bullet accordingly. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
…an it feeds (#670) Recovers work left uncommitted on 2026-09-22 in the Codex worktree for codex/evaluate-decoupled-track-processing (after #582): - docs/plans/lidar-vrlog-recording-contract-review.md (new): settles the vocabulary and ownership needed before VRLOG becomes the durable L4 observation boundary. One processed-observation recording with an explicit evidence contract; L5+ results and annotations in the package catalogue keyed to capture and run identities; web scenes as derived exports; PCAP as an optional earlier L1 source. Separates the six questions currently all called "profile". Proposed; no wire format or runtime change. - The asynchronous tracking plan takes the review's terms and cross-links it (+168/-43), applied three-way over #559's later edits to the same plan. Prettier-formatted (whitespace only), and the review's Related line gains the shared VRLOG plan, which landed the day after it was written. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…paths (#671) Carries the two still-valid items from the retired SQLite schema v2 freeze plan (removed in 510415a, before #559 merged) into the two-phase programme; nothing else in it survives the shared-VRLOG design. - The 000055-pattern bullet: every new or rebuilt foreign key declares its ON DELETE action and has a delete-path test, with PRAGMA foreign_key_check after the move. - P2-B acceptance: deleting a clip has a declared, tested effect on its labels (lidar_replay_annotations.replay_case_id cascades today, so it deletes reviewed labels), and nothing writes pcap_file once it is a read-only projection. - P2-F acceptance: deleting a site has a declared, tested effect on its deployments (site_config_periods.site_id cascades today). Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>


The measurement and evidence foundation the state estimation plan calls for, the heading fixes that could be validated without it, and the plans that put headway (tailgating) first. On every device it changes how the box heading behaves; it leaves the tracked position, track population and speeds where main has them.
Shipped, on by default
obb_heading_lock_max_rejections: 5): a heading stuck behind repeated guard rejections now releases and snaps to measurement instead of holding forever.min_associable_extent_metres: 0.5): a sub-metre scrap can no longer overwrite a vehicle-sized track's dimensions or seed a duplicate on it. It applies only to tracks whose extent belief is 2 m or more, so pedestrians and cyclists are unaffected.deleted_track_render_fade: 500ms): a deleted track's frozen box fades in about half a second instead of five.Replayed against main on the whole Columbus Broadway capture at 0.25x (record):
Bumper-to-bumper gap is measured along the box, so a box that points the way the vehicle is going is a precondition for the headway metric. Frames with two overlapping boxes rose from 11,843 to 12,849 of 20,599: a proximity signal, partly from longer-lived tracks, to check against annotated truth before headway pairing relies on it.
Position measurement: the medoid stays
D2 proposed filtering the OBB centre instead of the medoid, on condition that it be validated on real data first. E1 could not do that (its reference path was fitted to OBB-centre estimates), so it was A/B tested against kirk0's annotated road users (30 objects, 771 frames). Recall was level (0.534 vs 0.527), but the OBB centre switched identity 146 times against 92 and fragmented more (64 vs 55). Headway pairing needs stable identity, so the medoid stays the production measurement and
obb_centre_v1is opt-in (state plan 21.1 D5). On kirk0 the default confirms 25 tracks, main's figure.Experimental, off by default
Reachable from the replay harness (
-experiment,-measurement-mode), not the live default:obb_axis_coherence_enabled) and D2.2 bounded association shape cost (association_extent_cost_weight), both awaiting physical-object acceptance.Evidence and replay infrastructure
lidar_observations,lidar_track_estimatesandlidar_track_residuals. Only the offline replay harness writes them, so they stay empty on live devices. The livelidar_track_observationstable gainsframe_unix_nanosandmeasurement_source.lidar-state-estimation-baseline,lidar-track-scorecard(label-free, plus per-frame MOTA/MOTP/IDSW/FM/HOTA against a reference run),lidar-e1-analysis,lidar-closeness-auditandlidar-evidence-oracle. These are also the two job kinds main's analysis worker ([ai][go][docs] velocity worker and velocity jobs: a job runner over verified captures #586) runs.pi,mac,ci) is now a baseline dimension, with the Mac baselines recaptured and a manual Pi runbook.Web
/app/lidar/tracksrenders through the shared three.js scene player. Track picking and missed-region marking moved into the 3D view, and the 932-line flat canvas map is gone./lidar/capturesno longer leaves the page stuck scanning.Plans and documentation
backup/state-est-experiment-data-20260924and the LiDAR volume. Its ground-truth scores are marked unreliable because they matched first-sighting rows (see [ai][go] store analysis-run tracks as they finished, not as first sighted #584).Before and after merge
MustLoadDefaultConfiglooks for the tuning file on disk. Main's [go] fix lidar worker crash #593 adds the embedded fallback, and merging brings it in. The branch is one commit behind main.--configfile needsconfig-migratebefore it will start.OnMain: truefor thestate_estimation_baselineandtrack_scorecardjob kinds (internal/lidar/jobs/kinds.go).--lidar-udp-port, and the request cannot override it.pyproject.toml. The root lockfile records the pinned pnpm (hashes and provenance verified).Known gaps
piperformance baseline exists yet.Checklist
DESIGN.md.README.mdis up to date.docs/DEVLOG.mdis up to date.🤖 Generated with Claude Code