Skip to content

Add model-declared ARC rollout controls and episode recording - #213

Open
ElmoPA wants to merge 9 commits into
graphite-base/213from
codex/graph-consolidation-stack-20261007/15-arc-token-rollout
Open

ElmoPA wants to merge 9 commits into
graphite-base/213from
codex/graph-consolidation-stack-20261007/15-arc-token-rollout

Conversation

@ElmoPA

@ElmoPA ElmoPA commented Oct 7, 2026 •

Copy link
Copy Markdown
Collaborator

Add the reviewed capabilities from aidan/arc-token-shapes-20261001 at b1ecba63 above #212: additional token layouts, motion/hold replay speed, fastest-stream prefix execution, episode recording, checkpoint substring selection, HDF5 tempo diagnostics, gripper controls and camera cleanup.

Controls bind through typed model-owned declarations and pure decoder interfaces. Current Cartesian profiles explicitly select canonical200; historical M28/PR193 and E1 codecs require explicit identities. Per-arm rotation clocks remain independent, and tri codecs are decode-only. ARC flow rollout defaults are 20 sampler iterations, 50% execution and fastest-stream mode; training neural/optimizer/scheduler settings stay identical to the BC parent.

Recording initializes files in its worker, keeps close/discard requests outside the bounded row queue and distinguishes planned from written commands. Gripper updates preflight all arms, cap force at 50 N, verify rollback and close newly opened drivers if initialization fails. Failed model replacement clears prior observation history.

Validation: 1,165 affected CPU tests passed; all 469 shipped configs resolve; Ruff 0.8.6 lint/format and both changed JavaScript syntax checks pass. A fresh checkout restored all 402 artifacts from companion pin c6b5febd and matched all 76 tested source hashes. Tests, exact donor dispositions and evidence are in #211. Physical robot performance and historical checkpoints were not evaluated.

Stack: #198 → #212 → #213. The read-only combined merge preview retains exactly the four preexisting archive-test conflicts with current main and adds none. No merge to main or station deployment is included.

ElmoPA and others added 9 commits September 23, 2026 22:12
…itted files) as port baseline

Verbatim copy of ~/dev/EgoVerse uncommitted edits on 2026-09-30, cmp-verified, so the
ARC replay-speed port (1ebeabdf) lands on exactly what the station runs.
Port of 1ebeabdf (aidan/arc-rollout-speedup-20260920) onto the current station
code. The decoder warp is the original: speed scales moving phases, hold_speed
the holds (both arms under hold_threshold), one monotone wall-time to
token-clock map shared by both arms, applied before the codec SLERP; 1.0 / 1.0
is the unmodified decode. Lab and the cartesian (2M, 14) layouts take a uniform
speed (a codec with a scaled control period).

Instead of the old custom dashboard panel, the tempo rides the dashboard's
typed inference controls: every policy whose decoder has set_speed exposes
arc_speed_percent (and arc_hold_speed_percent for E1 tokens), 25-400 step 5,
whatever produced its profile (YAML, sidecar or config-derived). They apply at
the next replan like every other inference setting; model swaps republish
them. execution_plan logs when a faster replay ends before Repredict every.

Tests: tests/test_arc_decoder_speed.py (9). Robot suites unchanged: same 5
pre-existing test_robot_arc_campaigns failures before and after.
Station check (~/port_checks/arc_speed_check_20260930.log): both stationery
arcdur bundles expose the controls, time bundles do not; on the model token
the 2x uniform plan equals 1x read at twice the rate, and moving-only speeds
stay within 0.2 mm of the 1x path. Not yet run on the arms.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The baseline for the ARC speed-up comparison. A time policy's canonical
(H, 14) Euler chunk is a path on a uniform clock (row k at k * dt), so the ARC
decoder's warp applies unchanged: TimeChunkRetimer reads the chunk through it,
SLERPs rotation (intrinsic ZYX, as pose_matrix reads it) and holds the last
pose past the chunk end. speed == hold_speed is the naive uniform speed-up
(2x = every other row); hold_speed < speed keeps holds at the demonstrated
tempo. 100 % returns the chunk untouched, so default rollouts are unchanged.

load_graph_policy puts it in the decoder slot of every Euler time policy
(decoder null, no ARC round trip), so time and DP models get the same
dashboard controls as ARC ones. set_speed, the percent views and _warp move to
a ReplayTempo base shared by BimanualArcDecoder and the retimer. Control labels
are now "Replay speed (%)" / "Hold speed (%)" (names unchanged).

Tests: 2 new in tests/test_arc_decoder_speed.py (11 total); robot suites
unchanged (same 5 pre-existing test_robot_arc_campaigns failures).
Station check (~/port_checks/time_speed_check_20260930.log): slowpace_time
(HPT flow), elmoaidan time-300M and the towels298 graph DP load with the
retimer and both controls; on each model's own chunk 200/200 equals every other
row (0 on position/gripper, 4e-16 on rotation). Inference: HPT 59 ms, HPT-300M
43 ms, DP 567 ms (DDIM 100). Not yet run on the arms.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CSaa5mnstFBVJi8329KLDc
The Elmo+Aidan stationery runs trained three token shapes the station could
not decode. All three now bind a decoder through auto_inference_config:

- arcdurhyb (100, 18) -> e1_durhyb. Decode-only durhyb mode in the E1 codec:
  per arm, xyz + gripper on the translation clock and ypr on its own rotation
  clock; rows 0..M-2 are interval seconds, row M-1 the stream start delay.
  Streams are read by waypoint index against their own clock, so a delayed
  wrist turn and a held arm gripper decode on time. valid_steps counts the
  rotation and gripper streams, not just translation.
- lab wide (100, 28) / stacked (200, 14) -> lab_pw_wide / lab_pw_stacked,
  decoded with the PR #193 codec vendored verbatim (67863ca, the base of
  aidan/stattempo-lab-layouts). The station copy of arc_length_tokenizer has
  drifted and keeps serving the older bundles.
- inference_config: the Yam transform_list decides the lab token, not
  e1.variant (the lab runs inherit a stale e1.variant=arcdur); an implicit
  velocity_layout fails closed. Static flow_arcdurhyb profile added.

Not yet checked: durhyb against the training codec (7694986 is local to
Phoenix, unreachable 2026-10-01).

Tests: 7 new in tests/test_arc_token_shapes.py; decoder-speed, inference and
graph-policy suites pass; the same 5 pre-existing test_robot_arc_campaigns
failures as 007104c. Station (~/port_checks/token_shapes_check_20261001.py,
CPU): auto profiles change only for these 3 of 40 bundles; all 5 Elmo+Aidan
bundles strictly load, predict at 50 steps and decode finite plans.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The rollout dashboard could only save a display-only MP4. It can now also
record a full episode (d / Record episode) next to the unchanged video-only
recording, and the operator saves it as success, failure or unlabeled.

- Same layout as a demo: rollout_episode.RolloutEpisodeRecorder wraps the
  unchanged collect_demo.EpisodeWriter. Per executed control tick it writes the
  measured joints/EEF pose, the joint command sent to each arm, its forward
  kinematics (as the GELLO collector builds actions/eepose), and every camera
  frame as uncompressed RGB. Rollout-only data stays under rollout/: per-row
  timestamps and plan index, the policy's Cartesian target, and each plan
  with its inference settings and decoder stats. Attrs carry the checkpoint,
  policy config, outcome and end reason; a JSON manifest sits beside the file.
- Never blocks control: a background thread owns the file and the loop only
  enqueues copies (bounded queue). A slow disk or a write error ends the
  episode (writer_error), never the rollout. Recording refuses to start, and
  stops, below episode_recording.min_free_gb (default 20).
- Episode end paths: save with an outcome (complete), r/restart (complete,
  unlabeled), model change / camera reconnect / q / error (kept, incomplete),
  and discard (file deleted).
- Dashboard: Record episode button and d key, a save/discard panel while
  recording, frame count in the recording indicator, and a read-only
  Recorded episodes list (/api/episodes, metadata only).
- yam_rl2_hptflow_rollout.yaml enables it under
  /home/rohan/rollouts/yam_hptflow/episodes (about 5 GB per minute).

Tests: tests/test_rollout_episode.py (8) and 2 new dashboard tests; the
rollout dashboard, robot runtime, graph policy, ARC speed/token-shape,
inference config and GELLO suites pass. test_robot_arc_campaigns has the same
5 failures as 6da2402. Station check (~/port_checks/episode_check_20261004):
the episode layout is identical to a GELLO demo, Hdf5ReplayPolicy replays it
and hdf5_tempo reads it; three 480x640 cameras at 30 Hz sustain 83 MB/s, the
control-loop enqueue costs 1.0 ms mean (2.7 ms max) and the writer queue
never exceeded 1 row. Not yet run on the arms.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…d / tri decoders, ARC defaults, action-dt fix

Station-side changes made 2026-10-04..06 as uncommitted edits in this tree, committed together:

- Fastest-stream termination (arc_decoder, graph_policy, rollout dashboard): every ARC model starts with it on; a
  chunk ends when its fastest stream finishes, then the policy replans. A tri gripper stream ends a chunk only if
  its waypoints span at least 0.05, so an idle gripper never cuts a plan short.
- Decoders for every slowpace366 token: e1_profhyb (arcvelhyb, M x 18) and the tri layouts e1_durtri /
  e1_proftri (arcdurtri / arcveltri, M x 20: own translation, rotation and gripper clock per arm), matching training
  commits 76bc7a4a and 257e1fe. Speed-column tokens are converted to durations on the whole token (hold time
  dt * (H - 1)) before the fastest-stream cut; tri tokens read the gripper on its own clock and start delay.
- arc_length_tokenizer_m28.py: the M28 hybrid multistream codec vendored verbatim from EgoVerse-graph 99be4af
  (per-arm translation and rotation clocks; wide (M, 28), stacked (2M, 14) and four-clock (M, 18) tokens), kept
  apart from the station's own arc_length_tokenizer.py.
- ARC defaults (inference_config, rollout YAML): ARC_FLOW_INFERENCE_STEPS = 20 and ARC_EXECUTE_PERCENT = 50 for
  every ARC-decoded flow model (time models keep their recorded default, diffusion 100); both stay dashboard
  controls.
- "ARC action dt must be a positive number" on slowpace366 models: runs scored by OpenLoopSimEval record the
  control period as evaluator.control_dt; the E1 inference contract now reads it.

Tests: test_arc_decoder_speed, test_inference_config, test_rollout_dashboard, test_arc_hybrid_variants,
test_first_stream_arc_decoder. The GELLO recording / teleop dashboard edits in this tree are not part of this commit.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CSaa5mnstFBVJi8329KLDc
Integrate the audited token-shapes donor above the BC recipes while retaining
canonical per-arm clocks and model-owned pipeline declarations. Validation
artifacts remain on the exact companion pin.
@ElmoPA ElmoPA changed the title feat(robot): add HDF5 tempo bucket checker Add model-declared ARC rollout controls and episode recording Oct 7, 2026
@aidang3019
aidang3019 changed the base branch from codex/graph-consolidation-stack-20261007/14-aidan-bc to graphite-base/213 October 9, 2026 00:05
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants