Skip to content

Add Aidan BC recipes and four-clock hybrid ARC tokens - #212

Open
ElmoPA wants to merge 20 commits into
codex/graph-consolidation-20261001from
codex/graph-consolidation-stack-20261007/14-aidan-bc
Open

ElmoPA wants to merge 20 commits into
codex/graph-consolidation-20261001from
codex/graph-consolidation-stack-20261007/14-aidan-bc

Conversation

@ElmoPA

@ElmoPA ElmoPA commented Oct 7, 2026 •

Copy link
Copy Markdown
Collaborator

Add the missing training recipes and 18-channel E1 hybrid tokens from aidan/arc-bc-consolidated at 8ff8da40, as a child of #198. The 12 new recipes cover Elmo/Aidan 218/11 splits, HPT300 time/duration twins, the DP180 time baseline, and five slow-pace YAM/Aria co-training variants with frozen episode lists.

The composed neural graphs, optimizer, scheduler and exact episode selections match the donor. Duration/velocity hybrids retain independent translation and rotation clocks for each arm, preserve holds and start delays, and bind the explicit 2π rotation budget into model-owned preprocessing/inference contracts. Shared builders use this checkout and PACE launchers execute Python through scheduled srun steps.

Validation: a fresh checkout restored the BC-only companion snapshot c3a86edd and passed 143 affected CPU tests; all 466 shipped YAML configs resolve. Source parity and receipts are in #211. Full Linux CI retains the existing config/component/CPU/wheel gates. No training or physical robot execution was performed.

Stack: #198 → #212 → #213. Donor branches and the existing 13 PR heads/bases are unchanged. The complete old stack still has four preexisting archive-test modify/delete conflicts against current main; this layer adds no new integration conflict.

Aidan Gao and others added 20 commits September 29, 2026 23:40
The 229 organize_stationary episodes (Elmo 200 on 9-24, Aidan 29 on 9-19), validated on a fixed
11-episode split drawn with seed 42 and written as explicit lists (scripts/e1/stationery_elmoaidan_split.json)
so the lab-token twins built on Aniketh's stack train and validate on the same episodes.
build_elmoaidan.py writes this branch's arcdur configs and the twins' configs.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Same fixed 218/11 split and recipe; the model is the lab HPT-300M dims (840, 19x10, flow 6x320) as in the
towels394 hpt300 cell. Adds model/e1/hpt300_flow_wrists_time (the 300M arcdur model with the action dims at 14)
and the 180M time base experiment for the Elmo + Aidan pool that the 300M time twin inherits.
Compose + instantiate check: 253.2M (time, 100x14) / 253.3M (arcdur, 100x16).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…n clock (M, 18)

[14 canonical | translation dt L, R | rotation dt L, R]. xyz and gripper are sampled along the
translation arc (budget D), ypr along each arm's own geodesic rotation arc, and each stream stores its
interval seconds in rows 0..M-2 plus its start delay in row M-1 (padding in `dur`), so a wrist turn that
starts after the arm stops translating decodes on time. A translation hold keeps the arm put and times
the gripper across the window instead of dropping it. Other variants are byte-for-byte unchanged.

The rotation budget defaults to a full turn: on the fixed 100-frame YAM window the wrist exceeds the lab
hybrid's 24 deg in 63 % of rl2 stationery arm-windows, and a 24 deg stream froze it (GT round-trip
geodesic error 11.7 deg vs 1.1 deg for arcdur). At 2*pi: 0.07 deg over the full window (1,500 val windows),
xyz and gripper unchanged. tests/test_e1_arcdur_hybrid.py covers the in-place turn and a held arm's
gripper; the existing e1 tests pass (65).

Configs: 180M Elmo + Aidan run (model e1/hpt_flow_wrists_ft_arcdurhyb, data/experiment stattempo_elmoaidan_arcdurhyb).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The arcdur run with only variant changed, as towels394 arcvel twins its arcdur cell.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Twin of scratch_rl2_stattempo_slowpace_time with the graph Diffusion Policy (pretrained ResNet-18
encoders per camera, UNet [488, 976, 1952], DDIM 100): 180.5 M params vs the HPT run 180.4 M.
Baseline for the rollout speed-up comparison.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CSaa5mnstFBVJi8329KLDc
…me / arcdur)

For when Elmo's Aria sort-stationery data is ingested. Robot side is the
slow-pace pool exactly as scratch_rl2_stattempo_slowpace_{time,arcdur} (217
train, shared 24-ep robot val); the human side is every RL2 Aria episode
(SQL lab=rl2, embodiment=human_bimanual) matching --tasks / --since, frozen as
an explicit hash list at launch.

- model/e1/hpt_flow_wrists_ft{,_arcdur}_cotrain: the 180M h640t8 models plus
  a human_bimanual domain (own 14-D ee_pose stem, shared front-image stem),
  as the lab's yam+human cotrain models do.
- experiment/yam_arc_grid/cotrain_rl2_stattempo_slowpace_aria_{time,arcdur}:
  inherit the slowpace twins, swap model + generated data config.
- scripts/e1/build_slowpace_aria_cotrain.py: --list / build / --check (CPU:
  compose, sync missing zarrs, load samples, assert 30 fps). Aria zarrs say
  attrs.embodiment=aria_bimanual, so the leaf uses the embodiment override.
- scripts/e1/launch_slowpace_aria_cotrain.sh: build -> CPU check -> sbatch
  both (MODE=dry / MODE=smoke); exports the venv bin so s5cmd is found.

Smoke with stand-in RL2 Aria pick_place (6 eps): check OK (human chunks
(100,14)/(100,16) like robot, 30 fps), 300-step GPU smokes 13792686/87
COMPLETED, separate norm stats for embodiments 3 and 7, BimanualTempoEval ran,
3.0-3.5 steps/s. Stand-in data configs removed afterwards.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CSaa5mnstFBVJi8329KLDc
…ment aria, task "organize stationary")

Elmo's first RL2 Aria sort-stationery upload (2026-10-01, 20 eps) registered as embodiment=aria,
task="organize stationary". Both labels and both task spellings are now the defaults, and the
filter lambda matches the embodiment labels actually selected. The launcher runs its data check
in place when already inside a Slurm job.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CSaa5mnstFBVJi8329KLDc
…s (Elmo 10-01, 17 eps)

Generated by build_slowpace_aria_cotrain.py for Phoenix jobs 13800528 (time) and 13800529 (arcdur):
the 17 converted RL2 Aria organize_stationary episodes Elmo recorded 2026-10-01 (52,334 frames =
29 min after the converter dropped untracked frames; 71.6 min raw), relabelled in SQL from
embodiment=aria / task=organize stationary to human_bimanual / organize_stationary.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CSaa5mnstFBVJi8329KLDc
…", (M, 18))

Same waypoints and layout as arcdurhyb (xyz + gripper on the translation arc, ypr on each arm's own
rotation arc, 4 timing columns), but rows 0..M-2 hold each interval's mean speed (waypoint segment
length / duration; m/s translation, rad/s rotation) instead of its duration; row M-1 keeps the start
delay, or a translation hold's duration. Decode converts back to durations with the same segment
lengths and reuses the durhyb clock, so arcvelhyb and arcdurhyb carry identical timing content --
the arcvel-vs-arcdur parameterization question with an independent rotation clock.

Tests: 2 new in tests/test_e1_arcdur_hybrid.py; 92 E1/arc tests pass.
Real stationery val windows (600, shared 24-ep set): round-trip rotation 0.08 deg / xyz 11.13 mm for
both hybrids (plain arcdur 2.21 deg, arcvel 4.75 deg); profhyb vs durhyb decode max diff 4.9e-4.
Speeds: translation median 0.10 m/s (p99 0.87), rotation median 0.43 rad/s (p99 4.5).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CSaa5mnstFBVJi8329KLDc
…el, arcvelhyb)

- model/e1/hpt_flow_wrists_ft_arcdurhyb_cotrain: the 18-dim hybrid action stack (byte-identical to
  hpt_flow_wrists_ft_arcdurhyb) on the cotrain model; shared by arcdurhyb and arcvelhyb.
- experiments cotrain_rl2_stattempo_slowpace_aria_{arcvel,arcdurhyb,arcvelhyb} on the arcdur recipe.
- builder: --variants (default all five); non-time robot leaves = the slowpace arcdur config with only
  the transform variant changed; the manifest records the variants.
- launcher: loops over the manifest variants; run names carry the Aria episode count
  (stattempo_slowpace_aria<N>_cotrain_<variant>).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CSaa5mnstFBVJi8329KLDc
…variant runs

Generated by build_slowpace_aria_cotrain.py for Phoenix jobs 13823284-88 (time, arcdur, arcdurhyb,
arcvel, arcvelhyb): 37 converted RL2 Aria organize_stationary episodes from Elmo (10-01: 19, 10-02: 6,
10-03: 12; 130,006 frames = 1.20 h kept), with the slow-pace robot pool. The 20 rows added since the
aria17 runs were relabelled in SQL to human_bimanual / organize_stationary afterwards (filter accepts both).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CSaa5mnstFBVJi8329KLDc
…2 h)

The aria37 jobs sat in the queue for 2 h because their 72 h wall time overlapped PACE maintenance
reservation md-26-10 (10-05 06:00 -> 10-09 06:00, all nodes): Slurm will not start a job that would
run into it. 240k steps took ~23 h for the aria17 runs, so 30 h keeps a 30 % margin.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CSaa5mnstFBVJi8329KLDc
…aining)

The first profhyb stored a translation hold's duration (3.3 s) in row M-1. On the Aria data that row
is ~always 0 (q1 = q99 = 0), so per-element quantile normalization (q99 - q1 + 1e-6) mapped a hold to
6.6e8 and cotrain job 13823288 diverged (train loss median 1e4-1e5, spikes 1e9; robot val paired MSE
1-29 vs ~0.02). A hold is now all zeros and decodes over the decode horizon, where durhyb's uniform
hold durations put it; GT decode still equals durhyb's. On 3,000 real windows per embodiment the worst
normalized element is 12 (human) / 41 (robot), vs durhyb 67 / 41; no sample above 100.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CSaa5mnstFBVJi8329KLDc
…teacher-forced

Aidan 2026-10-03: replace the teacher-forced scores with the open-loop ones for future runs.

- open_loop_sim: E1 hybrid layouts e1_durhyb / e1_profhyb ((M, 18)); truncating to the executed
  25 % keeps each stream's start delay (row M-1) in the kept last row. Test: on a slow token the
  open-loop prefix decode equals the full (deployed) decode.
- cotrain_rl2_stattempo_slowpace_aria_{time,arcdur,arcvel,arcdurhyb,arcvelhyb}: now standalone (the
  teacher-forced parent's inline evaluator would merge into OpenLoopSimEval) with
  evaluator=eval_open_loop_sim (execute_fraction 0.25, DTW off, GT actions_time, baseline for time,
  e1_<layout> for arc), trainer.limit_val_batches 1.0 (whole ordered episodes). Recipe unchanged.
- scripts/e1/cotrain_ol_eval.sbatch: post-hoc open-loop scoring of cotrain checkpoints with the same
  experiments (mode=eval, the run's own norm stats).
- launcher: default WALL 32 h (open-loop validation adds ~1 h per run).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CSaa5mnstFBVJi8329KLDc
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CSaa5mnstFBVJi8329KLDc
… yam_source_frames pin) into local stationery work

Brings in 3c90bc9 and bef887c, pushed 2026-09-29 from another checkout,
under the 15 local commits (0bd4a3b..597fde7). A merge rather than a rebase
so the commits the stationery runs were trained from keep their hashes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@ElmoPA ElmoPA changed the title stationery: Elmo + Aidan arcdur run on a fixed 5 % split (218 / 11) Add Aidan BC recipes and four-clock hybrid ARC tokens Oct 7, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants