Skip to content

YAM arc BC: canonical ARC stack + stationery/towels grid, consolidated on current main - #160

Open
aidang3019 wants to merge 70 commits into
mainfrom
aidan/arc-bc-consolidated
Open

aidang3019 wants to merge 70 commits into
mainfrom
aidan/arc-bc-consolidated

Conversation

@aidang3019

@aidang3019 aidang3019 commented Sep 25, 2026 •

Copy link
Copy Markdown

What this is

The code, configs and launchers that every YAM arc BC grid run was trained from, on one branch off current main (161e3a0c):

  • Phoenix stationery / towels grid (9-17 → 9-20): fine-tunes off Elmo's yam97 checkpoint, ABC bases, RL2-only from-scratch cells, ABC+RL2 mixes (uniform and lab-balanced 50/50), RL2 towels-298 time / arcdur / arcvel.
  • RL2 towels-394 campaign on the Lambda loaner (9-25): time / arcdur / arcvel, the w400 window test, the ABC+RL2 40/60 mixes, and the HPT-300M / DP-300M arcdur twins (9-27).
  • Stationery tempo buckets (9-25 / 9-27): {slow, medium, fast, mixed, slowpace, all} x {time, arcdur}, plus per-bucket open-loop eval.

All experiments live under experiment/yam_arc_grid/ (launch with +experiment=yam_arc_grid/<name>), and their data under data/abc_arc/.

What to review here

Only the 37 commits authored by Aidan. The rest of the branch is other people's open PRs, carried because the runs depend on them:

Part Commits Where to review it
Aniketh's canonical ARC stack 14, cherry-picked with -x 12 are patch-identical to #161–#172. ebb422e0 is an older version of his c731afea. e227a094 (Lambda launcher settings) comes from his cluster/* branches and is in none of his PRs. His stack has since grown (#173–#193).
Ryan's weighted dataset mixtures 93980990 + merge 766c1f49 #141. The towels mix4060 runs need it.
Grid runs 37 (Aidan) This PR.

Once those land, this branch rebases down to the Aidan commits. Expect one conflict in bimanual_arc.get_keymap: keep the yam_source_frames argument (see below).

Where the +22.8k lines are

Share Lines added What
Aniketh's stack + Ryan's #141 ~8.4k Reviewed in their own PRs.
Run configs 8.1k in 149 files One experiment + data config per run, so each run can be relaunched by name. Twins (variant, bucket, task, Lambda copy, eval set) inherit from a sibling in the same family and hold only what differs.
Episode lists and split manifests 4.1k in 8 files The exact pools the runs used (rl2_towels394_episodes.txt, stationery_tempo_manifest.json, ...). The configs pin episode counts and name hashes against these, so a drifted pool fails at construction.
Launchers and tools 1.6k in 28 files Phoenix / Lambda sbatch launchers, data sync, harvest and eval helpers under scripts/e1, scripts/pi05 and scripts/clusters/lambda.
Library code ~350 lines + ~95 of tests The part that needs a careful read. Listed next.

Library changes

  1. init_weights_from (trainHydra.py, eval/checkpoint_loading.py): weights-only initialisation for fine-tunes. Tensors whose shape changed (e.g. a 14-D cartesian head going to a 16-D arc token) keep their fresh init and are logged. A key-set mismatch is still an error. cfg.ckpt_path and Slurm requeues always win, so a resume is never re-initialised.
  2. group_balance_sampler (rldb/zarr/group_balance_sampler.py, a hook in pl_data_utils.py): per-group balanced sampling set under train_dataloader_params.<embodiment>. It is used for the lab-balanced RL2 50 % / ABC 50 % runs, and logs the realised share every 320k draws. By frames, RL2 is only 2.8 % of stationery, so a uniform shuffle would train almost entirely on ABC.
  3. Two train sources of one embodiment (trainHydra.py, for Add weighted dataset mixtures and homogeneous graph training #141 mixes): norm stats are keyed by embodiment, so a second yam_bimanual source used to overwrite the first source's stats (or crash with KeyError). That now raises unless norm_stats.precomputed_norm_path points at pooled stats from a norm_stats_only pass over the union.
  4. yam_source_frames (rldb/embodiment/bimanual_arc.py): an explicit YAM window for keymap_mode: cartesian. See the next section.
  5. open_loop_sim: decodes E1 wide (M, 16) tokens (token_layout: e1_dur | e1_logdur | e1_profile, default lab, so existing configs are unaffected). Adds an opt-in per-episode dump of the executed prediction against ground truth.

⚠️ The YAM source window is now explicit

With keymap_mode: cartesian, bimanual_arc.get_keymap gives YAM a fixed 100-frame (3.3 s) window whatever horizon the yaml sets (canonical ba4509b1). Every grid run trained on that window. Until now it was a code default, so the grid yamls said horizon: 200 while the norm stats showed the last arc interval peaking at 3.27 s. Elmo's Lambda twin of the towels-298 arcdur run ran on code that honours horizon (400 frames, 13.27 s) and rolls out worse.

bef887c4 writes yam_source_frames: 100 into every YAM cartesian key_map of the grid's data configs. The window now shows in the saved and W&B configs, and these configs keep reproducing their runs if the default ever changes. The w400 cell sets 400 on purpose; it is the root-cause test for Elmo's twin, and it came out worse on 19/19 validation episodes. Aniketh's current stack also keeps cartesian YAM at 100. His 200-frame cap applies to arc_tokenizer_cartesian, which no grid config uses.

Not portable as-is

  • *_lambda configs read data from /workspace/users/agao81/datasets on the Lambda loaner, and the Lambda launchers assume that layout.
  • Phoenix launchers write runs under /storage/project/r-dxu345-0/agao81/.
  • Phoenix configs read the shared mirror /storage/project/r-dxu345-0/shared/egoverseS3ZarrDatasets, so they work for anyone on Phoenix.

Verification

  • Consolidation (9-24): all 26 grid experiments composed identically on the old tree and this branch. For towels-298 time / arcdur / arcvel, the training tensors are bit-identical on 40 fixed indices: actions_time and the camera SHA-256s exactly, actions_cartesian within 4e-15, the difference coming from main's NumPy pose transform. A 300-step GPU smoke (Phoenix job 13553687, 1×H200) finished COMPLETED 0:0 with Valid/E1/* logged.
  • Towels-394 + tempo merge (9-29): 111 targeted tests passed, and 64/64 yam_arc_grid experiments compose.
  • This update (9-29): after the inheritance rewrite (3c90bc9e), all 64 experiments and 103 data/abc_arc configs compose byte-identically to before. After the pin (bef887c4), the only difference in any composed config is the added yam_source_frames key: 238 key_map blocks read 100, and the w400 cell's 4 read 400.
  • Tests: pytest tests/ gave 1,382 passed and 6 failed at consolidation. All 6 fail without this branch too: 3 on plain main (test_planar_configs.py ×2, test_shared_robot_components.py::test_vendor_metadata_is_canonicalized_without_editing_episode) and 3 on the canonical stack's own tip (test_ice_wandb_resume.py).
  • The PR Auto Review check fails before it reviews anything, because the workflow has no Anthropic API key. It is unrelated to this diff.
Left out on purpose
  • aniketh/arc-configs-graph (4481e3d9, b430aa0c, 88b3786e): superseded by the canonical stack (ba4509b1 re-lands 88b3786e).
  • standalone/hpt180-stationery-configs: re-homed under experiment/abc_arc/robot_bc/ by the canonical stack.
  • Elmo's Lambda w400 runs (source_commit 56b4d1c9): that commit is not on GitHub. scratch_rl2_towels394_arcdur_w400_lambda reproduces its window here.
  • aidan/towels394-eval (Aniketh's feat(eval): standalone open-loop validation with distance DTW #174 + the E1 port) used for the open-loop DTW comparisons: it belongs with feat(eval): standalone open-loop validation with distance DTW #174.
  • make_stationery_{tempo,midtempo}_configs.py: they generated the tempo configs and were removed once the runs finished (re-running them would write the full copies back). They remain in git history.
Commit map for the Phoenix grid (old → new), for the run notes
old (aidan/stationery-ft-20260917) new commit
ed120944 bdf04a62 stationery mid-tempo holdout + E1 wide open-loop eval
b2f35635 d8da16c6 sync tooling for the 2026-09-17 rl2 stationery re-upload
ec32507b 447241bd stationery fine-tune off the yam97 checkpoint
902de75b 52a7f7c8 stationery ft: rl2 station episodes only
f189f60f 57382247 persistent workers fix epoch-boundary deadlock
d810ec98 673d0243 stationery pretrains
a07d692d 4fd2cd8a paths.output_dir null → hydra runtime dir
9bfee8d1 4a802898 launcher passes paths.output_dir explicitly
19f18b91 4082c766 grid configs under the ABC/RL2 separation rule
1dd56aec 5ee1bdb4 launcher: H-class pool + GPU-class guard
3d97d5ef 0e019749 mixed ABC+RL2 Arc+Vel runs
ca5d2f1f b295cc9e mixed TIME counterparts
5af693db 25d57830 lab-balanced 50/50 sampler
88e2e5a3 1d47e09b sampler realised-share logging + SQL test
6ddf673e 3fc7e657 arcdur twins of the balanced runs
97827d6a 7c9e2655 RL2 towels-298 time / arcdur / arcvel

The old tip is tagged on Phoenix as pre-consolidate/stationery-ft-20260917. The towels-394 line came in through the --no-ff merge b6e3ee40, with the pre-merge tip tagged pre-towels394-merge-20260929.

🤖 Generated with Claude Code

rpunamiya and others added 30 commits September 19, 2026 22:30
…val for E1 (M,16) tokens

Aidan, 2026-09-15: fold the rl2 YAM stationery uploads into the ABC stationery set,
hold the middle Step 0 tempo tertile out of training, and train two rollout-ready
models -- the lab time-indexed recipe and the arcdur arc row -- scored with
Aniketh's open-loop segment evaluator (open_loop_sim, 4481e3d).

EVALUATOR
open_loop_sim only knew the lab (M+1, 14) / (2M, 14) ARC layouts. Add token_layout
(lab | e1_dur | e1_logdur | e1_profile) and truncate_e1_wide_token so an E1 wide
(M, 16) prediction is truncated by waypoint rows (each row carries its own timing)
and detokenized with TokenizeBimanualArcLengthE1 built exactly as
robot/arc_decoder.BimanualArcDecoder builds that layout -- the scored decode is the
deployed decode. The scoring loop is unchanged; lab-layout behaviour is unchanged
(token_layout defaults to lab). Tests: truncation keeps waypoint/timing rows, and a
truncated e1_dur decode equals the robot decoder's first 25 steps.

CONFIGS (generated from the split manifest, not hand-edited)
data/abc_arc/stationery_midtempo_{time,arcdur}.yaml and eval-only configs for
val / test_mid / test_in, plus a *_smoke_* set on a 4-episode subset. Episode-hash
frozensets; the task clause admits 'sort the stationery into containers' (ABC) and
'organize_stationary' (rl2 YAM station) -- no SQL relabel. time = Ryan's
shorts_extreme_bc data block (Yam lab keymap, raw 100-frame eef_frame chunk);
arcdur = shorts_extreme_arcdur (bimanual_arc, D 0.40 m, M 100, (100, 16)).
experiment/abc_arc/stationery_midtempo_{time,arcdur}.yaml: the ported shorts-extreme
recipe (e1/hpt_flow_wrists, 30k steps, batch 32, 1 GPU, checkpoint every 5k) with
evaluator eval_open_loop_sim (10 val episodes every 5k in training, whole episodes).

SCRIPTS (scripts/e1/)
build_stationery_midtempo.py   split v2: train = slow + fast tertiles minus val and
                               test_in; val / test_mid identical to v1
make_stationery_midtempo_configs.py, make_smoke_manifest.py
stationery_midtempo_{train,eval}.sbatch, stationery_midtempo_smoke.sh,
stationery_midtempo_tests.sbatch   Phoenix gpu-h200 / cpu-small launchers
harvest_stationery_midtempo.py     select on val open-loop paired MSE, report
                                   selected + final on test_in / test_mid
package_stationery_midtempo.py     rollout artefact: checkpoint (sha256), norm stats
                                   with normalizer_state, Hydra config, filled
                                   yam_rollout.yaml (base_T_model from the rl2
                                   calibration, e1_dur decoder for arcdur), manifest
replay_check_stationery_midtempo.py  load the package via load_graph_policy and
                                   replay a recorded rl2 episode offline

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit ed120944eab5a485e934f104b1047143a19f2e62)
(cherry picked from commit b2f356359784fe68640d2edde31ca88f91aae7cd)
(cherry picked from commit ec32507b3eba2bb02c2680e4d32a8fc39f28d0d7)
(cherry picked from commit 902de75b813f792040239f1ed71b04d82272de45)
…C arc pretrain configs

Loader workers are re-forked at every epoch boundary by default and this stack deadlocks there: job 13307279 wedged at step 99 of a 100-step epoch, workers idle in do_poll, main thread in futex_wait, after limping at ~19 s/step. persistent_workers=true crosses three boundaries cleanly at 4.2 it/s (job 13308744). Also adds the ABC-only arc pretrain (Elmo ckpt is variant=time, so it has no arc pretraining) and a per-job hydra run dir.

(cherry picked from commit f189f60fc73499078c1bf3fe2ffe17203f58e0d8)
…ent base (ABC+RL2)

Fills the {time,arc} x {scratch,base} matrix. timepre_abc is the ABC-only time twin of arcpre_abc (Elmo ckpt is not a substitute: it trained on ABC + old RL2, 2264 eps). arcpre_abcrl2 mirrors the DATA POOL of Elmo checkpoint run (ABC + RL2 organize_stationary_updated, 2287 eps) with the arc-duration target, so a later RL2 fine-tune from it parallels the time arm.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit d810ec9861a089447705becafdee7c39d1c6327d)
…re all landing in a shared None/checkpoints)

Copied from Elmo resolved config where output_dir had already resolved to null; dirpath then rendered as literal None/checkpoints relative to the repo root, shared by every run. Lightning versioned the collisions (-v1/-v2) so nothing was lost, but ownership was only recoverable by mtime. Each job now checkpoints under its own hydra run dir.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit a07d692d3e4d097a02c83209ce5413727bf1e194)
… land in None/checkpoints

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 9bfee8d17c32aaf784ae6180c545a73da8af1b9e)
… cells, ABC-only towel bases; towel sync

Aidan 09-18: never mix ABC and RL2 unless told to. From-scratch = RL2 YAM data of the task only; pretraining bases = ABC only. Adds scratch_rl2_{stationery,towels}_{time,arcdur}, base_abc_towels_{time,arcdur}, towel data configs, and sync_towels (rl2 task is fold_towels after the 09-18 relabel). arcpre_abcrl2 (ABC+RL2) violated the rule and was cancelled before it ran.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 19f18b9151fb25eab9a10d37a63ce3ec543a6aba)
…one submit helper

Untyped gres lets PACE expand the partition list (a job once landed on a V100, no bf16), so grid_submit.sh excludes every non-H100/H200 GPU node and the job exits 42 if the GPU is not H100/H200. norm_stats.save_cache_dir was null (inherited from a resolved config), so no run was saving the stats a later fine-tune must reuse; now saved under the run dir.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 1dd56aecc1a8bc48e222294c5d84bbe3f24a5c8a)
… tasks; pinned-pool norm-stats recompute

Aidan 09-18 asked for two from-scratch runs on BOTH abc and rl2 data with the velocity arc tokenizer (variant arcvel -> profile, (100,16) = 14 canonical + per-arm path speed). This is the explicit exception to the never-mix rule and applies to these two configs only. Also commits the pinned 2089-episode data config + sbatch used to recompute arcpre_abc norm stats exactly.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 3d97d5ef827430ca797f0109711291c682197889)
…licitly authorised), both tasks; exclude hung H200 node

Same pool filter and valid_ratio as the *_mix_arcvel configs; only the action target differs (14-D time chunk).
atl1-1-03-018-14-0 hung job 13325462 in its first backward for 8.5 h (GPU 100 %, zero steps) -> added to the exclude list.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit ca5d2f1ff052bf64aa381ebdb198cabdb81b4e3c)
…h runs; arc evaluator on

group_balance_sampler: WeightedRandomSampler over MultiDataset.index_map, each named group gets a fixed share of
draws, uniform over frames within a group; membership from explicit episode lists (198 RL2 episodes per task).
Hooked into MultiDataModuleWrapper.train_dataloader next to anchor_sampler. Four scratch_mix5050_* experiments
supersede scratch_mix_* (cancelled); the arcvel pair now runs BimanualTempoEval so Valid/* is logged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 5af693db2487361b010c77135fe223372e38ce51)
…aws; end-to-end test against SQL ground truth

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 88e2e5a345383eb41e68fb6eb5ee96d5bff3bd57)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 6ddf673eadd9e3103b7d1cd7c4e559922295d9f9)
Aidan Gao and others added 18 commits September 25, 2026 13:23
- mix4060_towels394_arcvel_lambda and its pooled norm-only pass, built exactly
  like the time/arcdur members (Aidan 2026-09-25: another mixed run with arc vel
  to use more of the 8 GPUs).
- pack_train: CPUS_PER_TASK and PORT_BASE overrides; run with bash and
  SLURM_JOB_ID=<job> to put runs on a running pack job's spare GPUs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The venv interpreter links to the image's /usr/bin/python3.12, so test -x fails
outside the container (the login node, where in-job steps are launched).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A relayed last.ckpt already in the run dir makes Lightning's rolling checkpoint save as last-v1.ckpt,
so resuming from last.ckpt by name would rewind a requeued run to the relayed step.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…launcher)

stattempo_eval_{time,arcdur}: the training experiment with its inline BimanualTempoEval block replaced
by eval_open_loop_sim at the stationery mid-tempo settings (execute_fraction 0.25, DTW off, no videos),
because Hydra merges an experiment's inline evaluator block over any evaluator group selected on the
command line and OpenLoopSimEval forwards unknown kwargs to a parent that rejects them.
stattempo_eval.sbatch scores checkpoints on val_slow / val_medium / val_fast with the run's own norm stats.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…truth; stattempo_eval EXTRA overrides

episode_dump_dir (default None = nothing written, metrics unchanged) saves one .npz per scored
episode: segment starts, per-step frame index, and the executed (T, 14) prediction and ground truth
exactly as scored. Used to plot per-model error/gripper timelines for the stationery tempo runs.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Aidan 2026-09-27: two more arcdur models on the RL2 fold data, HPT at the larger
size and a DP at the same size. model/e1/hpt300_flow_wrists_arcdur = our arcdur
model at the lab HPT-300M dims (840 / 19 blocks x 10 heads / flow 6x320);
model/e1/dp300_wrists_arcdur = the lab graph DP-300M (UNet [512,1024,2048],
64-kp encoders) with front + both wrist cameras and the E1 arcdur (100, 16)
action. Experiments scratch_rl2_towels394_arcdur_{hpt300,dp300}_lambda are the
arcdur_lambda twin with only the model changed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… pools

train_slowpace = the 229 organize_stationary episodes (Elmo 200 + Aidan 29, collected at an instructed
slow pace) minus the shared held-out set; train_all = every usable rl2 stationery episode (427) minus the
same 24 held-out episodes, so both stay comparable with the tempo-bucket runs. Recipe unchanged.
Generator gains --pools and takes the episode count from the manifest; check script takes a pool list
(4/4 train + valid sets match the manifest).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…coders

dp300pt_wrists_arcdur = dp300_wrists_arcdur with pretrained: true and norm_layer: batch
(the group swap would re-initialise the pretrained norm layers; no EMA, so BN is safe).
Experiment twin scratch_rl2_towels394_arcdur_dp300pt_lambda, W&B id ..._20260928_s42.
Aidan 2026-09-28: all models should be inited from the pretrained resnet weights.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Brings aidan/towels394-lambda (the RL2 fold_towels 394-episode campaign on the Lambda loaner) onto the
consolidated branch next to the stationery tempo-bucket line, so one branch carries every run's code:
- Ryan Punamiya's weighted dataset mixtures (PR #141, merged as-is) for the 40/60 ABC/RL2 mix runs
- towels394 time / arcdur / arcvel / arcdur-w400 configs (Phoenix + Lambda twins), 40/60 mix configs
- the yam_source_frames window knob, pooled two-source norm stats (trainHydra _source_embodiment)
- Lambda pack_train.sbatch / setup_runtime.sbatch
- HPT-300M, DP-300M and DP-300M-pretrained arcdur model + experiment configs

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The 12 towels394 Lambda experiment twins and 8 data twins become defaults-inheritance children holding only
their delta (name, W&B id, tags, variant, model/data choice, dims, window). Parents: scratch arcdur, mix4060 arcdur,
norm_mix4060 arcdur; data: towels_rl2_394 arcdur, towels_mix4060 arcdur, normpool arcdur. All 64 yam_arc_grid
experiments compose byte-identical to before (sorted-key dump of every composed config, 0 differences).
dp300pt stays a full file: its encoders sit in pipeline.stages[0], a list, which OmegaConf replaces on merge.

trainHydra: Counter for the shared-embodiment check; bimanual_arc: one expression for the YAM window;
pack_train: one-line resume. 111 tests pass.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The 229 organize_stationary episodes (Elmo 200 on 9-24, Aidan 29 on 9-19), validated on a fixed
11-episode split drawn with seed 42 and written as explicit lists (scripts/e1/stationery_elmoaidan_split.json)
so the lab-token twins built on Aniketh's stack train and validate on the same episodes.
build_elmoaidan.py writes this branch's arcdur configs and the twins' configs.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
76 near-duplicate configs under experiment/yam_arc_grid and data/abc_arc
(bucket, variant, task and Lambda twins, eval and smoke sets) are now
`defaults: [<sibling>, _self_]` plus only the keys that differ, the pattern
a56cf34 used for towels394. Parents are always in the same run family (the
names differ only in variant / bucket / task / eval-set tokens), and a smoke or
Lambda config is never the parent of one that is not. Headers are unchanged.

The two config generators (make_stationery_tempo_configs.py,
make_stationery_midtempo_configs.py) are removed: those runs are finished, the
configs are the record, and re-running them would write the full copies back.
The data headers that said "do not edit by hand" now point at git history.

Verified: every yam_arc_grid experiment (64) and every abc_arc data config
(103) composes byte-identically before and after (hydra compose, unresolved).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…key_map

Under keymap_mode: cartesian the YAM window has been a code default
(Yam.ACTION_HORIZON = 100 frames, 3.3 s) that ignores the yaml `horizon: 200`.
Every grid run was trained on that default. Writing it into the configs makes
the window visible in the saved and W&B configs, and keeps these configs
reproducing their runs if the default or the horizon handling changes. The
w400 cell keeps its explicit 400.

Verified: 116 composed configs change, and in each one the only difference is
the added key (238 YAM cartesian key_map blocks now read 100, the w400 cell's
4 read 400). Training tensors are unchanged because 100 is what the code
already used.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Same fixed 218/11 split and recipe; the model is the lab HPT-300M dims (840, 19x10, flow 6x320) as in the
towels394 hpt300 cell. Adds model/e1/hpt300_flow_wrists_time (the 300M arcdur model with the action dims at 14)
and the 180M time base experiment for the Elmo + Aidan pool that the 300M time twin inherits.
Compose + instantiate check: 253.2M (time, 100x14) / 253.3M (arcdur, 100x16).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…n clock (M, 18)

[14 canonical | translation dt L, R | rotation dt L, R]. xyz and gripper are sampled along the
translation arc (budget D), ypr along each arm's own geodesic rotation arc, and each stream stores its
interval seconds in rows 0..M-2 plus its start delay in row M-1 (padding in `dur`), so a wrist turn that
starts after the arm stops translating decodes on time. A translation hold keeps the arm put and times
the gripper across the window instead of dropping it. Other variants are byte-for-byte unchanged.

The rotation budget defaults to a full turn: on the fixed 100-frame YAM window the wrist exceeds the lab
hybrid's 24 deg in 63 % of rl2 stationery arm-windows, and a 24 deg stream froze it (GT round-trip
geodesic error 11.7 deg vs 1.1 deg for arcdur). At 2*pi: 0.07 deg over the full window (1,500 val windows),
xyz and gripper unchanged. tests/test_e1_arcdur_hybrid.py covers the in-place turn and a held arm's
gripper; the existing e1 tests pass (65).

Configs: 180M Elmo + Aidan run (model e1/hpt_flow_wrists_ft_arcdurhyb, data/experiment stattempo_elmoaidan_arcdurhyb).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The arcdur run with only variant changed, as towels394 arcvel twins its arcdur cell.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Twin of scratch_rl2_stattempo_slowpace_time with the graph Diffusion Policy (pretrained ResNet-18
encoders per camera, UNet [488, 976, 1952], DDIM 100): 180.5 M params vs the HPT run 180.4 M.
Baseline for the rollout speed-up comparison.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CSaa5mnstFBVJi8329KLDc
…me / arcdur)

For when Elmo's Aria sort-stationery data is ingested. Robot side is the
slow-pace pool exactly as scratch_rl2_stattempo_slowpace_{time,arcdur} (217
train, shared 24-ep robot val); the human side is every RL2 Aria episode
(SQL lab=rl2, embodiment=human_bimanual) matching --tasks / --since, frozen as
an explicit hash list at launch.

- model/e1/hpt_flow_wrists_ft{,_arcdur}_cotrain: the 180M h640t8 models plus
  a human_bimanual domain (own 14-D ee_pose stem, shared front-image stem),
  as the lab's yam+human cotrain models do.
- experiment/yam_arc_grid/cotrain_rl2_stattempo_slowpace_aria_{time,arcdur}:
  inherit the slowpace twins, swap model + generated data config.
- scripts/e1/build_slowpace_aria_cotrain.py: --list / build / --check (CPU:
  compose, sync missing zarrs, load samples, assert 30 fps). Aria zarrs say
  attrs.embodiment=aria_bimanual, so the leaf uses the embodiment override.
- scripts/e1/launch_slowpace_aria_cotrain.sh: build -> CPU check -> sbatch
  both (MODE=dry / MODE=smoke); exports the venv bin so s5cmd is found.

Smoke with stand-in RL2 Aria pick_place (6 eps): check OK (human chunks
(100,14)/(100,16) like robot, 30 fps), 300-step GPU smokes 13792686/87
COMPLETED, separate norm stats for embodiments 3 and 7, BimanualTempoEval ran,
3.0-3.5 steps/s. Stand-in data configs removed afterwards.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CSaa5mnstFBVJi8329KLDc
aidang3019 and others added 2 commits October 1, 2026 22:22
…ment aria, task "organize stationary")

Elmo's first RL2 Aria sort-stationery upload (2026-10-01, 20 eps) registered as embodiment=aria,
task="organize stationary". Both labels and both task spellings are now the defaults, and the
filter lambda matches the embodiment labels actually selected. The launcher runs its data check
in place when already inside a Slurm job.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CSaa5mnstFBVJi8329KLDc
…s (Elmo 10-01, 17 eps)

Generated by build_slowpace_aria_cotrain.py for Phoenix jobs 13800528 (time) and 13800529 (arcdur):
the 17 converted RL2 Aria organize_stationary episodes Elmo recorded 2026-10-01 (52,334 frames =
29 min after the converter dropped untracked frames; 71.6 min raw), relabelled in SQL from
embodiment=aria / task=organize stationary to human_bimanual / organize_stationary.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CSaa5mnstFBVJi8329KLDc
aidang3019 and others added 8 commits October 3, 2026 02:32
…", (M, 18))

Same waypoints and layout as arcdurhyb (xyz + gripper on the translation arc, ypr on each arm's own
rotation arc, 4 timing columns), but rows 0..M-2 hold each interval's mean speed (waypoint segment
length / duration; m/s translation, rad/s rotation) instead of its duration; row M-1 keeps the start
delay, or a translation hold's duration. Decode converts back to durations with the same segment
lengths and reuses the durhyb clock, so arcvelhyb and arcdurhyb carry identical timing content --
the arcvel-vs-arcdur parameterization question with an independent rotation clock.

Tests: 2 new in tests/test_e1_arcdur_hybrid.py; 92 E1/arc tests pass.
Real stationery val windows (600, shared 24-ep set): round-trip rotation 0.08 deg / xyz 11.13 mm for
both hybrids (plain arcdur 2.21 deg, arcvel 4.75 deg); profhyb vs durhyb decode max diff 4.9e-4.
Speeds: translation median 0.10 m/s (p99 0.87), rotation median 0.43 rad/s (p99 4.5).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CSaa5mnstFBVJi8329KLDc
…el, arcvelhyb)

- model/e1/hpt_flow_wrists_ft_arcdurhyb_cotrain: the 18-dim hybrid action stack (byte-identical to
  hpt_flow_wrists_ft_arcdurhyb) on the cotrain model; shared by arcdurhyb and arcvelhyb.
- experiments cotrain_rl2_stattempo_slowpace_aria_{arcvel,arcdurhyb,arcvelhyb} on the arcdur recipe.
- builder: --variants (default all five); non-time robot leaves = the slowpace arcdur config with only
  the transform variant changed; the manifest records the variants.
- launcher: loops over the manifest variants; run names carry the Aria episode count
  (stattempo_slowpace_aria<N>_cotrain_<variant>).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CSaa5mnstFBVJi8329KLDc
…variant runs

Generated by build_slowpace_aria_cotrain.py for Phoenix jobs 13823284-88 (time, arcdur, arcdurhyb,
arcvel, arcvelhyb): 37 converted RL2 Aria organize_stationary episodes from Elmo (10-01: 19, 10-02: 6,
10-03: 12; 130,006 frames = 1.20 h kept), with the slow-pace robot pool. The 20 rows added since the
aria17 runs were relabelled in SQL to human_bimanual / organize_stationary afterwards (filter accepts both).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CSaa5mnstFBVJi8329KLDc
…2 h)

The aria37 jobs sat in the queue for 2 h because their 72 h wall time overlapped PACE maintenance
reservation md-26-10 (10-05 06:00 -> 10-09 06:00, all nodes): Slurm will not start a job that would
run into it. 240k steps took ~23 h for the aria17 runs, so 30 h keeps a 30 % margin.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CSaa5mnstFBVJi8329KLDc
…aining)

The first profhyb stored a translation hold's duration (3.3 s) in row M-1. On the Aria data that row
is ~always 0 (q1 = q99 = 0), so per-element quantile normalization (q99 - q1 + 1e-6) mapped a hold to
6.6e8 and cotrain job 13823288 diverged (train loss median 1e4-1e5, spikes 1e9; robot val paired MSE
1-29 vs ~0.02). A hold is now all zeros and decodes over the decode horizon, where durhyb's uniform
hold durations put it; GT decode still equals durhyb's. On 3,000 real windows per embodiment the worst
normalized element is 12 (human) / 41 (robot), vs durhyb 67 / 41; no sample above 100.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CSaa5mnstFBVJi8329KLDc
…teacher-forced

Aidan 2026-10-03: replace the teacher-forced scores with the open-loop ones for future runs.

- open_loop_sim: E1 hybrid layouts e1_durhyb / e1_profhyb ((M, 18)); truncating to the executed
  25 % keeps each stream's start delay (row M-1) in the kept last row. Test: on a slow token the
  open-loop prefix decode equals the full (deployed) decode.
- cotrain_rl2_stattempo_slowpace_aria_{time,arcdur,arcvel,arcdurhyb,arcvelhyb}: now standalone (the
  teacher-forced parent's inline evaluator would merge into OpenLoopSimEval) with
  evaluator=eval_open_loop_sim (execute_fraction 0.25, DTW off, GT actions_time, baseline for time,
  e1_<layout> for arc), trainer.limit_val_batches 1.0 (whole ordered episodes). Recipe unchanged.
- scripts/e1/cotrain_ol_eval.sbatch: post-hoc open-loop scoring of cotrain checkpoints with the same
  experiments (mode=eval, the run's own norm stats).
- launcher: default WALL 32 h (open-loop validation adds ~1 h per run).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CSaa5mnstFBVJi8329KLDc
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CSaa5mnstFBVJi8329KLDc
… yam_source_frames pin) into local stationery work

Brings in 3c90bc9 and bef887c, pushed 2026-09-29 from another checkout,
under the 15 local commits (0bd4a3b..597fde7). A merge rather than a rebase
so the commits the stationery runs were trained from keep their hashes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants