Skip to content

Sweep ARC stream grouping and component timing on LIBERO - #202

Draft
rpuns wants to merge 8 commits into
codex/libero-arc-mot-20260930from
codex/libero-arc-streams-20261001
Draft

rpuns wants to merge 8 commits into
codex/libero-arc-mot-20260930from
codex/libero-arc-streams-20261001

Conversation

@rpuns

@rpuns rpuns commented Oct 2, 2026 •

Copy link
Copy Markdown

Adds a controlled LIBERO sweep separating ARC support grouping, component velocity labels, and shared versus independent component clocks. Eleven variants run with STK and duration on Spatial, Object, Goal and LIBERO-10: 88 policy cells at fixed R=384°, D=1.6 m and M=36. Gripper streams preserve command events; angular-command controls distinguish rotation-coordinate changes from scalar grouping.

Every arm retains the U-Net DP, decoded loader, seed 42, 5001 epochs/global batch 1024, AdamW/EMA, predict-32/execute-16 replanning and 2500 paired evaluation episodes. Replay receipts must match source, codec and controls before training. GPU preflight and checkpoint verification enforce the original representation and full optimizer/EMA budget. Poor replay candidates are not silently pruned.

The October 5 20:44 PT artifact audit has 46/88 trained policies, 44 final evaluations and two partial evaluations. Twelve evaluations finished since the previous audit. All 22 Object policies are trained; 20 have final scores and two are still evaluating. Every scored policy reached its full training budget; partial refers only to rollout coverage.

Representation Mode Spatial Object Goal LIBERO-10
Reference STK 68.64% 47.16% 85.52% 54.64%
Reference DUR 69.48% 41.96% 89.36% 53.80%
Separate gripper STK 65.96% 54.20% 87.04% 55.60%
Separate gripper DUR 67.04% 63.72% 88.92% Pending
X/Y/Z streams STK 51.68% 40.60% 68.04% Pending
X/Y/Z streams DUR 66.76% 45.84% 87.84% Pending
Component timing, shared clock STK 63.60% 49.56% 85.40% Pending
Component timing, shared clock DUR 67.32% 52.64% 88.32% Pending
Independent component clocks STK 48.24% 43.88% 75.68% Pending
Independent component clocks DUR 68.24% 52.36% 88.12% 54.68%
X / YZ STK Pending 44.92% Pending Pending
X / YZ DUR Pending 51.16% Pending Pending
Y / XZ STK Pending 35.00% Pending Pending
Y / XZ DUR Pending 51.84% Pending Pending
Z / XY STK Pending 38.08% Pending Pending
Z / XY DUR Pending 50.16% Pending Pending
Grouped angular commands STK Pending 46.76% Pending Pending
Grouped angular commands DUR Pending 52.08% Pending Pending
Separate angular-command axes STK Pending 52.32% Pending Pending
Separate angular-command axes DUR Pending 67.92% Pending Pending
Separate translation and angular-command axes STK Pending 18.10%* Pending Pending
Separate translation and angular-command axes DUR Pending 42.49%* Pending Pending

Unmarked scores are final: 2500 episodes over ten tasks. Only Object all-scalar remains partial: STK* covers 1845 episodes/eight tasks; DUR* covers 1692 episodes/seven tasks. Partial comparisons use identical episode IDs and initial-state hashes. Those arms trail their matched angular-scalar controls by 28.78 and 31.32 points respectively on completed episodes.

The completed separate-gripper comparison is mixed across suites. STK changes Spatial / Object / Goal / LIBERO-10 by -2.68 / +7.04 / +1.52 / +0.96 points versus their references. DUR changes Spatial / Object / Goal by -2.44 / +21.76 / -0.44 points; LIBERO-10 DUR remains pending. The strongest completed gain is Object separate-gripper DUR at 63.72%, compared with reference DUR at 41.96%.

Completed controlled comparisons favor grouped XYZ under STK: separate X/Y/Z streams trail the separate-gripper control by 14.28 / 13.60 / 19.00 points on Spatial / Object / Goal. Duration improves scalar XYZ to 66.76% on Spatial and 87.84% on Goal, within 0.28 and 1.08 points of their grouped controls. Object scalar-DUR remains at 45.84%, 17.88 points below grouped-XYZ gripper-DUR. Component-rate STK labels trail group-rate STK on all three suites by 2.36 / 4.64 / 1.64 points.

Completed independent-clock versus shared-clock component-DUR differences are +0.92 / -0.28 / -0.20 points on Spatial / Object / Goal, with no consistent advantage. Independent STK clocks trail shared clocks by 15.36 / 5.68 / 9.72 points, now with full coverage on all three suites.

All three Object axis/pair groupings have now finished. X/YZ, Y/XZ and Z/XY score 44.92 / 35.00 / 38.08% STK and 51.16 / 51.84 / 50.16% DUR, all below their grouped-XYZ controls (54.20% STK, 63.72% DUR). Separating integrated angular-command axes while keeping XYZ grouped scores 52.32% STK / 67.92% DUR. The matched grouped-angular-command control scores 46.76% / 52.08%, so the isolated angular-grouping differences are +5.56 / +15.84 points. Angular-scalar DUR is also 4.20 points above the separate-gripper SO(3) control and is the strongest completed Object arm in this sweep. This comparison is not yet available on the other suites. The earlier one-task angular-driver DUR result of 1.33% became 52.08% at full coverage; it was not representative of the suite.

Shared-pool quota enforcement preempted the replacement workflows on October 4 at 16:52 PT. L40S-03 canceled the October 5 replacements again during bootstrap despite reporting free capacity. After verifying their artifact prefixes were empty, the three suites moved to L40S-01 with the original immutable checkpoint/episode requests. The accepted Object evaluation workflow remains in place. Four canonical workflows now use L40S-01 at NORMAL priority. The recovery preserved 45 completed policies, 11 then-final scores, 18 interrupted checkpoints, ten then-partial evaluations, and the remaining 25 unstarted policies. Completed-policy evaluations and checkpoint resumes are ordered first. There are up to 18 training lanes / 72 GPUs and 24 evaluation lanes / 24 GPUs. At the latest scheduler audit, 18 training jobs on 72 L40S GPUs and two evaluations were running. No training task is currently stuck in SCHEDULING; 24 further training tasks wait on lane dependencies. All 18 active training runs have uploaded new periodic checkpoints. The six active LIBERO-10 runs have checkpointed 57–63% of their budgets; the closest Spatial run has checkpointed 93%. These are conservative upload-based progress bounds, not finish-time guarantees. No canonical task is currently failed.

Checkpoint recovery verifies original runtime/replay/preflight and checkpoint hashes, restores optimizer/normalizer/EMA state through the unchanged trainer, and checks the exact remaining budget. It never falls back to fresh training. CPU readiness gates validate checksummed full-budget policy receipts, and the native evaluator independently validates the downloaded checkpoint. Model/codec/trainer source remains 69ff3fab69c2d322b590809461c80d70878ba71f; host orchestration source is recorded separately.

See the current 88-cell manifest, compact score table, all available scores, paired comparisons, per-task comparisons, and uploaded training progress. The snapshot includes all 88 cell states, native-validated episode evidence, checksums, analysis scripts and scheduler states.

Tests, documentation, notebooks and reports remain on the validation companion, restored through an immutable checksummed pin. See the protocol and matrix.

Validation:

  • All 329 Linux CI tests passed on runtime 9c36037a, including codecs, all 88 configurations, training/checkpoint/inference, optimizer/EMA resume, replay and evaluation. CI. No runtime code changed for these recoveries.
  • The seven October 5 submitted specifications passed OSMO validation. Four final canonical workflows were accepted on L40S-01; three preempted L40S-03 attempts are retained as history. Canonical assignments preserve 88 unique policy/evaluation cells and matching reference IDs.
  • Native episode validation passed for 54 directories in the current snapshot (44 complete suites plus ten partial-repetition directories); paired initial-state mismatches were zero. Full scores were recomputed from episode records, and artifact checksums were verified.
  • CI builds the runtime wheel before restoring companion files. Model, codec and trainer files have no differences from the pinned training source.

Replay covers 120 previously used demonstrations across the four suites, not learned-policy success. Independent STK component clocks succeed on 64 demonstrations versus 108 for shared clocks; independent duration succeeds on 107. Rotation carry still needs a distinct distance quantum and causal-state definition: the codec retains sub-R rotation, and R is a horizon cap. Active channel counts change parameters by less than 0.03% and alter timing's share of mean epsilon loss. One training seed is used; five evaluation repetitions are not independent trainings. Recovery is explicit, not automatic requeue or protected GPU capacity.

Stacked on #195 to retain the proven LIBERO runtime. Consolidation PR #198 and other agents' worktrees remain separate.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant