Skip to content

Integrate canonical ARC semantics with graph lifecycle contracts - #200

Open
ElmoPA wants to merge 41 commits into
mainfrom
codex/graph-consolidation-stack-20261001/12-canonical-arc
Open

ElmoPA wants to merge 41 commits into
mainfrom
codex/graph-consolidation-stack-20261001/12-canonical-arc

Conversation

@ElmoPA

@ElmoPA ElmoPA commented Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator

Integrate the canonical ARC stack from #197 while retaining the model-owned inference, data, checkpoint, and evaluation contracts from #150–#159. Preserve per-arm rotation clocks, explicit ARC layouts, timed reconstruction and mode-aware evaluation alongside retained graph recipes.

This layer preserves the original history-preserving merge 47b94298, including source tip 20507c68. Its diff is against the config-alias layer immediately below it; the weighted/E1 reconciliation follows in the next PR. No source branch or commit was rewritten.

Validation: stack construction checked parent ancestry and exact commit identity. #198 carries the assembled validation receipt for code commit 96df46f1 (1,993 CPU tests and 439 YAML/constructor contexts); it does not establish independent acceptance of this intermediate snapshot. See the ARC contract and conflict-resolution ledger in #198 for final integration decisions.

Keep this PR in draft pending the assembled real-weight/data OSMO gate and applicable fresh CI/review.

Stack order: #150–#159 → #199 → #200 → #201 → #198.

rpunamiya and others added 30 commits September 24, 2026 15:04
Every hpt_stems.ResNet image stem in the visual recipes passed weights:
null, so resnet18 was built with random weights -- the configs opted out
of the pretrained init the class itself declares ("DEFAULT"). All 9 stems
across yam_bimanual_hpt, hpt_yam_visual, and hpt_bimanual_visual now
request DEFAULT.

Pretrained weights are only correct if the input distribution matches, and
ResNet.forward applied no normalization while images arrive in [0, 1]. Add
an opt-in imagenet_normalize flag that applies the ImageNet channel stats
inside forward, and enable it on the stems that now load DEFAULT.

The flag defaults to False on purpose: E1ImageStem normalizes outside the
encoder, so defaulting to True would double-normalize every e1 run. It is
the only Python-side ResNet caller, and it does not pass the flag.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Ports the vectorized interpolation from the arc-speedup lineage onto the
three-way arc_chunking_mode tokenizer. The lineage was cherry-picked rather
than merged: it carries a second, independent race implementation
(translation_horizon_mode plus _race_hybrid_* methods) that duplicates what
arc_chunking_mode already does, and merging both would have left the tokenizer
with two switches for one behavior -- with these configs setting
arc_chunking_mode=race while translation_horizon_mode stayed at its "joint"
default, the wrong path would have fired.

_bracket_segments could not simply be reused for the per-arm modes. It clamps a
target at or past the end onto the final segment, whereas the per-arm rule takes
the first frame that reaches the distance and treats every later frame as a
hold, so _first_crossing_brackets vectorizes that second rule and the mode-aware
wrappers route each mode to the vectorization of its own scalar rule.

The hybrid path also stops calling tokenize(): it replaces every waypoint and
recomputes every velocity row, so that token was built and discarded. The guard
is narrower than the hybrid branch below it because the "duration" and "mean"
velocity modes still consume it.

Gates, both run under Slurm against 257c4f8 as the reference implementation
loaded from git rather than a monkeypatched copy of the current one:

  exact parity  17 corpus cases across all three chunking modes, byte-identical
                tokens and preserved rows under np.array_equal, zero mismatches
  throughput    74.4 -> 5.1 ms joint_distance, 76.8 -> 5.2 ms race,
                77.0 -> 5.2 ms multistream; ratios 0.067-0.069 against a
                required 0.70

The frozen corpus is rekeyed onto arc_chunking_mode and extended: the cases
inherited from the other lineage covered joint and race only, and their "joint"
names actually resolved to multistream here. GOLDEN is regenerated from the
reference implementation, so it records the pre-optimization bytes.

116 tests pass across parity, tokenizer, horizon dedup, open-loop sim eval and
YAM hybrid ARC.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Two differences from main's default ARC that only the hybrid chunking modes
see, both in the path that skips ArcLengthTokenizer.tokenize():

- A stationary arm was resampled on a zero-length distance coordinate, which
  freezes its gripper at frame 0. tokenize_at's kind="zero" token repeats the
  position and zeroes translational velocity, but still carries the gripper to
  the end of the window. Restate that rule per arm in the hybrid paths.

- The distance-resolved source window was bounded at 600 frames, three times
  main's fixed 200-frame ARC window. Frames past an arm's D crossing feed no
  clock, so inside 200 the two policies agree; past it the trainer sees chunks
  several times longer in wall time than main's ARC ever produced. multistream
  waits for both arms and ran a median of 372 source frames. Bound the buffer
  at 200.

max_steps_per_chunk is deliberately not restated: an arm that moves but cannot
reach D inside the window resamples its available motion, which is both main's
behaviour and the contract pinned by
test_multistream_short_arm_resamples_all_100_waypoints_then_holds_on_decode.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
(cherry picked from commit 22a73ab342af284678f28b84e8b3f9b69b1651b1)
AnikethCheluva and others added 8 commits September 29, 2026 03:31
The stacked layout puts the velocity rows under the waypoints, so one bimanual
arc token is (2M, 14) and every column carries both a shape value and a rate.
A per-column normalizer therefore pools metres with metres per second, and the
sequence the model reads is twice as long as the number of waypoints.

The wide layout puts the velocity rows beside the waypoints instead: (M, 28),
columns 0..13 shape and 14..27 velocity. Same numbers, half the sequence
length, and each stream gets its own normalization statistics.

Wide is now the default wherever it is defined. It needs one velocity row per
waypoint, so it has no meaning for velocity_mode='mean', which emits a single
row for the whole token; default_bimanual_velocity_layout names that one
exception rather than hiding it in a branch. bimanual_arc_token_shape is the
single source of truth for the resulting shape, and _join_token is the only
place the two halves are joined, so the layouts cannot drift apart.

The knob is called velocity_layout, not token_layout: token_layout already
means the E1 deploy layouts ('lab', 'e1_profile', 'e1_dur') in arc_decoder.py
and inference_config.py, and two different meanings for one name is a trap.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The default bimanual arc token is now wide -- (M, 28), velocity beside each
waypoint -- but both evaluators still asserted the stacked (2M, 14) shape, so
an arc run raised on its own predictions and a shape check meant to catch
misconfiguration fired on the correct configuration.

Detection and conversion are split:

  bimanual_arc_token_shapes(M, velocity_mode) lists every shape the mode can
  produce, so _is_arc_prediction and _is_arc recognize a token in either
  layout. One evaluator serves both arms of the ablation, and wide is
  undefined for velocity_mode="mean", which is why this is a list of shapes
  rather than one.

  stack_arc_token(token) restacks wide to stacked and passes a stacked token
  through. Everything downstream indexes token ROWS -- the truncation helpers,
  the arcmatch waypoint slice -- so each evaluator converts once at its
  boundary and pins its own codec to velocity_layout="stacked". The shape says
  which layout arrived, so no run has to declare its layout twice.

The abc evaluator's mode-mismatch check now runs before its width check: a
token for the other velocity mode can itself be wide, and "wrong
velocity_mode" is the more precise diagnosis than "not a bimanual chunk".

Verified locally: the two new cases in tests/test_bimanual_velocity_layout.py
pass (15 total). The evaluator paths import torch, so the full suite still has
to run through Slurm.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The arc experiments carry 100 token rows of 28 columns, so the flow stages
read ${oc.select:hpt.token_dim,${hpt.action_dim}} while the ee_pose proprio
stem keeps hpt.action_dim: ee_pose is 14-dim permanently. The oc.select
fallback leaves every non-arc run byte-identical.

The config contracts pin the wide shape, including the token_dim knob, so a
silent fall back to the stacked default fails here instead of in training.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@ElmoPA ElmoPA changed the title feat(arc): preserve native source windows and timed reconstruction Integrate canonical ARC semantics with graph lifecycle contracts Oct 2, 2026
@ElmoPA
ElmoPA marked this pull request as ready for review October 2, 2026 19:02

ElmoPA commented Oct 7, 2026 •

Copy link
Copy Markdown
Collaborator Author

Merge activity

  • Oct 7, 9:07 PM UTC: A user started a stack merge that includes this pull request via Graphite.
  • Oct 7, 9:28 PM UTC: Graphite couldn't merge this PR because it had merge conflicts.

@ElmoPA
ElmoPA changed the base branch from codex/graph-consolidation-stack-20261001/11-config-aliases to graphite-base/200 October 7, 2026 21:26
ElmoPA added a commit that referenced this pull request Oct 7, 2026
…#199)

Replace five identical model copies with thin aliases and express the two Qwen 180M variants through explicit flow-head overrides. Existing callers retain their resolved graph behavior.

This extends the existing graph-contract stack ending at #159. Review the config deduplication and compatibility layer independently of the ARC and weighted-training integration above it. The branch points at the original consolidation commit `67357575`; no commit or code was rewritten to create this stack.

Validation: the committed deduplication evidence records unchanged composition for all 254 contexts in this layer. Stack construction verified exact commit identity and parent ancestry. The complete assembled CPU validation and source dispositions are attached to #198; those results apply to the assembled code commit, not separately to every intermediate layer.

Keep this PR in draft pending the assembled real-weight/data OSMO gate and applicable fresh CI/review.

Stack order: #150–#159 → #199 → #200 → #201 → #198.
@ElmoPA
ElmoPA changed the base branch from graphite-base/200 to main October 7, 2026 21:27
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants