An auditable, framework-neutral reproduction of the publicly specified algorithm in CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents (Li et al., 2026), with a thin integration path for THUDM/slime.
This repository reproduces the algorithmic core. It does not claim to reproduce the paper's 30B/106B benchmark scores: the exact prompts, resume template, training subset/order, reward implementation, PPO settings, and several infrastructure details are not public.
For the full experiment narrative—including the 96-step long-horizon result, the value-function design, all Phase 1–17 outcomes, failure analyses, and cost decisions—read Reproducing CompactRL: What Worked, What Failed, and Why We Did Not Scale. A Chinese version is also available.
- Fixed-budget compaction trigger:
C - |history| < T_comp. - Atomic assistant-action + environment-observation history steps.
- Same-policy summary prompt and
system + resume(summary) + recent tailreconstruction, with the paper's defaultk=2tail behavior. - Shared terminal task reward for execution and summary segments.
- Token-normalized clipped PPO objective.
- Local token-level GAE plus the paper's cross-trajectory correction:
A_hat[s,i] = (gamma * lambda_s) ** N_after[s] * A_local[s,i]. - VAPO-style length-adaptive
lambda = 1 - 1 / (1.5 * response_length). - Separate actor/critic summary masks for a clean direct-summary-loss ablation.
- Paper-reported 50-step critic warm-up and
2 critic : 1 actorupdate support through a minimal patch against the pinned public slime commit. - Optional tokenwise-lambda and stitched-bootstrap value-target ablations, plus compaction-boundary value-gap and explained-variance diagnostics.
- Exact zero-critic full-trajectory reference used to test reward-distance correction.
- A dependency-free toy actor-critic experiment comparing joint summary training with a frozen-summary ablation.
- A 96-step memory-chain experiment with three context resets, multi-seed learning curves, a frozen-summary control, and a readable successful trace.
- A deterministic 360-task synthetic memory environment with exact rewards, balanced train/dev/test splits, and pinned regeneration hashes.
- Optional common-random-number SGLang sampling with stable per-task and per-action/summary seeds, auditable seed metadata, and exact saved-trajectory comparison gates for paired training arms.
- Optional actor-only occurrence credit, bounded post-coverage stopping pressure, and marker-gated causal-prefix redistribution. These preserve task GAE/critic targets and include exact sampled-token mass audits.
- Versioned reported-paper, eight-GPU debug, and causal-ablation settings.
- An out-of-tree integration for pinned public
slime: native compactlist[Sample]rollout generation, exact token provenance, verifier rewards, chronological segment recovery, a custom advantage hook, and an 8-GPU PPO launcher. Source/API compatibility, the official slime container runtime, a four-A10 one-update actor/critic/rollout job, and a bounded five-update Full-vs-summary-loss-masked pilot are validated live. - A publication-blocking live E2E audit: immutable run identity, independent advantage/return recomputation, actor/critic gradient artifacts, checkpoint inventory, lightweight evidence download, and a final report that cannot be generated before every real-run gate passes.
No third-party package is required for the core tests:
PYTHONPATH=src python3 -m unittest discover -s tests -v
PYTHONPATH=src python3 examples/toy_joint_training.py
PYTHONPATH=src python3 experiments/long_horizon_memory_chain.pyOr run both:
bash scripts/run_checks.shGenerate or validate the simulation data:
PYTHONPATH=src python3 scripts/generate_sim_data.py
PYTHONPATH=src python3 scripts/summarize_sim_data.py
PYTHONPATH=src:. python3 scripts/build_plan_grammar_sft.py \
--source data/synthetic/train.jsonl \
--sft-output data/synthetic/plan_grammar_sft.jsonl \
--rl-output data/synthetic/ordered_plan_train.jsonlValidate against a pinned slime checkout:
python3 scripts/check_slime_compat.py --slime-root /path/to/slimeAfter an authorized Modal GPU run, download only its lightweight evidence and build the gated final report:
bash scripts/download_modal_evidence.shThe report builder intentionally fails while the live e2e_audit.json is
missing or unsuccessful.
The committed final E2E report passed both the
96-step causal-learning gate and the live pinned-slime integration gate.
The newer Modal scale-pilot report
records nonzero post-compaction rewards and the single-seed arm comparison.
Its task-type audit is also the current LLM blocker: both arms scored 0/23
on reward-critical early_fact memory, so the pilot is not presented as proof
of learned LLM summary retention. Its historical masked arm also omitted the
summary from critic training; the value-function fix below supersedes that
ablation contract.
The value-function fix report records
the corrected mask/critic wiring and a passing bounded 4×A10 runtime audit.
The subsequent three-arm scale comparison
completed 60 updates per arm with 50 critic-only warmup updates. It verifies
the corrected value path at scale, but its saved-checkpoint capability result
is superseded: actor and critic inherited one save directory, and the later
critic save replaced the actor shards. The
deterministic checkpoint evaluation
documents the diagnosis, role-specific checkpoint fix, passing fixed-run audit,
and a paired dev-60 result. After two actor updates the fixed model scored
12/60 versus 14/60 for the base model, so there is not yet evidence of held-out
LLM capability improvement.
The follow-up seed replication and checkpoint-resume report
records a complete 60-rollout continuation and a second deterministic dev-60
comparison. Seed 17 scored 20/60 while seed 23 scored 12/60 versus the same
14/60 base, so the positive seed-17 capability result did not replicate. The
implementation is operational, but stable LLM improvement remains an open gate.
The controlled critic learning-rate ablation
shows that reducing the resumed critic LR from 3e-6 to 1e-6 halves mean
value loss and materially reduces actor/critic gradient excursions. Its dev-60
score changes only from 12/60 to 13/60 (p=1.0) while invalid commands and
truncation regress, so the stability gain is not presented as capability gain.
The follow-on actor reference-KL ablation
verifies slime's actor-only k3 KL path on the same continuation. KL 0.01
partially recovers invalid-command and truncation rates and ties the base model
at 14/60, but early-fact memory falls to 1/20. It remains an engineering
stability control, not a report-reported setting or a capability result.
The KL dose-response follow-up
records a positive KL 0.05 candidate: 21/60 overall and 12/20 early-fact
versus 13/60 and 2/20 for the matched no-KL continuation (+10/-2,
p=0.0386). It is the first positive long-horizon memory signal in the scaled
LLM experiment, but it remains provisional until common-random-number and
second-seed replication because stochastic training trajectories differ.
The subsequent deterministic sampling validation
uses two independent 4×A10 runs to verify exact equality of all eight saved
temperature-1 token trajectories and all 183 derived request seeds. This
enables common-random-number KL replication while retaining an explicit
boundary: GPU optimizer updates are not bitwise deterministic.
The controlled KL replication
then rejects the provisional KL 0.05 advantage: under common-random-number
training, no KL scores 26/60 while KL 0.05 scores 19/60 (+3/-10 from the KL
arm, p=0.0923). The no-KL checkpoint itself scores 26/60 versus 14/60 for the
frozen base (+15/-3, p=0.00754), including 10/20 versus 2/20 early-fact
tasks. This is a positive single-run candidate, not yet a replicated capability
claim; an identical independent rerun is the next gate.
That identical no-KL rerun
scores 25/60, including 16/20 early-fact, versus the original 26/60 and frozen
base 14/60. The original and repeat each beat base in paired tests
(p=0.00754 and p=0.0192), so the aggregate gain is robust to one independent
same-seed GPU rerun. Their early/latest composition and value diagnostics still
vary materially; cross-model-seed replication was the remaining gate at that
stage.
The follow-up deterministic training validation
adds the pinned slime Megatron reproducibility recipe and runs two independent
2-update 4×A10 jobs. Both updates' 16 token trajectories, actor gradients, four
critic updates, and value diagnostics match exactly. This validates controlled
optimizer experiments through two updates; it is infrastructure evidence, not
a model-quality result.
The full deterministic second-seed KL replication
extends that control to eight updates from the seed-17 checkpoint. Update 60
matches all sampled tokens, value/advantage diagnostics, and both critic
updates exactly before the KL objective separates the arms. No KL scores 29/60
versus frozen-base 14/60 (+17/-2, p=0.000729), giving positive no-KL
results on both seed 17 and seed 23. KL 0.05 scores 32/60 on seed 17 but its
direct advantage over no KL is not significant and reverses on seed 23, so KL
0 remains the default. Every arm remains 0/20 on ordered-plan; the next gate
is a 4096-token no-compaction learnability control.
That full-context learnability control
scores 40/60: early-fact and latest-update are both 20/20, with zero
compactions and zero truncations, but ordered-plan remains 0/20. Trace analysis
shows a grammar failure: all plan responses invent an audit_records array
instead of the required {"plan":"a>b>c"} string. The repository therefore
keeps the default prompt unchanged and adds a versioned, non-leaking plan-schema
hint as the next calibration gate before grammar SFT or more sparse-reward RL.
The first schema-hint calibration
keeps ordered-plan at 0/20. It eliminates invalid commands, but the model emits
placeholder ordinals or distractor records instead of the three milestone
tokens. A fixed fake-value demonstration is the final prompt-only gate; failure
there routes to a small grammar SFT rather than a compacted RL comparison.
That fixed-example gate reaches
only 1/20 under full context. The repository now includes a pinned-slime,
one-GPU grammar SFT warm-start. The first terminal-only recipe is sharply
non-monotonic: update 10 scores 0/20, update 20 scores 20/20, and update 30
returns to 0/20 as the action protocol collapses. The passing update-20
checkpoint is a valid gated candidate, while a lower-LR corrected recipe
behavior-clones legal next actions and terminal JSON to seek a wider stable
window. Compaction summaries remain entirely outside SFT targets. See the
grammar-SFT diagnosis.
The corrected run passes the full-context gate at updates 10, 20, and 30
(20/20, zero invalid commands). Its selected earliest checkpoint, update 10,
scores 8/20 with compaction versus 20/20 without compaction on the exact
same task IDs. This clean 12-task paired regression (p=0.000488) establishes
the phase-2 summary-state-transfer target while retaining nonzero exploration
reward for CompactRL.
The first live phase-2 smoke then exposed an important warm-start value bug:
slime's random missing critic head generated nonzero returns and a policy
gradient on an all-zero-reward batch. The integration now zero-initializes only
that missing scalar head. A same-task, byte-identical trajectory replay reduces
values, returns, advantages, both critic updates, and the actor gradient to
exactly zero; the automated safety audit passes 19/19 checks. See the
phase-2 warm-start report.
The first 12-rollout phase-2 pilot
then passes the full slime E2E audit and receives seven real sparse rewards,
but held-out compacted accuracy moves from 8/20 to 7/20. Paired traces show
three gains and four regressions; regressed summaries repeat the first
milestone and grow by 28.5 tokens on average. Because the nominal four-step
critic warm-up contained no reward, the next controlled setting uses 12
critic-only batches and halves actor LR before making any improvement claim.
That 20-rollout
reward-effective warm-up
is the first positive long-horizon result: the critic sees three positive tasks
before the actor is enabled, the final actor scores 16/20 versus the frozen
baseline's 8/20, and paired outcomes are nine improvements to one regression
(p=0.0215). Summary milestone recall rises from 56.10% to 93.90% with zero
invented plan tokens. The subsequent
independent-seed replication and locked test
initially passes the two remaining gates. Seed 29 scores 15/20 on dev, while
the fresh test20 comparison is 6/20 for the SFT baseline, 16/20 for RL seed
23, and 13/20 for RL seed 29. Those checkpoint measurements remain valid,
but the later
seed-sensitivity and reward-gate experiment
changes the conclusion: an identical seed-37 run scores only 6/20 on dev,
2/20 on ordered-plan test, and 30/60 on full test. Its test summary recall
falls to 21.43%, with increased invention and repetition. The fixed ungated
recipe succeeds on two of three observed training seeds and is not robust
enough to scale as a default.
The opt-in raw-reward gate fixes the concrete zero-positive-batch failure mode.
It continues critic training but skips an actor update unless the batch has at
least one positive verifier reward. Under the exact seed-37 counterfactual all
eight actor-phase batches remain zero-positive, so actor/critic update counts
are 0/40 and deterministic dev20 is exactly unchanged from SFT at 8/20.
The remote skip audit passes all 11 applicable checks. This is a verified safe
no-op, not learning: the next experiment must create a new dev/test protocol
and make reward-bearing actor batches sufficiently likely before model
scale-up. Across all runs, the evidence is small-model synthetic long-horizon
wiring and capability evidence through pinned slime, not a GLM-scale benchmark
claim or a universal Pareto improvement.
That next experiment is now complete. The
phase-3 protocol
uses new phase3 task IDs, 480 train-only ordered-plan rows, and fresh 60-task
ordered-plan dev/test slices. It selects rollout/global batch 24 from the
frozen-actor reward estimate (12/288) and requires at least four
reward-bearing actor batches plus a 10-point dev gain on two independent seeds
before opening the new locked test. All three seeds improve fresh dev from the
SFT baseline's 22/60: seed 41 scores 31/60, seed 53 41/60, and diagnostic
seed 67 36/60. Seeds 41 and 67 pass all 12 per-seed training/dev checks; seed
53 fails only the preregistered invention guardrail because summary invention
incidence rises by 6.45 points versus a five-point maximum. Under the strict
first-two-seed rule, locked test remains unopened. The exact evidence and
interpretation are in the
phase-3 result.
The next stability experiment is preregistered on another fresh namespace.
Phase 4 compares
matched no-KL and actor reference-KL 0.01 arms while leaving rewards, GAE,
returns, critic targets, and the raw-reward actor gate unchanged. It adds an
explicit repetition guardrail and requires the KL arm to reduce invention or
repetition without losing more than five capability points in two fixed seed
pairs before any new locked test can open.
Phase 4 is now complete and the KL control fails that gate. In pair 1, no-KL
and KL score 28/60 and 25/60; in pair 2 they score 52/60 and 39/60.
KL consistently reduces invention, repetition, and summary length, but loses
5.0 and 21.7 exact-match points plus 6.7 and 23.4 milestone-recall
points against its matched no-KL arm. All GAE/value/KL objective audits pass,
so this is an actor-objective result rather than a changed critic target. The
phase-4 locked test remains unopened. See the
phase-4 result.
Across the four phase-4 arms, zero-reward trajectories still account for
26.0%--41.0% of absolute actor advantage mass once a batch-level gate has
admitted an update. Phase 5 therefore isolates an opt-in, actor-only control:
compute the same GAE/returns for every trajectory and keep critic training
unchanged, but zero policy advantages for zero raw-reward logical trajectories
inside a reward-bearing batch. This engineering ablation is not part of the
report-faithful default. The fixed data, two seed pairs, and open-test rule are
in the phase-5 preregistration.
Phase 5 is now complete and does not pass its two-pair gate. Pair 1 passes and
favors positive-trajectory advantages: exact match rises from 42/60 to
44/60, milestone recall from 86.9% to 92.3%, invention falls from 6%
to 2%, and repetition from 13% to 10%. Pair 2 has zero reward-bearing
actor batches after the 12-step critic warmup, so both arms correctly perform
zero actor updates, continue all 40 critic updates, and exactly reproduce the
SFT dev result (19/60). The second pair therefore cannot exercise the mask;
the phase-5 locked test remains unopened. The audits, value-function behavior,
and interpretation are in the
phase-5 result.
Phase 6 is preregistered on a fresh namespace to replicate the pair-1 effect
without the pair-2 opportunity failure. It fixes 36 updates (12
critic-only plus 24 actor-phase), batch size 48, exactly 1728 unique
ordered-plan train tasks, a 2048-token dynamic microbatch cap, two new seed
pairs, and a minimum of six
reward-bearing actor batches per arm. The scale is chosen using only already
open phase-5 train outcomes; neither fresh dev nor locked test participates.
See the phase-6 protocol.
Phase 6 is now complete. Scaling removes the actor-opportunity failure and all
four arms learn the long-horizon task: final dev exact match is 55/60 and
50/60 in pair 1, and 52/60 and 48/60 in pair 2, against the fresh SFT
baseline's 19/60. But the positive-only actor mask loses to the batch-gate
control in both pairs, and every RL arm violates the preregistered summary
invention/repetition guardrails. All critic, GAE, return, mask, and optimizer
audits pass. The phase-6 locked test remains unopened; see the
phase-6 result.
Phase 7 addresses the exposed objective misspecification without changing the
critic architecture. Its candidate keeps one shared scalar return but sets it
to terminal exact match x minimum observed-summary plan-token F1. Missing
milestones lower recall; invented, future, or repeated milestone tokens lower
precision. The scalar critic, local and cross-trajectory GAE, masks, failed
trajectory critic training, and optimizer schedule remain intact. It uses a
fresh 1728-task train namespace and two fixed seed pairs; a remote auditor must
reconstruct every shaped reward before any dev claim or locked-test opening.
See the
phase-7 preregistration.
Phase 7 is now complete and the multiplicative fidelity reward fails its
two-pair gate. It retains the long-horizon capability gain: the candidates
score 51/60 and 52/60 against the fresh SFT baseline's 24/60. But neither
candidate reduces invention or repetition by the required ten points versus
its matched terminal-reward control; both also exceed the SFT-relative quality
limits and make summaries 12.7--13.2 tokens longer than control. All scalar
reward reconstruction, critic, GAE, return, mask, optimizer, and checkpoint
audits pass. The failure is therefore attributed to sparse multiplicative
credit (R=0 for every failed task), not to the scalar critic architecture.
The phase-7 locked test remains unopened. See the
phase-7 result.
Phase 8 is preregistered on another fresh namespace. It keeps the successful
phase-7 return unchanged but gives failed trajectories a dense quality cost:
R = task_exact + minimum_summary_F1 - 1. Thus faithful failures receive zero
and degenerate failures receive a negative scalar. Negative trajectories stay
in both actor and critic training whenever a positive trajectory admits the
batch; the scalar head, GAE, cross-trajectory correction, masks, schedule, and
control remain unchanged. The fixed hashes, seed pairs, negative-return audit,
dev gates, and locked-test rule are in the
phase-8 preregistration.
Phase 8 is now complete and fails in the same way in both seed pairs. The
dense candidates score 0/60 and 2/60, versus 53/60 and 55/60 for the
matched terminal controls; their summaries collapse to 2.9 and 5.6 mean
tokens. All reward reconstruction, critic, value-mask, GAE, return, gradient,
schedule, and checkpoint audits pass, and the dense critics track the supplied
negative returns. The failure is therefore actor-side: admitted batches retain
predominantly negative trajectories, whose signed policy gradients suppress
almost every sampled summary. The locked test remains closed. See the
phase-8 result.
Phase 9 freezes the dense critic target and changes only actor routing. Both
arms use R = task_exact + minimum_summary_F1 - 1 and train the critic on every
trajectory. After GAE, only the candidate zeroes actor advantages for logical
returns that are not positive. Fresh data, two fixed seed pairs, strong
SFT-relative quality gates, matched rescue gates, and the locked-test policy
are frozen in the
phase-9 protocol.
Phase 9 is complete and replicates the mechanism in both fixed seed pairs. The signed controls collapse to 5/60 and 0/60 dev, while actor-only any-success routing reaches 51/60 and 55/60. Every reward, value, GAE, return, mask, schedule, checkpoint, matched-stream, and cross-arm tensor audit passes. The two candidates nevertheless reach 18.5%/40.2% invention and 98.9%/94.6% repetition, failing the same two frozen quality guardrails. The locked test remains closed. See the phase-9 result.
Phase 10 is preregistered on another fresh namespace to isolate the remaining
actor-admission issue. Both arms keep Phase 9's dense critic and post-GAE actor
routing; the control admits any success (R > 0) while the candidate admits
only a task success whose minimum sampled-summary F1 is exactly one
(R > 0.999). The threshold is shared by the batch gate and trajectory mask,
and every below-threshold trajectory remains critic data. See the
phase-10 protocol.
Phase 10 is complete and the locked test remains closed. Pair 1's perfect-fidelity arm keeps 54/60 dev exact versus 57/60 for any-success while cutting repetition from 93.75% to 37.89%, but it still fails the frozen SFT-relative invention and repetition limits. Pair 2 supplies the missing strict causal opportunity: at an identical update-12 tensor batch, the candidate masks exactly one additional imperfect-success actor trajectory while retaining 5,474 critic tokens and all return targets. Both Pair 2 arms complete 72 critic updates and reach final value explained variance above 0.73. Pair 2 dev was not run after a low-Modal-balance warning because Pair 1 had already made the locked-test decision irreversible. See the phase-10 result.
Phase 11 passed its bounded preflight and both training gates. Stage 2a
completed in 1,556.24 seconds and used at most 103.75 A10 GPU-minutes. Stage
2b resumed that exact checkpoint at update 13, completed through update 35 in
2,673.45 seconds, and used at most 178.23 A10 GPU-minutes (about $3.27 at the
recorded A10 rate). Across the two stages the candidate made 19 actor and all
72 critic optimizer updates. Runtime, fidelity, matched task-stream,
strict-update-12, summary-KL coverage, and objective-parity audits all passed.
The training authorization is consumed and closed. The candidate-only dev
screen was fail-closed before GPU allocation and capped at 600 seconds on
2xA10 (20 A10 GPU-minutes per attempt). It is now complete: the candidate kept
54/60 exact and 99.19% summary recall, but failed invention (7.53% > 5%)
and repetition (65.59% > 20%). Fresh dev and locked test remain closed. It
nevertheless improved the exact matched no-KL control from 50 to 54 exact,
26.88% to 7.53% invention, and 88.17% to 65.59% repetition. The
direction works; coefficient 0.01 with the full actor-token denominator is
not strong enough. It applies k3
reference KL only to sampled summary tokens while retaining the original full
actor-token denominator; PPO advantages, returns, actor admission, and the
dense critic path are unchanged. To conserve quota, the completed Phase-10
Pair-2 perfect-fidelity run is the matched no-KL control, so only one candidate
would be trained. To avoid paying the full candidate cost before the first
strictly comparable actor update is known-good, candidate training is split:
Stage 2a runs updates 0--12 from the frozen SFT warm start and performs the
strict update-12 tensor audit; only a separately authorized Stage 2b may resume
that exact checkpoint for updates 13--35. Both stages emit a persistent,
fail-closed gate artifact. The preflight completed one actor update and four
critic updates with every built-in audit passing; cumulative training usage
across the failed device-placement attempt, passing retry, Stage 2a, and Stage
2b was at most 340.99 A10 GPU-minutes (about $6.26 at the recorded A10 rate).
See the
phase-11 protocol.
The toy is a behavioral smoke test, not evidence of LLM improvement. It checks the causal idea: a summary decision and a post-compaction execution decision share one terminal reward, so updating the summary policy can beat freezing it.
For history h_t = (system, user, (action_1, observation_1), ...):
- If remaining context is below
T_comp, sample a summaryS_tfrom the same trainable policy using a fixed summary instruction. - Continue from
system + resume(S_t) + last_k_atomic_steps, reducingkfrom two if needed to fit the peak context. - Export every execution and summary response as a chronological segment. Optimize only tokens actually sampled by the policy; mask prompts, observations, copied summaries, and copied tail turns. In the summary-loss-masked ablation, keep the summary segment and its critic mask but set only its actor mask to zero.
- Give every segment from one logical rollout the same verifier reward.
- Compute local GAE inside each segment. If
N_after[s]optimized tokens occur in later segments, multiply all local advantages in segmentsby(gamma * lambda_s) ** N_after[s]. - Optimize PPO by summing over all enabled tokens and dividing once by the total enabled-token count in the global batch.
The crucial invariant is token provenance: summary tokens get loss only at the turn where the policy sampled them. Their copied appearance in the resumed prompt must be masked.
The public slime main branch audited for this project (commit
fb42ae456fac8166afb604f13b30d22bb3c75053, 2026-07-15) provides the required
infrastructure primitives:
- PPO with a critic;
--calculate-per-token-loss;--custom-advantage-function-path;list[Sample]fan-out with sharedrollout_id;- loss masks and token-in/token-out agent adapters;
- compacted/sub-agent trajectory grouping.
It does not expose a built-in CompactionRL estimator or the paper's rollout,
so this repository supplies both through slime's supported customization
interfaces. Keeping the research logic here makes the equations testable
without a large framework deployment.
See the slime integration guide for the concrete contract and the technical analysis for report claims, experimental results, ambiguities, and a staged reproduction plan.
src/compactionrl/advantages.py paper equations and exact reference
src/compactionrl/compaction.py trigger and context reconstruction
src/compactionrl/losses.py token-level PPO reducer and ablation
src/compactionrl/adapters/slime.py optional public-slime advantage hook
integrations/slime/rollout.py slime custom compact rollout
integrations/slime/run_*.sh pinned configurable-GPU PPO launcher
integrations/slime/modal_e2e.py official-image Modal preflight and E2E
integrations/slime/patches/ minimal pinned-slime critic-loop patch
scripts/audit_slime_e2e.py independent live-artifact E2E gate
scripts/audit_reward_gated_skip.py zero-positive actor-skip audit
scripts/summarize_training_curve.py logical reward curve from slime dumps
scripts/summarize_value_diagnostics.py critic and boundary-value log summary
scripts/summarize_eval_rollout.py fixed held-out logical-task metrics
scripts/compare_eval_runs.py paired checkpoint comparison and exact test
scripts/render_megatron_roles.py isolated actor/critic checkpoint paths
scripts/build_final_e2e_report.py publication-blocking evidence report
scripts/download_modal_evidence.sh lightweight Modal evidence retrieval
examples/toy_joint_training.py joint-vs-frozen summary smoke test
experiments/long_horizon_memory_chain.py 96-step multi-compaction experiment
src/compactionrl/simulation.py deterministic environment and verifier
data/synthetic/*.jsonl 240/60/60 generated task splits
data/synthetic_phase3/*.jsonl fresh 1440/180/180 preregistered splits
configs/ reported settings and debug presets
results/long_horizon_memory_chain.* measured multi-seed result and path trace
results/modal/image_inspection.json measured official-image runtime preflight
tests/ deterministic algorithm tests
docs/TECHNICAL_ANALYSIS.md deep report analysis and disclosure gaps
docs/SLIME_INTEGRATION.md integration contract and launch sketch
docs/SIMULATION.md data design and first experiment settings
blog/compactionrl-reproduction.zh-CN.md complete Chinese experiment narrative
blog/compactionrl-reproduction.md complete English experiment narrative
reports/modal_scale_pilot_20260718.md bounded live comparison and claim boundary
reports/modal_scale_value_comparison_20260718.md corrected three-arm scale report
reports/modal_checkpoint_evaluation_20260718.md checkpoint diagnosis and dev eval
reports/modal_kl_dose_response_20260719.md KL 0/0.01/0.05 candidate result
reports/modal_deterministic_sampling_20260719.md exact SGLang trajectory gate
reports/modal_controlled_kl_replication_20260719.md common-random-number KL result
reports/modal_deterministic_no_kl_replication_20260719.md same-seed robustness result
reports/modal_deterministic_training_validation_20260719.md exact optimizer repeat gate
reports/modal_deterministic_seed17_kl_replication_20260719.md eight-update second-seed gate
reports/modal_full_context_learnability_20260719.md no-compaction grammar gate
reports/modal_phase2_seed_sensitivity_reward_gate_20260719.md current phase-2 conclusion
reports/modal_phase3_preregistered_protocol_20260719.md fresh-data scale-up gate
reports/modal_plan_schema_hint_20260719.md non-leaking plan-schema calibration
reports/modal_plan_example_gate_20260719.md fixed-example gate and SFT decision
reports/phase16_offline_credit_redesign_20260724.md coverage-stop live result
reports/phase17_causal_prefix_postmortem_20260724.md causal-credit mechanism result and rejected pairing
results/modal/scale_20260718/ manifests, audits, curves, value diagnostics
results/modal/checkpoint_eval_20260718/ fixed checkpoint/evaluation summary
There are three separate evidence levels:
- Verified here: equations, trigger/reconstruction state transition, token-loss weighting, deterministic tests, the controlled 96-step learning experiment, and live Qwen/slime actor+critic updates with independently audited gradients, GAE, masks, compact paths, checkpoint artifacts, and nonzero exact-match rewards after compaction.
- Specified by the paper but not run here: SWE-Dev training through slime, GLM-4.7-Flash / GLM-4.5-Air-SFT, Harbor, Terminus-KIRA, and paper-scale ablations.
- Not publicly reproducible exactly: the internal GLM-5.2 training run and undisclosed recipe details.