Skip to content

Update public inference release, runtime contracts and documentation - #11

Merged
iChubai merged 226 commits into
OpenEnvision:mainfrom
iChubai:pr/upstream-public-inference-20261005
Oct 5, 2026
Merged

iChubai merged 226 commits into
OpenEnvision:mainfrom
iChubai:pr/upstream-public-inference-20261005

Conversation

@iChubai

@iChubai iChubai commented Oct 5, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Update upstream to the public inference release with unified inference/evaluation entrypoints, native model runtimes, scoped diffusion acceleration, strict regression tooling, and bilingual documentation. This is a release-wide source update.

The candidate's tracked files match release commit 62b624f8 exactly, including README. The candidate commit is 4e4c015a; upstream main is included in its history.

User-Visible Behavior

  • Run models through typed worldfoundry-eval run commands, the TUI, or Studio; inspect readiness before loading weights.
  • Use reversible acceleration with rounding, cache, streaming-state, and checkpoint contracts.
  • Compare raw numerical outputs and decoded exported artifacts against accepted references. Missing output, changed shape, non-finite values, and numerical drift fail validation.
  • Model pages distinguish runtime evidence, static checkpoint checks, and pending validation; generated launch commands are checked against executable schemas.

Affected Pipeline / Benchmark

  • Pipelines: packaged visual generation, navigation/video, 3D/geometry, and embodied-action runtimes under worldfoundry/synthesis and worldfoundry/base_models.
  • Benchmarks: packaged evaluation catalogs and runners under worldfoundry/data/benchmarks and worldfoundry/evaluation.
  • Runtime profiles: the release inventory under worldfoundry/data/models/runtime and benchmark runtime profiles.
  • Public entrypoints: worldfoundry-eval run, zoo, tui, model asset preparation, and Studio inference.

Change Type

  • Bug fix
  • Model integration
  • Benchmark integration
  • Pipeline/runtime change
  • Documentation only
  • Test/QA tooling

Asset / API / GPU Requirements

  • CPU contracts use the declared test dependencies; docs require Node 20.19+ and the locked npm dependencies.
  • Real inference requires model-specific checkpoints and CUDA environments. The local validation host uses H100 GPUs.
  • Hosted providers, gated checkpoints, datasets, and simulators require their declared credentials or assets when selected.

Commands Run

The passing source checks below correspond to the release tree, which is identical to this candidate.

Command Result Notes
make test PYTHON=python3 Pass 805 tests, zero skips; complete selection and source checks.
make test-docs-contracts PYTHON=python3 Pass 29 tests, zero skips; published commands and catalog claims.
make docs-check PYTHON=python3 Pass Generated API, model, benchmark, and coverage data.
bash scripts/docs/build.sh --skip-bootstrap Pass Type checking and static site build.
make lint PYTHON=python3 Pass Formatting, source, manifests, and registries.
python3 tests/manual/inference_regression_coverage.py --base origin/main --output /tmp/worldfoundry-release-certification-20261005/upstream-coverage.json Fail The full upstream comparison affects 238 native variants; 218 lack a replay recipe.

Release CI passed geometry impact, inference replay, CPU tests, inference tensors, public surface, and packaging/license checks for 62b624f8.

Checkpoint / API Key Needs

  • Checkpoints and weights: per-model manifests; staged outside the source tree.
  • Environment variables: model asset roots and provider credentials declared by the selected runtime. Replay suites use WORLDFOUNDRY_REGRESSION_* fixture settings.
  • API providers: optional, selected through runtime profiles.
  • Local cache: existing checkpoints and fixture inputs are required for real GPU replay.

Sample Artifact Evidence

  • Public source evidence: release CI results and artifacts.
  • The native replay matrix contains 42 cases, including camera-conditioned DreamX, two-chunk spatial Alaya, and conditioned DIAMOND CS:GO.
  • Candidate GPU certification and complete coverage of the upstream comparison remain pending. Real inference evidence is tied to its executed source revision, inputs, weights, and environment.

Validation Matrix Status

Area Status Evidence / Command
Unit or focused regression Pass Public CPU and documentation contracts; identical release tree.
Real inference validation Not run for candidate Exact-commit GPU certification pending.
Streaming or multi-turn path Pass for CPU contracts Release inference tensor checks; candidate checkpoint trajectories pending.
Benchmark runner / metric path Not run at full benchmark scale Declared assets and official runtime requirements apply.
Docs build or link check Pass Docs build and generated-data checks.
GPU/API-key dependent path Not run for candidate Requires model-specific assets and environments.

Leaderboard Validity Impact

  • Readiness remains scoped to the evidence recorded for each model and benchmark.
  • Demo, static, and partial validation evidence does not establish full benchmark or leaderboard validity.
  • Source tests, preflight declarations, and numerical parity are separate from official benchmark scoring.

Compatibility And Risk

  • The public source layout and entrypoints are updated across the release; downstream users should follow the maintained CLI and model pages.
  • Complete GPU replay requires local assets and model-specific environments.
  • The upstream comparison has 218 native variants without replay recipes; complete candidate GPU certification remains pending.

Checklist

  • I kept the change narrowly scoped. This is a complete release update.
  • I did not commit secrets, API keys, large checkpoints, or generated cache files.
  • I reused packaged fixtures where practical.
  • I updated docs for public behavior, installation, and entry commands.
  • I documented external asset and runtime requirements.
  • I kept readiness and leaderboard claims scoped to their evidence.

iChubai and others added 30 commits September 9, 2026 01:14
Publish the current CLI, TUI, Studio, and model configuration layout without
merging private development history. Exclude private training, local tests,
runtime artifacts, and restricted source trees already gated for this release.

Keep shared inference helpers independent of removed training modules. Tighten
ignore and packaging rules, audit wheel and sdist contents in CI, and refresh
public documentation. Retain MoVerse, MVDiffusion, and WorldOlympiad runtimes.

Preserve runtime environment compatibility exports and package the Studio
logo. Remove the Studio preset that depended on a private test fixture.
…pth Pro acknowledgements.

The builder heading and controls were a flat dump; grouping them as a card makes the pipeline identity and command tabs readable. The vendored ACKNOWLEDGEMENTS.md was leftover license text, not runtime.

Co-authored-by: Cursor <cursoragent@cursor.com>
…igures, variants, and harvested media with the command builder.

Keep the shared docs-site widgets and benchmark-hub pages that use the same MDX wiring.

Co-authored-by: Cursor <cursoragent@cursor.com>
…ot README.

Dispatch, routing, and scoring membership now load from catalog shards instead of handwritten registries. Leftover READMEs and evaluation shims move into fumadocs or shared modules so ids cannot drift across tables.

Co-authored-by: Cursor <cursoragent@cursor.com>
… tab-panel code blocks.

A raw script in the React tree trips React 19; head injection keeps the first paint on the stored theme. Drop leftover LFS exceptions for deleted readme-demo assets.

Co-authored-by: Cursor <cursoragent@cursor.com>
Relative lvdm paths break after configs moved under worldfoundry/data; one helper keeps val lists and data_config.yaml next to the runtime configs.

Co-authored-by: Cursor <cursoragent@cursor.com>
Default YAML and eval-id lists now live under worldfoundry/data/models/runtime/configs rather than next to vendored source.

Co-authored-by: Cursor <cursoragent@cursor.com>
…g them.

find_spec can initialize parent packages and optional CUDA extensions; PathFinder only locates modules so missing flash-attn still falls back to PyTorch SDPA.

Co-authored-by: Cursor <cursoragent@cursor.com>
…rts.

Wrappers now fail closed when optional GPU stacks are absent, and vendored fidelity modules import sibling helpers instead of the public torch_fidelity package.

Co-authored-by: Cursor <cursoragent@cursor.com>
…ersioned checkpoints.

Contract and existing-results paths now call one scoring helper, so metric identities, batch checkpoints, and runner parameters decide reuse instead of wrapping inference at import time.

Co-authored-by: Cursor <cursoragent@cursor.com>
Add native Zing extensions over shared Wan components and official inference
adapters for Astronex-World, SolarWM, and AlayaWorld v1.1. Refresh Echo-WM
configuration and input validation. Keep runtime configs in package data
and document asset requirements, licenses, and separate backend environments.

Exclude training lifecycle, training configs, and test artifacts. Validate
CPU model parity, rollout contracts, CLI routes, subprocess failures, registry
references, and scoped generated docs. GPU checkpoint inference remains pending.
Keep Cosmos sampling state in FP32 and generate complete temporal windows before trimming short outputs. Match released MAGI-2 sampling geometry, CFG and audio export, restore reusable component residency, and use shared fused inference kernels. Validate Cosmos GPU demos and a full 10-second MAGI-2 preview; retain diagnostics outside the release.
Fix public model loading and inference paths found during checkpoint-backed
validation, remove the unsupported ThinkSound integration, and correct the
FastVideo I2V catalog/binding. Record per-model GPU and blocker evidence
without treating interface checks or incomplete runs as successful inference.

Keep local tests out of the public Git tree, remove the three requested
Warp-as-History license/provenance files, and retain upstream attribution in
the consolidated third-party notices. SANA-WM Streaming refiner mask handling
is repaired and CPU-checked; its GPU retest remains pending in the ledger.
iChubai added 26 commits October 4, 2026 01:13
Bind policy payloads before execution and use the supported distributed control-group signature. Preserve initialization errors and lifecycle cleanup.
Report profile import and parsing failures while keeping the metadata fallback for models without a profile.
…lures

Skip staging checks when disabled and download-space checks on HF cache hits. Rebuild merged safetensors caches only for malformed files while preserving device, memory, and permission errors.
Preserve the existing default stage-three seed while honoring explicit public seed and deterministic requests across captioning, depth reconstruction, rendering, and causal inference. Keep distributed seed offsets inside the NumPy range. Add public API, subprocess RNG, geometry ordering, and strict CUDA rendering contracts to the required CPU and CUDA test selections.
@iChubai
iChubai marked this pull request as ready for review October 5, 2026 13:56
Copilot AI balanced review requested due to automatic review settings October 5, 2026 13:56
@iChubai
iChubai merged commit 322afce into OpenEnvision:main Oct 5, 2026
6 of 8 checks passed

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants