Skip to content

About

No description, website, or topics provided.

Resources

Stars

144 stars

Watchers

6 watching

Forks

Repository files navigation

Patch Policy: Efficient Embodied Control via Dense Visual Representations

[Project Website] [Paper] [Dataset]

Gaoyue Zhou*, Zichen Jeff Cui*, Ada Langford, Bowen Tan, Yann LeCun and Lerrel Pinto, New York University, Meta AI, AMI Labs

demo.mp4

Method overview

This repo contains code for training and reproducing sim environment experiments across four simulation environments: Push-T, Block Pushing, LIBERO Goal, and Cube.

Getting Started

  1. Setup
  2. Datasets
  3. Training a policy
  4. Choosing the visual encoder

Setup

Create the conda environment (this installs everything, including the CUDA build of PyTorch):

conda env create -f conda_env.yml
conda activate patch-policy

Tested on Ubuntu 22.04 with CUDA 12.8. To log training runs, log in to Weights & Biases with wandb login (or set export WANDB_MODE=disabled to turn logging off). In ./configs/env_vars/env_vars.yaml, set wandb_entity to your wandb username.

Intel XPU Support

Patch Policy training and LIBERO Goal evaluation are supported on Intel GPUs through the PyTorch XPU backend. The simulator runs on the CPU, while the frozen visual encoder and policy run on xpu:0 through Accelerate.

Requirements

  • Python 3.12
  • PyTorch 2.13.0+xpu
  • A supported Intel GPU with the matching Intel GPU runtime and oneAPI drivers
  • The LIBERO assets configured as described below

The repository's conda_env.yml installs CUDA PyTorch. Keep that environment for CUDA and create the separate XPU environment from conda_env_xpu.yml:

conda env create -f conda_env_xpu.yml
conda activate patch-policy-xpu

The validated XPU environment uses Python 3.12, PyTorch 2.13.0+xpu, Accelerate 1.14.0, MuJoCo 3.2.7, and robosuite 1.4.1.

Verify the backend before starting a run:

python -c "import torch; print(torch.__version__, torch.xpu.is_available())"

Download and unpack the datasets as described in Datasets, then run the LIBERO Goal configuration with device=xpu:

MUJOCO_GL=egl WANDB_MODE=disabled python -u train_policy.py \
   --config-name train_libero_goal_1gpu \
   device=xpu \
   env_vars.dataset_root=/path/to/patch_policy_datasets \
   env_vars.save_path=/path/to/patch_policy_outputs \
   epochs=10 eval_on_env_freq=5 num_env_evals=10 num_final_evals=50 num_envs=5 \
   +dataset.subset_fraction=1.0

The XPU path was validated on the LIBERO Goal task with 10 tasks and one episode per final evaluation. The run produced loadable model_final.pt checkpoints and finite actions, with 50% final-evaluation success; an epoch-10 evaluation reached 100%. The validation run used subset_fraction=1.0 and 5 parallel environments. Simulation remains CPU/EGL, and WANDB_MODE=disabled can be used when W&B logging is not configured.

Datasets

The datasets for all four simulation environments are hosted on the Hugging Face Hub: gaoyuezhou/patch-policy-datasets.

Environment Folder (after unzip) Zip size
Push-T pusht_dataset 61 MB
Cube cube_dataset 2.2 GB
LIBERO Goal libero_dataset 7.7 GB
Block Pushing block_push_dataset 5.9 GB

Download

  1. Install the Hugging Face Hub CLI:
    pip install "huggingface_hub==0.36.2"
    
  2. Download the dataset repo to a local directory (this is the directory all four datasets will live in):
    huggingface-cli download gaoyuezhou/patch-policy-datasets \
      --repo-type dataset --local-dir patch_policy_datasets
    
    To download only a subset, add e.g. --include "pusht_dataset.zip".
  3. Unzip each dataset in place:
    cd patch_policy_datasets
    for f in *.zip; do unzip -q "$f"; done
    

Point the code at the data

The dataloading code reads from a single dataset_root directory — no code changes are needed, you only set this path.

  • In ./configs/env_vars/env_vars.yaml, set dataset_root to the unzipped directory (e.g. the absolute path to patch_policy_datasets), and set save_path to where you want training/rollout results saved (e.g. the root directory of this repo).

The expected layout under dataset_root is:

patch_policy_datasets/
├── pusht_dataset/
├── cube_dataset/
├── libero_dataset/
└── block_push_dataset/

Note: The .pth/.npy/.pkl files are loaded with torch.load / numpy, which unpickle data. Only use datasets you trust.

Training a policy

Policy training and online evaluation both run through train_policy.py, driven by the configs in configs/. A run trains the policy on top of a frozen visual encoder and periodically rolls it out in the simulator.

python train_policy.py --config-name train_pusht        # Push-T
python train_policy.py --config-name train_blockpush    # Block Pushing
python train_policy.py --config-name train_cube         # Cube
MUJOCO_GL=egl python train_policy.py --config-name train_libero_goal   # LIBERO Goal
  • Diffusion policy variants are available for every environment — append _diffusion to the config name (e.g. train_pusht_diffusion). The default configs use a VQ-BeT policy head.
  • Checkpoints are written under save_path.

Single-GPU configs

The config names above assume a node of 8 GPUs. We also provide _1gpu variants of every config, tuned to fit within 32 GB of VRAM:

python train_policy.py --config-name train_pusht_1gpu
python train_policy.py --config-name train_blockpush_1gpu
python train_policy.py --config-name train_cube_1gpu
MUJOCO_GL=egl python train_policy.py --config-name train_libero_goal_1gpu

These runs use DINOv2 ViT-S (dino_patch) and precompute the frozen encoder's features once at startup. Observation windows are shortened where memory requires it — LIBERO Goal VQ-BeT drops from 10 to 2, and Cube from 5 to 2 — and batch sizes are set per environment. _diffusion variants are available here too (e.g. train_pusht_diffusion_1gpu).

Launcher configs for submitting to a SLURM cluster are in configs/cluster/.

Choosing the visual encoder

The encoder is frozen and selected via the encoder config group (configs/encoder/); off-the-shelf encoders need no training. The default is DINOv2 patch features (dino_patch). Override it on the command line:

python train_policy.py --config-name train_pusht encoder=webssl_patch
python train_policy.py --config-name train_pusht encoder=vjepa2_patch
python train_policy.py --config-name train_pusht encoder=dinov3_patch
python train_policy.py --config-name train_pusht encoder=siglip2_patch

Each of these uses dense patch features. Two pooled variants are also provided for comparison: *_patch_avg_pool (patch tokens mean-pooled into a single vector) is available for every encoder above as well as dino, and *_cls (the CLS token instead of patch tokens) is available for dino, dinov3, and webssl. ResNet-18 baselines (resnet18_imagenet, resnet18_random) and pretrained DynaMo encoders are also supported — see configs/encoder/ for the full list.

Note (DINOv3): the DINOv3 weights are hosted in a gated Hugging Face repo. To use the dinov3_* encoders, request access at facebook/dinov3-vits16plus-pretrain-lvd1689m and log in with hf auth login before launching.

Citation

If you find our work useful, please consider citing:

@misc{zhou2026patchpolicyefficientembodied,
      title={Patch Policy: Efficient Embodied Control via Dense Visual Representations}, 
      author={Gaoyue Zhou and Zichen Jeff Cui and Ada Langford and Bowen Tan and Yann LeCun and Lerrel Pinto},
      year={2026},
      eprint={2607.18236},
      archivePrefix={arXiv},
      primaryClass={cs.RO},
      url={https://arxiv.org/abs/2607.18236}, 
}

About

No description, website, or topics provided.

Resources

Stars

144 stars

Watchers

6 watching

Forks

Releases

Packages

Contributors

Languages