[Project Website] [Paper] [Dataset]
Gaoyue Zhou*, Zichen Jeff Cui*, Ada Langford, Bowen Tan, Yann LeCun and Lerrel Pinto, New York University, Meta AI, AMI Labs
demo.mp4
This repo contains code for training and reproducing sim environment experiments across four simulation environments: Push-T, Block Pushing, LIBERO Goal, and Cube.
Create the conda environment (this installs everything, including the CUDA build of PyTorch):
conda env create -f conda_env.yml
conda activate patch-policy
Tested on Ubuntu 22.04 with CUDA 12.8. To log training runs, log in to Weights & Biases with wandb login (or set export WANDB_MODE=disabled to turn logging off). In ./configs/env_vars/env_vars.yaml, set wandb_entity to your wandb username.
Patch Policy training and LIBERO Goal evaluation are supported on Intel GPUs
through the PyTorch XPU backend. The simulator runs on the CPU, while the
frozen visual encoder and policy run on xpu:0 through Accelerate.
Requirements
- Python 3.12
- PyTorch
2.13.0+xpu - A supported Intel GPU with the matching Intel GPU runtime and oneAPI drivers
- The LIBERO assets configured as described below
The repository's conda_env.yml installs CUDA PyTorch. Keep that environment
for CUDA and create the separate XPU environment from conda_env_xpu.yml:
conda env create -f conda_env_xpu.yml
conda activate patch-policy-xpuThe validated XPU environment uses Python 3.12, PyTorch 2.13.0+xpu, Accelerate 1.14.0,
MuJoCo 3.2.7, and robosuite 1.4.1.
Verify the backend before starting a run:
python -c "import torch; print(torch.__version__, torch.xpu.is_available())"Download and unpack the datasets as described in Datasets, then
run the LIBERO Goal configuration with device=xpu:
MUJOCO_GL=egl WANDB_MODE=disabled python -u train_policy.py \
--config-name train_libero_goal_1gpu \
device=xpu \
env_vars.dataset_root=/path/to/patch_policy_datasets \
env_vars.save_path=/path/to/patch_policy_outputs \
epochs=10 eval_on_env_freq=5 num_env_evals=10 num_final_evals=50 num_envs=5 \
+dataset.subset_fraction=1.0The XPU path was validated on the LIBERO Goal task with 10 tasks and one
episode per final evaluation. The run produced loadable model_final.pt
checkpoints and finite actions, with 50% final-evaluation success; an
epoch-10 evaluation reached 100%. The validation run used subset_fraction=1.0 and
5 parallel environments. Simulation remains CPU/EGL, and WANDB_MODE=disabled
can be used when W&B logging is not configured.
The datasets for all four simulation environments are hosted on the Hugging Face Hub: gaoyuezhou/patch-policy-datasets.
| Environment | Folder (after unzip) | Zip size |
|---|---|---|
| Push-T | pusht_dataset |
61 MB |
| Cube | cube_dataset |
2.2 GB |
| LIBERO Goal | libero_dataset |
7.7 GB |
| Block Pushing | block_push_dataset |
5.9 GB |
- Install the Hugging Face Hub CLI:
pip install "huggingface_hub==0.36.2" - Download the dataset repo to a local directory (this is the directory all four datasets will live in):
To download only a subset, add e.g.
huggingface-cli download gaoyuezhou/patch-policy-datasets \ --repo-type dataset --local-dir patch_policy_datasets--include "pusht_dataset.zip". - Unzip each dataset in place:
cd patch_policy_datasets for f in *.zip; do unzip -q "$f"; done
The dataloading code reads from a single dataset_root directory — no code changes are needed, you only set this path.
- In
./configs/env_vars/env_vars.yaml, setdataset_rootto the unzipped directory (e.g. the absolute path topatch_policy_datasets), and setsave_pathto where you want training/rollout results saved (e.g. the root directory of this repo).
The expected layout under dataset_root is:
patch_policy_datasets/
├── pusht_dataset/
├── cube_dataset/
├── libero_dataset/
└── block_push_dataset/
Note: The
.pth/.npy/.pklfiles are loaded withtorch.load/numpy, which unpickle data. Only use datasets you trust.
Policy training and online evaluation both run through train_policy.py, driven by the configs in configs/. A run trains the policy on top of a frozen visual encoder and periodically rolls it out in the simulator.
python train_policy.py --config-name train_pusht # Push-T
python train_policy.py --config-name train_blockpush # Block Pushing
python train_policy.py --config-name train_cube # Cube
MUJOCO_GL=egl python train_policy.py --config-name train_libero_goal # LIBERO Goal
- Diffusion policy variants are available for every environment — append
_diffusionto the config name (e.g.train_pusht_diffusion). The default configs use a VQ-BeT policy head. - Checkpoints are written under
save_path.
The config names above assume a node of 8 GPUs. We also provide _1gpu variants of
every config, tuned to fit within 32 GB of VRAM:
python train_policy.py --config-name train_pusht_1gpu
python train_policy.py --config-name train_blockpush_1gpu
python train_policy.py --config-name train_cube_1gpu
MUJOCO_GL=egl python train_policy.py --config-name train_libero_goal_1gpu
These runs use DINOv2 ViT-S (dino_patch) and precompute the frozen encoder's
features once at startup. Observation windows are shortened where memory requires
it — LIBERO Goal VQ-BeT drops from 10 to 2, and Cube from 5 to 2 — and batch sizes
are set per environment. _diffusion variants are available here too (e.g.
train_pusht_diffusion_1gpu).
Launcher configs for submitting to a SLURM cluster are in configs/cluster/.
The encoder is frozen and selected via the encoder config group (configs/encoder/); off-the-shelf encoders need no training. The default is DINOv2 patch features (dino_patch). Override it on the command line:
python train_policy.py --config-name train_pusht encoder=webssl_patch
python train_policy.py --config-name train_pusht encoder=vjepa2_patch
python train_policy.py --config-name train_pusht encoder=dinov3_patch
python train_policy.py --config-name train_pusht encoder=siglip2_patch
Each of these uses dense patch features. Two pooled variants are also provided for comparison: *_patch_avg_pool (patch tokens mean-pooled into a single vector) is available for every encoder above as well as dino, and *_cls (the CLS token instead of patch tokens) is available for dino, dinov3, and webssl. ResNet-18 baselines (resnet18_imagenet, resnet18_random) and pretrained DynaMo encoders are also supported — see configs/encoder/ for the full list.
Note (DINOv3): the DINOv3 weights are hosted in a gated Hugging Face repo. To use the
dinov3_*encoders, request access at facebook/dinov3-vits16plus-pretrain-lvd1689m and log in withhf auth loginbefore launching.
If you find our work useful, please consider citing:
@misc{zhou2026patchpolicyefficientembodied,
title={Patch Policy: Efficient Embodied Control via Dense Visual Representations},
author={Gaoyue Zhou and Zichen Jeff Cui and Ada Langford and Bowen Tan and Yann LeCun and Lerrel Pinto},
year={2026},
eprint={2607.18236},
archivePrefix={arXiv},
primaryClass={cs.RO},
url={https://arxiv.org/abs/2607.18236},
}