[Project Website] [Paper] [Dataset]
Gaoyue Zhou, Zichen Jeff Cui, Ada Langford, Bowen Tan, Yann LeCun and Lerrel Pinto, New York University, Meta AI, AMI Labs
This repo contains code for training and reproducing sim environment experiments across four simulation environments: Push-T, Block Pushing, LIBERO Goal, and Cube.
Fork Note: I have extended this repository to load and interact with the environments from a custom simulator. This simulator will not be provided.
Dependencies are managed with uv. It installs Python 3.12 and everything else, including the CUDA build of PyTorch:
uv sync # Push-T only
uv sync --all-extras # all four environments
Extras are per environment: --extra blockpush, --extra libero, --extra cube.
Run commands through uv run (e.g. uv run python train_policy.py ...), or activate the
environment with source .venv/bin/activate.
Tested on Ubuntu 22.04 with CUDA 12.8. Training runs log to TensorBoard in each run directory; Weights & Biases logging is off by default.
To turn it on, log in with wandb login, set wandb_entity to your wandb username in ./configs/env_vars/env_vars.yaml, and run with WANDB_MODE=online.
The datasets for all four simulation environments are hosted on the Hugging Face Hub: gaoyuezhou/patch-policy-datasets.
| Environment | Folder (after unzip) | Zip size |
|---|---|---|
| Push-T | pusht_dataset |
61 MB |
| Cube | cube_dataset |
2.2 GB |
| LIBERO Goal | libero_dataset |
7.7 GB |
| Block Pushing | block_push_dataset |
5.9 GB |
All datasets go to ~/data, the dataset directory of this machine. Keep data out of the
repository. This repo uses:
~/data/patch_policy_datasets/
- The Hugging Face Hub CLI is part of the environment already (
uv sync). - Download the dataset repo to that directory (all four datasets live in it):
To download only a subset, add e.g.
uv run hf download gaoyuezhou/patch-policy-datasets \ --repo-type dataset --local-dir ~/data/patch_policy_datasets--include "pusht_dataset.zip". - Unzip each dataset in place:
cd ~/data/patch_policy_datasets for f in *.zip; do unzip -q "$f" && rm "$f"; done
The dataloading code reads from a single dataset_root directory — no code changes are needed, you only set this path.
- In
./configs/env_vars/env_vars.yaml, setdataset_rootto the unzipped directory (here~/data/patch_policy_datasets, written as an absolute path), and setsave_pathto where you want training/rollout results saved (e.g. the root directory of this repo). - You can also override it per run:
uv run python train_policy.py --config-name train_pusht env_vars.dataset_root=/other/path.
The expected layout under dataset_root is:
patch_policy_datasets/
├── pusht_dataset/
├── cube_dataset/
├── libero_dataset/
└── block_push_dataset/
Note: The
.pth/.npy/.pklfiles are loaded withtorch.load/numpy, which unpickle data. Only use datasets you trust.
Policy training and online evaluation both run through train_policy.py, driven by the configs in configs/. A run trains the policy on top of a frozen visual encoder and periodically rolls it out in the simulator.
- Diffusion policy variants are available for every environment — append
_diffusionto the config name (e.g.train_pusht_diffusion). The default configs use a VQ-BeT policy head. - Checkpoints are written under
save_path.
The config names above assume a node of 8 GPUs. We also provide _1gpu variants of
every config, tuned to fit within 32 GB of VRAM:
Fork Note: The code in this repository, specifically the dataloader logic, was optimized for training and inference on a single 3090 RTX.
Specifically, this repository uses Diffusion-Policy as its main driver, though VQ-BeT does work too.
python train_policy.py --config-name train_pusht_1gpu
python train_policy.py --config-name train_blockpush_1gpu
python train_policy.py --config-name train_cube_1gpu
MUJOCO_GL=egl python train_policy.py --config-name train_libero_goal_1gpu
These runs use DINOv2 ViT-S (dino_patch) and precompute the frozen encoder's
features once at startup. Observation windows are shortened where memory requires
it — LIBERO Goal VQ-BeT drops from 10 to 2, and Cube from 5 to 2 — and batch sizes
are set per environment. _diffusion variants are available here too (e.g.
train_pusht_diffusion_1gpu).
Launcher configs for submitting to a SLURM cluster are in configs/cluster/.
The encoder is frozen and selected via the encoder config group (configs/encoder/); off-the-shelf encoders need no training. The default is DINOv2 patch features (dino_patch). Override it on the command line:
python train_policy.py --config-name train_pusht encoder=webssl_patch
python train_policy.py --config-name train_pusht encoder=vjepa2_patch
python train_policy.py --config-name train_pusht encoder=dinov3_patch
python train_policy.py --config-name train_pusht encoder=siglip2_patch
Each of these uses dense patch features. Two pooled variants are also provided for comparison: *_patch_avg_pool (patch tokens mean-pooled into a single vector) is available for every encoder above as well as dino, and *_cls (the CLS token instead of patch tokens) is available for dino, dinov3, and webssl. ResNet-18 baselines (resnet18_imagenet, resnet18_random) and pretrained DynaMo encoders are also supported — see configs/encoder/ for the full list.
Note (DINOv3): the DINOv3 weights are hosted in a gated Hugging Face repo. To use the
dinov3_*encoders, request access at facebook/dinov3-vits16plus-pretrain-lvd1689m and log in withhf auth loginbefore launching.
If you find our work useful, please consider citing:
@misc{zhou2026patchpolicyefficientembodied,
title={Patch Policy: Efficient Embodied Control via Dense Visual Representations},
author={Gaoyue Zhou and Zichen Jeff Cui and Ada Langford and Bowen Tan and Yann LeCun and Lerrel Pinto},
year={2026},
eprint={2607.18236},
archivePrefix={arXiv},
primaryClass={cs.RO},
url={https://arxiv.org/abs/2607.18236},
}