Skip to content
 
 

Repository files navigation

Patch Policy: Efficient Embodied Control via Dense Visual Representations

[Project Website] [Paper] [Dataset]

Gaoyue Zhou, Zichen Jeff Cui, Ada Langford, Bowen Tan, Yann LeCun and Lerrel Pinto, New York University, Meta AI, AMI Labs

Method overview

This repo contains code for training and reproducing sim environment experiments across four simulation environments: Push-T, Block Pushing, LIBERO Goal, and Cube.

Fork Note: I have extended this repository to load and interact with the environments from a custom simulator. This simulator will not be provided.

Getting Started

  1. Setup
  2. Datasets
  3. Training a policy
  4. Choosing the visual encoder

Setup

Dependencies are managed with uv. It installs Python 3.12 and everything else, including the CUDA build of PyTorch:

uv sync                     # Push-T only
uv sync --all-extras        # all four environments

Extras are per environment: --extra blockpush, --extra libero, --extra cube.

Run commands through uv run (e.g. uv run python train_policy.py ...), or activate the environment with source .venv/bin/activate.

Tested on Ubuntu 22.04 with CUDA 12.8. Training runs log to TensorBoard in each run directory; Weights & Biases logging is off by default. To turn it on, log in with wandb login, set wandb_entity to your wandb username in ./configs/env_vars/env_vars.yaml, and run with WANDB_MODE=online.

Datasets

The datasets for all four simulation environments are hosted on the Hugging Face Hub: gaoyuezhou/patch-policy-datasets.

Environment Folder (after unzip) Zip size
Push-T pusht_dataset 61 MB
Cube cube_dataset 2.2 GB
LIBERO Goal libero_dataset 7.7 GB
Block Pushing block_push_dataset 5.9 GB

Where the data lives

All datasets go to ~/data, the dataset directory of this machine. Keep data out of the repository. This repo uses:

~/data/patch_policy_datasets/

Download

  1. The Hugging Face Hub CLI is part of the environment already (uv sync).
  2. Download the dataset repo to that directory (all four datasets live in it):
    uv run hf download gaoyuezhou/patch-policy-datasets \
      --repo-type dataset --local-dir ~/data/patch_policy_datasets
    
    To download only a subset, add e.g. --include "pusht_dataset.zip".
  3. Unzip each dataset in place:
    cd ~/data/patch_policy_datasets
    for f in *.zip; do unzip -q "$f" && rm "$f"; done
    

Point the code at the data

The dataloading code reads from a single dataset_root directory — no code changes are needed, you only set this path.

  • In ./configs/env_vars/env_vars.yaml, set dataset_root to the unzipped directory (here ~/data/patch_policy_datasets, written as an absolute path), and set save_path to where you want training/rollout results saved (e.g. the root directory of this repo).
  • You can also override it per run: uv run python train_policy.py --config-name train_pusht env_vars.dataset_root=/other/path.

The expected layout under dataset_root is:

patch_policy_datasets/
├── pusht_dataset/
├── cube_dataset/
├── libero_dataset/
└── block_push_dataset/

Note: The .pth/.npy/.pkl files are loaded with torch.load / numpy, which unpickle data. Only use datasets you trust.

Training a policy

Policy training and online evaluation both run through train_policy.py, driven by the configs in configs/. A run trains the policy on top of a frozen visual encoder and periodically rolls it out in the simulator.

  • Diffusion policy variants are available for every environment — append _diffusion to the config name (e.g. train_pusht_diffusion). The default configs use a VQ-BeT policy head.
  • Checkpoints are written under save_path.

Single-GPU configs

The config names above assume a node of 8 GPUs. We also provide _1gpu variants of every config, tuned to fit within 32 GB of VRAM:

Fork Note: The code in this repository, specifically the dataloader logic, was optimized for training and inference on a single 3090 RTX.

Specifically, this repository uses Diffusion-Policy as its main driver, though VQ-BeT does work too.

python train_policy.py --config-name train_pusht_1gpu
python train_policy.py --config-name train_blockpush_1gpu
python train_policy.py --config-name train_cube_1gpu
MUJOCO_GL=egl python train_policy.py --config-name train_libero_goal_1gpu

These runs use DINOv2 ViT-S (dino_patch) and precompute the frozen encoder's features once at startup. Observation windows are shortened where memory requires it — LIBERO Goal VQ-BeT drops from 10 to 2, and Cube from 5 to 2 — and batch sizes are set per environment. _diffusion variants are available here too (e.g. train_pusht_diffusion_1gpu).

Launcher configs for submitting to a SLURM cluster are in configs/cluster/.

Choosing the visual encoder

The encoder is frozen and selected via the encoder config group (configs/encoder/); off-the-shelf encoders need no training. The default is DINOv2 patch features (dino_patch). Override it on the command line:

python train_policy.py --config-name train_pusht encoder=webssl_patch
python train_policy.py --config-name train_pusht encoder=vjepa2_patch
python train_policy.py --config-name train_pusht encoder=dinov3_patch
python train_policy.py --config-name train_pusht encoder=siglip2_patch

Each of these uses dense patch features. Two pooled variants are also provided for comparison: *_patch_avg_pool (patch tokens mean-pooled into a single vector) is available for every encoder above as well as dino, and *_cls (the CLS token instead of patch tokens) is available for dino, dinov3, and webssl. ResNet-18 baselines (resnet18_imagenet, resnet18_random) and pretrained DynaMo encoders are also supported — see configs/encoder/ for the full list.

Note (DINOv3): the DINOv3 weights are hosted in a gated Hugging Face repo. To use the dinov3_* encoders, request access at facebook/dinov3-vits16plus-pretrain-lvd1689m and log in with hf auth login before launching.

Citation

If you find our work useful, please consider citing:

@misc{zhou2026patchpolicyefficientembodied,
      title={Patch Policy: Efficient Embodied Control via Dense Visual Representations}, 
      author={Gaoyue Zhou and Zichen Jeff Cui and Ada Langford and Bowen Tan and Yann LeCun and Lerrel Pinto},
      year={2026},
      eprint={2607.18236},
      archivePrefix={arXiv},
      primaryClass={cs.RO},
      url={https://arxiv.org/abs/2607.18236}, 
}

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages