Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 

Repository files navigation

Awesome VLA Quantization

Awesome Status Topic

A curated paper list for quantization, compression, action discretization, and deployment of Vision-Language-Action (VLA) models and robot foundation models.

Aim

This repository aims to build a focused and continuously updated technical map for Vision-Language-Action (VLA) model quantization, compression, and efficient robotic deployment. It is not just a general VLA/WAM paper collection; it organizes papers around the question: how can VLA models be compressed and deployed on real robots under latency, memory, and compute constraints?

The repository serves two goals:

  1. Academic survey leadership: provide a structured literature foundation for VLA quantization surveys, method comparison, taxonomy design, and open-problem analysis.
  2. Industrial R&D reference: support method selection for edge robotics, low-latency inference, low-memory deployment, embedded control, and real-robot evaluation.

The seed list is built from arXiv quantization VLA search results and local paper files, then manually organized into core VLA quantization, action-side quantization, low-bit compression, post-compression recovery, VLA/WAM targets, and general quantization background.

Difference from General VLA/WAM Lists

Compared with general VLA/WAM awesome lists, this repository is more vertical and deployment-oriented:

  • More focused: centered on VLA quantization, compression, and efficient deployment.
  • More technical-roadmap oriented: not only lists papers, but asks what is quantized, why performance drops, and how to recover it.
  • Stronger action-side perspective: action tokenizers, VQ action codebooks, and action discretization are treated as first-class VLA quantization topics.
  • Better suited for survey writing: includes taxonomy, core papers, background papers, benchmarks, and open problems.
  • Better suited for industrial deployment: tracks bit-width, latency, memory, edge robots, task success rate, and real-robot deployment.

Who This Repository Serves

  • Researchers: quickly understand the core papers, technical logic, and open problems of VLA quantization.
  • Survey authors: reuse the taxonomy, grouping logic, and method-comparison dimensions.
  • Robotics engineers: identify compression methods for OpenVLA, Qwen-VLA, RT-2, Octo, π0, and other VLA/WAM backbones.
  • Edge deployment teams: compare 1-bit, sub-4-bit, mixed-bit, quantization-aware pruning, and recovery methods.

Open Problems

  • Which modules should be quantized first: LLM backbone, vision encoder, projector, action head, or action tokenizer?
  • Can LLM/VLM quantization metrics predict robot task-success degradation?
  • How should action-aware calibration be designed to directly preserve action accuracy and task success?
  • How can action-side quantization and model-side quantization be co-designed?
  • Are 1-bit and sub-4-bit VLA models stable on real robots?
  • When is RL recovery, distillation, or adapter-based recovery required after compression?
  • Do quantization methods generalize across LIBERO, CALVIN, SimplerEnv, and real-robot setups?
  • How can hardware-aware VLA quantization be adapted to edge robot chips?

中文版本

Roadmap: From DuQuant to Test-Time Training

The suggested evolution from DuQuant to Test-Time Training is used as a forward-looking roadmap: from static distribution-aware quantization to deployment-time adaptive quantized VLA systems.

DuQuant / LLM-VLM quantization
        ↓
VLA-specific PTQ
        ↓
Action-aware low-bit quantization
        ↓
Post-compression recovery
        ↓
Test-time calibration / test-time training
        ↓
World-model-guided adaptive quantized VLA

This roadmap connects:

  • DuQuant / SmoothQuant / AWQ: make transformer weights and activations easier to quantize.
  • QuantVLA / DA-PTQ / Q-QVLA / QVLA: adapt PTQ to VLA models and action quality.
  • ActQuant / BitVLA / HBVLA / DyQ-VLA: move toward low-bit, 1-bit, action-aware, and temporal-aware quantization.
  • RLRC / SQAP-VLA: recover task success after compression.
  • Test-Time Training / Adaptation: adapt the quantized VLA during deployment under new environments, cameras, objects, embodiments, or action distributions.
  • World-model-guided recovery: use a world model or WAM to evaluate next-state/action feasibility and guide test-time recovery.

See the detailed roadmap: papers/roadmap-duquant-to-ttt.md.

Contents

Overview

VLA quantization is becoming important because modern robot foundation models combine large language models, vision encoders, action heads, action tokenizers, and policy decoders. Efficient deployment requires quantization methods that preserve both semantic reasoning and low-level control accuracy.

This repository focuses on papers related to:

  • post-training quantization for VLA models;
  • low-bit, 1-bit, sub-4-bit, and mixed-bit VLA compression;
  • saliency/action-guided quantization for imitation learning and robot control;
  • vector-quantized action tokenizers and action discretization;
  • recovery, calibration, and robustness after compression;
  • related world-action-model and behavior-cloning background.

Coverage note: the initial seed list is built from the local folder D:\六个月论文\六个月论文\VLA量化综述. Public arXiv/GitHub metadata should be checked before release because network access was blocked during this generation step.

VLA Quantization Logic

VLA quantization is not simply “LLM quantization plus robot policy compression”. The survey should follow the VLA inference pipeline:

Language instruction
        ↓
LLM / multimodal reasoning backbone
        ↓
Vision encoder + projector / cross-modal fusion
        ↓
Action representation: continuous action, action chunk, action token, VQ codebook
        ↓
Action head / policy decoder
        ↓
Robot execution and closed-loop feedback

The core logic is:

  1. What is quantized? LLM backbone, vision encoder, projector, action head, action tokenizer, or the full VLA model.
  2. What failure is addressed? scale mismatch, quantization drift, low-bit accuracy loss, action distribution shift, task-success degradation, or post-compression recovery.

Taxonomy

Category Core question Representative topics
VLA model quantization How to compress a complete VLA model or its key modules? PTQ, scale calibration, drift-aware quantization, robust quantization
Low-bit / extreme quantization How to push VLA models below 4-bit or even to 1-bit? 1-bit VLA, sub-4-bit, mixed-bit allocation
Action-aware quantization How to make quantization sensitive to robot actions? action-guided calibration, saliency-aware imitation learning
Action tokenization How to quantize the action space itself? VQ action tokenizer, multimodal action tokenizer, action discretization
Recovery / robustness How to recover task success after compression? RL recovery, post-quantization adaptation, robust objectives
Background / benchmarks Which VLA models and tasks are used for evaluation? Qwen-VLA, OpenVLA, RT-2, LIBERO, CALVIN, SimplerEnv

Entry Format

Use the following format, matching the requested style:

- **Qwen-VLA**, Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments. arXiv. [arXiv] [Website]

Append [Code] [Notes] when available.

Confirmed arXiv quantization VLA Search Results

I read the local export Arxiv-VLA-Quant合集.txt and excluded the user-marked non-quantization indices:

1, 3, 5, 7, 8, 10, 11, 12, 13, 16, 18, 20, 22, 23, 24, 25, 27, 36, 37

The remaining core quantization/compression papers are:

  • Mix-QVLA, Mix-QVLA: Task-Evidence-Aware Mixed-Precision Quantization of Vision-Language-Action Models. arXiv:2606.19565. [arXiv]
  • Q-QVLA, Q-QVLA: Robust Quantization for Vision-Language-Action Models via Composite Rotation and Perstep Scaling. arXiv:2605.28803. [arXiv]
  • ActQuant, ActQuant: Sub-4-bit Action-Guided Quantization for Vision-Language-Action Models. arXiv:2605.24011. [arXiv]
  • DA-PTQ, DA-PTQ: Drift-Aware Post-Training Quantization for Efficient Vision-Language-Action Models. arXiv:2604.11572. [arXiv]
  • DyQ-VLA, DyQ-VLA: Temporal-Dynamic-Aware Quantization for Embodied Vision-Language-Action Models. arXiv:2603.07904. [arXiv]
  • LiteVLA-Edge, LiteVLA-Edge: Quantized On-Device Multimodal Control for Embedded Robotics. arXiv:2603.03380. [arXiv]
  • QuantVLA, QuantVLA: Scale-Calibrated Post-Training Quantization for Vision-Language-Action Models. arXiv:2602.20309. [arXiv]
  • HBVLA, HBVLA: Pushing 1-Bit Post-Training Quantization for Vision-Language-Action Models. arXiv:2602.13710. [arXiv]
  • QVLA, QVLA: Not All Channels Are Equal in Vision-Language-Action Model's Quantization. arXiv:2602.03782. [arXiv]
  • QDepth-VLA, QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models. arXiv:2510.14836. [arXiv]
  • SQAP-VLA, SQAP-VLA: A Synergistic Quantization-Aware Pruning Framework for High-Performance Vision-Language-Action Models. arXiv:2509.09090. [arXiv]
  • VQ-VLA, VQ-VLA: Improving Vision-Language-Action Models Via Scaling Vector-Quantized Action Tokenizers. arXiv:2507.01016. [arXiv]
  • RLRC, RLRC: Reinforcement Learning-based Recovery for Compressed Vision-Language-Action Models. arXiv:2506.17639. [arXiv]
  • BitVLA, BitVLA: 1-bit Vision-Language-Action Models for Robotics Manipulation. arXiv:2506.07530. [arXiv]
  • EaqVLA, EaqVLA: Encoding-aligned Quantization for Vision-Language-Action Models. arXiv:2505.21567. [arXiv]
  • SAQIL, Saliency-Aware Quantized Imitation Learning for Efficient Robotic Control. arXiv:2505.15304. [arXiv]
  • QAIL, Quantization-Aware Imitation-Learning for Resource-Efficient Robotic Control. arXiv:2412.01034. [arXiv]

HybridVLA and OpenVLA are kept as background / quantization targets rather than core quantization papers.

See the full structured list: papers/paper-list.md.

Highlighted Papers

VLA Model Quantization

  • QuantVLA, QuantVLA: Scale-Calibrated Post-Training Quantization for Vision-Language-Action Models. arXiv. [arXiv] [Code] [Notes]

    • Logic: VLA PTQ; focuses on scale calibration.
    • Local source: QuantVLA Scale-Calibrated Post-Training Quantization for Vision-Language-Action Models.pdf.
  • DA-PTQ, DA-PTQ: Drift-Aware Post-Training Quantization for Efficient Vision-Language-Action Models. arXiv. [arXiv] [Code] [Notes]

    • Logic: VLA PTQ; focuses on quantization drift and action drift.
    • Local source: DA-PTQ Drift-Aware Post-Training Quantization for Efficient Vision-Language-Action Models.pdf.
  • Ω-QVLA, Ω-QVLA: Robust Quantization for Vision-Language-Action Models. arXiv. [arXiv] [Code] [Notes]

    • Logic: robust VLA quantization and deployment stability.
    • Local source: Ω-QVLA Robust Quantization for Vision-Language-Action Models via.pdf.
  • MIx-QVLA, MIx-QVLA: Mixed-bit Quantization for Vision-Language-Action Models. arXiv. [arXiv] [Code] [Notes]

    • Logic: mixed-bit allocation across VLA modules/layers.
    • Local source: MIx-QVLA 17 Jun 2606.19565v1.pdf.
  • BitVLA, BitVLA: 1-bit Vision-Language-Action Models for Robotics Manipulation. arXiv. [arXiv] [Code] [Notes]

    • Logic: extreme 1-bit VLA compression.
    • Local source: BitVLA 1-bit Vision-Language-Action Models for Robotics Manipulation.pdf.
  • ActQuant, ActQuant: Sub-4-bit Action-Guided Quantization for Vision-Language-Action Models. arXiv. [arXiv] [Code] [Notes]

    • Logic: action-guided sub-4-bit VLA quantization.
    • Local source: ActQuant Sub-4-bit Action-Guided Quantization for Vision-Language-Action Models.pdf.
  • RLRC, RLRC: Reinforcement Learning-based Recovery for Compressed Vision-Language-Action Models. arXiv. [arXiv] [Code] [Notes]

    • Logic: RL-based post-compression recovery for VLA models.
    • Local source: RLRC Reinforcement Learning-based Recovery for Compressed Vision-Language-Action Models.pdf.

Action Quantization and Tokenization

  • VQ-VLA, VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers. ICCV 2025. [arXiv] [Website] [Code] [Notes]

    • Logic: quantizes the action space through vector-quantized action tokenizers.
    • Local source: VQ-VLA Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers(ICCV2025).pdf.
  • X-Tokenizer, X-Tokenizer: A Multimodal Action Tokenizer for Vision-Language-Action Pretraining. arXiv. [arXiv] [Website] [Code] [Notes]

    • Logic: multimodal action tokenizer for VLA pretraining.
    • Local source: X-Tokenizer A Multimodal Action Tokenizer for Vision-Language-Action Pretraining.pdf.
  • Behavior Cloning with Action Quantization, Understanding Behavior Cloning with Action Quantization. arXiv. [arXiv] [Code] [Notes]

    • Logic: explains how action quantization affects behavior cloning.
    • Local source: Understanding Behavior Cloning with Action Quantization.pdf.

Robot Policy / Imitation Learning Quantization

  • SAQIL, Saliency-Aware Quantized Imitation Learning for Efficient Robotic Control. arXiv. [arXiv] [Code] [Notes]
    • Logic: saliency-aware quantized imitation learning; adjacent to VLA quantization.
    • Local source: Saliency-Aware Quantized Imitation Learning for Efficient Robotic Control.pdf.

VLA Backbones / Quantization Targets

  • Qwen-VLA, Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments. arXiv. [arXiv] [Website]

  • OpenVLA, OpenVLA: An Open-Source Vision-Language-Action Model. arXiv 2024. [arXiv] [Website] [Code]

  • RT-2, RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control. arXiv 2023. [arXiv] [Website]

  • Octo, Octo: An Open-Source Generalist Robot Policy. arXiv 2024. [arXiv] [Website] [Code]

  • π0, π0: A Vision-Language-Action Flow Model for General Robot Control. arXiv. [arXiv] [Website]

Detailed List

See the 40–50 paper seed list: papers/paper-list.md

Survey and Background

See also: papers/survey.md

  • Survey / Review Notes.
    Local source: 综述.pdf.

  • MotuBrain: An Advanced World Action Model.
    Focus: world action model background related to embodied foundation models.
    Local source: MotuBrain An Advanced World Action Model.pdf.

Chinese Notes

The local folder also contains Chinese notes/translations that can be linked from a private documentation space or summarized in README.zh-CN.md:

  • 7.QuantVLA_CN.pdf
  • 8.QVLA中文.pdf
  • MixQVLA_CN.pdf
  • X-Tokenizer-ZH.pdf
  • RLRC_CN.pdf
  • ACT_Quant_CN.pdf
  • J-bit_zh_CN.pdf
  • MotuBrain_CN.pdf
  • Understanding _CN.pdf

Contributing

New entries are welcome. Please include:

  1. title, authors, year, venue or arXiv ID;
  2. links to paper, code, project page, and notes if available;
  3. method category and quantization target;
  4. bit-width/compression setting;
  5. evaluated VLA model and robot benchmark.

TODO before public release

  • Re-check the arXiv search result for all VLA quantization papers.
  • Add official paper/code/project links.
  • Verify years, venues, authors, arXiv IDs, benchmarks, and bit-widths.
  • Add a compact timeline similar to the reference awesome list.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors