World action models often feed pretrained video VAE encoder latents directly into downstream action policies, so quantization must preserve the latent representation expected by the frozen policy as well as reconstruction quality. The authors find that direct NVFP4 quantization leaves W4A4 error uncompensated, while joint quantization-aware training can restore reconstruction but may severely damage control: on Wan2.1, VBench-7 reaches 0.7403 versus 0.7409 for FP32, yet LIBERO success falls from 95.5% to 10.5%. Decoder-only experiments attribute the failure to activation quantize-dequantize operations, which change the reconstruction signal and decoder Jacobian, redirect encoder gradients, and cause persistent latent drift. LatentQuant addresses this with two stages: it first aligns the quantized encoder with its high-precision counterpart, then freezes the encoder and adapts the decoder. Across Wan2.1 and Wan2.2, the method preserves near-baseline control and high reconstruction quality, reporting 95.75% LIBERO success and 68.8% RoboTwin success. On NVIDIA B300 GPUs, NVFP4 execution also delivers 1.17x-1.26x end-to-end VAE speedups over BF16 cuDNN.
AI News
The latest AI releases, research, products, and industry updates.
Loading...