The article explains how vision-language-action models (VLAs) use transformer-based AI to control robots. Unlike a language model that maps text to text, a VLA receives camera images, text instructions and information about the robot’s state, then outputs actions such as joint-angle or gripper-position targets. The discussion focuses on Physical Intelligence’s open-weight π0.5 VLA, while noting that it is widely used but not the company’s most advanced model. The article first introduces the linear algebra, neural-network and training concepts needed to understand the system: matrix multiplication, embeddings, nonlinearities, loss functions, backpropagation, gradient descent and minibatches. It then describes how text is tokenized and embedded, how image patches are encoded, and how attention and transformer blocks combine information from different inputs. In π0.5, a SigLIP vision encoder, the PaliGemma vision-language model, the FAST action-token system and a new 300-million-parameter action transformer are trained together for robotic manipulation and image-description tasks. The action transformer starts with roughly 50 random actions and repeatedly refines them over 10 iterations using image, language and robot-state information. Producing action chunks avoids rerunning the entire model for every individual movement; a conventional robot controller then converts the predicted targets into motor torques. The author argues that the substantial overlap between robot AI and chatbot architectures makes rapid future progress plausible, while emphasizing that current robotic capabilities remain limited and that it is uncertain whether VLAs will remain the dominant approach.
AI News
The latest AI releases, research, products, and industry updates.
Loading...