Visualizing attention maps for VLA models to understand how the model focuses on different parts of the input.
This project demonstrates how robotic arms can learn using Visual Language Action (VLA) models like SmolVLA & GR00T N1.5 and imitation learning like Action Chunking Transformer(ACT) .
This project demonstrates how robotic arms can learn using Visual Language Action (VLA) models and imitation learning.
The system uses ROS2 for robot control and PyTorch for model inference.