How to transfer the imitation learning techniques to a more complex robotic arm like Franka Emika Panda.
In previous work, we demonstrated how robotic arms can learn from human demonstrations through Action Chunking Transformer(ACT). In this part, we demonsrate how we shift to Franka emika panda .
The first step is collecting high-quality demonstration data from human teleoperation.
We use the same LeRobot dataset format, which stores episodes as Parquet files and mp4 for front and wrist camera observation.
# Robot Joint States (7-DOF)
# Camera Observations
LeRobot dataset structure with end effector positions, quaternions and camera observations
Human demonstrations are collected via teleoperation using a leader-follower setup, where the operator controls a leader arm and the follower arm mimics the movements.
Leader-follower teleoperation for data collection
After collecting demonstration data, we train imitation learning models to predict robot actions from visual observations.
A pure imitation learning approach that predicts action sequences("chunks") rather than single actions . It uses a transformer encoder-decoder architecture with a CVAE (Conditional Variational Autoencoder) for modeling action distributions.

ACT Architecture (Source: ACT Paper)
The trained model is deployed on the robot for real-time inference and autonomous task execution.
The model runs at ~30Hz, predicting action chunks that are executed by the robot controller in real-time.
Autonomous task execution after 50 episodes training
Real-time action classification using the trained ACT model