VLA (vision-language-action model)
- In French:
- VLA (modèle vision-langage-action)
- Abbreviation:
- VLA
Definition
A model that takes images and a natural-language instruction as input and directly produces a robot’s actions. It starts from a vision-language model pretrained on web data, then is trained on robot trajectories, to benefit from the general knowledge acquired online. The term was introduced in 2023 with the RT-2 model.
Checked on against the references below.