Skip to main content
FREN
Menu

VLA (vision-language-action model)

In French:
VLA (modèle vision-langage-action)
Abbreviation:
VLA

Definition

A model that takes images and a natural-language instruction as input and directly produces a robot’s actions. It starts from a vision-language model pretrained on web data, then is trained on robot trajectories, to benefit from the general knowledge acquired online. The term was introduced in 2023 with the RT-2 model.

Checked on against the references below.

Related terms

References

Back to the glossary