Vision-language model
- In French:
- Modèle vision-langage
- Abbreviation:
- VLM
Definition
An AI model trained on large quantities of image-text pairs, which links what it sees to descriptions in natural language. It can recognise objects described to it in words without task-specific training. It is the basis of the vision-language-action models used in robotics.
Checked on against the references below.