Skip to main content
FREN
Menu

Vision-language model

In French:
Modèle vision-langage
Abbreviation:
VLM

Definition

An AI model trained on large quantities of image-text pairs, which links what it sees to descriptions in natural language. It can recognise objects described to it in words without task-specific training. It is the basis of the vision-language-action models used in robotics.

Checked on against the references below.

Related terms

References

Back to the glossary