For the complete documentation index, see llms.txt. Markdown versions of documentation pages are available by appending .md to the page URL.
Primary navigation

Multimodal

Multimodality refers to a model's ability to understand and generate content using various input types—such as text, images, audio, and video.

VisionImagesSpeech