A multimodal model can accept or generate more than text, typically images and sometimes audio or video. Capability flags in the Model Registry mark which models support vision, image generation, or audio, so routing can filter to models that can actually handle the input.

Why it matters

Real applications rarely receive text alone. Multimodal models let one request carry images, and sometimes audio, without a separate pipeline.

In Infere

Capability flags in the Model Registry mark which models support vision, image generation, or audio. The Auto-Router can filter candidates by required capability, so a request with an image is only routed to models that can process it.

Example

Send an image plus a question to a vision model and get a description, a set of extracted fields, or a classification back in the same OpenAI-compatible response shape.

← Back to the full glossary

Put the platform behind the terms

Route, evaluate, and monitor every AI request from one OpenAI-compatible platform.

Start Free → Explore the Features