A multimodal model can accept or generate more than text, typically images and sometimes audio or video. Capability flags in the Model Registry mark which models support vision, image generation, or audio, so routing can filter to models that can actually handle the input.
Why it matters
Real applications rarely receive text alone. Multimodal models let one request carry images, and sometimes audio, without a separate pipeline.
In Infere
Capability flags in the Model Registry mark which models support vision, image generation, or audio. The Auto-Router can filter candidates by required capability, so a request with an image is only routed to models that can process it.