An embedding is a numeric vector that represents the meaning of a piece of text, so similar meanings sit close together in vector space. Embeddings power semantic search, retrieval, deduplication, and clustering. Infere exposes embedding models alongside chat models through the same catalog and API.

Why it matters

Embeddings turn meaning into numbers, which is what makes semantic search, retrieval, and deduplication possible at scale.

How it works

Text is passed through an embedding model to produce a vector; similar meanings land close together in vector space. Comparing vectors by cosine distance ranks content by relevance to a query.

Example

"how do I reset my password" and "password reset steps" have nearly identical embeddings even with no shared keywords.

← Back to the full glossary

Put the platform behind the terms

Route, evaluate, and monitor every AI request from one OpenAI-compatible platform.

Start Free → Explore the Features