A benchmark is a standardized test used to compare models. Benchmarks are useful for a first-pass ranking, but they are public, saturate over time, and rarely match your workload. Treat them as a starting filter, then evaluate candidates on your own dataset.

Why it matters

Benchmarks give a fast, shared starting point for comparing models, but they are public and saturate, so they rarely predict performance on your specific workload.

How it works

A benchmark is a fixed test set with a defined scoring method. Use it to shortlist candidates, then evaluate the shortlist on your own dataset, because a model can lead a public benchmark and still underperform on your domain.

Example

A model topping a general reasoning benchmark may still lose to a smaller, cheaper model on your support-ticket task.

← Back to the full glossary

Put the platform behind the terms

Route, evaluate, and monitor every AI request from one OpenAI-compatible platform.

Start Free → Explore the Features