A span is a single unit of work inside a trace - one model call, one retrieval step, or one tool invocation. Spans carry their own timing, cost, and attributes, so a trace becomes a waterfall you can read to find which step is slow or expensive.

Why it matters

Knowing a trace was slow is less useful than knowing which step was slow. Spans localize time and cost inside a workflow.

How it works

Each model call, retrieval step, or tool invocation is a span with its own timing, cost, and attributes. Stacked in order, spans form the waterfall that makes a trace readable.

Example

A trace shows the retrieval span taking 2 seconds while the model span is fast - so you optimize retrieval, not the model.

Related terms

← Back to the full glossary

Put the platform behind the terms

Route, evaluate, and monitor every AI request from one OpenAI-compatible platform.

Start Free → Explore the Features