A dataset is a curated collection of inputs - and usually expected outputs - used to run evaluations. Good datasets cover the edge cases that actually occur in production. Infere keeps two surfaces: a Data Hub for curating production requests, and frozen dataset versions for reproducible runs.
Why it matters
Evaluation is only as good as its examples. A dataset that misses production edge cases will pass while the real product fails.
In Infere
Infere keeps two surfaces: a Data Hub for curating collections of production requests (sampling, branching, deduplication, export), and frozen dataset versions that make evaluation runs reproducible.