Experiments & Variant Comparison
The standalone Experiments (A/B testing) feature - hypotheses, statistical significance, and automatic winner declaration - is not part of the current release. There is no Experiments page in the application.
The supported way to compare prompt, model, or configuration variants is the Evaluator Runs workflow:
- Curate a dataset (Evals → Datasets) with the test inputs you care about.
- Configure evaluators that score outputs on the criteria that matter to you.
- Start one run per variant - baseline, then candidate (new prompt version or model).
- Select the two runs and click Compare to review per-evaluator mean deltas and verdicts, then make the promotion decision yourself.
Change one variable at a time and use a large enough dataset so deltas are meaningful - treat
within_noise results as "no evidence of a difference".