The standalone Experiments (A/B testing) feature - hypotheses, statistical significance, and automatic winner declaration - is not part of the current release. There is no Experiments page in the application.

The supported way to compare prompt, model, or configuration variants is the Evaluator Runs workflow:

  1. Curate a dataset (Evals → Datasets) with the test inputs you care about.
  2. Configure evaluators that score outputs on the criteria that matter to you.
  3. Start one run per variant - baseline, then candidate (new prompt version or model).
  4. Select the two runs and click Compare to review per-evaluator mean deltas and verdicts, then make the promotion decision yourself.

Change one variable at a time and use a large enough dataset so deltas are meaningful - treat within_noise results as "no evidence of a difference".