Evaluators are automated quality-assessment tools that score AI outputs against defined criteria. Infere ships with a library of ~60 built-in evaluator templates, and you can create your own - as Python code run in a secure sandbox, as LLM-as-judge evaluators built in a guided wizard, or as .eml files versioned in Git.

  • Evaluator types - CODE (sandboxed Python), LLM judge (boolean, range, categorical, comparative), text similarity (BLEU, ROUGE, cosine, fuzzy), heuristics, conversational, agent, pipeline, and group.
  • Template library - answer quality, RAG (faithfulness, groundedness), safety (toxicity, bias, PII leak), security/red-team (prompt-injection, jailbreak), agent & tools, and format checks.
  • Datasets - evaluation datasets with train/dev/test partitions, frozen versions, human annotations, and import from production traffic.
  • Runs & comparison - run evaluators against datasets with multiple trials, then compare two runs side-by-side with mean deltas, pass@k, significance bands, and verdicts (improvement / regression / within noise).
  • Online & gated evaluation - continuously sample live traffic, cluster failures into a taxonomy, and define regression policies that gate prompt deployments on evaluation results.

Everything evaluator-related lives under Observe & Improve → Evaluators in the sidebar: Manage Evaluators, Datasets, Runs, Online, Failures, and Regression.