Use several eval types, from cheap code checks to careful human review:
Unit-style assertions: deterministic code checks on outputs. Examples include "response is valid JSON," "contains the refund policy link," "under 200 tokens," and "no personal data regex matches."
Golden dataset evals: run the system on curated inputs with known reference outputs or labels. Compare via exact match, string similarity, embedding similarity, or a grader.
LLM-as-judge: a strong model grades outputs against criteria, either standalone (pointwise) or against a reference.
Pairwise or preference evals: a judge or human picks the better of two outputs (A vs B). Comparing is easier and more reliable than assigning absolute scores.
Rubric-based evals: scoring against explicit written criteria, one dimension at a time (accuracy, completeness, tone), rather than a single fuzzy "quality" score.
Human review: expert annotators label outputs. The gold standard for correctness, used to bootstrap datasets and calibrate automated judges.
Mature systems layer all of these: assertions everywhere, judges on samples, humans on a small calibrated slice.
Rewriting in plainer words…
This answer doesn't lend itself to a diagram - it reads best . No credits were charged.