LearnThatStack Ace your next interview
AI Evals & Observability · question
Question 2 of 55

What are the main types of evals used for LLM applications?

beginner
← All AI Evals & Observability questions
Re-explain

Use several eval types, from cheap code checks to careful human review:

  • Unit-style assertions: deterministic code checks on outputs. Examples include "response is valid JSON," "contains the refund policy link," "under 200 tokens," and "no personal data regex matches."
  • Golden dataset evals: run the system on curated inputs with known reference outputs or labels. Compare via exact match, string similarity, embedding similarity, or a grader.
  • LLM-as-judge: a strong model grades outputs against criteria, either standalone (pointwise) or against a reference.
  • Pairwise or preference evals: a judge or human picks the better of two outputs (A vs B). Comparing is easier and more reliable than assigning absolute scores.
  • Rubric-based evals: scoring against explicit written criteria, one dimension at a time (accuracy, completeness, tone), rather than a single fuzzy "quality" score.
  • Human review: expert annotators label outputs. The gold standard for correctness, used to bootstrap datasets and calibrate automated judges.

Mature systems layer all of these: assertions everywhere, judges on samples, humans on a small calibrated slice.

Rewriting in plainer words…

This answer doesn't lend itself to a diagram - it reads best . No credits were charged.

Why there's no diagram: “”

The interactive diagram is below the answer - jump to diagram ↓ · Below it, the related concept . Jump to it ↓

The diagram below the answer is the concept . Jump to it ↓

Tailored explanation · switch back to · ·
What should the new diagram focus on?
How well did you know this?
AI:

Saved in this browser - sign in to keep your review list.

How should your speech become text?

Listening… your words appear above as you speak - tap Stop when you're done.

Recording · cr - tap Stop & transcribe when you're done.

Transcribing with AI…

Voice:

Keep going - a few more words and AI can grade it.

Interview lens

Likely follow-ups, what you can say, and the weak answers to avoid.

Sign in free to open it Free account - the lens opens as soon as you're back.

Want a quick review of the fundamentals? See the AI Evals & Observability cheatsheet.

← Back to all AI Evals & Observability questions
Pro · $10/mo

48 of 55 AI Evals & Observability answers are in Pro.

Full answers, code samples, and AI explanations that go simpler or deeper. Cancel anytime.

  • Full answers + code
  • AI explanations, simpler or deeper
  • 1,000 AI credits / month
  • Cancel anytime