Evals definition
LLM evaluation, often called evals, is the practice of systematically testing a large language model or LLM application against defined tasks and quality criteria, such as accuracy, faithfulness to sources, safety, format compliance, latency and cost. Evals combine test datasets, automated scoring and human review to compare models, prompts and releases with evidence.
Why LLM evaluation matters
LLM applications are easy to demo and hard to trust. Outputs vary between runs, a prompt fix for one case can break three others, and a provider's model update can change behavior overnight. Without evals, teams rely on informal spot checks, a few hand-picked questions tried before release, and discover regressions from users. Evals replace that with a repeatable test that runs on every change, the same way unit tests protect ordinary code.
Public benchmarks and leaderboards measure general model ability, which helps with shortlisting but says little about your use case. Application evals use your own tasks, documents and users' real questions, and they are the evidence that decides whether a release ships.
LLM evaluation methods
Most teams combine methods. Cheap code checks run on every case, LLM judges score qualities such as tone and faithfulness, and humans review a sample plus anything the automated checks flag. Each method covers blind spots in the others, and together they produce a score the team can trust.
- Code-based checks: exact match, regular expressions, JSON schema validation and unit tests for generated code.
- Reference comparison: similarity to known good answers, useful for extraction and classification.
- LLM-as-a-judge: a model scores outputs against a written rubric, calibrated against human ratings.
- Pairwise comparison: judges pick the better of two answers from different prompts or models.
- Human review: domain experts grade samples, essential for high-stakes content.
- Online signals: user ratings, task completion, escalation rates and A/B tests in production.
- Red teaming: deliberate attacks that probe for unsafe or policy-breaking behavior.
What to measure
Pick a small set of criteria that match the product. A RAG assistant needs retrieval recall, faithfulness to sources, answer relevance and correct refusals when information is missing. An extraction pipeline needs field-level accuracy and schema validity. An agent needs task success rate, tool selection accuracy and number of steps. Every application should track latency and cost per request alongside quality, since a better answer that takes twice as long may not be better for users.
How to build an eval set
- Collect real inputs from logs, support tickets or domain experts, not invented examples.
- Include edge cases, ambiguous requests and adversarial inputs.
- Write clear pass criteria or reference answers for each case.
- Start with 50 to 200 cases and grow the set from production failures.
- Check LLM judges against human labels before trusting their scores.
- Run the suite automatically in CI on every prompt, model or retrieval change.
LLM evaluation tools
Open-source and commercial options include promptfoo, DeepEval, Ragas, OpenAI Evals, LangSmith, Langfuse, Braintrust and Arize Phoenix, covering test runs, tracing and dashboards. The tool matters less than the dataset and criteria. Nexzem builds the eval set with each client during the first weeks of an LLM project and wires it into the release pipeline, so every change ships with a before and after score.