Skip to lesson content

THE ILLUSTRATED LLM TUTORIAL / 11 OF 15

Evaluation of LLMs

Build a task-specific evaluation rubric and compute simple outcome metrics.

Try the example ↓
Evaluation of LLMs: Test set: Answerable + unanswerable cases; Run: Same inputs for each version; Score: Correctness, support, latency; Review: Inspect failures before release
Lesson 11 visual guide · Read the four steps, then explore the explanation below.
  1. 01Test setAnswerable + unanswerable cases
  2. 02RunSame inputs for each version
  3. 03ScoreCorrectness, support, latency
  4. 04ReviewInspect failures before release

Evaluate several dimensions

Check accuracy, relevance, coherence, helpfulness, safety, robustness, groundedness, and efficiency. A system can sound helpful while using unsupported evidence. Define what success means for the actual use case.

Match metrics to the task

Exact match suits some structured outputs. Precision and recall suit extracted items. BLEU and ROUGE measure forms of text overlap; METEOR includes additional matching, and BERTScore compares contextual representations. None alone proves factual correctness.

Language-model and human evaluation

Perplexity measures predictive performance on token sequences. Human reviewers can judge usefulness and evidence support. LLM judges are scalable but can show bias; calibrate them against human judgments and do not treat their scores as ground truth.

Evaluate the whole system

For RAG, test both retrieval and the final answer. Freeze a representative test set, include unanswerable and adversarial cases, and compare changes using the same criteria. Record latency and resource use alongside quality.

Worked example

Three structured predictions are compared with reference outputs. A second calculation scores extracted sets; it deliberately distinguishes false positives from missed items.

Download lesson 11 Python example

Python 3 / standard library
expected = ["30", "5", "unknown"]
predicted = ["30", "7", "unknown"]
accuracy = sum(a == b for a, b in zip(expected, predicted))/3
truth = {"invoice_id", "total", "currency"}
found = {"invoice_id", "total", "customer_age"}
tp = len(truth & found)
precision, recall = tp/len(found), tp/len(truth)
print("Exact-match accuracy:", round(accuracy, 3))
print("Precision:", round(precision, 3))
print("Recall:", round(recall, 3))

Expected output

Exact-match accuracy: 0.667
Precision: 0.667
Recall: 0.667

All three scores happen to be 0.667 in this example, but they measure different things. Adding irrelevant extracted fields reduces precision; omitting required fields reduces recall.

Practice and self-check

Common mistake

Averages can hide serious failures. Report the count and nature of critical errors, plus performance on important subgroups or difficult cases.

Student tasks

  1. Add a fourth correct test case and recalculate exact-match accuracy.
  2. Remove customer_age from found and calculate precision and recall.
  3. Write a three-level evidence-support rubric for a document Q&A assistant.
Checkpoint — open after attempting the tasks

With the spurious field removed, precision is 1.0 and recall remains 2/3. A rubric can distinguish fully supported, partly supported, and unsupported answers.