Evaluate several dimensions
Check accuracy, relevance, coherence, helpfulness, safety, robustness, groundedness, and efficiency. A system can sound helpful while using unsupported evidence. Define what success means for the actual use case.
Match metrics to the task
Exact match suits some structured outputs. Precision and recall suit extracted items. BLEU and ROUGE measure forms of text overlap; METEOR includes additional matching, and BERTScore compares contextual representations. None alone proves factual correctness.
Language-model and human evaluation
Perplexity measures predictive performance on token sequences. Human reviewers can judge usefulness and evidence support. LLM judges are scalable but can show bias; calibrate them against human judgments and do not treat their scores as ground truth.
Evaluate the whole system
For RAG, test both retrieval and the final answer. Freeze a representative test set, include unanswerable and adversarial cases, and compare changes using the same criteria. Record latency and resource use alongside quality.
Worked example
Three structured predictions are compared with reference outputs. A second calculation scores extracted sets; it deliberately distinguishes false positives from missed items.
Download lesson 11 Python example
expected = ["30", "5", "unknown"]
predicted = ["30", "7", "unknown"]
accuracy = sum(a == b for a, b in zip(expected, predicted))/3
truth = {"invoice_id", "total", "currency"}
found = {"invoice_id", "total", "customer_age"}
tp = len(truth & found)
precision, recall = tp/len(found), tp/len(truth)
print("Exact-match accuracy:", round(accuracy, 3))
print("Precision:", round(precision, 3))
print("Recall:", round(recall, 3))Expected output
Exact-match accuracy: 0.667
Precision: 0.667
Recall: 0.667All three scores happen to be 0.667 in this example, but they measure different things. Adding irrelevant extracted fields reduces precision; omitting required fields reduces recall.
Practice and self-check
Common mistake
Averages can hide serious failures. Report the count and nature of critical errors, plus performance on important subgroups or difficult cases.
Student tasks
- Add a fourth correct test case and recalculate exact-match accuracy.
- Remove customer_age from found and calculate precision and recall.
- Write a three-level evidence-support rubric for a document Q&A assistant.
Checkpoint — open after attempting the tasks
With the spurious field removed, precision is 1.0 and recall remains 2/3. A rubric can distinguish fully supported, partly supported, and unsupported answers.
