Skip to lesson content

THE ILLUSTRATED LLM TUTORIAL / 15 OF 15

Key Takeaways and Next Steps

Combine the lessons into a small, testable document-assistant project.

Try the example ↓
Key Takeaways and Next Steps: Define: User need + measurable targets; Build: Evidence, answer, validation; Test: Normal, missing, conflicting input; Improve: Review failures and rerun tests
Lesson 15 visual guide · Read the four steps, then explore the explanation below.
  1. 01DefineUser need + measurable targets
  2. 02BuildEvidence, answer, validation
  3. 03TestNormal, missing, conflicting input
  4. 04ImproveReview failures and rerun tests

A complete development workflow

Define the user problem and success criteria; prepare permitted data; choose a model and retrieval method; design prompts; evaluate; deploy with monitoring. Return to earlier steps when failures reveal missing evidence or unclear requirements.

Capstone: a policy assistant

Use five short fictional company policies. Answer with source IDs, return unknown when evidence is missing, and never invent live account data. Separate retrieval, generation, validation, and presentation so failures can be diagnosed.

Acceptance and review

Create answerable, unanswerable, conflicting, and instruction-in-document test cases. Check support and correctness separately from formatting. Track response time and failures. Keep deployment decisions tied to explicit gates and human review where appropriate.

A learning plan

First reproduce the arithmetic and toy examples. Next build retrieval and evaluation. Then connect a real model using its official documentation and verify its behavior. Explore fine-tuning only after you can measure a concrete limitation of the baseline.

Worked example

The release gate below is an illustrative rule for a small prototype. The thresholds are chosen for this exercise, not universal production standards.

Download lesson 15 Python example

Python 3 / standard library
results = {
    "correct": 18, "total": 20,
    "unsupported": 0, "schema_failures": 0,
    "p95_seconds": 2.4
}
accuracy = results["correct"] / results["total"]
release = (accuracy >= 0.90 and
           results["unsupported"] == 0 and
           results["schema_failures"] == 0 and
           results["p95_seconds"] <= 3.0)
print("Accuracy:", round(accuracy, 2))
print("Prototype gate:", "PASS" if release else "FAIL")

Expected output

Accuracy: 0.9
Prototype gate: PASS

The prototype passes these example checks. Twenty cases are not sufficient evidence for broad reliability; expand coverage to match the risks and variability of the intended use.

Practice and self-check

Common mistake

Passing a small test set is the start of evidence, not proof that the system will always work. Preserve failures as regression cases.

Student tasks

  1. Build the five-policy corpus and write at least 20 test questions with expected evidence.
  2. Create a report showing accuracy, unsupported-answer count, and latency; document every failure.
  3. Change unsupported to 1 and explain why the gate fails even if average accuracy remains high.
Checkpoint — open after attempting the tasks

The gate fails because unsupported must equal zero. Submit your corpus, prompts, test set, result table, failure analysis, and one proposed improvement.