Skip to lesson content

THE ILLUSTRATED LLM TUTORIAL / 09 OF 15

Inference and Text Generation

Compare decoding strategies and explain the generation loop.

Try the example ↓
Inference and Text Generation: Prefill: Read the prompt once; Scores: Produce next-token logits; Decode: Greedy or filtered sampling; Loop: Append until a stop condition
Lesson 09 visual guide · Read the four steps, then explore the explanation below.
  1. 01PrefillRead the prompt once
  2. 02ScoresProduce next-token logits
  3. 03DecodeGreedy or filtered sampling
  4. 04LoopAppend until a stop condition

Prefill and decoding

Prefill processes the prompt. Decoding extends the answer token by token. A key/value cache stores earlier attention keys and values so they need not be recomputed in full for every step; it still consumes memory.

Choosing the next token

Greedy decoding selects the largest score. Sampling draws from a probability distribution. Top-k keeps a fixed number of candidates. Top-p keeps the smallest high-probability set whose cumulative mass meets the threshold, then renormalizes for sampling.

Temperature and stopping

Temperature rescales logits before softmax: lower positive values sharpen the distribution. It does not guarantee correctness or full reproducibility. Stop at EOS, a supported stop sequence, or a configured generation limit.

Efficiency trade-offs

Batching, efficient kernels, quantization, and speculative decoding can improve serving performance. Benefits depend on hardware, workload, and implementation. Quantization can change quality. Beam search retains multiple candidate sequences but is not universally best for chat.

Worked example

Given a next-token distribution, we build a top-p candidate set at 0.7. The first two candidates total 0.75, so both are retained.

Download lesson 09 Python example

Python 3 / standard library
probabilities = {"blue": 0.50, "gray": 0.25, "clear": 0.15, "dark": 0.10}
p = 0.70
kept, mass = [], 0.0
for token, probability in sorted(probabilities.items(), key=lambda x: -x[1]):
    kept.append((token, probability))
    mass += probability
    if mass >= p:
        break
print("Greedy:", max(probabilities, key=probabilities.get))
print("Top-p:", [(t, round(v/mass, 3)) for t, v in kept])

Expected output

Greedy: blue
Top-p: [('blue', 0.667), ('gray', 0.333)]

The retained probabilities become about 0.667 and 0.333 after renormalization. A subsequent sampling step would choose between them; the code only constructs the distribution.

Practice and self-check

Common mistake

Do not confuse maximum output tokens with total context capacity. A larger output allowance can raise latency without fixing a poor prompt.

Student tasks

  1. Run with top-p=0.9 and identify the retained tokens.
  2. Explain how top-k=2 differs from top-p=0.7 when the distribution changes.
  3. Name two useful latency measurements for a chat application.
Checkpoint — open after attempting the tasks

At 0.9 the first three tokens are retained. Useful measurements include time to first token, generation tokens per second, and total response time.