Three architecture families
Encoder-only models commonly create bidirectional representations for classification or retrieval.
Decoder-only models generate with causal attention. Encoder-decoder models encode an input and decode an output; T5 belongs to this third family.
Inside a decoder block
Attention mixes information across allowed token positions. A feed-forward network transforms features at each position. Residual connections add sublayer inputs back to outputs. Normalization stabilizes activations; exact placement and type vary across models.
From IDs to predictions
Input IDs select embedding vectors. Positional information makes order available. Repeated transformer blocks produce hidden states. An output projection maps a state to vocabulary logits, and softmax converts logits into probabilities.
Parallelism has limits
Training can process many sequence positions together under a causal mask. Autoregressive generation still depends on previously generated tokens. Dense attention also becomes expensive as sequences grow; transformers are not universally faster for every workload.
Worked example
The code demonstrates a residual addition and a final softmax. It deliberately omits learned attention, normalization, and the MLP; it is not a complete transformer.
Download lesson 04 Python example
from math import exp
hidden = [1.0, 2.0, 3.0]
sublayer = [0.2, -0.3, 0.1]
residual = [x+y for x, y in zip(hidden, sublayer)]
logits = [2.0, 1.0, 0.0]
shifted = [exp(x-max(logits)) for x in logits]
probabilities = [x/sum(shifted) for x in shifted]
print("Residual:", [round(x, 1) for x in residual])
print("Probabilities:", [round(x, 3) for x in probabilities])Expected output
Residual: [1.2, 1.7, 3.1]
Probabilities: [0.665, 0.245, 0.09]The residual path preserves the original hidden state while adding an update. Subtracting the largest logit before exponentiation improves numerical stability without changing softmax probabilities.
Practice and self-check
Common mistake
An architecture sketch is not an exact specification of every model. Pre-norm/post-norm, activation functions, normalization, and positional methods differ.
Student tasks
- Draw the path from token IDs to next-token probabilities.
- Explain why a decoder must hide future training tokens.
- Change all logits by adding 100 and verify that the probabilities stay the same.
Checkpoint — open after attempting the tasks
A causal mask prevents a position from seeing the answer it is being trained to predict. Softmax is invariant to adding the same constant to every logit.
