Skip to lesson content

THE ILLUSTRATED LLM TUTORIAL / 13 OF 15

Fine-Tuning and Alignment

Distinguish model adaptation methods and explain LoRA with a small matrix example.

Try the example ↓
Fine-Tuning and Alignment: Base: Keep original matrix W; Adapter: Learn low-rank B and A; Update: Delta W = B x A; Evaluate: Task gain + regression checks
Lesson 13 visual guide · Read the four steps, then explore the explanation below.
  1. 01BaseKeep original matrix W
  2. 02AdapterLearn low-rank B and A
  3. 03UpdateDelta W = B x A
  4. 04EvaluateTask gain + regression checks

When to adapt a model

Fine-tuning continues training on selected data to improve behavior or domain performance. Prompting or RAG may be enough for some tasks. Fine-tuning is not a dependable substitute for retrieving current facts or enforcing access control.

Full tuning and PEFT

Full fine-tuning updates all selected model weights. Parameter-efficient fine-tuning updates a smaller parameter set. LoRA learns low-rank updates to selected matrices; adapters insert trainable modules. Soft prompt and prefix tuning optimize trainable representations, not ordinary typed prompts.

Instruction and preference training

Supervised instruction tuning learns from input/target pairs. RLHF uses human preference information in a reinforcement-learning pipeline. DPO uses preference pairs with a different optimization objective.

Constitutional approaches use explicit principles in feedback or training. None guarantees that every output is safe or correct.

QLoRA and evaluation

QLoRA combines a frozen quantized base with trainable low-rank adapters and supporting memory-saving techniques. Validate task gains, general capability retention, robustness, and evidence fidelity on held-out examples; monitor overfitting and regression.

Worked example

For a two-by-two base matrix, a rank-one update is the outer product of B and A. Real LoRA also uses a scaling factor and targets selected model layers.

Download lesson 13 Python example

Python 3 / standard library
W = [[1, 0], [0, 1]]
B = [[1], [2]]
A = [[0.1, 0.2]]
delta = [[B[i][0]*A[0][j] for j in range(2)] for i in range(2)]
adapted = [[round(W[i][j]+delta[i][j], 2) for j in range(2)] for i in range(2)]
print("Adapted:", adapted)
d_in, d_out, rank = 1000, 1000, 8
print("Full parameters:", d_in*d_out)
print("Adapter parameters:", rank*(d_in+d_out))

Expected output

Adapted: [[1.1, 0.2], [0.2, 1.4]]
Full parameters: 1000000
Adapter parameters: 16000

The adapted matrix is [[1.1,0.2],[0.2,1.4]]. A 1,000-by-1,000 matrix has one million parameters; rank-eight A and B together have 16,000. This counts trainable adapter parameters, not total model memory.

Practice and self-check

Common mistake

A small training set of preferred answers can cause memorization or regressions. Keep validation/test cases separate and compare with the original model.

Student tasks

  1. Change the adapter rank in the parameter-count calculation to 16.
  2. Decide whether a frequently changing price list is better handled through retrieval or weight updates, and explain.
  3. Specify three held-out tests for an instruction-tuned support assistant.
Checkpoint — open after attempting the tasks

Rank 16 uses 32,000 adapter parameters. A changing price list is generally better read from an authoritative current data source.