Skip to the lessons

Path

How post-training works

Understand where worked examples, human preferences, reward signals and evaluation fit. No coding is needed.

9 lessons186 minutes

Prepares you for
It gives you the words the other paths use. A criterion — one named quality you judge by. A rubric — the checklist those criteria sit in. Then a project brief stops being a wall of unknown words.
What you learn
You follow how written examples, human preferences, reward signals and evaluation fit together to change a model. You see how a score can go up while the real task quietly gets worse. No coding is needed.
Why this order
It follows the training steps in the order they really run. So each idea explains the next one.
  1. Worked examples and SFT18 minutes
  2. From preferences to a training signal20 minutes
  3. Evaluation closes the loop20 minutes
  4. What an RL environment is18 minutes
  5. Spot reward hacking20 minutes
  6. What a reward model is20 minutes
  7. Direct preference training22 minutes
  8. Where your labels go22 minutes
  9. Audit a preference batch26 minutes

Curriculum reviewed: