Path
How post-training works
Understand where worked examples, human preferences, reward signals and evaluation fit. No coding is needed.
- Prepares you for
- It gives you the words the other paths use. A criterion — one named quality you judge by. A rubric — the checklist those criteria sit in. Then a project brief stops being a wall of unknown words.
- What you learn
- You follow how written examples, human preferences, reward signals and evaluation fit together to change a model. You see how a score can go up while the real task quietly gets worse. No coding is needed.
- Why this order
- It follows the training steps in the order they really run. So each idea explains the next one.
- Worked examples and SFT18 minutes
- From preferences to a training signal20 minutes
- Evaluation closes the loop20 minutes
- What an RL environment is18 minutes
- Spot reward hacking20 minutes
- What a reward model is20 minutes
- Direct preference training22 minutes
- Where your labels go22 minutes
- Audit a preference batch26 minutes
Curriculum reviewed: