Path: How post-training works
From preferences to a training signal
Understand what pairwise choices capture, and what they can miss when guidelines or reviewers disagree.
- Starting level
- Some basics
- Time
- 20 minutes
- Public lesson guide
- EN + हिंदी
You will learn to
Explain what preference choices capture and where disagreement needs adjudication.
Why this matters
A preference is only as useful as the task, rubric, reviewer context, and disagreement process behind it.
Three things to remember
- Pair the same prompt with comparable responses.
- Record the decisive criterion, not just A or B.
- Treat valid disagreement as task-design evidence.
See one example
Two medical reviewers interpret an underspecified prompt differently.
Average their labels silently.
Record the ambiguity, clarify the prompt or rubric, and adjudicate with relevant expertise.
Hiding disagreement creates noisy labels; resolving its cause improves the signal.
Make something yourself
Draw a four-step preference-data flow and name one place bias or ambiguity can enter.
Read further
The source is optional. The complete teaching and practice are available inside AI2Bharat.
Read the primary or official source: Hugging Face Alignment Course