Skip to lesson

Path: How post-training works

From preferences to a training signal

Understand what pairwise choices capture, and what they can miss when guidelines or reviewers disagree.

Starting level
Some basics
Time
20 minutes
Public lesson guide
EN + हिंदी

You will learn to

Explain what preference choices capture and where disagreement needs adjudication.

Why this matters

A preference is only as useful as the task, rubric, reviewer context, and disagreement process behind it.

Three things to remember

  • Pair the same prompt with comparable responses.
  • Record the decisive criterion, not just A or B.
  • Treat valid disagreement as task-design evidence.

See one example

Task

Two medical reviewers interpret an underspecified prompt differently.

Weak approach

Average their labels silently.

Stronger approach

Record the ambiguity, clarify the prompt or rubric, and adjudicate with relevant expertise.

Why the stronger approach works

Hiding disagreement creates noisy labels; resolving its cause improves the signal.

Make something yourself

Draw a four-step preference-data flow and name one place bias or ambiguity can enter.

A useful starting pointInclude prompt, response pair, reviewer decision, and training use.
Practise this lesson

Read further

The source is optional. The complete teaching and practice are available inside AI2Bharat.

Read the primary or official source: Hugging Face Alignment Course