Skip to lesson
  1. Learn
  2. See
  3. Try
  4. Make

Direct preference training

What this lesson is about

Compare the two routes: train a scorer and use RL, or move the model straight from the pair.

A firm tells your circle it will no longer build a scorer for this work. Nobody in the room knows what that means.

Why this matters

The older route trains a scorer from your pairs, then improves the model against it with RL. Direct preference training, published in 2023, drops the scorer and moves the model straight from the chosen and rejected reply. Your marking does not change. One habit now costs more: a pair marked as a tie gives this route nothing at all.

Lesson 7 of 9 on How post-training worksnextWhere your labels goPart 2 ends in a work sample you can send

Read the primary or official source: Rafailov et al., Direct Preference Optimization (arXiv)