All learning paths Learn See Try Make Audit a preference batch What this lesson is about Open a batch of preference marks and count the faults a scorer would learn from it, then write the note.
A batch of nine hundred marked pairs lands on your desk before it travels to a lab. What do you count first?
A How many pairs each person finished, so a slow desk shows up before the lab complains. B Three faults poison a scorer: length winning, the reply shown first winning, and one reviewer marking alone. C Whether the chosen replies read well, since good chosen replies make a good batch.
Check my recall Why this matters A published assistant run measured agreement between a lab's own researchers and its paid reviewers at about sixty-three percent. That is what careful work on hard cases looks like. So a batch that agrees with itself almost perfectly is the one to open first. What you are looking for is a habit the scorer would copy.
Audit a batch of preference marks for the faults a scorer would learn. Write the note that goes with it.
A published assistant run measured agreement between a lab's own researchers and its paid reviewers at about sixty-three percent. That is what careful work on hard cases looks like. So a batch that agrees with itself almost perfectly is the one to open first. What you are looking for is a habit the scorer would copy.
Three faults poison a scorer: length winning, the reply shown first winning, and one reviewer marking alone. Near-perfect agreement is a warning, not a prize, because hard cases split careful people. Report the counts with the size of the sample, and say which fault you could not test for. Task Audit sixty of nine hundred preference pairs from a health helpline before they go for training.
Weak approach Reads the sixty pairs, finds every mark sensible, and writes that the batch looks clean and may go. Counts nothing and names no reviewer.
Stronger approach Counts: the longer reply won forty-eight of sixty, and the reply shown first won forty-one. Reports both with the sample size, and asks for thirty pairs to be marked again with the order flipped.
Why the stronger approach works Both read the same sixty pairs. The strong audit turns that reading into two counts a buyer can check.
Try a changed situation Two reviewers agreed on fifty-nine of sixty hard pairs in your sample. The buyer calls it excellent work.
Make something yourself Audit a sample of one preference batch. Count the length wins and the shown-first wins, then write the note that travels with the batch.
Lesson 9 of 9 on How post-training works next Ground answers in image and video Part 2 ends in a work sample you can send The rules this work sample is read against
Read the primary or official source: Bai et al., Training a Helpful and Harmless Assistant with RLHF (arXiv)