Skip to lesson
  1. Learn
  2. See
  3. Try
  4. Make

Where your labels go

What this lesson is about

Follow a rating batch from your screen to a training run, and find the checks it passes on the way.

You finish two hundred rows on Friday and press send. Your cousin asks whether a model changes that evening.

Why this matters

A published assistant run describes the loop plainly. Batches are pooled, a part is kept aside, the model is trained, and fresh batches arrive on the changed model. Two things decide whether the run can be believed later. A held-out set — items the model never trains on — must stay unread during repairs. And every batch must carry who marked it and when.

Lesson 8 of 9 on How post-training worksnextAudit a preference batchPart 2 ends in a work sample you can send

Read the primary or official source: Bai et al., Training a Helpful and Harmless Assistant with RLHF (arXiv)