All learning paths Learn See Try Make Check the model that marks What this lesson is about Measure a model grader against your own marks, and test it for the three habits that make its figure worthless.
Your buyer wants two thousand replies marked every week and offers a model to do it. Where do you start?
A Run the model on all of them, then read a handful to see whether its marks look sensible. B Mark a set by hand first, then count how often the marking model agrees with you. C Ask the buyer which model it is, and decide from how that model was built and tested.
Check my recall Why this matters A model can mark far more answers than a circle can. A published study measured a strong model agreeing with people about eighty percent of the time. Two people agree about that often. The same study names the habits to test for. Longer answers win, the answer shown first wins, and a model likes its own writing.
Check a model that marks answers against your own marks before you trust a single figure from it.
A model can mark far more answers than a circle can. A published study measured a strong model agreeing with people about eighty percent of the time. Two people agree about that often. The same study names the habits to test for. Longer answers win, the answer shown first wins, and a model likes its own writing.
Mark a set by hand first, then count how often the marking model agrees with you. Test three habits: the longer answer winning, the first answer winning, and the model liking its own writing. Publish the agreement figure with the count beside it, never the model's score on its own. Task A team wants a model to mark two thousand Hindi answers a week.
Weak approach Runs the model on all two thousand, reports the pass rate, and says the marks looked sensible on a few. Marks none of them by hand.
Stronger approach Marks eighty by hand first. The model agrees on sixty-two, and fourteen of the eighteen splits go to the longer answer. Reports both counts before the rest is run.
Why the stronger approach works Both used the model. Only the strong one can say what a figure from it is worth.
Try a changed situation The model puts the answer shown first on top in seven of ten pairs. Your own marks split evenly.
Make something yourself Mark forty answers by hand, then have a model mark the same forty. Report the agreement and name the habit behind the splits.
Lesson 7 of 8 on Coding & reasoning evaluation next Turn a wrong answer into a test Part 2 ends in a work sample you can send
Read the primary or official source: Zheng et al., Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (arXiv)