All learning paths Learn See Try Make Adapt a rubric to a robotics domain What this lesson is about Take a client's rubric — the scoring guide that says helpful, safe, complete — and rewrite it for one robotics domain. Two raters should then decide twenty cases the same way twice.
A client's brief for scoring a warehouse arm says: mark each run as safe and useful. Twenty runs arrive tomorrow. What do you send back?
A Rewrite helpful, safe and complete as four criteria, each anchored to a thing a camera or a clock shows. B Send back safe and useful with a paragraph under each word explaining what the client means by it. C Send the brief back as it is, and sit with the second rater for five runs so you score alike.
Check my recall Why this matters A published robot evaluation framework asks each rater for a progress score, a preference, and a reason. It leaves how to decide the preference to the rater. So raters often give two runs the same progress and still prefer one, for speed or confidence. That gap is where a rubric — the scoring guide — earns its place. Each criterion — one named quality you judge by — needs an anchor. The anchor is the line that says what a pass looks like in this domain.
Adapt a general scoring brief to one robotics domain, so two raters decide twenty cases the same way twice.
A published robot evaluation framework asks each rater for a progress score, a preference, and a reason. It leaves how to decide the preference to the rater. So raters often give two runs the same progress and still prefer one, for speed or confidence. That gap is where a rubric — the scoring guide — earns its place. Each criterion — one named quality you judge by — needs an anchor. The anchor is the line that says what a pass looks like in this domain.
Rewrite helpful, safe and complete as four criteria, each anchored to a thing a camera or a clock shows. Test it on twenty cases with two raters, twice; the splits that remain are the lines still to rewrite. A split is a fault in the rubric, not in a rater. The case goes into the anchor, never into a note about the person. Task A client sends a brief for a floor-cleaning robot in a hospital ward: helpful, safe, complete. Adapt it.
Weak approach Adds the word ward to each line: helpful in the ward, safe in the ward, complete in the ward. Two raters then split on eleven of twenty runs, and the note blames the newer rater.
Stronger approach Writes four anchors. Complete: every marked tile passed over once. Safe: never within half a metre of a bed with a patient in it. Helpful: a cleaned tile is dry inside two minutes. Time: the ward done inside the twenty-minute slot. Two raters split on three of twenty, and each of the three goes into an anchor.
Why the stronger approach works The weak rubric changed the words and left the raters to judge. The strong one moved every judgement into a line both can see.
Try a changed situation Your rubric for a delivery robot went to two raters. They split on nine of twenty runs, and the client asks which rater to keep.
Make something yourself Take a generic helpful, safe, complete brief and adapt it to one named robotics domain. Then run twenty cases past two raters, twice.
Lesson 10 of 10 on Multimodal & physical AI next Read the whole submission first Part 3 ends in a work sample you can send The rules this work sample is read against
Read the primary or official source: Atreya et al., RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies (arXiv)