All learning paths Learn See Try Make Judge a generated image against its prompt What this lesson is about Judge a generated image or clip in four buckets. What was asked, artefacts, style across the set, and fit for the place.
A printer in Meerut shows you a model-made wedding card image. The mangalsutra is drawn as a plain chain. What do you write?
A Write that the necklace is an artefact, since the model drew it wrong, and ask for that region repainted. B Put each fault in one bucket: asked, artefact, style across the set, or fit for the place. C Write that the card passes the asked-for check, because a bride with a necklace is what the prompt said.
Check my recall Why this matters A published study collected rich feedback on eighteen thousand generated images. Raters marked two things separately: the regions that looked wrong, and the prompt words the image missed. The two are kept apart because they are fixed differently. A wrong region can be painted over. A missed word needs a new image, or a better model. That study did not name two more buckets: style across a set, and fit for the place. The listings ask for both.
Judge a generated image or clip against its prompt, and file each fault in the bucket it belongs to.
A published study collected rich feedback on eighteen thousand generated images. Raters marked two things separately: the regions that looked wrong, and the prompt words the image missed. The two are kept apart because they are fixed differently. A wrong region can be painted over. A missed word needs a new image, or a better model. That study did not name two more buckets: style across a set, and fit for the place. The listings ask for both.
Put each fault in one bucket: asked, artefact, style across the set, or fit for the place. An artefact is something no real camera would show: extra fingers, melted text, a repeated tile, a smeared edge. A cultural mismatch is drawn cleanly and is still wrong for the place. It needs a person from that place, not a repaint. Task Six images for a Diwali sweet-shop poster came back from the model. Review them for the client.
Weak approach Writes that image three has bad hands and image five looks off. Says the set is fine otherwise and sends it.
Stronger approach Files image three's six fingers as an artefact. Files image five's snow on the rooftop as a cultural mismatch, not an artefact. Notes that images one and four use two different lamp styles, which is a style fault across the set.
Why the stronger approach works Both saw the same faults. Only the strong review tells the client which fix each fault needs.
Try a changed situation A generated clip for a tractor advert shows the driver's hand with six fingers for a moment. The shop board's spelling changes mid-clip. What do you file?
Make something yourself Take six generated images for one prompt. File every fault in a bucket, and write the fix each bucket points to.
Lesson 8 of 10 on Multimodal & physical AI next Write a robot task somebody else can score Part 3 ends in a work sample you can send
Read the primary or official source: Liang et al., Rich Human Feedback for Text-to-Image Generation (arXiv)