Skip to the lessons

Path

Multimodal & physical AI

Practise reviewing images, documents, video and robot actions. Say what you can see, and say what you are unsure of.

10 lessons244 minutes

Prepares you for
It prepares you for image, document, video and robot-action work on multimodal AI projects. The first two quests are annotation — marking up what a file shows, so a model can learn from it. The last quest is the expert end of that work. You judge generated media, then write the robot task and its scoring guide.
What you learn
You describe only what an image or a video really shows. You check the writing and the reading order on documents. You break a robot's action into steps you can see. You mark exactly when an event starts and ends. Then you judge a generated image against its prompt, and write the robot task and its scoring guide yourself.
Why this order
Still images come before video, and video before robot actions. Each one adds time, and then a real-world result. Judging made images and writing the task come last, once you can say what a file really shows.
  1. Ground answers in image and video24 minutes
  2. Review documents and OCR22 minutes
  3. Review robot-action traces28 minutes
  4. Mark when an event starts and ends22 minutes
  5. Line up video with other sensors24 minutes
  6. Mark where a machine must stop24 minutes
  7. Read a table without inventing a cell22 minutes
  8. Judge a generated image against its prompt24 minutes
  9. Write a robot task somebody else can score26 minutes
  10. Adapt a rubric to a robotics domain28 minutes

Curriculum reviewed: