Path
Multimodal & physical AI
Practise reviewing images, documents, video and robot actions. Say what you can see, and say what you are unsure of.
- Prepares you for
- It prepares you for image, document, video and robot-action work on multimodal AI projects. The first two quests are annotation — marking up what a file shows, so a model can learn from it. The last quest is the expert end of that work. You judge generated media, then write the robot task and its scoring guide.
- What you learn
- You describe only what an image or a video really shows. You check the writing and the reading order on documents. You break a robot's action into steps you can see. You mark exactly when an event starts and ends. Then you judge a generated image against its prompt, and write the robot task and its scoring guide yourself.
- Why this order
- Still images come before video, and video before robot actions. Each one adds time, and then a real-world result. Judging made images and writing the task come last, once you can say what a file really shows.
- Ground answers in image and video24 minutes
- Review documents and OCR22 minutes
- Review robot-action traces28 minutes
- Mark when an event starts and ends22 minutes
- Line up video with other sensors24 minutes
- Mark where a machine must stop24 minutes
- Read a table without inventing a cell22 minutes
- Judge a generated image against its prompt24 minutes
- Write a robot task somebody else can score26 minutes
- Adapt a rubric to a robotics domain28 minutes
Curriculum reviewed: