Path: Multimodal & physical AI
Ground answers in image and video
Describe only what the media supports, identify the relevant region or moment, and flag ambiguity.
- Starting level
- New learner
- Time
- 24 minutes
- Public lesson guide
- EN + हिंदी
You will learn to
Separate visible evidence from inference in image and video tasks.
Why this matters
A confident label can still be wrong when an object, action, or moment is partly hidden.
Three things to remember
- Name the exact region or moment supporting the label.
- Describe occlusion, blur, or missing context.
- Use the project uncertainty option instead of guessing.
See one example
Task
Review whether a hand has picked up a cup.
Weak approach
The cup was picked up.
Stronger approach
The hand approaches the cup; contact is hidden and lift is not visible, so pickup is unconfirmed.
Why the stronger approach works
The stronger review anchors every claim to a visible frame and preserves uncertainty.
Make something yourself
Write a grounded review with observation, uncertain inference, and the next view needed.
A useful starting pointUse three labels: visible, unclear, and not shown.
Practise this lessonRead further
The source is optional. The complete teaching and practice are available inside AI2Bharat.
Read the primary or official source: NVIDIA Isaac Sim