4Implementing Computer Vision Solutions
- 4.1Image and Video Generation and Editing
In AI-103, computer vision centers on generation more than recognition. Understand—from a developer’s view—deploying image- and video-generation models in Microsoft Foundry, generating images and videos from text prompts and reference media, and editing them via mask-based inpainting and prompt-driven changes.
- 4.2Multimodal Understanding and Responsible AI
Understand using multimodal models to comprehend visual context (captions, visual question answering, accessibility alt-text), extracting visual characteristics with Azure Content Understanding, and Responsible AI specific to multimodal content (filters for unsafe visual content, indirect prompt injection via text embedded in images, and visual policy such as watermarks).

