6Implementing Text, Speech, Computer Vision, and Image Generation
- 6.1Implementing Text and Speech
Understand the two routes to implementing text and speech in Foundry—task-focused Foundry Tools (Azure AI Language / Azure AI Speech) and multimodal models that take audio directly—plus text analysis, speech-to-text (STT), text-to-speech (TTS), speech translation, and how to respond to spoken prompts.
- 6.2Implementing Computer Vision and Image Generation
Understand how to implement two capabilities in Foundry—"understanding images" (visual input) and "creating images" (image generation). Cover image interpretation/OCR via multimodal models or Azure AI Vision, and text-to-image generation via generative models, with selection guidance.

