What's changed: Added per-section figures (cert-figure-retrofit). New AI-901 Chapter 5 (Domain 2: Content Understanding = multimodal information extraction (docs/forms/images/audio/video → structured), schema/field definition, prebuilt vs custom, positioning vs Document Intelligence/Vision/Speech, build flow = analyzer → schema → analyze → results, combining with RAG/agents)
5.1Overview of Azure Content Understanding
Understand the role of Azure Content Understanding—extracting structured fields/insights from diverse content (documents, forms, images, audio, video)—and the idea of a schema (fields) that defines what to extract.
Azure Content Understanding is a service that extracts needed information in structured form from diverse content—documents, forms, images, audio, and video. For example, "biller, amount, due date" from an invoice PDF; "summary, speaker, topics" from a meeting recording; "a description of what is shown" from an image. Input is messy "unstructured data," output is "structured data" that fits fields/tables—that conversion is its role. It is the central service for AI-901’s information-extraction workload.
5.1.1Supported inputs (multimodal)
- Documents / forms: extract fields (amounts, dates, parties) from invoices, receipts, contracts.
- Images: extract a description, tags, and specified fields of what is shown.
- Audio: extract transcription, summary, speakers, key points from calls/meetings.
- Video: extract scenes, summaries, and elements from footage.
Continue reading — free sign-up
You're reading the free preview. Sign up free to read this section in full, plus every chapter (including 4+) and all questions.

