What's changed: Added per-section figures (cert-figure-retrofit). New AI-300 Chapter 4 (Domain 4 "GenAI quality assurance and observability": evaluation = test datasets/data mapping/AI quality metrics (groundedness/relevance/coherence/fluency)/risk-safety evaluation/automated evaluation workflows; observability = Foundry continuous monitoring/performance (latency/throughput/response time)/cost (token consumption/resource usage)/logging-tracing-debugging)
4.1Evaluating and validating generative AI apps and agents
Understand creating test datasets and data mapping, AI quality metrics (groundedness/relevance/coherence/fluency), risk and safety evaluations for harmful content detection, and configuring automated evaluation workflows using built-in/custom metrics.
Because generative AI output is non-deterministic, evaluation is central to operations. AI-300 asks you to evaluate quality and safety with measurable metrics and to automate it.
4.1.1Test datasets and quality metrics
For evaluation, prepare a test dataset (inputs and expected outputs/ground truth) and data mapping (defining column correspondence). Measure generative-AI quality with AI quality metrics: groundedness (is it based on the provided context = the opposite of hallucination), relevance (does it answer the question), coherence (is it logically consistent), and fluency (is it natural). For RAG, groundedness and relevance especially matter.
4.1.2Risk/safety evaluation and automated evaluation
Continue reading — free sign-up
You're reading the free preview. Sign up free to read this section in full, plus every chapter (including 4+) and all questions.

