4Implementing observability practices and troubleshooting issues
- 4.1Instrumenting telemetry and managing logs
Understand collecting logs/metrics (Ops Agent, OpenTelemetry, Cloud Audit Logs, VPC Flow Logs, Google Cloud Managed Service for Prometheus), log optimization (filter/sampling/exclusions/cost), synthetic monitors, custom/log-based metrics, the Logs Explorer and query language, log export/retention (BigQuery/Pub/Sub/Cloud Storage), and redacting PII/PHI.
- 4.2Metrics, dashboards, alerts, and distributed tracing
Understand metric analysis via the Metrics Explorer, dashboards (PromQL, sharing, playbooks), alerting policies (SLI/SLO, cost) with third-party integration (e.g., PagerDuty), distributed tracing (OpenTelemetry, Cloud Trace, trace-log correlation), and troubleshooting infrastructure/CI-CD/app/performance/latency issues.

