2Ingesting and processing the data
- 2.1Planning and building pipelines
Understand defining data sources and sinks, transformation/orchestration logic, batch and streaming (windowing, late data) processing, and choosing the right processing services (Dataflow, Apache Beam, Dataproc, Cloud Data Fusion, BigQuery, Pub/Sub, Spark, Kafka).
- 2.2Deploying and operationalizing pipelines
Understand job automation and orchestration (Cloud Composer, Workflows), data cleansing and AI data enrichment, integrating new data sources, and continuous delivery of pipelines via CI/CD.

