Google Cloud Associate Data PractitionerStudy guide
The associate certification for data ingestion, analysis, pipelines, and management on Google Cloud (Associate Data Practitioner).
About Google Cloud Associate Data Practitioner (GCP-ADP)
Google Cloud Associate Data Practitioner (GCP-ADP) is a Associate-level certification from Google Cloud. This page organizes the exam scope into a 4-chapter, 11-section study guide and lets you check your understanding with exam-style practice questions. A good flow is to read the chapters below in order, then test yourself via "Practice questions."
Exam domains (approximate weighting)
- Data preparation and ingestion~30%
- Data analysis and presentation~27%
- Data pipeline orchestration~18%
- Data management~25%
Weights are approximate guidance for the live exam. Each domain is covered in detail in the chapters and sections below.
Official exam information: https://cloud.google.com/learn/certification/data-practitioner
1Data preparation and ingestion
- 1.1Data processing methodologies and quality
Understand the difference and selection among ETL, ELT, and ETLT data-manipulation methodologies, assessing data quality, data cleaning with Cloud Data Fusion, BigQuery, and Dataflow, and choosing the right data transfer tool (Storage Transfer Service, Transfer Appliance).
- 1.2Data formats and extract/load
Understand data formats (CSV, JSON, Apache Parquet, Apache Avro, structured tables), choosing extraction tools (Dataflow, BigQuery Data Transfer Service, Database Migration Service, Cloud Data Fusion), and loading into Google Cloud storage with the gcloud / bq CLI, Storage Transfer Service, and client libraries.
- 1.3Choosing storage and data location
Understand choosing among storage destinations (Cloud Storage, BigQuery, Cloud SQL, Firestore, Bigtable, Spanner, AlloyDB), classifying data requirements as structured, semi-structured, or unstructured, and selecting data location types: regional, dual-regional, multi-regional, and zonal.
2Data analysis and presentation
- 2.1Analysis with BigQuery and notebooks
Understand running SQL queries in BigQuery to produce reports and key insights, analyzing and visualizing data with Jupyter notebooks (Colab Enterprise), and conducting data analysis to answer business questions.
- 2.2Dashboards with Looker
Understand creating, modifying, and sharing dashboards to answer business questions, choosing between Looker and Looker Studio, and manipulating basic LookML parameters that define the data model.
- 2.3Using machine learning models
Understand creating ML models with BigQuery ML and AutoML, using pretrained large language models (LLMs) via remote connections in BigQuery, the steps of a standard ML project (data collection, training, evaluation, prediction), and organizing models in the Model Registry.
3Data pipeline orchestration
- 3.1Designing data pipelines
Understand choosing a data transformation tool (Dataproc, Dataflow, Cloud Data Fusion, Cloud Composer, Dataform) by business requirements, evaluating ELT vs ETL use cases, and combining the products needed to implement basic transformation pipelines.
- 3.2Scheduling, automation, and monitoring
Understand automating data processing with scheduled queries and Cloud Scheduler/Cloud Composer, monitoring pipelines via the Dataflow job UI and Cloud Logging/Cloud Monitoring, choosing an orchestration solution (Cloud Composer, Workflows, Dataproc Workflow Templates), and event-driven ingestion from Pub/Sub to BigQuery with Eventarc triggers.
4Data management
- 4.1Access control and governance
Understand enforcing least privilege with IAM, the difference among basic roles, predefined roles, and permissions for data services (BigQuery, Cloud Storage), Cloud Storage access control (public/private, uniform access), and when to share data securely with Analytics Hub.
- 4.2Lifecycle management
Understand choosing Cloud Storage classes by access frequency and retention requirements, configuring rules to auto-delete objects after a period to reduce storage cost (Cloud Storage, BigQuery), and evaluating archiving services by business requirements.
- 4.3High availability, disaster recovery, and security
Understand Google-managed backup/recovery for Cloud Storage and Cloud SQL, when to use replication, primary/secondary location types for redundancy, the choice among CMEK/CSEK/GMEK encryption keys, the role of Cloud KMS, and encryption in transit vs at rest.

