What's changed: Created Professional Data Engineer Chapter 3 (Domain 3 "Storing": storage choice and DWH = BigQuery/BigLake/AlloyDB/Bigtable/Spanner/Cloud SQL/Cloud Storage/Firestore/Memorystore, normalization/denormalization, partitioning/clustering; data lakes and platforms = data lake (Cloud Storage) discovery/access/cost, Knowledge Catalog, federated governance).
3.2Data lakes and data platforms
Understand managing data lakes (discovery, access, cost controls), building data platforms with Knowledge Catalog (one product for both governance and metadata discovery), federated governance for distributed data systems, and choosing between lakes and warehouses.
Organizational data spans not only the structured data warehouse but also the data lake that holds diverse raw data. A data platform governs across them.
3.2.1Managing the data lake
A data lake (often on Cloud Storage) cheaply stores diverse raw data. Operationally, set up data discovery (what exists), access control (who can use it), and cost controls (lifecycle). Lake data can be processed/analyzed, and lake+warehouse combinations (lakehouse-like) are possible. Map "cheaply store raw data = data lake (Cloud Storage)" and "manage discovery/access/cost."
3.2.2Knowledge Catalog and federated governance
Knowledge Catalog is a data platform that centrally manages quality and governance across distributed data (lakes/warehouses) and also discovers and organizes metadata (search, schema, lineage). Governance and cataloging are one product, not two (renamed from Dataplex Universal Catalog in April 2026; the older Data Catalog is the legacy product you migrate from). Understand where data lives and what it means. As organizations grow, federated governance (domains own responsibility while sharing common rules) works better than central control. Map both "cross-data governance" and "metadata discovery" to Knowledge Catalog.
Common: requirement → means. E.g., "cheaply store diverse raw data" = data lake (Cloud Storage); both "govern quality/governance across distributed data" and "discover/organize metadata" = Knowledge Catalog (one product covers both); "distributed governance at scale" = federated governance.
Watch the mix-ups: (1) Data lake (diverse raw) and warehouse (structured analytics) differ (combinable). (2) Knowledge Catalog is one product covering both cross-data governance and metadata discovery - do not treat them as two products (the older Data Catalog is the legacy one you migrate from). (3) Choose central vs federated governance by scale.
3.2.3Section summary
- Data lakes (Cloud Storage) cheaply store raw data; manage discovery/access/cost
- Knowledge Catalog handles both cross-data governance and metadata discovery
- At scale, use federated governance for distributed control
Sign in to track progress — Log in.
Quick check
(just a quick review)Q1. Which is commonly the foundation of a data lake storing diverse raw data cheaply?
Q2. To centrally manage quality/catalog/governance across distributed data (lakes/warehouses), which is best?
Q3. To discover/organize metadata and understand where data lives and what it means, what do you use?
Q4. At scale, to let domains own responsibility while sharing common rules, which approach best fits?
Q5. Which correctly contrasts data lakes and data warehouses?
Q6. Which best lists what to set up when operating a data lake?
Keep track of your progress
The full study guide is free to read. Sign up free to practice with the question bank, track what you have read, review your mistakes, and highlight passages.

