2Collaborating within and across teams to manage data and models
- 2.1Exploring and preprocessing organization-wide data
Understand exploring organization-wide data (Cloud Storage, BigQuery, Spanner, Cloud SQL, Apache Spark, Apache Hadoop), organizing data types (tabular/text/speech/image/video), managing datasets in Agent Platform, preprocessing (Dataflow, TensorFlow Extended [TFX], BigQuery), creating/consolidating features in Feature Store on Gemini Enterprise Agent Platform, privacy of data usage (PII/PHI), and ingesting data into Agent Platform for inference.
- 2.2Notebooks and experiment tracking
Understand choosing the Jupyter backend on Google Cloud (Gemini Enterprise Agent Platform Workbench, Colab Enterprise, notebooks on Dataproc), security best practices in Gemini Enterprise Agent Platform Workbench, Spark kernels, code-repository integration, developing with frameworks (TensorFlow/PyTorch/sklearn/Spark/JAX), leveraging foundation/open-source models in Model Garden, tracking ML experiments (Experiments on Agent Platform, Kubeflow Pipelines, Vertex AI TensorBoard), and evaluating generative AI solutions.

