📋 Table of Contents
This page contains all 279 GCP Professional Machine Learning Engineer exam questions with answers.
How to use: Review all questions below. The correct answer is highlighted in green. You can print this page or save as PDF for offline studying.
Question 1
Question 1: You are building an ML model to detect anomalies in real-time
sensor data. You will use Pub/Sub to handle incoming requests. You want to store the results for
analytics and visualization. How should you configure the pipeline?
• A. 1 = Dataflow, 2 = AI Platform, 3 = BigQuery
• B. 1 = DataProc, 2 = AutoML, 3 = Cloud Bigtable
• C. 1 = BigQuery, 2 = AutoML, 3 = Cloud Functions
• D. 1 = BigQuery, 2 = AI Platform, 3 = Cloud Storage
Question 2
Question 2: Your organization wants to make its internal shuttle service
route more efficient. The shuttles currently stop at all pick-up points across the city every 30
minutes between 7 am and 10 am. The development team has already built an application on Google
Kubernetes Engine that requires users to confirm their presence and shuttle station one day in
advance. What approach should you take?
• A. 1. Build a tree-based regression model that predicts how many passengers will be picked up at
each shuttle station. 2. Dispatch an appropriately sized shuttle and provide the map with the
required stops based on the prediction.
• B. 1. Build a tree-based classification model that predicts whether the shuttle should pick up
passengers at each shuttle station. 2. Dispatch an available shuttle and provide the map with the
required stops based on the prediction.
• C. 1. Define the optimal route as the shortest route that passes by all shuttle stations with
confirmed attendance at the given time under capacity constraints. 2. Dispatch an appropriately
sized shuttle and indicate the required stops on the map.
• D. 1. Build a reinforcement learning model with tree-based classification models that predict the
presence of passengers at shuttle stops as agents and a reward function around a distance-based
metric. 2. Dispatch an appropriately sized shuttle and provide the map with the required stops based
on the simulated outcome.
Question 3
Question 3: You were asked to investigate failures of a production line
component based on sensor readings. After receiving the dataset, you discover that less than 1% of
the readings are positive examples representing failure incidents. You have tried to train several
classification models, but none of them converge. How should you resolve the class imbalance
problem?
• A. Use the class distribution to generate 10% positive examples.
• B. Use a convolutional neural network with max pooling and softmax activation.
• C. Downsample the data with upweighting to create a sample with 10% positive examples.
• D. Remove negative examples until the numbers of positive and negative examples are equal.
Question 4
Question 4: You want to rebuild your ML pipeline for structured data on
Google Cloud. You are using PySpark to conduct data transformations at scale, but your pipelines are
taking over 12 hours to run. To speed up development and pipeline run time, you want to use a
serverless tool and SQL syntax. You have already moved your raw data into Cloud Storage. How should
you build the pipeline on Google Cloud while meeting the speed and processing requirements?
Question 5
Question 5: You manage a team of data scientists who use a cloud-based
backend system to submit training jobs. This system has become very difficult to administer, and you
want to use a managed service instead. The data scientists you work with use many different
frameworks, including Keras, PyTorch, theano, Scikit-learn, and custom libraries. What should you
do?
• A. Use the AI Platform custom containers feature to receive training jobs using any framework.
• B. Configure Kubeflow to run on Google Kubernetes Engine and receive training jobs through TF Job.
• C. Create a library of VM images on Compute Engine, and publish these images on a centralized
repository.
• D. Set up Slurm workload manager to receive jobs that can be scheduled to run on your cloud
infrastructure.
Question 6
Question 6: You work for an online retail company that is creating a visual
search engine. You have set up an end-to-end ML pipeline on Google Cloud to classify whether an
image contains your company's product. Expecting the release of new products in the near
future, you configured a retraining functionality in the pipeline so that new data can be fed into
your ML models. You also want to use AI Platform's continuous evaluation service to ensure that
the models have high accuracy on your test dataset. What should you do?
• A. Keep the original test dataset unchanged even if newer products are incorporated into
retraining.
• B. Extend your test dataset with images of the newer products when they are introduced to
retraining.
• C. Replace your test dataset with images of the newer products when they are introduced to
retraining.
• D. Update your test dataset with images of the newer products when your evaluation metrics drop
below a pre-decided threshold.
Question 7
Question 7: You need to build classification workflows over several
structured datasets currently stored in BigQuery. Because you will be performing the classification
several times, you want to complete the following steps without writing code: exploratory data
analysis, feature selection, model building, training, and hyperparameter tuning and serving. What
should you do?
Question 8
You work for a public transportation company and need to build a model to
estimate delay times for multiple transportation routes. Predictions are served directly to users in
an app in real time. Because different seasons and population increases impact the data relevance,
you will retrain the model every month. You want to follow Google-recommended best practices. How
should you configure the end-to-end architecture of the predictive model?
Question 9
You are developing ML models with AI Platform for image segmentation on CT
scans. You frequently update your model architectures based on the newest available research papers,
and have to rerun training on the same dataset to benchmark their performance. You want to minimize
computation costs and manual intervention while having version control for your code. What should
you do?
Question 10
Our team needs to build a model that predicts whether images contain a
driver's license, passport, or credit card. The data engineering team already built the
pipeline and generated a dataset composed of 10,000 images with driver's licenses, 1,000 images
with passports, and 1,000 images with credit cards. You now have to train a model with the following
label map: [`˜drivers_license', `˜passport', `˜credit_card']. Which loss function
should you use?
Question 11
You are designing an ML recommendation model for shoppers on your
company's ecommerce website. You will use Recommendations AI to build, test, and deploy your
system. How should you develop recommendations that increase revenue while following best practices?
Question 12
Question 12: You are designing an architecture with a serverless ML system to
enrich customer support tickets with informative metadata before they are routed to a support agent.
You need a set of models to predict ticket priority, predict ticket resolution time, and perform
sentiment analysis to help agents make strategic decisions when they process support requests.
Tickets are not expected to have any domain-specific terms or jargon. The proposed architecture has
the following flow: Which endpoints should the Enrichment Cloud Functions call? • A. 1 = AI
Platform, 2 = AI Platform, 3 = AutoML Vision • B. 1 = AI Platform, 2 = AI Platform, 3 = AutoML
Natural Language • C. 1 = AI Platform, 2 = AI Platform, 3 = Cloud Natural Language API • D. 1 =
Cloud Natural Language API, 2 = AI Platform, 3 = Cloud Vision API
Figure: Enrichment architecture
Question 13
You have trained a deep neural network model on Google Cloud. The model has
low loss on the training data, but is performing worse on the validation data. You want the model to
be resilient to overfitting. Which strategy should you use when retraining the model?
Question 14
Question 14: You built and manage a production system that is responsible for
predicting sales numbers. Model accuracy is crucial, because the production model is required to
keep up with market changes. Since being deployed to production, the model hasn't changed;
however the accuracy of the model has steadily deteriorated. What issue is most likely causing the
steady decline in model accuracy?
Question 15
You have been asked to develop an input pipeline for an ML training model
that processes images from disparate sources at a low latency. You discover that your input data
does not fit in memory. How should you create a dataset following Google-recommended best practices?
Question 16
You are an ML engineer at a large grocery retailer with stores in multiple
regions. You have been asked to create an inventory prediction model. Your model's features
include region, location, historical demand, and seasonal popularity. You want the algorithm to
learn from new inventory data on a daily basis. Which algorithms should you use to build the model?
Question 17
You are building a real-time prediction engine that streams files which may
contain Personally Identifiable Information (PII) to Google Cloud. You want to use the Cloud Data
Loss Prevention (DLP) API to scan the files. How should you ensure that the PII is not accessible by
unauthorized individuals?
Question 18
You work for a large hotel chain and have been asked to assist the marketing
team in gathering predictions for a targeted marketing strategy. You need to make predictions about
user lifetime value (LTV) over the next 20 days so that marketing can be adjusted accordingly. The
customer dataset is in BigQuery, and you are preparing the tabular data for training with AutoML
Tables. This data has a time signal that is spread across multiple columns. How should you ensure
that AutoML fits the best model to your data?
Question 19
You have written unit tests for a Kubeflow Pipeline that require custom
libraries. You want to automate the execution of unit tests with each new push to your development
branch in Cloud Source Repositories. What should you do?
Question 20
You are training an LSTM-based model on AI Platform to summarize text using
the following job submission script: You want to ensure that training time is minimized without
significantly compromising the accuracy of your model. What should you do?
gcloud ai-platform jobs submit training $JOB_NAME \
--package-path $TRAINER_PACKAGE_PATH \
--module-name $MAIN_TRAINER_MODULE \
--job-dir $JOB_DIR \
--region $REGION \
--scale-tier basic \
-- \
--epochs 20 \
--batch_size=32 \
--learning_rate=0.001
Question 21
Question 21: You have deployed multiple versions of an image classification
model on AI Platform. You want to monitor the performance of the model versions over time. How
should you perform this comparison?
• A. Compare the loss performance for each model on a held-out dataset.
• B. Compare the loss performance for each model on the validation data.
• C. Compare the receiver operating characteristic (ROC) curve for each model using the What-If
Tool.
• D. Compare the mean average precision across the models using the Continuous Evaluation feature.
Question 22
Question 22: You trained a text classification model. You have the following
SignatureDefs:
You started a TensorFlow-serving component server and tried to send an HTTP request to get a
prediction using: headers = {"content-type": "application/json"} json_response =
requests.post('http: //localhost:8501/v1/models/text_model:predict', data=data,
headers=headers)
What is the correct way to write the predict request?
• A. data = json.dumps({ג€signature_name ג :€ג€seving_defaultג ,€ג€instances ג€ [['ab',
'bc',
'cd']]})
• B. data = json.dumps({ג€signature_name ג :€ג€serving_defaultג ,€ג€instances ג€ [['a',
'b', 'c',
'd', 'e', 'f']]})
• C. data = json.dumps({ ג€signature_nameג :€ג€serving_default ג ,€ג€instances ג€ [['a',
'b', 'c'],
['d', 'e', 'f']]})
• D. data = json.dumps({ג€signature_name ג :€ג€serving_defaultג ,€ג€instances ג€ [['a',
'b'], ['c',
'd'], ['e', 'f']])
Question 23
Your organization's call center has asked you to develop a model that
analyzes customer sentiments in each call. The call center receives over one million calls daily,
and data is stored in Cloud Storage. The data collected must not leave the region in which the call
originated, and no Personally Identifiable Information (PII) can be stored or analyzed. The data
science team has a third-party tool for visualization and access which requires a SQL ANSI-2011
compliant interface. You need to select components for data processing and for analytics. How should
the data pipeline be designed?
Question 24
You are an ML engineer at a global shoe store. You manage the ML models for
the company's website. You are asked to build a model that will recommend new products to the
user based on their purchase behavior and similarity with other users. What should you do?
Question 25
Question 25: You work for a social media company. You need to detect whether
posted images contain cars. Each training example is a member of exactly one class. You have trained
an object detection neural network and deployed the model version to AI Platform Prediction for
evaluation. Before deployment, you created an evaluation job and attached it to the AI Platform
Prediction model version. You notice that the precision is lower than your business requirements
allow. How should you adjust the model's final layer softmax threshold to increase precision?
• A. Increase the recall.
• B. Decrease the recall.
• C. Increase the number of false positives.
• D. Decrease the number of false negatives.
Question 26
Question 26: You are responsible for building a unified analytics environment
across a variety of on-premises data marts. Your company is experiencing data quality and security
challenges when integrating data across the servers, caused by the use of a wide range of
disconnected tools and temporary solutions. You need a fully managed, cloud-native data integration
service that will lower the total cost of work and reduce repetitive work. Some members on your team
prefer a codeless interface for building Extract, Transform, Load (ETL) process. Which service
should you use?
• A. Dataflow
• B. Dataprep
• C. Apache Flink
• D. Cloud Data Fusion
Question 27
Question 27: You are an ML engineer at a regulated insurance company. You are
asked to develop an insurance approval model that accepts or rejects insurance applications from
potential customers. What factors should you consider before building the model?
• A. Redaction, reproducibility, and explainability
• B. Traceability, reproducibility, and explainability
• C. Federated learning, reproducibility, and explainability
• D. Differential privacy, federated learning, and explainability
Question 28
Question 28: You are training a Resnet model on AI Platform using TPUs to
visually categorize types of defects in automobile engines. You capture the training profile using
the Cloud TPU profiler plugin and observe that it is highly input-bound. You want to reduce the
bottleneck and speed up your model training process. Which modifications should you make to the
tf.data dataset? (Choose two.)
Question 29
You have trained a model on a dataset that required computationally expensive
preprocessing operations. You need to execute the same preprocessing at prediction time. You
deployed the model on AI Platform for high-throughput online prediction. Which architecture should
you use?
Question 30
Question 30: Your team trained and tested a DNN regression model with good
results. Six months after deployment, the model is performing poorly due to a change in the
distribution of the input data. How should you address the input differences in production?
Question 31
You need to train a computer vision model that predicts the type of
government ID present in a given image using a GPU-powered virtual machine on Compute Engine. You
use the following parameters:
✑ Optimizer: SGD
✑ Image shape = 224224—ֳ
✑ Batch size = 64
✑ Epochs = 10
✑ Verbose =2
During training you encounter the following error: ResourceExhaustedError: Out Of Memory (OOM) when
allocating tensor. What should you do?
Question 32
You developed an ML model with AI Platform, and you want to move it to
production. You serve a few thousand queries per second and are experiencing latency issues.
Incoming requests are served by a load balancer that distributes them across multiple Kubeflow
CPU-only pods running on Google Kubernetes Engine (GKE). Your goal is to improve the serving latency
without changing the underlying infrastructure. What should you do?
Question 33
You have a demand forecasting pipeline in production that uses Dataflow to
preprocess raw data prior to model training and prediction. During preprocessing, you employ Z-score
normalization on data stored in BigQuery and write it back to BigQuery. New training data is added
every week. You want to make the process more efficient by minimizing computation time and manual
intervention. What should you do?
Question 34
Question 34: You need to design a customized deep neural network in Keras
that will predict customer purchases based on their purchase history. You want to explore model
performance using multiple model architectures, store training data, and be able to compare the
evaluation metrics in the same dashboard. What should you do?
• A. Create multiple models using AutoML Tables.
• B. Automate multiple training runs using Cloud Composer.
• C. Run multiple training jobs on AI Platform with similar job names.
• D. Create an experiment in Kubeflow Pipelines to organize multiple runs.
Question 35
You are developing a Kubeflow pipeline on Google Kubernetes Engine. The first
step in the pipeline is to issue a query against BigQuery. You plan to use the results of that query
as the input to the next step in your pipeline. You want to achieve this in the easiest way
possible. What should you do?
Question 36
You are building a model to predict daily temperatures. You split the data
randomly and then transformed the training and test datasets. Temperature data for model training is
uploaded hourly. During testing, your model performed with 97% accuracy; however, after deploying to
production, the model's accuracy dropped to 66%. How can you make your production model more
accurate?
Question 37
You are developing models to classify customer support emails. You created
models with TensorFlow Estimators using small datasets on your on-premises system, but you now need
to train the models using large datasets to ensure high performance. You will port your models to
Google Cloud and want to minimize code refactoring and infrastructure overhead for easier migration
from on-prem to cloud. What should you do?
Question 38
You have trained a text classification model in TensorFlow using AI Platform.
You want to use the trained model for batch predictions on text data stored in BigQuery while
minimizing computational overhead. What should you do?
Question 39
You work with a data engineering team that has developed a pipeline to clean
your dataset and save it in a Cloud Storage bucket. You have created an ML model and want to use the
data to refresh your model as soon as new data is available. As part of your CI/CD workflow, you
want to automatically run a Kubeflow Pipelines training job on Google Kubernetes Engine (GKE). How
should you architect this workflow?
Question 40
Question 40: You have a functioning end-to-end ML pipeline that involves
tuning the hyperparameters of your ML model using AI Platform, and then using the best-tuned
parameters for training. Hypertuning is taking longer than expected and is delaying the downstream
processes. You want to speed up the tuning job without significantly compromising its effectiveness.
Which actions should you take? (Choose two). • A. Decrease the number of parallel trials. • B.
Decrease the range of floating-point values. • C. Set the early stopping parameter to TRUE. • D.
Change the search algorithm from Bayesian search to random search. • E. Decrease the maximum number
of trials during subsequent training phases.
Question 41
Your team is building an application for a global bank that will be used by
millions of customers. You built a forecasting model that predicts customers' account balances
3 days in the future. Your team will use the results in a new feature that will notify users when
their account balance is likely to drop below $25. How should you serve your predictions?
• A. 1. Create a Pub/Sub topic for each user. 2. Deploy a Cloud Function that sends a notification
when your model predicts that a user's account balance will drop below the $25 threshold.
• B. 1. Create a Pub/Sub topic for each user. 2. Deploy an application on the App Engine standard
environment that sends a notification when your model predicts that a user's account balance
will drop below the $25 threshold.
• C. 1. Build a notification system on Firebase. 2. Register each user with a user ID on the
Firebase Cloud Messaging server, which sends a notification when the average of all account balance
predictions drops below the $25 threshold.
• D. 1. Build a notification system on Firebase. 2. Register each user with a user ID on the
Firebase Cloud Messaging server, which sends a notification when your model predicts that a
user's account balance will drop below the $25 threshold.
Question 42
You are an ML engineer at a global car manufacture. You need to build an ML
model to predict car sales in different cities around the world. Which features or feature crosses
should you use to train city-specific relationships between car type and number of sales?
Question 43
You work for a large technology company that wants to modernize their contact
center. You have been asked to develop a solution to classify incoming calls by product so that
requests can be more quickly routed to the correct support team. You have already transcribed the
calls using the Speech-to-Text API. You want to minimize data preprocessing and development time.
How should you build the model?
Question 44
Question 45: You are training a TensorFlow model on a structured dataset with
100 billion records stored in several CSV files. You need to improve the input/output execution
performance. What should you do?
Question 45
As the lead ML Engineer for your company, you are responsible for building ML
models to digitize scanned customer forms. You have developed a TensorFlow model that converts the
scanned images into text and stores them in Cloud Storage. You need to use your ML model on the
aggregated data collected at the end of each day with minimal manual intervention. What should you
do?
Question 46
Question 47: You recently joined an enterprise-scale company that has
thousands of datasets. You know that there are accurate descriptions for each table in BigQuery, and
you are searching for the proper BigQuery table to use for a model you are building on AI Platform.
How should you find the data that you need?
Question 47
You started working on a classification problem with time series data and
achieved an area under the receiver operating characteristic curve (AUC ROC) value of 99% for
training data after just a few experiments. You haven't explored using any sophisticated
algorithms or spent any time on hyperparameter tuning. What should your next step be to identify and
fix the problem?
Question 48
Question 49: You work for an online travel agency that also sells advertising
placements on its website to other companies. You have been asked to predict the most relevant web
banner that a user should see next. Security is important to your company. The model latency
requirements are 300ms@p99, the inventory is thousands of web banners, and your exploratory analysis
has shown that navigation context is a good predictor. You want to Implement the simplest solution.
How should you configure the prediction pipeline?
Question 49
Question 50: Your team is building a convolutional neural network (CNN)-based
architecture from scratch. The preliminary experiments running on your on-premises CPU-only
infrastructure were encouraging, but have slow convergence. You have been asked to speed up model
training to reduce time-to-market. You want to experiment with virtual machines (VMs) on Google
Cloud to leverage more powerful hardware. Your code does not include any manual device placement and
has not been wrapped in Estimator model-level abstraction. Which environment should you train your
model on?
Question 50
Question 51: You work on a growing team of more than 50 data scientists who
all use AI Platform. You are designing a strategy to organize your jobs, models, and versions in a
clean and scalable way. Which strategy should you choose?
• A. Set up restrictive IAM permissions on the AI Platform notebooks so that only a single user or
group can access a given instance.
• B. Separate each data scientist's work into a different project to ensure that the jobs,
models, and versions created by each data scientist are accessible only to that user.
• C. Use labels to organize resources into descriptive categories. Apply a label to each created
resource so that users can filter the results by label when viewing or monitoring the resources.
• D. Set up a BigQuery sink for Cloud Logging logs that is appropriately filtered to capture
information about AI Platform resource usage. In BigQuery, create a SQL view that maps users to the
resources they are using
Question 51
You are training a deep learning model for semantic image segmentation with
reduced training time. While using a Deep Learning VM Image, you receive the following error: The
resource
'projects/deeplearning-platform/zones/europe-west4-c/acceleratorTypes/nvidia-tesla-k80'
was not found. What should you do?
Question 52
Your team is working on an NLP research project to predict political
affiliation of authors based on articles they have written. You have a large training dataset that
is structured like this: You followed the standard 80%-10%-10% data distribution across the
training, testing, and evaluation subsets. How should you distribute the training examples across
the train-test-eval subsets while maintaining the 80-10-10 proportion?
• A. Distribute texts randomly across the train-test-eval subsets: Train set: [TextA1, TextB2, ...]
Test set: [TextA2, TextC1, TextD2, ...] Eval set: [TextB1, TextC2, TextD1, ...]
• B. Distribute authors randomly across the train-test-eval subsets: Train set: [TextA1, TextA2,
TextD1, TextD2, ...] Test set: [TextB1, TextB2, ...] Eval set: [TextC1, TextC2, ...]
• C. Distribute sentences randomly across the train-test-eval subsets: Train set: [SentenceA11,
SentenceA21, SentenceB11, SentenceB21, SentenceC11, SentenceD21 ...] Test set: [SentenceA12,
SentenceA22, SentenceB12, SentenceC22, SentenceC12, SentenceD22 ...] Eval set: [SentenceA13,
SentenceA23, SentenceB13, SentenceC23, SentenceC13, SentenceD31 ...]
• D. Distribute paragraphs of texts (i.e., chunks of consecutive sentences) across the
train-test-eval subsets: Train set: [SentenceA11, SentenceA12, SentenceD11, SentenceD12 ...] Test
set: [SentenceA13, SentenceB13, SentenceB21, SentenceD23, SentenceC12, SentenceD13 ...] Eval set:
[SentenceA11, SentenceA22, SentenceB13, SentenceD22, SentenceC23, SentenceD11 ...]
Question 53
Question 54: Your team has been tasked with creating an ML solution in Google
Cloud to classify support requests for one of your platforms. You analyzed the requirements and
decided to use TensorFlow to build the classifier so that you have full control of the model's
code, serving, and deployment. You will use Kubeflow pipelines for the ML platform. To save time,
you want to build on existing resources and use managed services instead of building a completely
new model. How should you build the classifier?
• A. Use the Natural Language API to classify support requests.
• B. Use AutoML Natural Language to build the support requests classifier.
• C. Use an established text classification model on AI Platform to perform transfer learning.
• D. Use an established text classification model on AI Platform as-is to classify support requests.
Question 54
Question 55: You recently joined a machine learning team that will soon
release a new project. As a lead on the project, you are asked to determine the production readiness
of the ML components. The team has already tested features and data, model development, and
infrastructure. Which additional readiness check should you recommend to the team?
Question 55
You work for a credit card company and have been asked to create a custom
fraud detection model based on historical data using AutoML Tables. You need to prioritize detection
of fraudulent transactions while minimizing false positives. Which optimization objective should you
use when training the model?
Question 56
Your company manages a video sharing website where users can watch and upload
videos. You need to create an ML model to predict which newly uploaded videos will be the most
popular so that those videos can be prioritized on your company's website. Which result should
you use to determine whether the model is successful?
Question 57
You are working on a Neural Network-based project. The dataset provided to
you has columns with different ranges. While preparing the data for model training, you discover
that gradient optimization is having difficulty moving weights to a good solution. What should you
do?
Question 58
Your data science team needs to rapidly experiment with various features,
model architectures, and hyperparameters. They need to track the accuracy metrics for various
experiments and use an API to query the metrics over time. What should they use to track and report
their experiments while minimizing manual effort?
Question 59
You work for a bank and are building a random forest model for fraud
detection. You have a dataset that includes transactions, of which 1% are identified as fraudulent.
Which data transformation strategy would likely improve the performance of your classifier?
Question 60
You are using transfer learning to train an image classifier based on a
pre-trained EfficientNet model. Your training dataset has 20,000 images. You plan to retrain the
model once per day. You need to minimize the cost of infrastructure. What platform components and
configuration environment should you use?
Question 61
While conducting an exploratory analysis of a dataset, you discover that
categorical feature A has substantial predictive power, but it is sometimes missing. What should you
do?
Question 62
You work for a large retailer and have been asked to segment your customers
by their purchasing habits. The purchase history of all customers has been uploaded to BigQuery. You
suspect that there may be several distinct customer segments, however you are unsure of how many,
and you don’t yet understand the commonalities in their behavior. You want to find the most
efficient solution. What should you do?
Question 63
Question 64: You recently designed and built a custom neural network that
uses critical dependencies specific to your organization’s framework. You need to train the model
using a managed training service on Google Cloud. However, the ML framework and related dependencies
are not supported by AI Platform Training. Also, both your model and your data are too large to fit
in memory on a single machine. Your ML framework of choice uses the scheduler, workers, and servers
distribution structure. What should you do?
Question 64
Question 65: While monitoring your model training’s GPU utilization, you
discover that you have a native synchronous implementation. The training data is split into multiple
files. You want to reduce the execution time of your input pipeline. What should you do? • A.
Increase the CPU load • B. Add caching to the pipeline • C. Increase the network bandwidth • D. Add
parallel interleave to the pipeline
Question 65
Question 66: Your data science team is training a PyTorch model for image
classification based on a pre-trained ResNet model. You need to perform hyperparameter tuning to
optimize for several parameters. What should you do?
Question 66
Question 67: You have a large corpus of written support cases that can be
classified into 3 separate categories: Technical Support, Billing Support, or Other Issues. You need
to quickly build, test, and deploy a service that will automatically classify future written
requests into one of the categories. How should you configure the pipeline?
Question 67
You need to quickly build and train a model to predict the sentiment of
customer reviews with custom categories without writing code. You do not have enough data to train a
model from scratch. The resulting model should have high predictive performance. Which service
should you use?
• A. AutoML Natural Language
• B. Cloud Natural Language API
• C. AI Hub pre-made Jupyter Notebooks
• D. AI Platform Training built-in algorithms
Question 68
Question 69: You need to build an ML model for a social media application to
predict whether a user’s submitted profile photo meets the requirements. The application will inform
the user if the picture meets the requirements. How should you build a model to ensure that the
application does not falsely accept a non-compliant picture? • A. Use AutoML to optimize the model’s
recall in order to minimize false negatives. • B. Use AutoML to optimize the model’s F1 score in
order to balance the accuracy of false positives and false negatives. • C. Use Vertex AI Workbench
user-managed notebooks to build a custom model that has three times as many examples of pictures
that meet the profile photo requirements. • D. Use Vertex AI Workbench user-managed notebooks to
build a custom model that has three times as many examples of pictures that do not meet the profile
photo requirements.
Question 69
You lead a data science team at a large international corporation. Most of
the models your team trains are large-scale models using high-level TensorFlow APIs on AI Platform
with GPUs. Your team usually takes a few weeks or months to iterate on a new version of a model. You
were recently asked to review your team’s spending. How should you reduce your Google Cloud compute
costs without impacting the model’s performance?
Question 70
You need to train a regression model based on a dataset containing 50,000
records that is stored in BigQuery. The data includes a total of 20 categorical and numerical
features with a target variable that can include negative values. You need to minimize effort and
training time while maximizing model performance. What approach should you take to train this
regression model?
Question 71
You are building a linear model with over 100 input features, all with values
between –1 and 1. You suspect that many features are non-informative. You want to remove the
non-informative features from your model while keeping the informative ones in their original form.
Which technique should you use?
Question 72
You work for a global footwear retailer and need to predict when an item will
be out of stock based on historical inventory data. Customer behavior is highly dynamic since
footwear demand is influenced by many different factors. You want to serve models that are trained
on all available data, but track your performance on specific subsets of data before pushing to
production. What is the most streamlined and reliable way to perform this validation?
Question 73
You have deployed a model on Vertex AI for real-time inference. During an
online prediction request, you get an “Out of Memory” error. What should you do?
Question 74
Question 75: You work at a subscription-based company. You have trained an
ensemble of trees and neural networks to predict customer churn, which is the likelihood that
customers will not renew their yearly subscription. The average prediction is a 15% churn rate, but
for a particular customer the model predicts that they are 70% likely to churn. The customer has a
product usage history of 30%, is located in New York City, and became a customer in 1997. You need
to explain the difference between the actual prediction, a 70% churn rate, and the average
prediction. You want to use Vertex Explainable AI. What should you do?
Question 75
You are working on a classification problem with time series data. After
conducting just a few experiments using random cross-validation, you achieved an Area Under the
Receiver Operating Characteristic Curve (AUC ROC) value of 99% on the training data. You haven’t
explored using any sophisticated algorithms or spent any time on hyperparameter tuning. What should
your next step be to identify and fix the problem?
Question 76
You need to execute a batch prediction on 100 million records in a BigQuery
table with a custom TensorFlow DNN regressor model, and then store the predicted results in a
BigQuery table. You want to minimize the effort required to build this inference pipeline. What
should you do?
Question 77
You are creating a deep neural network classification model using a dataset
with categorical input values. Certain columns have a cardinality greater than 10,000 unique values.
How should you encode these categorical values as input into the model?
Question 78
You need to train a natural language model to perform text classification on
product descriptions that contain millions of examples and 100,000 unique words. You want to
preprocess the words individually so that they can be fed into a recurrent neural network. What
should you do?
Question 79
You work for an online travel agency that also sells advertising placements
on its website to other companies. You have been asked to predict the most relevant web banner that
a user should see next. Security is important to your company. The model latency requirements are
300ms@p99, the inventory is thousands of web banners, and your exploratory analysis has shown that
navigation context is a good predictor. You want to Implement the simplest solution. How should you
configure the prediction pipeline?
Question 80
Question 81: Your data science team has requested a system that supports
scheduled model retraining, Docker containers, and a service that supports autoscaling and
monitoring for online prediction requests. Which platform components should you choose for this
system?
• A. Vertex AI Pipelines and App Engine
• B. Vertex AI Pipelines, Vertex AI Prediction, and Vertex AI Model Monitoring
• C. Cloud Composer, BigQuery ML, and Vertex AI Prediction
• D. Cloud Composer, Vertex AI Training with custom containers, and App Engine
Question 81
You are profiling the performance of your TensorFlow model training time and
notice a performance issue caused by inefficiencies in the input data pipeline for a single 5
terabyte CSV file dataset on Cloud Storage. You need to optimize the input pipeline performance.
Which action should you try first to increase the efficiency of your pipeline?
Question 82
You need to design an architecture that serves asynchronous predictions to
determine whether a particular mission-critical machine part will fail. Your system collects data
from multiple sensors from the machine. You want to build a model that will predict a failure in the
next N minutes, given the average of each sensor’s data from the past 12 hours. How should you
design the architecture?
Question 83
Question 84: Your company manages an application that aggregates news
articles from many different online sources and sends them to users. You need to build a
recommendation model that will suggest articles to readers that are similar to the articles they are
currently reading. Which approach should you use?
Question 84
Question 85: You work for a large social network service provider whose users
post articles and discuss news. Millions of comments are posted online each day, and more than 200
human moderators constantly review comments and flag those that are inappropriate. Your team is
building an ML model to help human moderators check content on the platform. The model scores each
comment and flags suspicious comments to be reviewed by a human. Which metric(s) should you use to
monitor the model’s performance?
• A. Number of messages flagged by the model per minute
• B. Number of messages flagged by the model per minute confirmed as being inappropriate by humans.
• C. Precision and recall estimates based on a random sample of 0.1% of raw messages each minute
sent to a human for review
• D. Precision and recall estimates based on a sample of messages flagged by the model as
potentially inappropriate each minute
Question 85
Question 86: You are a lead ML engineer at a retail company. You want to
track and manage ML metadata in a centralized way so that your team can have reproducible
experiments by generating artifacts. Which management solution should you recommend to your team?
• A. Store your tf.logging data in BigQuery.
• B. Manage all relational entities in the Hive Metastore.
• C. Store all ML metadata in Google Cloud’s operations suite.
• D. Manage your ML workflows with Vertex ML Metadata.
Question 86
You have been given a dataset with sales predictions based on your company’s
marketing activities. The data is structured and stored in BigQuery, and has been carefully managed
by a team of data analysts. You need to prepare a report providing insights into the predictive
capabilities of the data. You were asked to run several ML models with different levels of
sophistication, including simple models and multilayered neural networks. You only have a few hours
to gather the results of your experiments. Which Google Cloud tools should you use to complete this
task in the most efficient and self-serviced way?
Question 87
Question 88: You are an ML engineer at a bank. You have developed a binary
classification model using AutoML Tables to predict whether a customer will make loan payments on
time. The output is used to approve or reject loan requests. One customer’s loan request has been
rejected by your model, and the bank’s risks department is asking you to provide the reasons that
contributed to the model’s decision. What should you do?
Question 88
Question 89: You work for a magazine distributor and need to build a model
that predicts which customers will renew their subscriptions for the upcoming year. Using your
company’s historical data as your training set, you created a TensorFlow model and deployed it to AI
Platform. You need to determine which customer attribute has the most predictive power for each
prediction served by the model. What should you do?
• A. Use AI Platform notebooks to perform a Lasso regression analysis on your model, which will
eliminate features that do not provide a strong signal.
• B. Stream prediction results to BigQuery. Use BigQuery’s CORR(X1, X2) function to calculate the
Pearson correlation coefficient between each feature and the target variable.
• C. Use the AI Explanations feature on AI Platform. Submit each prediction request with the
‘explain’ keyword to retrieve feature attributions using the sampled Shapley method.
• D. Use the What-If tool in Google Cloud to determine how your model will perform when individual
features are excluded. Rank the feature importance in order of those that caused the most
significant performance drop when removed from the model.
Question 89
Question 90: You are working on a binary classification ML algorithm that
detects whether an image of a classified scanned document contains a company’s logo. In the dataset,
96% of examples don’t have the logo, so the dataset is very skewed. Which metrics would give you the
most confidence in your model?
Question 90
Question 91: You work on the data science team for a multinational beverage
company. You need to develop an ML model to predict the company’s profitability for a new line of
naturally flavored bottled waters in different locations. You are provided with historical data that
includes product types, product sales volumes, expenses, and profits for all regions. What should
you use as the input and output for your model? • A. Use latitude, longitude, and product type as
features. Use profit as model output. • B. Use latitude, longitude, and product type as features.
Use revenue and expenses as model outputs. • C. Use product type and the feature cross of latitude
with longitude, followed by binning, as features. Use profit as model output. • D. Use product type
and the feature cross of latitude with longitude, followed by binning, as features. Use revenue and
expenses as model outputs.
Question 91
Question 92: You work on the data science team for a multinational beverage
company. You need to develop an ML model to predict the company’s profitability for a new line of
naturally flavored bottled waters in different locations. You are provided with historical data that
includes product types, product sales volumes, expenses, and profits for all regions. What should
you use as the input and output for your model?
Question 92
You have been asked to build a model using a dataset that is stored in a
medium-sized (~10 GB) BigQuery table. You need to quickly determine whether this data is suitable
for model development. You want to create a one-time report that includes both informative
visualizations of data distributions and more sophisticated statistical analyses to share with other
ML engineers on your team. You require maximum flexibility to create your report. What should you
do? • A. Use Vertex AI Workbench user-managed notebooks to generate the report. • B. Use the Google
Data Studio to create the report. • C. Use the output from TensorFlow Data Validation on Dataflow to
generate the report. • D. Use Dataprep to create the report.
Question 93
You work on an operations team at an international company that manages a
large fleet of on-premises servers located in few data centers around the world. Your team collects
monitoring data from the servers, including CPU/memory consumption. When an incident occurs on a
server, your team is responsible for fixing it. Incident data has not been properly labeled yet.
Your management team wants you to build a predictive maintenance solution that uses monitoring data
from the VMs to detect potential failures and then alerts the service desk team. What should you do
first?
Question 94
You are developing an ML model that uses sliced frames from video feed and
creates bounding boxes around specific objects. You want to automate the following steps in your
training pipeline: ingestion and preprocessing of data in Cloud Storage, followed by training and
hyperparameter tuning of the object model using Vertex AI jobs, and finally deploying the model to
an endpoint. You want to orchestrate the entire pipeline with minimal cluster management. What
approach should you use?
• A. Use Kubeflow Pipelines on Google Kubernetes Engine.
• B. Use Vertex AI Pipelines with TensorFlow Extended (TFX) SDK.
• C. Use Vertex AI Pipelines with Kubeflow Pipelines SDK.
• D. Use Cloud Composer for the orchestration.
Question 95
You are training an object detection machine learning model on a dataset that
consists of three million X-ray images, each roughly 2 GB in size. You are using Vertex AI Training
to run a custom training application on a Compute Engine instance with 32-cores, 128 GB of RAM, and
1 NVIDIA P100 GPU. You notice that model training is taking a very long time. You want to decrease
training time without sacrificing model performance. What should you do?
• A. Increase the instance memory to 512 GB and increase the batch size.
• B. Replace the NVIDIA P100 GPU with a v3-32 TPU in the training job.
• C. Enable early stopping in your Vertex AI Training job.
• D. Use the tf.distribute.Strategy API and run a distributed training job.
Question 96
Question 97: You are a data scientist at an industrial equipment
manufacturing company. You are developing a regression model to estimate the power consumption in
the company’s manufacturing plants based on sensor data collected from all of the plants. The
sensors collect tens of millions of records every day. You need to schedule daily training runs for
your model that use all the data collected up to the current date. You want your model to scale
smoothly and require minimal development work. What should you do?
Question 97
Question 98: You built a custom ML model using scikit-learn. Training time is
taking longer than expected. You decide to migrate your model to Vertex AI Training, and you want to
improve the model’s training time. What should you try out first?
Question 98
You are an ML engineer at a travel company. You have been researching
customers’ travel behavior for many years, and you have deployed models that predict customers’
vacation patterns. You have observed that customers’ vacation destinations vary based on seasonality
and holidays; however, these seasonal variations are similar across years. You want to quickly and
easily store and compare the model versions and performance statistics across years. What should you
do?
Question 99
Question 100: You are an ML engineer at a manufacturing company. You need to
build a model that identifies defects in products based on images of the product taken at the end of
the assembly line. You want your model to preprocess the images with lower computation to quickly
extract features of defects in products. Which approach should you use to build the model?
Question 100
Question 101: You are developing an ML model intended to classify whether
X-ray images indicate bone fracture risk. You have trained a ResNet architecture on Vertex AI using
a TPU as an accelerator, however you are unsatisfied with the training time and memory usage. You
want to quickly iterate your training code but make minimal changes to the code. You also want to
minimize impact on the model’s accuracy. What should you do?
Question 101
Question 102: You have successfully deployed to production a large and
complex TensorFlow model trained on tabular data. You want to predict the lifetime value (LTV) field
for each subscription stored in the BigQuery table named subscription.subscriptionPurchase in the
project named my-fortune500-company-project. You have organized all your training code, from
preprocessing data from the BigQuery table up to deploying the validated model to the Vertex AI
endpoint, into a TensorFlow Extended (TFX) pipeline. You want to prevent prediction drift, i.e., a
situation when a feature data distribution in production changes significantly over time. What
should you do?
Question 102
Question 103: You recently developed a deep learning model using Keras, and
now you are experimenting with different training strategies. First, you trained the model using a
single GPU, but the training process was too slow. Next, you distributed the training across 4 GPUs
using tf.distribute.MirroredStrategy (with no other changes), but you did not observe a decrease in
training time. What should you do?
Question 103
Question 104: You work for a gaming company that has millions of customers
around the world. All games offer a chat feature that allows players to communicate with each other
in real time. Messages can be typed in more than 20 languages and are translated in real time using
the Cloud Translation API. You have been asked to build an ML system to moderate the chat in real
time while assuring that the performance is uniform across the various languages and without
changing the serving infrastructure. You trained your first model using an in-house word2vec model
for embedding the chat messages translated by the Cloud Translation API. However, the model has
significant differences in performance across the different languages. How should you improve it? •
A. Add a regularization term such as the Min-Diff algorithm to the loss function. • B. Train a
classifier using the chat messages in their original language. • C. Replace the in-house word2vec
with GPT-3 or T5. • D. Remove moderation for languages for which the false positive rate is too
high.
Question 104
Question 105: You work for a gaming company that develops massively
multiplayer online (MMO) games. You built a TensorFlow model that predicts whether players will make
in-app purchases of more than $10 in the next two weeks. The model’s predictions will be used to
adapt each user’s game experience. User data is stored in BigQuery. How should you serve your model
while optimizing cost, user experience, and ease of management?
Question 105
You are building a linear regression model on BigQuery ML to predict a
customer’s likelihood of purchasing your company’s products. Your model uses a city name variable as
a key predictive component. In order to train and serve the model, your data must be organized in
columns. You want to prepare your data using the least amount of coding while maintaining the
predictable variables. What should you do?
Question 106
Question 107: You are an ML engineer at a bank that has a mobile application.
Management has asked you to build an ML-based biometric authentication for the app that verifies a
customer’s identity based on their fingerprint. Fingerprints are considered highly sensitive
personal information and cannot be downloaded and stored into the bank databases. Which learning
strategy should you recommend to train and deploy this ML mode?
Question 107
Question 108: You are experimenting with a built-in distributed XGBoost model
in Vertex AI Workbench user-managed notebooks. You use BigQuery to split your data into training and
validation sets using the following queries:
CREATE OR REPLACE TABLE ‘myproject.mydataset.training‘ AS
(SELECT * FROM ‘myproject.mydataset.mytable‘ WHERE RAND() <= 0.8);
CREATE OR REPLACE TABLE ‘myproject.mydataset.validation‘ AS
(SELECT * FROM ‘myproject.mydataset.mytable‘ WHERE RAND() <= 0.2);
After training the model, you achieve an area under the receiver operating characteristic curve (AUC
ROC) value of 0.8, but after deploying the model to production, you notice that your model
performance has dropped to an AUC ROC value of 0.65. What problem is most likely occurring?
Question 108
During batch training of a neural network, you notice that there is an
oscillation in the loss. How should you adjust your model to ensure that it converges?
Question 109
Question 110: You work for a toy manufacturer that has been experiencing a
large increase in demand. You need to build an ML model to reduce the amount of time spent by
quality control inspectors checking for product defects. Faster defect detection is a priority. The
factory does not have reliable Wi-Fi. Your company wants to implement the new ML model as soon as
possible. Which model should you use?
• A. AutoML Vision Edge mobile-high-accuracy-1 model
• B. AutoML Vision Edge mobile-low-latency-1 model
• C. AutoML Vision model
• D. AutoML Vision Edge mobile-versatile-1 model
Question 110
Question 111: You need to build classification workflows over several
structured datasets currently stored in BigQuery. Because you will be performing the classification
several times, you want to complete the following steps without writing code: exploratory data
analysis, feature selection, model building, training, and hyperparameter tuning and serving. What
should you do?
Question 111
Question 112: You are an ML engineer in the contact center of a large
enterprise. You need to build a sentiment analysis tool that predicts customer sentiment from
recorded phone conversations. You need to identify the best approach to building a model while
ensuring that the gender, age, and cultural differences of the customers who called the contact
center do not impact any stage of the model development pipeline and results. What should you do? •
A. Convert the speech to text and extract sentiments based on the sentences. • B. Convert the speech
to (text) and build a model based on the words. • C. Extract sentiment directly from the voice
recordings. • D. Convert the speech to text and extract sentiment using syntactical analysis.
Question 112
You need to analyze user activity data from your company’s mobile
applications. Your team will use BigQuery for data analysis, transformation, and experimentation
with ML algorithms. You need to ensure real-time ingestion of the user activity data into BigQuery.
What should you do?
Question 113
Question 114: You work for a gaming company that manages a popular online
multiplayer game where teams with 6 players play against each other in 5-minute battles. There are
many new players every day. You need to build a model that automatically assigns available players
to teams in real time. User research indicates that the game is more enjoyable when battles have
players with similar skill levels. Which business metrics should you track to measure your model’s
performance?
Question 114
You are building an ML model to predict trends in the stock market based on a
wide range of factors. While exploring the data, you notice that some features have a large range.
You want to ensure that the features with the largest magnitude don’t overfit the model. What should
you do?
Question 115
You work for a biotech startup that is experimenting with deep learning ML
models based on properties of biological organisms. Your team frequently works on early-stage
experiments with new architectures of ML models, and writes custom TensorFlow ops in C++. You train
your models on large datasets and large batch sizes. Your typical batch size has 1024 examples, and
each example is about 1 MB in size. The average size of a network with all weights and embeddings is
20 GB. What hardware should you choose for your models?
Question 116
Question 117: You are an ML engineer at an ecommerce company and have been
tasked with building a model that predicts how much inventory the logistics team should order each
month. Which approach should you take?
• A. Use a clustering algorithm to group popular items together. Give the list to the logistics team
so they can increase inventory of the popular items.
• B. Use a regression model to predict how much additional inventory should be purchased each month.
Give the results to the logistics team at the beginning of the month so they can increase inventory
by the amount predicted by the model.
• C. Use a time series forecasting model to predict each item's monthly sales. Give the results
to the logistics team so they can base inventory on the amount predicted by the model.
• D. Use a classification model to classify inventory levels as UNDER_STOCKED, OVER_STOCKED, and
CORRECTLY_STOCKED. Give the report to the logistics team each month so they can fine-tune inventory
levels.
Question 117
Question 118: You are building a TensorFlow model for a financial institution
that predicts the impact of consumer spending on inflation globally. Due to the size and nature of
the data, your model is long-running across all types of hardware, and you have built frequent
checkpointing into the training process. Your organization asked you to minimize cost. What hardware
should you choose?
Question 118
You work for a company that provides an anti-spam service that flags and
hides spam posts on social media platforms. Your company currently uses a list of 200,000 keywords
to identify suspected spam posts. If a post contains more than a few of these keywords, the post is
identified as spam. You want to start using machine learning to flag spam posts for human review.
What is the main advantage of implementing machine learning for this business case?
Question 119
One of your models is trained using data provided by a third-party data
broker. The data broker does not reliably notify you of formatting changes in the data. You want to
make your model training pipeline more robust to issues like this. What should you do?
Question 120
Question 121: You work for a company that is developing a new video streaming
platform. You have been asked to create a recommendation system that will suggest the next video for
a user to watch. After a review by an AI Ethics team, you are approved to start development. Each
video asset in your company’s catalog has useful metadata (e.g., content type, release date,
country), but you do not have any historical user event data. How should you build the
recommendation system for the first version of the product?
• A. Launch the product without machine learning. Present videos to users alphabetically, and start
collecting user event data so you can develop a recommender model in the future.
• B. Launch the product without machine learning. Use simple heuristics based on content metadata to
recommend similar videos to users, and start collecting user event data so you can develop a
recommender model in the future.
• C. Launch the product with machine (sic) machine learning. Use a publicly available dataset such
as MovieLens to train a model using the Recommendations AI, and then apply this trained model to
your data.
• D. Launch the product with machine learning. Generate embeddings for each video by training an
autoencoder on the content metadata using TensorFlow. Cluster content based on the similarity of
these embeddings, and then recommend videos from the same cluster.
Question 121
Question 122:
You recently built the first version of an image segmentation model for a self-driving car. After
deploying the model, you observe a decrease in the area under the curve (AUC) metric. When analyzing
the video recordings, you also discover that the model fails in highly congested traffic but works
as expected when there is less traffic. What is the most likely reason for this result?
• A. The model is overfitting in areas with less traffic and underfitting in areas with more
traffic.
• B. AUC is not the correct metric to evaluate this classification model.
• C. Too much data representing congested areas was used for model training.
• D. Gradients become small and vanish while backpropagating from the output to input nodes.
Question 122
Question 123: You are developing an ML model to predict house prices. While
preparing the data, you discover that an important predictor variable, distance from the closest
school, is often missing and does not have high variance. Every instance (row) in your data is
important. How should you handle the missing data?
• A. Delete the rows that have missing values.
• B. Apply feature crossing with another column that does not have missing values.
• C. Predict the missing values using linear regression.
• D. Replace the missing values with zeros.
Question 123
Question 124: You are an ML engineer responsible for designing and
implementing training pipelines for ML models. You need to create an end-to-end training pipeline
for a TensorFlow model. The TensorFlow model will be trained on several terabytes of structured
data. You need the pipeline to include data quality checks before training and model quality checks
after training but prior to deployment. You want to minimize development time and the need for
infrastructure maintenance. How should you build and orchestrate your training pipeline?
• A. Create the pipeline using Kubeflow Pipelines domain-specific language (DSL) and predefined
Google Cloud components. Orchestrate the pipeline using Vertex AI Pipelines.
• B. Create the pipeline using TensorFlow Extended (TFX) and standard TFX components. Orchestrate
the pipeline using Vertex AI Pipelines.
• C. Create the pipeline using Kubeflow Pipelines domain-specific language (DSL) and predefined
Google Cloud components. Orchestrate the pipeline using Kubeflease Pipelines deployed on Google
Kubernetes Engine.
• D. Create the pipeline using TensorFlow Extended (TFX) and standard TFX components. Orchestrate
the pipeline using Kubeflow Pipelines deployed on Google Kubernetes Engine.
Question 124
You manage a team of data scientists who use a cloud-based backend system to
submit training jobs. This system has become very difficult to administer, and you want to use a
managed service instead. The data scientists you work with use many different frameworks, including
Keras, PyTorch, theano, scikit-learn, and custom libraries. What should you do?
Question 125
Question 126: You are training an object detection model using a Cloud TPU
v2. Training time is taking longer than expected. Based on this simplified trace obtained with a
Cloud TPU profile, what action should you take to decrease training time in a cost-efficient way?
Question 126
While performing exploratory data analysis on a dataset, you find that an
important categorical feature has 5% null values. You want to minimize the bias that could result
from the missing values. How should you handle the missing values?
Question 127
Question 128: You are an ML engineer on an agricultural research team working
on a crop disease detection tool to detect leaf rust spots in images of crops to determine the
presence of a disease. These spots, which can vary in shape and size, are correlated to the severity
of the disease. You want to develop a solution that predicts the presence and severity of the
disease with high accuracy. What should you do?
Question 128
You have been asked to productionize a proof-of-concept ML model built using
Keras. The model was trained in a Jupyter notebook on a data scientist’s local machine. The notebook
contains a cell that performs data validation and a cell that performs model analysis. You need to
orchestrate the steps contained in the notebook and automate the execution of these steps for weekly
retraining. You expect much more training data in the future. You want your solution to take
advantage of managed services while minimizing cost. What should you do?
Question 129
You are working on a system log anomaly detection model for a cybersecurity
organization. You have developed the model using TensorFlow, and you plan to use it for real-time
prediction. You need to create a Dataflow pipeline to ingest data via Pub/Sub and write the results
to BigQuery. You want to minimize the serving latency as much as possible. What should you do?
Question 130
Question 131: You work on a data science team at a bank and are creating an
ML model to predict loan default risk. You have collected and cleaned hundreds of millions of
records worth of training data in a BigQuery table, and you now want to develop and compare multiple
models on this data using TensorFlow and Vertex AI. You want to minimize any bottlenecks during the
data ingestion state while considering scalability. What should you do?
• A. Use the BigQuery client library to load data into a dataframe, and use
tf.data.Dataset.from_tensor_slices() to read it.
• B. Export data to CSV files in Cloud Storage, and use tf.data.TextLineDataset() to read them.
• C. Convert the data into TFRecords, and use tf.data.TFRecordDataset() to read them.
• D. Use TensorFlow I/O’s BigQuery Reader to directly read the data.
Question 131
You work on a data science team at a bank and are creating an ML model to
predict loan default risk. You have collected and cleaned hundreds of millions of records worth of
training data in a BigQuery table, and you now want to develop and compare multiple models on this
data using TensorFlow and Vertex AI. You want to minimize any bottlenecks during the data ingestion
state while considering scalability. What should you do?
Question 132
Question 133: You have recently created a proof-of-concept (POC) deep
learning model. You are satisfied with the overall architecture, but you need to determine the value
for a couple of hyperparameters. You want to perform hyperparameter tuning on Vertex AI to determine
both the appropriate embedding dimension for a categorical feature used by your model and the
optimal learning rate. You configure the following settings: • For the embedding dimension, you set
the type to INTEGER with a minValue of 16 and maxValue of 64. • For the learning rate, you set the
type to DOUBLE with a minValue of 10e-05 and maxValue of 10e-02. You are using the default Bayesian
optimization tuning algorithm, and you want to maximize model accuracy. Training time is not a
concern. How should you set the hyperparameter scaling for each hyperparameter and the
maxParallelTrials? • A. Use UNIT_LINEAR_SCALE for the embedding dimension, UNIT_LOG_SCALE for the
learning rate, and a large number of parallel trials. • B. Use UNIT_LINEAR_SCALE for the embedding
dimension, UNIT_LOG_SCALE for the learning rate, and a small number of parallel trials. • C. Use
UNIT_LOG_SCALE for the embedding dimension, UNIT_LINEAR_SCALE for the learning rate, and a large
number of parallel trials. • D. Use UNIT_LOG_SCALE for the embedding dimension, UNIT_LINEAR_SCALE
for the learning rate, and a small number of parallel trials.
Question 133
Question 134: You are the Director of Data Science at a large company, and
your Data Science team has recently begun using the Kubeflow Pipelines SDK to orchestrate their
training pipelines. Your team is struggling to integrate their custom Python code into the Kubeflow
Pipelines SDK. How should you instruct them to proceed in order to quickly integrate their code with
the Kubeflow Pipelines SDK?
Question 134
Question 135: You work for the AI team of an automobile company, and you are
developing a visual defect detection model using TensorFlow and Keras. To improve your model
performance, you want to incorporate some image augmentation functions such as translation,
cropping, and contrast tweaking. You randomly apply these functions to each training batch. You want
to optimize your data processing pipeline for run time and compute resources utilization. What
should you do?
Question 135
You work for an online publisher that delivers news articles to over 50
million readers. You have built an AI model that recommends content for the company’s weekly
newsletter. A recommendation is considered successful if the article is opened within two days of
the newsletter’s published date and the user remains on the page for at least one minute. All the
information needed to compute the success metric is available in BigQuery and is updated hourly. The
model is trained on eight weeks of data, on average its performance degrades below the acceptable
baseline after five weeks, and training time is 12 hours. You want to ensure that the model’s
performance is above the acceptable baseline while minimizing cost. How should you monitor the model
to determine when retraining is necessary?
Question 136
Question 137: You deployed an ML model into production a year ago. Every
month, you collect all raw requests that were sent to your model prediction service during the
previous month. You send a subset of these requests to a human labeling service to evaluate your
model’s performance. After a year, you notice that your model's performance sometimes degrades
significantly after a month, while other times it takes several months to notice any decrease in
performance. The labeling service is costly, but you also need to avoid large performance
degradations. You want to determine how often you should retrain your model to maintain a high level
of performance while minimizing cost. What should you do?
• A. Train an anomaly detection model on the training dataset, and run all incoming requests through
this model. If an anomaly is detected, send the most recent serving data to the labeling service.
• B. Identify temporal patterns in your model’s performance over the previous year. Based on these
patterns, create a schedule for sending serving data to the labeling service for the next year.
• C. Compare the cost of the labeling service with the lost revenue due to model performance
degradation over the past year. If the lost revenue is greater than the cost of the labeling
service, increase the frequency of model retraining; otherwise, decrease the model retraining
frequency.
• D. Run training-serving skew detection batch jobs every few days to compare the aggregate
statistics of the features in the training dataset with recent serving data. If skew is detected,
send the most recent serving data to the labeling service.
Question 137
You work for a company that manages a ticketing platform for a large chain of
cinemas. Customers use a mobile app to search for movies they’re interested in and purchase tickets
in the app. Ticket purchase requests are sent to Pub/Sub and are processed with a Dataflow streaming
pipeline configured to conduct the following steps: 1. Check for availability of the movie tickets
at the selected cinema. 2. Assign the ticket price and accept payment. 3. Reserve the tickets at the
selected cinema. 4. Send successful purchases to your database. Each step in this process has low
latency requirements (less than 50 milliseconds). You have developed a logistic regression model
with BigQuery ML that predicts whether offering a promo code for free popcorn increases the chance
of a ticket purchase, and this prediction should be added to the ticket purchase process. You want
to identify the simplest way to deploy this model to production while adding minimal latency. What
should you do?
Question 138
Question 139: You work on a team in a data center that is responsible for
server maintenance. Your management team wants you to build a predictive maintenance solution that
uses monitoring data to detect potential server failures. Incident data has not been labeled yet.
What should you do first?
Question 139
Question 140: You work for a retailer that sells clothes to customers around
the world. You have been tasked with ensuring that ML models are built in a secure manner.
Specifically, you need to protect sensitive customer data that might be used in the models. You have
identified four fields containing sensitive data that are being used by your data science team: AGE,
IS_EXISTING_CUSTOMER, LATITUDE_LONGITUDE, and SHIRT_SIZE. What should you do with the data before it
is made available to the data science team for training purposes?
Question 140
You work for a magazine publisher and have been tasked with predicting
whether customers will cancel their annual subscription. In your exploratory data analysis, you find
that 90% of individuals renew their subscription every year, and only 10% of individuals cancel
their subscription. After training a NN Classifier, your model predicts those who cancel their
subscription with 99% accuracy and predicts those who renew their subscription with 82% accuracy.
How should you interpret these results?
Question 141
Question 142: You have built a model that is trained on data stored in
Parquet files. You access the data through a Hive table hosted on Google Cloud. You preprocessed
these data with PySpark and exported it as a CSV file into Cloud Storage. After preprocessing, you
execute additional steps to train and evaluate your model. You want to parametrize this model
training in Kubeflow Pipelines. What should you do?
Question 142
You have developed an ML model to detect the sentiment of users’ posts on
your company's social media page to identify outages or bugs. You are using Dataflow to provide
real-time predictions on data ingested from Pub/Sub. You plan to have multiple training iterations
for your model and keep the latest two versions live after every run. You want to split the traffic
between the versions in an 80:20 ratio, with the newest model getting the majority of the traffic.
You want to keep the pipeline as simple as possible, with minimal management required. What should
you do?
Question 143
You are developing an image recognition model using PyTorch based on ResNet50
architecture. Your code is working fine on your local laptop on a small subsample. Your full dataset
has 200k labeled images. You want to quickly scale your training workload while minimizing cost. You
plan to use 4 V100 GPUs. What should you do?
Question 144
You have trained a DNN regressor with TensorFlow to predict housing prices
using a set of predictive features. Your default precision is tf.float64, and you use a standard
TensorFlow estimator. Your model performs well, but just before deploying it to production, you
discover that your current serving latency is 10ms @ 90 percentile and you currently serve on CPUs.
Your production requirements expect a model latency of 8ms @ 90 percentile. You're willing to
accept a small decrease in performance in order to reach the latency requirement. Therefore, your
plan is to improve latency while evaluating how much the model's prediction decreases. What
should you first try to quickly lower the serving latency?
Question 145
Question 146: You work on the data science team at a manufacturing company.
You are reviewing the company’s historical sales data, which has hundreds of millions of records.
For your exploratory data analysis, you need to calculate descriptive statistics such as mean,
median, and mode; conduct complex statistical tests for hypothesis testing; and plot variations of
the features over time. You want to use as much of the sales data as possible in your analyses while
minimizing computational resources. What should you do?
Question 146
Question 147: Your data science team needs to rapidly experiment with various
features, model architectures, and hyperparameters. They need to track the accuracy metrics for
various experiments and use an API to query the metrics over time. What should they use to track and
report their experiments while minimizing manual effort?
Question 147
You are training an ML model using data stored in BigQuery that contains
several values that are considered Personally Identifiable Information (PII). You need to reduce the
sensitivity of the dataset before training your model. Every column is critical to your model. How
should you proceed?
Question 148
You recently deployed an ML model. Three months after deployment, you notice
that your model is underperforming on certain subgroups, thus potentially leading to biased results.
You suspect that the inequitable performance is due to class imbalances in the training data, but
you cannot collect more data. What should you do? (Choose two.)
Question 149
Question 150: You are working on a binary classification ML algorithm that
detects whether an image of a classified scanned document contains a company’s logo. In the dataset,
96% of examples don’t have the logo, so the dataset is very skewed. Which metric would give you the
most confidence in your model?
Question 150
While running a model training pipeline on Vertex Al, you discover that the
evaluation step is failing because of an out-of-memory error. You are currently using TensorFlow
Model Analysis (TFMA) with a standard Evaluator TensorFlow Extended (TFX) pipeline component for the
evaluation step. You want to stabilize the pipeline without downgrading the evaluation quality while
minimizing infrastructure overhead. What should you do?
Question 151
Question 152: You are developing an ML model using a dataset with categorical
input variables. You have randomly split half of the data into training and test sets. After
applying one-hot encoding on the categorical variables in the training set, you discover that one
categorical variable is missing from the test set. What should you do?
Question 152
Question 153: You work for a bank and are building a random forest model for
fraud detection. You have a dataset that includes transactions, of which 1% are identified as
fraudulent. Which data transformation strategy would likely improve the performance of your
classifier?
Question 153
Question 154: You are developing a classification model to support
predictions for your company’s various products. The dataset you were given for model development
has class imbalance. You need to minimize false positives and false negatives. What evaluation
metric should you use to properly train the model?
Question 154
You are training an object detection machine learning model on a dataset that
consists of three million X-ray images, each roughly 2 GB in size. You are using Vertex AI Training
to run a custom training application on a Compute Engine instance with 32-cores, 128 GB of RAM, and
1 NVIDIA P100 GPU. You notice that model training is taking a very long time. You want to decrease
training time without sacrificing model performance. What should you do?
Question 155
Question 156: You need to build classification workflows over several
structured datasets currently stored in BigQuery. Because you will be performing the classification
several times, you want to complete the following steps without writing code: exploratory data
analysis, feature selection, model building, training, and hyperparameter tuning and serving. What
should you do?
• A. Train a TensorFlow model on Vertex AI.
• B. Train a classification Vertex AutoML model.
• C. Run a logistic regression job on BigQuery ML.
• D. Use scikit-learn in Vertex AI Workbench user-managed notebooks with pandas library.
Question 156
Question 157: You recently developed a deep learning model. To test your new
model, you trained it for a few epochs on a large dataset. You observe that the training and
validation losses barely changed during the training run. You want to quickly debug your model. What
should you do first?
• A. Verify that your model can obtain a low loss on a small subset of the dataset
• B. Add handcrafted features to inject your domain knowledge into the model
• C. Use the Vertex AI hyperparameter tuning service to identify a better learning rate
• D. Use hardware accelerators and train your model for more epochs
Question 157
You are a data scientist at an industrial equipment manufacturing company.
You are developing a regression model to estimate the power consumption in the company’s
manufacturing plants based on sensor data collected from all of the plants. The sensors collect tens
of millions of records every day. You need to schedule daily training runs for your model that use
all the data collected up to the current date. You want your model to scale smoothly and require
minimal development work. What should you do?
Question 158
Your organization manages an online message board. A few months ago, you
discovered an increase in toxic language and bullying on the message board. You deployed an
automated text classifier that flags certain comments as toxic or harmful. Now some users are
reporting that benign comments referencing their religion are being misclassified as abusive. Upon
further inspection, you find that your classifier's false positive rate is higher for comments
that reference certain underrepresented religious groups. Your team has a limited budget and is
already overextended. What should you do?
Question 159
Question 160: You work for a magazine distributor and need to build a model
that predicts which customers will renew their subscriptions for the upcoming year. Using your
company’s historical data as your training set, you created a TensorFlow model and deployed it to
Vertex AI. You need to determine which customer attribute has the most predictive power for each
prediction served by the model. What should you do?
• A. Stream prediction results to BigQuery. Use BigQuery’s CORR(X1, X2) function to calculate the
Pearson correlation coefficient between each feature and the target variable.
• B. Use Vertex Explainable AI. Submit each prediction request with the explain' keyword to
retrieve feature attributions using the sampled Shapley method.
• C. Use Vertex AI Workbench user-managed notebooks to perform a Lasso regression analysis on your
model, which will eliminate features that do not provide a strong signal.
• D. Use the What-If tool in Google Cloud to determine how your model will perform when individual
features are excluded. Rank the feature importance in order of those that caused the most
significant performance drop when removed from the model.
Question 160
Question 161: You work for a company that is developing a new video streaming
platform. You have been asked to create a recommendation system that will suggest the next video for
a user to watch. After a review by an AI Ethics team, you are approved to start development. Each
video asset in your company’s catalog has useful metadata (e.g., content type, release date,
country), but you do not have any historical user event data. How should you build the
recommendation system for the first version of the product?
Question 161
Question 162: You built a custom ML model using scikit-learn. Training time
is taking longer than expected. You decide to migrate your model to Vertex AI Training, and you want
to improve the model’s training time. What should you try out first?
Question 162
Question 163: You are an ML engineer at a retail company. You have built a
model that predicts a coupon to offer an ecommerce customer at checkout based on the items in their
cart. When a customer goes to checkout, your serving pipeline, which is hosted on Google Cloud,
joins the customer's existing cart with a row in a BigQuery table that contains the
customers' historic purchase behavior and uses that as the model's input. The web team is
reporting that your model is returning predictions too slowly to load the coupon offer with the rest
of the web page. How should you speed up your model's predictions?
Question 163
You work for a small company that has deployed an ML model with autoscaling
on Vertex AI to serve online predictions in a production environment. The current model receives
about 20 prediction requests per hour with an average response time of one second. You have
retrained the same model on a new batch of data, and now you are canary testing it, sending ~10% of
production traffic to the new model. During this canary test, you notice that prediction requests
for your new model are taking between 30 and 180 seconds to complete. What should you do?
Question 164
You want to train an AutoML model to predict house prices by using a small
public dataset stored in BigQuery. You need to prepare the data and want to use the simplest, most
efficient approach. What should you do?
Question 165
Question 166: You developed a Vertex AI ML pipeline that consists of
preprocessing and training steps and each set of steps runs on a separate custom Docker image. Your
organization uses GitHub and GitHub Actions as CI/CD to run unit and integration tests. You need to
automate the model retraining workflow so that it can be initiated both manually and when a new
version of the code is merged in the main branch. You want to minimize the steps required to build
the workflow while also allowing for maximum flexibility. How should you configure the CI/CD
workflow?
• A. Trigger a Cloud Build workflow to run tests, build custom Docker images, push the images to
Artifact Registry, and launch the pipeline in Vertex AI Pipelines.
• B. Trigger GitHub Actions to run the tests, launch a job on Cloud Run to build custom Docker
images, push the images to Artifact Registry, and launch the pipeline in Vertex AI Pipelines.
• C. Trigger GitHub Actions to run the tests, build custom Docker images, push the images to
Artifact Registry, and launch the pipeline in Vertex AI Pipelines.
• D. Trigger GitHub Actions to run the tests, launch a Cloud Build workflow to build custom Docker
images, push the images to Artifact Registry, and launch the pipeline in Vertex AI Pipelines.
Question 166
Question 167: You are working with a dataset that contains customer
transactions. You need to build an ML model to predict customer purchase behavior. You plan to
develop the model in BigQuery ML, and export it to Cloud Storage for online prediction. You notice
that the input data contains a few categorical features, including product category and payment
method. You want to deploy the model as quickly as possible. What should you do?
• A. Use the TRANSFORM clause with the ML.ONE_HOT_ENCODER function on the categorical features at
model creation and select the categorical and non-categorical features.
• B. Use the ML.ONE_HOT_ENCODER function on the categorical features and select the encoded
categorical features and non-categorical features as inputs to create your model.
• C. Use the CREATE MODEL statement and select the categorical and non-categorical features.
• D. Use the ML.MULTI_HOT_ENCODER function on the categorical features, and select the encoded
categorical features and non-categorical features as inputs to create your model.
Question 167
You need to develop an image classification model by using a large dataset
that contains labeled images in a Cloud Storage bucket. What should you do?
Question 168
Question 169: You are developing a model to detect fraudulent credit card
transactions. You need to prioritize detection, because missing even one fraudulent transaction
could severely impact the credit card holder. You used AutoML to train a model on users'
profile information and credit card transaction data. After training the initial model, you notice
that the model is failing to detect many fraudulent transactions. How should you adjust the training
parameters in AutoML to improve model performance? (Choose two.) • A. Increase the score threshold •
B. Decrease the score threshold. • C. Add more positive examples to the training set • D. Add more
negative examples to the training set • E. Reduce the maximum number of node hours for training
Question 169
You need to deploy a scikit-learn classification model to production. The
model must be able to serve requests 24/7, and you expect millions of requests per second to the
production application from 8 am to 7 pm. You need to minimize the cost of deployment. What should
you do?
Question 170
You need to deploy a scikit-learn classification model to production. The
model must be able to serve requests 24/7, and you expect millions of requests per second to the
production application from 8 am to 7 pm. You need to minimize the cost of deployment. What should
you do?
Question 171
You created an ML pipeline with multiple input parameters. You want to
investigate the tradeoffs between different parameter combinations. The parameter options are •
Input dataset • Max tree depth of the boosted tree regressor • Optimizer learning rate. You need to
compare the pipeline performance of the different parameter combinations measured in F1 score, time
to train, and model complexity. You want your approach to be reproducible, and track all pipeline
runs on the same platform. What should you do?
Question 172
Question 173: You received a training-serving skew alert from a Vertex AI
Model Monitoring job running in production. You retrained the model with more recent training data,
and deployed it back to the Vertex AI endpoint, but you are still receiving the same alert. What
should you do?
Question 173
You developed a custom model by using Vertex AI to forecast the sales of your
company’s products based on historical transactional data. You anticipate changes in the feature
distributions and the correlations between the features in the near future. You also expect to
receive a large volume of prediction requests. You plan to use Vertex AI Model Monitoring for drift
detection and you want to minimize the cost. What should you do?
Question 174
You have recently trained a scikit-learn model that you plan to deploy on
Vertex AI. This model will support both online and batch prediction. You need to preprocess input
data for model inference. You want to package the model for deployment while minimizing additional
code. What should you do?
Question 175
You have developed an ML model to detect the sentiment of users’ posts on
your company's social media page to identify outages or bugs. You are using Dataflow to provide
real-time predictions on data ingested from Pub/Sub. You plan to have multiple training iterations
for your model and keep the latest two versions live after every run. You want to split the traffic
between the versions in an 80:20 ratio, with the newest model getting the majority of the traffic.
You want to keep the pipeline as simple as possible, with minimal management required. What should
you do?
Question 176
Question 177: You have created a Vertex AI pipeline that includes two steps.
The first step preprocesses 10 TB data completes in about 1 hour, and saves the result in a Cloud
Storage bucket. The second step uses the processed data to train a model. You need to update the
model’s code to allow you to test different algorithms. You want to reduce pipeline execution time
and cost while also minimizing pipeline changes. What should you do?
• A. Add a pipeline parameter and an additional pipeline step. Depending on the parameter value, the
pipeline step conducts or skips data preprocessing, and starts model training.
• B. Create another pipeline without the preprocessing step, and hardcode the preprocessed Cloud
Storage file location for model training.
• C. Configure a machine with more CPU and RAM from the compute-optimized machine family for the
data preprocessing step.
• D. Enable caching for the pipeline job, and disable caching for the model training step.
Question 177
Question 178: You work for a bank. You have created a custom model to predict
whether a loan application should be flagged for human review. The input features are stored in a
BigQuery table. The model is performing well, and you plan to deploy it to production. Due to
compliance requirements the model must provide explanations for each prediction. You want to add
this functionality to your model code with minimal effort and provide explanations that are as
accurate as possible. What should you do?
Question 178
You recently used XGBoost to train a model in Python that will be used for
online serving. Your model prediction service will be called by a backend service implemented in
Golang running on a Google Kubernetes Engine (GKE) cluster. Your model requires pre and
postprocessing steps. You need to implement the processing steps so that they run at serving time.
You want to minimize code changes and infrastructure maintenance, and deploy your model into
production as quickly as possible. What should you do?
Question 179
Question 180: You recently deployed a pipeline in Vertex AI Pipelines that
trains and pushes a model to a Vertex AI endpoint to serve real-time traffic. You need to continue
experimenting and iterating on your pipeline to improve model performance. You plan to use Cloud
Build for CI/CD. You want to quickly and easily deploy new pipelines into production, and you want
to minimize the chance that the new pipeline implementations will break in production. What should
you do?
Question 180
Question 181: You work for a bank with strict data governance requirements.
You recently implemented a custom model to detect fraudulent transactions. You want your training
code to download internal data by using an API endpoint hosted in your project’s network. You need
the data to be accessed in the most secure way, while mitigating the risk of data exfiltration. What
should you do?
• A. Enable VPC Service Controls for peerings, and add Vertex AI to a service perimeter.
• B. Create a Cloud Run endpoint as a proxy to the data. Use Identity and Access Management (IAM)
authentication to secure access to the endpoint from the training job.
• C. Configure VPC Peering with Vertex AI, and specify the network of the training job.
• D. Download the data to a Cloud Storage bucket before calling the training job.
Question 181
Question 182: You are profiling the performance of your TensorFlow model
training time and notice a performance issue caused by inefficiencies in the input data pipeline for
a single 5 terabyte CSV file dataset on Cloud Storage. You need to optimize the input pipeline
performance. Which action should you try first to increase the efficiency of your pipeline?
Question 182
Question 183: You are training an ML model on a large dataset. You are using
a TPU to accelerate the training process. You notice that the training process is taking longer than
expected. You discover that the TPU is not reaching its full capacity. What should you do? • A.
Increase the learning rate • B. Increase the number of epochs • C. Decrease the learning rate • D.
Increase the batch size
Question 183
Question 184: You work for a retail company. You have a managed tabular
dataset in Vertex AI that contains sales data from three different stores. The dataset includes
several features, such as store name and sale timestamp. You want to use the data to train a model
that makes sales predictions for a new store that will open soon. You need to split the data between
the training, validation, and test sets. What approach should you use to split the data?
Question 184
You have developed a BigQuery ML model that predicts customer chum, and
deployed the model to Vertex AI Endpoints. You want to automate the retraining of your model by
using minimal additional code when model feature values change. You also want to minimize the number
of times that your model is retrained to reduce training costs. What should you do?
Question 185
You have been tasked with deploying prototype code to production. The feature
engineering code is in PySpark and runs on Dataproc Serverless. The model training is executed by
using a Vertex AI custom training job. The two steps are not connected, and the model training must
currently be run manually after the feature engineering step finishes. You need to create a scalable
and maintainable production process that runs end-to-end and tracks the connections between steps.
What should you do?
Question 186
You recently deployed a scikit-learn model to a Vertex AI endpoint. You are
now testing the model on live production traffic. While monitoring the endpoint, you discover twice
as many requests per hour than expected throughout the day. You want the endpoint to efficiently
scale when the demand increases in the future to prevent users from experiencing high latency. What
should you do?
Question 187
Question 188: You work at a bank. You have a custom tabular ML model that was
provided by the bank’s vendor. The training data is not available due to its sensitivity. The model
is packaged as a Vertex AI Model serving container, which accepts a string as input for each
prediction instance. In each string, the feature values are separated by commas. You want to deploy
this model to production for online predictions and monitor the feature distribution over time with
minimal effort. What should you do?
Question 188
You recently deployed a model to a Vertex AI endpoint. Your data drifts
frequently, so you have enabled request-response logging and created a Vertex AI Model Monitoring
job. You have observed that your model is receiving higher traffic than expected. You need to reduce
the model monitoring cost while continuing to quickly detect drift. What should you do?
Question 189
You work for a retail company. You have been asked to develop a model to
predict whether a customer will purchase a product on a given day. Your team has processed the
company’s sales data, and created a table with the following rows: • Customer_id • Product_id • Date
• Days_since_last_purchase (measured in days) • Average_purchase_frequency (measured in 1/days) •
Purchase (binary class, if customer purchased product on the Date) You need to interpret your
model’s results for each individual prediction. What should you do?
Question 190
You work for a company that captures live video footage of checkout areas in
their retail stores. You need to use the live video footage to build a model to detect the number of
customers waiting for service in near real time. You want to implement a solution quickly and with
minimal effort. How should you build the model?
Question 191
Question 197: You work as an analyst at a large banking firm. You are
developing a robust scalable ML pipeline to train several regression and classification models. Your
primary focus for the pipeline is model interpretability. You want to productionize the pipeline as
quickly as possible. What should you do?
• A. Use Tabular Workflow for Wide & Deep through Vertex AI Pipelines to jointly train wide
linear models and deep neural networks
• B. Use Google Kubernetes Engine to build a custom training pipeline for XGBoost-based models
• C. Use Tabular Workflow for TabNet through Vertex AI Pipelines to train attention-based models
• D. Use Cloud Composer to build the training pipelines for custom deep learning-based models
Question 192
You developed a Transformer model in TensorFlow to translate text. Your
training data includes millions of documents in a Cloud Storage bucket. You plan to use distributed
training to reduce training time. You need to configure the training job while minimizing the effort
required to modify code and to manage the cluster’s configuration. What should you do?
Question 193
Question 199: You are developing a process for training and running your
custom model in production. You need to be able to show lineage for your model and predictions. What
should you do?
Question 194
Question 200: You work for a hotel and have a dataset that contains
customers’ written comments scanned from paper-based customer feedback forms, which are stored as
PDF files. Every form has the same layout. You need to quickly predict an overall satisfaction score
from the customer comments on each form. How should you accomplish this task?
Question 195
You developed a Vertex AI pipeline that trains a classification model on data
stored in a large BigQuery table. The pipeline has four steps, where each step is created by a
Python function that uses the KubeFlow v2 API. The components have the following names: [Note:
Component names are not specified in the question text, so they are omitted from the JSON]. You
launch your Vertex AI pipeline as the following: [Note: Launch command/reference is not specified,
so omitted]. You perform many model iterations by adjusting the code and parameters of the training
step. You observe high costs associated with the development, particularly the data export and
preprocessing steps. You need to reduce model development costs. What should you do?
• A. Change the components’ YAML filenames to export.yaml, preprocess.yaml,
f"train-{dt}.yaml", f"calibrate-{dt}.yaml".
• B. Add the {"kubeflow.v1.caching": True} parameter to the set of params provided to your
PipelineJob.
• C. Move the first step of your pipeline to a separate step, and provide a cached path to Cloud
Storage as an input to the main pipeline.
• D. Change the name of the pipeline to f"my-awesome-pipeline-{dt}".
Question 196
Question 202: You work for a startup that has multiple data science
workloads. Your compute infrastructure is currently on-premises, and the data science workloads are
native to PySpark. Your team plans to migrate their data science workloads to Google Cloud. You need
to build a proof of concept to migrate one data science job to Google Cloud. You want to propose a
migration process that requires minimal cost and effort. What should you do first?
• A. Create a n2-standard-4 VM instance and install Java, Scala, and Apache Spark dependencies on
it.
• B. Create a Google Kubernetes Engine cluster with a basic node pool configuration, install Java,
Scala, and Apache Spark dependencies on it.
• C. Create a Standard (1 master, 3 workers) Dataproc cluster, and run a Vertex AI Workbench
notebook instance on it.
• D. Create a Vertex AI Workbench notebook with instance type n2-standard-4.
Question 197
Question 203: A TensorFlow machine learning model on Compute Engine virtual
machines (n2-standard-32) takes two days to complete training. The model has custom TensorFlow
operations that must run partially on a CPU. You want to reduce the training time in a
cost-effective manner. What should you do?
• A. Change the VM type to n2-highmem-32.
• B. Change the VM type to e2-standard-32.
• C. Train the model using a VM with a GPU hardware accelerator.
• D. Train the model using a VM with a TPU hardware accelerator.
Question 198
Question 204: You work for an auto insurance company. You are preparing a
proof-of-concept ML application that uses images of damaged vehicles to infer damaged parts. Your
team has assembled a set of annotated images from damage claim documents in the company’s database.
The annotations associated with each image consist of a bounding box for each identified damaged
part and the part name. You have been given a sufficient budget to train models on Google Cloud. You
need to quickly create an initial model. What should you do? • A. Download a pre-trained object
detection model from TensorFlow Hub. Fine-tune the model in Vertex AI Workbench by using the
annotated image data. • B. Train an object detection model in AutoML by using the annotated image
data. • C. Create a pipeline in Vertex AI Pipelines and configure the AutoMLTrainingJobRunOp
component to train a custom object detection model by using the annotated image data. • D. Train an
object detection model in Vertex AI custom training by using the annotated image data.
Question 199
You are analyzing customer data for a healthcare organization that is stored
in Cloud Storage. The data contains personally identifiable information (PII). You need to perform
data exploration and preprocessing while ensuring the security and privacy of sensitive fields. What
should you do?
Question 200
You are building a predictive maintenance model to preemptively detect part
defects in bridges. You plan to use high definition images of the bridges as model inputs. You need
to explain the output of the model to the relevant stakeholders so they can take appropriate action.
How should you build the model?
• A. Use scikit-learn to build a tree-based model, and use SHAP values to explain the model output.
• B. Use scikit-learn to build a tree-based model, and use partial dependence plots (PDP) to explain
the model output.
• C. Use TensorFlow to create a deep learning-based model, and use Integrated Gradients to explain
the model output.
• D. Use TensorFlow to create a deep learning-based (PDP) to explain the model output.
Question 201
Question 207: You work for a hospital that wants to optimize how it schedules
operations. You need to create a model that uses the relationship between the number of surgeries
scheduled and beds used. You want to predict how many beds will be needed for patients each day in
advance based on the scheduled surgeries. You have one year of data for the hospital organized in
365 rows. The data includes the following variables for each day: • Number of scheduled surgeries •
Number of beds occupied • Date You want to maximize the speed of model development and testing. What
should you do? • A. Create a BigQuery table. Use BigQuery ML to build a regression model, with
number of beds as the target variable, and number of scheduled surgeries and date features (such as
day of week) as the predictors. • B. Create a BigQuery table. Use BigQuery ML to build an ARIMA
model, with number of beds as the target variable, and date as the time variable. • C. Create a
Vertex AI tabular dataset. Train an AutoML regression model, with number of beds as the target
variable, and number of scheduled minor surgeries and date features (such as day of the week) as the
predictors. • D. Create a Vertex AI tabular dataset. Train a Vertex AI AutoML Forecasting model,
with number of beds as the target variable, number of scheduled surgeries as a covariate and date as
the time variable.
Question 202
You recently developed a wide and deep model in TensorFlow. You generated
training datasets using a SQL script that preprocessed raw data in BigQuery by performing
instance-level transformations of the data. You need to create a training pipeline to retrain the
model on a weekly basis. The trained model will be used to generate daily recommendations. You want
to minimize model development and training time. How should you develop the training pipeline?
Question 203
Question 209: You are training a custom language model for your company using
a large dataset. You plan to use the Reduction Server strategy on Vertex AI. You need to configure
the worker pools of the distributed training job. What should you do?
• A. Configure the machines of the first two worker pools to have GPUs, and to use a container image
where your training code runs. Configure the third worker pool to have GPUs, and use the
reductionserver container image.
• B. Configure the machines of the first two worker pools to have GPUs and to use a container image
where your training code runs. Configure the third worker pool to use the reductionserver container
image without accelerators, and choose a machine type that prioritizes bandwidth.
• C. Configure the machines of the first two worker pools to have TPUs and to use a container image
where your training code runs. Configure the third worker pool without accelerators, and use the
reductionserver container image without accelerators, and choose a machine type that prioritizes
bandwidth.
• D. Configure the machines of the first two pools to have TPUs, and to use a container image where
your training code runs. Configure the third pool to have TPUs, and use the reduction
(reductionserver) container image.
Question 204
Question 210: You have trained a model by using data that was preprocessed in
a batch Dataflow pipeline. Your use case requires real-time inference. You want to ensure that the
data preprocessing logic is applied consistently between training and serving. What should you do?
Question 205
You need to develop a custom TensorFlow model that will be used for online
predictions. The training data is stored in BigQuery. You need to apply instance-level data
transformations to the data for model training and serving. You want to use the same preprocessing
routine during model training and serving. How should you configure the preprocessing routine?
Question 206
You are pre-training a large language model on Google Cloud. This model
includes custom TensorFlow operations in the training loop. Model training will use a large batch
size, and you expect training to take several weeks. You need to configure a training architecture
that minimizes both training time and compute costs. What should you do?
Question 207
Question 213: You are building a TensorFlow text-to-image generative model by
using a dataset that contains billions of images with their respective captions. You want to create
a low maintenance, automated workflow that reads the data from a Cloud Storage bucket collects
statistics, splits the dataset into training/validation/test datasets performs data transformations
trains the model using the training/validation datasets, and validates the model by using the test
dataset. What should you do?
• A. Use the Apache Airflow SDK to create multiple operators that use Dataflow and Vertex AI
services. Deploy the workflow on Cloud Composer.
• B. Use the MLFlow SDK and deploy it on a Google Kubernetes Engine cluster. Create multiple
components that use Dataflow and Vertex AI services.
• C. Use the Kubeflow Pipelines (KFP) SDK to create multiple components that use Dataflow and Vertex
AI services. Deploy the workflow on Vertex AI Pipelines.
• D. Use the TensorFlow Extended (TFX) SDK to create multiple components that use Dataflow and
Vertex AI services. Deploy the workflow on Vertex AI Pipelines.
Question 208
Question 214: You are developing an ML pipeline using Vertex AI Pipelines.
You want your pipeline to upload a new version of the XGBoost model to Vertex AI Model Registry and
deploy it to Vertex AI Endpoints for online inference. You want to use the simplest approach. What
should you do? • A. Use the Vertex AI REST API within a custom component based on a
vertex-ai/prediction/xgboost-cpu image • B. Use the Vertex AI ModelEvaluationOp component to
evaluate the model • C. Use the Vertex AI SDK for Python within a custom component based on a
python:3.10 image • D. Chain the Vertex AI ModelUploadOp and ModelDeployOp components together
Question 209
You work for an online retailer. Your company has a few thousand short
lifecycle products. Your company has five years of sales data stored in BigQuery. You have been
asked to build a model that will make monthly sales predictions for each product. You want to use a
solution that can be implemented quickly with minimal effort. What should you do?
Question 210
You are creating a model training pipeline to predict sentiment scores from
text-based product reviews. You want to have control over how the model parameters are tuned, and
you will deploy the model to an endpoint after it has been trained. You will use Vertex AI Pipelines
to run the pipeline. You need to decide which Google Cloud pipeline components to use. What
components should you choose?
Question 211
Question 217: Your team frequently creates new ML models and runs
experiments. Your team pushes code to a single repository hosted on Cloud Source Repositories. You
want to create a continuous integration pipeline that automatically retrains the models whenever
there is any modification of the code. What should be your first step to set up the CI pipeline?
Question 212
You have built a custom model that performs several memory-intensive
preprocessing tasks before it makes a prediction. You deployed the model to a Vertex AI endpoint,
and validated that results were received in a reasonable amount of time. After routing user traffic
to the endpoint, you discover that the endpoint does not autoscale as expected when receiving
multiple requests. What should you do?
Question 213
Question 219: Your company manages an ecommerce website. You developed an ML
model that recommends additional products to users in near real time based on items currently in the
user’s cart. The workflow will include the following processes:
1. The website will send a Pub/Sub message with the relevant data and then receive a message with
the prediction from Pub/Sub
2. Predictions will be stored in BigQuery
3. The model will be stored in a Cloud Storage bucket and will be updated frequently
You want to minimize prediction latency and the effort required to update the model. How should you
reconfigure the architecture?
• A. Write a Cloud Function that loads the model into memory for prediction. Configure the function
to be triggered when messages are sent to Pub/Sub.
• B. Create a pipeline in Vertex AI Pipelines that performs preprocessing, prediction, and
postprocessing. Configure the pipeline to be triggered by a Cloud Function when messages are sent to
Pub/Sub.
• C. Expose the model as a Vertex AI endpoint. Write a custom DoFn in a Dataflow job that calls the
endpoint for prediction.
• D. Use the RunInference API with WatchFilePattern in a Dataflow job that wraps around the model
and serves predictions.
Question 214
Question 220: You are collaborating on a model prototype with your team. You
need to create a Vertex AI Workbench environment for the members of your team and also limit access
to other employees in your project. What should you do?
Question 215
Question 221: You work at a leading healthcare firm developing
state-of-the-art algorithms for various use cases. You have unstructured textual data with custom
labels. You need to extract and classify various medical phrases with these labels. What should you
do?
Question 216
Question 222: You developed a custom model by using Vertex AI to predict your
application's user churn rate. You are using Vertex AI Model Monitoring for skew detection. The
training data stored in BigQuery contains two sets of features - demographic and behavioral. You
later discover that two separate models trained on each set perform better than the original model.
You need to configure a new model monitoring pipeline that splits traffic among the two models. You
want to use the same prediction-sampling-rate and monitoring-frequency for each model. You also want
to minimize management effort. What should you do?
Question 217
Question 223: You work for a pharmaceutical company based in Canada. Your
team developed a BigQuery ML model to predict the number of flu infections for the next month in
Canada. Weather data is published weekly, and flu infection statistics are published monthly. You
need to configure a model retraining policy that minimizes cost. What should you do?
• A. Download the weather and flu data each week. Configure Cloud Scheduler to execute a Vertex AI
pipeline to retrain the model weekly.
• B. Download the weather and flu data each month. Configure Cloud Scheduler to execute a Vertex AI
pipeline to retrain the model monthly.
• C. Download the weather and flu data each week. Configure Cloud Scheduler to execute a Vertex AI
pipeline to retrain the model every month.
• D. Download the weather data each week, and download the flu data each month. Deploy the model to
a Vertex AI endpoint with feature drift monitoring, and retrain the model if a monitoring alert is
detected.
Question 218
You are building a MLOps platform to automate your company’s ML experiments
and model retraining. You need to organize the artifacts for dozens of pipelines. How should you
store the pipelines’ artifacts?
Question 219
Question 225: You work for a telecommunications company. You’re building a
model to predict which customers may fail to pay their next phone bill. The purpose of this model is
to proactively offer at-risk customers assistance such as service discounts and bill deadline
extensions. The data is stored in BigQuery and the predictive features that are available for model
training include: - Customer_id - Age - Salary (measured in local currency) - Sex - Average bill
value (measured in local currency) - Number of phone calls in the last month (integer) - Average
duration of phone calls (measured in minutes) You need to investigate and mitigate potential bias
against disadvantaged groups, while preserving model accuracy. What should you do? • A. Determine
whether there is a meaningful correlation between the sensitive features and the other features.
Train a BigQuery ML boosted trees classification model and exclude the sensitive features and any
meaningfully correlated features. • B. Train a BigQuery (corrected from 'tram' to
'train') BigQuery ML boosted trees classification model with all features. Use the
ML.GLOBAL_EXPLAIN method to calculate the global attribution values for each feature of the model.
If the feature importance value for any of the sensitive features exceeds a threshold, discard the
model and train without this feature. • C. Train a BigQuery ML boosted trees classification model
with all features. Use the ML.EXPLAIN_PREDICT method to calculate the attribution values for each
feature for each customer in a test set. If for any individual customer, the importance value for
any feature exceeds a predefined threshold, discard the model and train the model again without this
feature. • D. Define a fairness metric that is represented by accuracy across the sensitive
features. Train a BigQuery ML boosted trees classification model with all features. Use the trained
model to make predictions on a test set. Join the data back with the sensitive features, and
calculate a fairness metric to investigate whether it meets your requirements.
Question 220
You recently trained a XGBoost model that you plan to deploy to production
for online inference. Before sending a predict request to your model’s binary, you need to perform a
simple data preprocessing step. This step exposes a REST API that accepts requests in your internal
VPC Service Controls and returns predictions. You want to configure this preprocessing step while
minimizing cost and effort. What should you do?
Question 221
You work at a bank. You need to develop a credit risk model to support loan
application decisions. You decide to implement the model by using a neural network in TensorFlow.
Due to regulatory requirements, you need to be able to explain the model’s predictions based on its
features. When the model is deployed, you also want to monitor the model’s performance over time.
You decided to use Vertex AI for both model development and deployment. What should you do?
Question 222
Question 228: You are investigating the root cause of a misclassification
error made by one of your models. You used Vertex AI Pipelines to train and deploy the model. The
pipeline reads data from BigQuery, creates a copy of the data in Cloud Storage in TFRecord format,
trains the model in Vertex AI Training on that copy, and deploys the model to a Vertex AI endpoint.
You have identified the specific version of that model that misclassified, and you need to recover
the data this model was trained on. How should you find that copy of the data?
Question 223
You work for a manufacturing company. You need to train a custom image
classification model to detect product defects at the end of an assembly line. Although your model
is performing well, some images in your holdout set are consistently mislabeled with high
confidence. You want to use Vertex AI to understand your model’s results. What should you do?
Question 224
Question 230:
ou are training models in Vertex AI by using data that spans across multiple Google Cloud projects.
You need to find, track, and compare the performance of the different versions of your models.
Which Google Cloud services should you include in your ML workflow?
Question 225
You are using Keras and TensorFlow to develop a fraud detection model.
Records of customer transactions are stored in a large table in BigQuery. You need to preprocess
these records in a cost-effective and efficient way before you use them to train the model. The
trained model will be used to perform batch inference in BigQuery. How should you implement the
preprocessing workflow?
Question 226
You need to use TensorFlow to train an image classification model. Your
dataset is located in a Cloud Storage directory and contains millions of labeled images. Before
training the model, you need to prepare the data. You want the data preprocessing and model training
workflow to be as efficient, scalable, and low maintenance as possible. What should you do?
Question 227
You are building a custom image classification model and plan to use Vertex
AI Pipelines to implement the end-to-end training. Your dataset consists of images that need to be
preprocessed before they can be used to train the model. The preprocessing steps include resizing
the images, converting them to grayscale, and extracting features. You have already implemented some
Python functions for the preprocessing tasks. Which components should you use in your pipeline?
Question 228
Question 234: You work for a retail company that is using a regression model
built with BigQuery ML to predict product sales. This model is being used to serve online
predictions. Recently you developed a new version of the model that uses a different architecture
(custom model). Initial analysis revealed that both models are performing as expected. You want to
deploy the new version of the model to production and monitor the performance over the next two
months. You need to minimize the impact to the existing and future model users. How should you
deploy the model?
Question 229
You are using Vertex AI and TensorFlow to develop a custom image
classification model. You need the model’s decisions and the rationale to be understandable to your
company’s stakeholders. You also want to explore the results to identify any issues or potential
biases. What should you do?
Question 230
Question 236: You work for a large retailer, and you need to build a model to
predict customer chum. The company has a dataset of historical customer data, including customer
demographics purchase history, and website activity. You need to create the model in BigQuery ML and
thoroughly evaluate its performance. What should you do?
• A. Create a linear regression model in BigQuery ML, and register the model in Vertex AI Model
Registry. Evaluate the model performance in Vertex AI .
• B. Create a logistic regression model in BigQuery ML and register the model in Vertex AI Model
Registry. Evaluate the model performance in Vertex AI .
• C. Create a linear regression model in BigQuery ML. Use the ML.EVALUATE function to evaluate the
model performance.
• D. Create a logistic regression model in BigQuery ML. Use the ML.CONFUSION_MATRIX function to
evaluate the model performance.
Question 231
You are developing a model to identify traffic signs in images extracted from
videos taken from the dashboard of a vehicle. You have a dataset of 100,000 images that were cropped
to show one out of ten different traffic signs. The images have been labeled accordingly for model
training, and are stored in a Cloud Storage bucket. You need to be able to tune the model during
each training run. How should you train the model?
Question 232
Question 238: You have deployed a scikit-team model to a Vertex AI endpoint
using a custom model server. You enabled autoscaling: however, the deployed model fails to scale
beyond one replica, which led to dropped requests. You notice that CPU utilization remains low even
during periods of high load. What should you do?
Question 233
Question 239: You work for a pet food company that manages an online forum.
Customers upload photos of their pets on the forum to share with others. About 20 photos are
uploaded daily. You want to automatically and in near real time detect whether each uploaded photo
has an animal. You want to prioritize time and minimize cost of your application development and
deployment. What should you do?
• A. Send user-submitted images to the Cloud Vision API. Use object localization to identify all
objects in the image and compare the results against a list of animals.
• B. Download an object detection model from TensorFlow Hub. Deploy the model to a Vertex AI
endpoint. Send new user-submitted images to the model endpoint to classify whether each photo has an
animal.
• C. Manually label previously submitted images with bounding boxes around any animals. Build an
AutoML object detection model by using Vertex AI. Deploy the model to a Vertex AI endpoint Send new
user-submitted images to your model endpoint to detect whether each photo has an animal.
• D. Manually label previously submitted images as having animals or not. Create an image dataset on
Vertex AI. Train a classification model by using Vertex AutoML to distinguish the two classes.
Deploy the model to a Vertex AI endpoint. Send new user-submitted images to your model endpoint to
classify whether each photo has an animal.
Question 234
You work at a mobile gaming startup that creates online multiplayer games.
Recently, your company observed an increase in players cheating in the games, leading to a loss of
revenue and a poor user experience. You built a binary classification model to determine whether a
player cheated after a completed game session, and then send a message to other downstream systems
to ban the player that cheated. Your model has performed well during testing, and you now need to
deploy the model to production. You want your serving solution to provide immediate classifications
after a completed game session to avoid further loss of revenue. What should you do?
Question 235
Question 241: You have created a Vertex AI pipeline that automates custom
model training. You want to add a pipeline component that enables your team to most easily
collaborate when running different executions and comparing metrics both visually and
programmatically. What should you do?
• A. Add a component to the Vertex AI pipeline that logs metrics to a BigQuery table. Query the
table to compare different executions of the pipeline. Connect BigQuery to Looker Studio to
visualize metrics.
• B. Add a component to the Vertex AI pipeline that logs metrics to a BigQuery table. Load the table
into a pandas DataFrame to compare different executions of the pipeline. Use Matplotlib to visualize
metrics.
• C. Add a component to the Vertex AI pipeline that logs metrics to Vertex ML Metadata. Use Vertex
AI Experiments to compare different executions of the pipeline. Use Vertex AI TensorBoard to
visualize metrics.
• D. Add a component to the Vertex AI pipeline that logs metrics to Vertex ML Metadata. Load the
Vertex ML Metadata into a pandas DataFrame to compare different executions of the pipeline. Use
Matplotlib to visualize metrics.
Question 236
Question 242: Your team is training a large number of ML models that use
different algorithms, parameters, and datasets. Some models are trained in Vertex AI Pipelines, and
some are trained on Vertex AI Workbench notebook instances. Your team wants to compare the
performance of the models across both services. You want to minimize the effort required to store
the parameters and metrics. What should you do?
Question 237
Question 243: You work on a team that builds state-of-the-art deep learning
models by using the TensorFlow framework. Your team runs multiple ML experiments each week, which
makes it difficult to track the experiment runs. You want a simple approach to effectively track,
visualize, and debug ML experiment runs on Google Cloud while minimizing any overhead code. How
should you proceed?
• A. Set up Vertex AI Experiments to track metrics and parameters. Configure Vertex AI TensorBoard
for visualization.
• B. Set up a Cloud Function to write and save metrics files to a Cloud Storage bucket. Configure a
Google Cloud VM to host TensorBoard locally for visualization.
• C. Set up a Vertex AI Workbench notebook instance. Use the instance to save metrics data in a
Cloud Storage bucket and to host TensorBoard locally for visualization.
• D. Set up a Cloud Function to write and save metrics files to a BigQuery table. Configure a Google
Cloud VM to host TensorBoard locally for visualization.
Question 238
Your work for a textile manufacturing company. Your company has hundreds of
machines, and each machine has many sensors. Your team used the sensory data to build hundreds of ML
models that detect machine anomalies. Models are retrained daily, and you need to deploy these
models in a cost-effective way. The models must operate 24/7 without downtime and make sub
millisecond predictions. What should you do?
Question 239
You are developing an ML model that predicts the cost of used automobiles
based on data such as location, condition, model type, color, and engine/battery efficiency. The
data is updated every night. Car dealerships will use the model to determine appropriate car prices.
You created a Vertex AI pipeline that reads the data splits the data into training/evaluation/test
sets performs feature engineering trains the model by using the training dataset and validates the
model by using the evaluation dataset. You need to configure a retraining workflow that minimizes
cost. What should you do?
Question 240
Question 246: You recently used BigQuery ML to train an AutoML regression
model. You shared results with your team and received positive feedback. You need to deploy your
model for online prediction as quickly as possible. What should you do?
Question 241
You recently used BigQuery ML to train an AutoML regression model. You shared
results with your team and received positive feedback. You need to deploy your model for online
prediction as quickly as possible. What should you do?
Question 242
You trained a model packaged it with a custom Docker container for serving,
and deployed it to Vertex AI Model Registry. When you submit a batch prediction job, it fails with
this error: "Error model server never became ready. Please validate that your model file or
container configuration are valid." There are no additional errors in the logs. What should you
do?
Question 243
You are developing an ML model to identify your company’s products in images.
You have access to over one million images in a Cloud Storage bucket. You plan to experiment with
different TensorFlow models by using Vertex AI Training. You need to read images at scale during
training while minimizing data I/O bottlenecks. What should you do?
Question 244
Question 250: You work at an ecommerce startup. You need to create a customer
churn prediction model. Your company’s recent sales records are stored in a BigQuery table. You want
to understand how your initial model is making predictions. You also want to iterate on the model as
quickly as possible while minimizing cost. How should you build your first model? • A. Export the
data to a Cloud Storage bucket. Load the data into a pandas DataFrame on Vertex AI Workbench and
train a logistic regression model with scikit-learn. • B. Create a tf.data.Dataset by using the
TensorFlow BigQueryClient. Implement a deep neural network in TensorFlow. • C. Prepare the data in
BigQuery and associate the data with a Vertex AI dataset. Create an AutoMLTabularTrainingJob to tram
a classification model. • D. Export the data to a Cloud Storage bucket. Create a tf.data.Dataset to
read the data from Cloud Storage. Implement a deep neural network in TensorFlow.
Question 245
You work at an ecommerce startup. You need to create a customer churn
prediction model. Your company’s recent sales records are stored in a BigQuery table. You want to
understand how your initial model is making predictions. You also want to iterate on the model as
quickly as possible while minimizing cost. How should you build your first model?
Question 246
Question 252: You are developing a training pipeline for a new XGBoost
classification model based on tabular data. The data is stored in a BigQuery table. You need to
complete the following steps: 1. Randomly split the data into training and evaluation datasets in a
65/35 ratio 2. Conduct feature engineering 3. Obtain metrics for the evaluation dataset 4. Compare
models trained in different pipeline executions How should you execute these steps?
Question 247
You work for a company that sells corporate electronic products to thousands
of businesses worldwide. Your company stores historical customer data in BigQuery. You need to build
a model that predicts customer lifetime value over the next three years. You want to use the
simplest approach to build the model and you want to have access to visualization tools. What should
you do?
Question 248
You work for a delivery company. You need to design a system that stores and
manages features such as parcels delivered and truck locations over time. The system must retrieve
the features with low latency and feed those features into a model for online prediction. The data
science team will retrieve historical data at a specific point in time for model training. You want
to store the features with minimal effort. What should you do?
Question 249
You are working on a prototype of a text classification model in a managed
Vertex AI Workbench notebook. You want to quickly experiment with tokenizing text by using a Natural
Language Toolkit (NLTK) library. How should you add the library to your Jupyter kernel?
Question 250
Question 255: You have recently used TensorFlow to train a classification
model on tabular data. You have created a Dataflow pipeline that can transform several terabytes of
data into training or prediction datasets consisting of TFRecords. You now need to productionize the
model, and you want the predictions to be automatically uploaded to a BigQuery table on a weekly
schedule. What should you do?
Question 251
You work for an online grocery store. You recently developed a custom ML
model that recommends a recipe when a user arrives at the website. You chose the machine type on the
Vertex AI endpoint to optimize costs by using the queries per second (QPS) that the model can serve,
and you deployed it on a single machine with 8 vCPUs and no accelerators. A holiday season is
approaching and you anticipate four times more traffic during this time than the typical daily
traffic. You need to ensure that the model can scale efficiently to the increased demand. What
should you do?
Question 252
Question 257: You recently trained an XGBoost model on tabular data. You plan
to expose the model for internal use as an HTTP microservice. After deployment, you expect a small
number of incoming requests. You want to productionize the model with the least amount of effort and
latency. What should you do?
Question 253
Question 258: You work for an international manufacturing organization that
ships scientific products all over the world. Instruction manuals for these products need to be
translated to 15 different languages. Your organization’s leadership team wants to start using
machine learning to reduce the cost of manual human translations and increase translation speed. You
need to implement a scalable solution that maximizes accuracy and minimizes operational overhead.
You also want to include a process to evaluate and fix incorrect translations. What should you do?
• A. Create a workflow using Cloud Function triggers. Configure a Cloud Function that is triggered
when documents are uploaded to an input Cloud Storage bucket. Configure another Cloud Function that
translates the documents using the Cloud Translation API, and saves the translations to an output
Cloud Storage bucket. Use human reviewers to evaluate the incorrect translations.
• B. Create a Vertex AI pipeline that processes the documents launches, an AutoML Translation
training job, evaluates the translations and deploys the model to a Vertex AI endpoint with
autoscaling and model monitoring. When there is a predetermined skew between training and live data,
re-trigger the pipeline with the latest data.
• C. Use AutoML Translation to train a model. Configure a Translation Hub project, and use the
trained model to translate the documents. Use human reviewers to evaluate the incorrect
translations.
• D. Use Vertex AI custom training jobs to fine-tune a state-of-the-art open source pretrained model
with your data. Deploy the model to a Vertex AI endpoint with autoscaling and model monitoring. When
there is a predetermined skew between the training and live data, configure a trigger to run another
training job with the latest data.
Question 254
You have developed an application that uses a chain of multiple scikit-learn
models to predict the optimal price for your company’s products. The workflow logic is shown in the
diagram. Members of your team use the individual models in other solution workflows. You want to
deploy this workflow while ensuring version control for each individual model and the overall
workflow. Your application needs to be able to scale down to zero. You want to minimize the compute
resource utilization and the manual effort required to manage this solution. What should you do?
Question 255
Question 260: You are developing a model to predict whether a failure will
occur in a critical machine part. You have a dataset consisting of a multivariate time series and
labels indicating whether the machine part failed. You recently started experimenting with a few
different preprocessing and modeling approaches in a Vertex AI Workbench notebook. You want to log
data and track artifacts from each run. How should you set up your experiments?
• A. 1. Use the Vertex AI SDK to create an experiment and set up Vertex ML Metadata.
2. Use the log_time_series_metrics function to track the preprocessed data, and use the log_merrics
function to log loss values.
• B. 1. Use the Vertex AI SDK to create an experiment and set up Vertex ML Metadata.
2. Use the log_time_series_metrics function to track the preprocessed data, and use the log_metrics
function to log loss values.
• C. 1. Create a Vertex AI TensorBoard instance and use the Vertex AI SDK to create an experiment
and associate the TensorBoard instance.
2. Use the assign_input_artifact method to track the preprocessed data and use the
log_time_series_metrics function to log loss values.
• D. 1. Create a Vertex AI TensorBoard instance, and use the Vertex AI SDK to create an experiment
and associate the TensorBoard instance.
2. Use the log_time (time series metrics) function to track the preprocessed data, and use the
log_metrics function to log loss values.
Question 256
Question 261: You are developing a recommendation engine for an online
clothing store. The historical customer transaction data is stored in BigQuery and Cloud Storage.
You need to perform exploratory data analysis (EDA), preprocessing and model training. You plan to
rerun these EDA, preprocessing, and training steps as you experiment with different types of
algorithms. You want to minimize the cost and development effort of running these steps as you
experiment. How should you configure the environment?
• A. Create a Vertex AI Workbench user-managed notebook using the default VM instance, and use the
%%bigquery magic commands in Jupyter to query the tables.
• B. Create a Vertex AI Workbench managed notebook to browse and query the tables directly from the
JupyterLab interface.
• C. Create a Vertex AI Workbench user-managed notebook on a Dataproc Hub, and use the %%bigquery
magic commands in Jupyter to query the tables.
• D. Create a Vertex AI Workbench managed notebook on a Dataproc cluster, and use the
spark-bigquery-connector to access the tables.
Question 257
Question 262: You recently deployed a model to a Vertex AI endpoint and set
up online serving in Vertex AI Feature Store. You have configured a daily batch ingestion job to
update your featurestore. During the batch ingestion jobs, you discover that CPU utilization is high
in your featurestore’s online serving nodes and that feature retrieval latency is high. You need to
improve online serving performance during the daily batch ingestion. What should you do?
Question 258
You are developing a custom TensorFlow classification model based on tabular
data. Your raw data is stored in BigQuery. contains hundreds of millions of rows, and includes both
categorical and numerical features. You need to use a MaxMin scaler on some numerical features, and
apply a one-hot encoding to some categorical features such as SKU names. Your model will be trained
over multiple epochs. You want to minimize the effort and cost of your solution. What should you do?
Question 259
You work for a retail company. You have been tasked with building a model to
determine the probability of churn for each customer. You need the predictions to be interpretable
so the results can be used to develop marketing campaigns that target at-risk customers. What should
you do?
Question 260
Question 265: You work for a company that is developing an application to
help users with meal planning. You want to use machine learning to scan a corpus of recipes and
extract each ingredient (e.g., carrot, rice, pasta) and each kitchen cookware (e.g., bowl, pot,
spoon) mentioned. Each recipe is saved in an unstructured text file. What should you do?
• A. Create a text dataset on Vertex AI for entity extraction Create two entities called
“ingredient” and “cookware”, and label at least 200 examples of each entity. Train an AutoML entity
extraction model to extract occurrences of these entity types. Evaluate performance on a holdout
dataset.
• B. Create a multi-label text classification dataset on Vertex AI. Create a test dataset, and label
each recipe that corresponds to its ingredients and cookware. Train a multi-class classification
model. Evaluate the model’s performance on a holdout dataset.
• C. Use the Entity Analysis method of the Natural Language API to extract the ingredients and
cookware from each recipe. Evaluate the model's performance on a prelabeled dataset.
• D. Create a text dataset on Vertex AI for entity extraction. Create as many entities as there are
different ingredients and cookware. Train an AutoML entity extraction model to extract those
entities. Evaluate the model’s performance on a holdout dataset.
Question 261
You work for an organization that operates a streaming music service. You
have a custom production model that is serving a 'next song' recommendation based on a
user's recent listening history. Your model is deployed on a Vertex AI endpoint. You recently
retrained the same model by using fresh data. The model received positive test results offline. You
now want to test the new model in production while minimizing complexity. What should you do?
Question 262
Question 267: You created a model that uses BigQuery ML to perform linear
regression. You need to retrain the model on the cumulative data collected every week. You want to
minimize the development effort and the scheduling cost. What should you do?
• A. Use BigQuery’s scheduling service to run the model retraining query periodically.
• B. Create a pipeline in Vertex AI Pipelines that executes the retraining query, and use the Cloud
Scheduler API to run the query weekly.
• C. Use Cloud Scheduler to trigger a Cloud Function every week that runs the query for retraining
the model.
• D. Use the BigQuery API Connector and Cloud Scheduler to trigger Workflows every week that
retrains the model.
Question 263
You want to migrate a scikit-learn classifier model to TensorFlow. You plan
to train the TensorFlow classifier model using the same training set that was used to train the
scikit-learn model, and then compare the performances using a common test set. You want to use the
Vertex AI Python SDK to manually log the evaluation metrics of each model and compare them based on
their F1 scores and confusion matrices. How should you log the metrics?
Question 264
Question 269: You are developing a model to help your company create more
targeted online advertising campaigns. You need to create a dataset that you will use to train the
model. You want to avoid creating or reinforcing unfair bias in the model. What should you do?
(Choose two.) • A. Include a comprehensive set of demographic features • B. Include only the
demographic groups that most frequently interact with advertisements • C. Collect a random sample of
production traffic to build the training dataset • D. Collect a stratified sample of production
traffic to build the training dataset • E. Conduct fairness tests across sensitive categories and
demographics on the trained model
Question 265
Question 270: You are developing an ML model in a Vertex AI Workbench
notebook. You want to track artifacts and compare models during experimentation using different
approaches. You need to rapidly and easily transition successful experiments to production as you
iterate on your model implementation. What should you do?
• A. 1. Initialize the Vertex SDK with the name of your experiment. Log parameters and metrics for
each experiment, and attach dataset and model artifacts as inputs and outputs to each execution.
2. After a successful experiment create a Vertex AI pipeline.
• B. 1. Initialize the Vertex SDK with the name of your experiment. Log parameters and metrics for
each experiment, save your dataset to a Cloud Storage bucket, and upload the models to Vertex AI
Model Registry.
2. After a successful experiment, create a Vertex AI pipeline.
• C. 1. Create a Vertex AI pipeline with parameters you want to track as arguments to your
PipelineJob. Use the Metrics, Model, and Dataset artifact types from the Kubeflow Pipelines DSL as
the inputs and outputs of the components in your pipeline.
2. Associate the pipeline with your experiment when you submit the job.
• D. 1. Create a Vertex AI pipeline. Use the Dataset and Model artifact types from the Kubeflow
Pipelines DSL as the inputs and (outputs of the components in your pipeline.
2. In your training component, use the Vertex AI SDK to create an experiment run. Configure the
log_params and log_metrics functions to track parameters and metrics of your experiment.
Question 266
Question 271: You recently created a new Google Cloud project. After testing
that you can submit a Vertex AI Pipeline job from the Cloud Shell, you want to use a Vertex AI
Workbench user-managed notebook instance to run your code from that instance. You created the
instance and ran the code but this time the job fails with an insufficient permissions error. What
should you do?
Question 267
Question 272: You work for a semiconductor manufacturing company. You need to
create a real-time application that automates the quality control process. High-definition images of
each semiconductor are taken at the end of the assembly line in real time. The photos are uploaded
to a Cloud Storage bucket along with tabular data that includes each semiconductor’s batch number,
serial number, dimensions, and weight. You need to configure model training and serving while
maximizing model accuracy. What should you do?
Question 268
Question 273: You work for a rapidly growing social media company. Your team
builds TensorFlow recommender models in an on-premises CPU cluster. The data contains billions of
historical user events and 100,000 categorical features. You notice that as the data increases, the
model training time increases. You plan to move the models to Google Cloud. You want to use the most
scalable approach that also minimizes training time. What should you do?
Question 269
You are training and deploying updated versions of a regression model with
tabular data by using Vertex AI Pipelines, Vertex AI Training, Vertex AI Experiments, and Vertex AI
Endpoints. The model is deployed in a Vertex AI endpoint, and your users call the model by using the
Vertex AI endpoint. You want to receive an email when the feature data distribution changes
significantly, so you can retrigger the training pipeline and deploy an updated version of your
model. What should you do?
Question 270
You have trained an XG.Boost model that you plan to deploy on Vertex AI for
online prediction. You are now uploading your model to Vertex AI Model Registry, and you need to
configure the explanation method that will serve online prediction requests to be returned with
minimal latency. You also want to be alerted when feature attributions of the model meaningfully
change over time. What should you do?
Question 271
Question 276: You work at a gaming startup that has several terabytes of
structured data in Cloud Storage. This data includes gameplay time data, user metadata, and game
metadata. You want to build a model that recommends new games to users that requires the least
amount of coding. What should you do?
Question 272
You work for a large bank that serves customers through an application hosted
in Google Cloud that is running in the US and Singapore. You have developed a PyTorch model to
classify transactions as potentially fraudulent or not. The model is a three-layer perceptron that
uses both numerical and categorical features as input, and hashing happens within the model. You
deployed the model to the us-central1 region on nl-highcpu-16 machines, and predictions are served
in real time. The model's current median response latency is 40 ms. You want to reduce latency,
especially in Singapore, where some customers are experiencing the longest delays. What should you
do?
Question 273
You need to train an XGBoost model on a small dataset. Your training code
requires custom dependencies. You want to minimize the startup time of your training job. How should
you set up your Vertex AI custom training job?
Question 274
You are creating an ML pipeline for data processing, model training, and
model deployment that uses different Google Cloud services. You have developed code for each
individual task, and you expect a high frequency of new files. You now need to create an
orchestration layer on top of these tasks. You only want this orchestration pipeline to run if new
files are present in your dataset in a Cloud Storage bucket. You also want to minimize the compute
node costs. What should you do?
Question 275
You are using Kubeflow Pipelines to develop an end-to-end PyTorch-based MLOps
pipeline. The pipeline reads data from BigQuery, processes the data, conducts feature engineering,
model training, model evaluation, and deploys the model as a binary file to Cloud Storage. You are
writing code for several different versions of the feature engineering and model training steps, and
running each new version in Vertex AI Pipelines. Each pipeline run is taking over an hour to
complete. You want to speed up the pipeline execution to reduce your development time, and you want
to avoid additional costs. What should you do?
Question 276
You work at an organization that maintains a cloud-based communication
platform that integrates conventional chat, voice, and video conferencing into one platform. The
audio recordings are stored in Cloud Storage. All recordings have an 8 kHz sample rate and are more
than one minute long. You need to implement a new feature in the platform that will automatically
transcribe voice call recordings into text for future applications, such as call summarization and
sentiment analysis. How should you implement the voice call transcription feature following
Google-recommended best practices?
Question 277
Question 283: You work for a multinational organization that has recently
begun operations in Spain. Teams within your organization will need to work with various Spanish
documents, such as business, legal, and financial documents. You want to use machine learning to
help your organization get accurate translations quickly and with the least effort. Your
organization does not require domain-specific terms or jargon. What should you do?
Question 278
You have a custom job that runs on Vertex AI on a weekly basis. The job is
implemented using a proprietary ML workflow that produces the datasets, models, and custom
artifacts, and sends them to a Cloud Storage bucket. Many different versions of the datasets and
models were created. Due to compliance requirements, your company needs to track which model was
used for making a particular prediction, and needs access to the artifacts for each model. How
should you configure your workflows to meet these requirements?
Question 279
Question 285: You have recently developed a custom model for image
classification by using a neural network. You need to automatically identify the values for learning
rate, number of layers, and kernel size. To do this, you plan to run multiple jobs in parallel to
identify the parameters that optimize performance. You want to minimize custom code development and
infrastructure management. What should you do?
• A. Train an AutoML image classification model.
• B. Create a custom training job that uses the Vertex AI Vizier SDK for parameter optimization.
• C. Create a Vertex AI hyperparameter tuning job.
• D. Create a Vertex AI pipeline that runs different model training jobs in parallel.