Skip to content
BytePatterns

MLA-C02 · Domain 4: Operating, Monitoring, and Securing ML and AI Solutions · 24% of the exam

Task 4.2: Optimize and manage ML and AI infrastructure costs and performance.

Choosing instance families and purchasing options for training and inference, dashboards and tracing, budgets and cost tools, and keeping token, embedding and vector storage costs of foundation model applications under control.

Study it

  • Observability for generative AI and agents: CloudWatch, AgentCore Observability and X-Ray

    Partly covered by: CloudWatch, Alarms & X-Ray

  • Choosing instances and purchasing options for training and inference

    Partly covered by: AWS Cost Levers

  • Foundation model costs: tokens, caching, batch inference and vector storage

    Partly covered by: Context Windows, Tokenization

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

A company serves a deep learning image model on SageMaker AI with high, steady request volume, and inference cost is now its largest ML bill. The model is built with PyTorch, and the team can compile it with the AWS Neuron SDK. Which instance family should the ML engineer evaluate first to lower the cost per inference?

  1. ALarger general purpose instances with more vCPUs
  2. BBurstable instances, to pay less for idle time
  3. CInf2 instances, based on AWS Inferentia2
  4. DThe same instance type with SageMaker AI managed warm pools
Show the answer and why
  • ALarger general purpose instances with more vCPUs

    Incorrect

    More CPU capacity can raise throughput, but general purpose CPUs are not built for deep learning inference, so cost per inference rarely drops this way.

  • BBurstable instances, to pay less for idle time

    Incorrect

    Burstable instances provide a baseline CPU level with the ability to burst, for workloads with low-to-moderate CPU use. A high, steady inference volume needs sustained performance.

  • CInf2 instances, based on AWS Inferentia2

    Correct

    Inferentia chips are designed for high performance at the lowest cost in Amazon EC2 for deep learning and generative AI inference, and Inf2 instances are built on Inferentia2. The Neuron SDK integrates with PyTorch to deploy models on them.

  • DThe same instance type with SageMaker AI managed warm pools

    Incorrect

    Warm pools keep training infrastructure ready between training jobs. They do not apply to endpoint inference cost.

For steady, high-volume deep learning inference, purpose-built accelerators such as Inferentia are a main cost lever. Use load tests, for example with Inference Recommender, to confirm the choice.

Question 2 · choose 1

An ML platform team runs SageMaker AI notebooks, training jobs and endpoints at a fairly steady level that it expects to keep for at least a year. Teams regularly switch between ml.m5 and ml.c5 instances and between two Regions. Which purchase option lowers the bill while keeping that flexibility?

  1. AA Compute Savings Plan
  2. BA SageMaker AI Savings Plan
  3. CAn EC2 Instance Savings Plan for the m5 family
  4. DManaged Spot capacity for every workload
Show the answer and why
  • AA Compute Savings Plan

    Incorrect

    Compute Savings Plans apply to EC2 instance usage, Fargate and Lambda. They do not cover SageMaker AI instance usage.

  • BA SageMaker AI Savings Plan

    Correct

    SageMaker AI Savings Plans offer up to 64% off On-Demand rates and apply automatically to SageMaker AI usage regardless of instance family, size, Region and component, such as notebooks or training.

  • CAn EC2 Instance Savings Plan for the m5 family

    Incorrect

    EC2 Instance Savings Plans apply to one EC2 instance family in one Region. They do not cover SageMaker AI usage, and they would not follow a move to c5 or another Region.

  • DManaged Spot capacity for every workload

    Incorrect

    Managed Spot training can cut training cost, but Spot capacity can be interrupted, which does not suit endpoints that must stay up.

Commitment discounts are service-specific: SageMaker AI Savings Plans for SageMaker AI usage, Compute and EC2 Instance Savings Plans for EC2 and related compute.

Question 3 · choose 1

Data scientists experiment in a separate AWS account. Leadership wants spending in that account capped: when costs reach the monthly limit, new SageMaker AI resources should be blocked automatically until an administrator reviews the situation. What should the ML engineer configure?

  1. AAn AWS Budgets budget action that applies an IAM policy or SCP
  2. BCost allocation tags on every SageMaker AI resource
  3. CAWS Cost Explorer forecasts, reviewed by the team each week
  4. DAWS Compute Optimizer recommendations for SageMaker AI
Show the answer and why
  • AAn AWS Budgets budget action that applies an IAM policy or SCP

    Correct

    AWS Budgets can run an action when a budget exceeds a cost or usage threshold, such as applying an IAM policy or a service control policy that stops new resources from being provisioned.

  • BCost allocation tags on every SageMaker AI resource

    Incorrect

    Activated cost allocation tags organize costs in Cost Explorer and reports. They describe spending but do not stop it.

  • CAWS Cost Explorer forecasts, reviewed by the team each week

    Incorrect

    Cost Explorer shows and forecasts costs. A weekly review is manual and does not block anything when the limit is reached.

  • DAWS Compute Optimizer recommendations for SageMaker AI

    Incorrect

    Compute Optimizer recommends better-sized resources. It does not enforce a spending cap.

To turn a budget into a control rather than a notification, attach a budget action that applies an IAM policy or SCP when the threshold is crossed.

Question 4 · choose 2

A company runs two Amazon Bedrock workloads. An interactive assistant sends the same 6,000-token policy document at the start of every request, followed by a short user question. A nightly job classifies 2 million support tickets, and its results are needed by the next morning. Which changes will reduce inference cost for these workloads? (Choose TWO.)

  1. APrompt caching for the repeated policy document
  2. BThe Priority service tier for the nightly classification job
  3. CNo-commitment Provisioned Throughput for the assistant
  4. DMoving the policy document to the end of each prompt
  5. EA Bedrock batch inference job for the nightly classification
Show the answer and why
  • APrompt caching for the repeated policy document

    Correct

    Prompt caching lets supported models reuse long, repeated prompt prefixes across requests, which reduces input token cost and latency. Static content should come first in the prompt.

  • BThe Priority service tier for the nightly classification job

    Incorrect

    Priority delivers faster responses for a price premium. The nightly job does not need speed, so this raises cost.

  • CNo-commitment Provisioned Throughput for the assistant

    Incorrect

    Provisioned Throughput is billed hourly for the capacity you buy, whether or not requests arrive. It adds throughput; it does not lower the token cost of a repeated prompt.

  • DMoving the policy document to the end of each prompt

    Incorrect

    Caching works on prompt prefixes, so static content belongs at the beginning. Moving it to the end makes reuse less likely.

  • EA Bedrock batch inference job for the nightly classification

    Correct

    Batch inference processes many prompts asynchronously from S3, and Bedrock offers select models for batch inference at a 50% lower price than on-demand inference.

Token cost falls by sending less new input (prompt caching for repeated prefixes) and by paying less per token for work that can wait (batch inference).

Question 5 · choose 1

A team is about to deploy a new model to a SageMaker AI real-time endpoint. It must meet a p99 latency target at the lowest hosting cost, and the team does not know which instance type and count to choose. What is the most efficient way to decide?

  1. ARun an Amazon SageMaker Inference Recommender job on the model
  2. BDeploy on the largest GPU instance and scale down later
  3. CUse AWS Compute Optimizer recommendations before the first deployment
  4. DCheck AWS Cost Explorer for the cheapest instance used last month
Show the answer and why
  • ARun an Amazon SageMaker Inference Recommender job on the model

    Correct

    Inference Recommender automates load testing across SageMaker AI instance types and configurations and recommends the deployment that delivers the best performance at the lowest cost.

  • BDeploy on the largest GPU instance and scale down later

    Incorrect

    Starting with the largest instance guarantees high cost and still requires manual load testing to find the right size.

  • CUse AWS Compute Optimizer recommendations before the first deployment

    Incorrect

    Compute Optimizer bases its recommendations on observed utilization, so it has nothing to analyze for a model that has never run.

  • DCheck AWS Cost Explorer for the cheapest instance used last month

    Incorrect

    Past spending on other workloads says nothing about whether an instance meets this model's latency target.

Right-size before launch with benchmark data. Inference Recommender runs the load tests and compares instance types and configurations for you.

Question 6 · choose 1

A company serves a customized model in Amazon Bedrock through Provisioned Throughput with no commitment. Traffic is steady around the clock, and the company plans to keep the model in production for at least a year. How can it lower the hourly price without reducing capacity?

  1. ABuy a 6-month commitment term
  2. BDelete the Provisioned Throughput each night
  3. CCut the number of model units in half
  4. DKeep no commitment and add a budget alert
Show the answer and why
  • ABuy a 6-month commitment term

    Correct

    Provisioned Throughput is billed hourly, and the longer the commitment term (no commitment, 1 month or 6 months), the more discounted the hourly price.

  • BDelete the Provisioned Throughput each night

    Incorrect

    Traffic runs around the clock, so deleting the Provisioned Throughput each night would take the model offline while requests still arrive.

  • CCut the number of model units in half

    Incorrect

    Model units set the throughput level, so halving them reduces capacity.

  • DKeep no commitment and add a budget alert

    Incorrect

    A budget alert reports spend; it does not lower the hourly price.

For steady, long-lived Provisioned Throughput, a longer commitment term is the main price lever.

Question 7 · choose 1

A data lake keeps 400 TB of training data in S3 Standard. Some datasets are read every week and others not for months, and the pattern is unknown and changes over time. Reads must stay immediate, without a restore step. Which storage choice lowers cost with no ongoing effort?

  1. AS3 Glacier Deep Archive for all of the datasets
  2. BS3 Intelligent-Tiering without the archive tiers
  3. CS3 Intelligent-Tiering with Deep Archive Access on
  4. DS3 Standard with the files compressed
Show the answer and why
  • AS3 Glacier Deep Archive for all of the datasets

    Incorrect

    Deep Archive objects must be restored before they can be read, which breaks the need for immediate reads.

  • BS3 Intelligent-Tiering without the archive tiers

    Correct

    Intelligent-Tiering moves objects between frequent, infrequent and Archive Instant Access tiers based on access, without performance impact; the optional archive tiers are the ones that need a restore.

  • CS3 Intelligent-Tiering with Deep Archive Access on

    Incorrect

    Objects in the Deep Archive Access tier must be restored first, so turning it on breaks the need for immediate reads.

  • DS3 Standard with the files compressed

    Incorrect

    Compression can help, but data that is rarely read still stays in the most expensive storage class.

Use Intelligent-Tiering for unknown or changing access patterns, and turn on its archive tiers only when a restore delay is acceptable.

Question 8 · choose 1

A support assistant sends every request to the largest model in a model family, although most questions are simple. The team wants a single serverless endpoint that sends each request to the model in that family predicted to answer it well, to balance quality and cost, without building its own model-selection logic. What should the ML engineer use?

  1. AA cross-Region inference profile
  2. BDistillation into a new, smaller model
  3. CAmazon Bedrock intelligent prompt routing
  4. DProvisioned Throughput for the large model
Show the answer and why
  • AA cross-Region inference profile

    Incorrect

    Cross-Region inference spreads requests for one model across Regions; it does not choose between models.

  • BDistillation into a new, smaller model

    Incorrect

    Distillation trains a new model, and the team would still need its own logic to decide which model gets each request.

  • CAmazon Bedrock intelligent prompt routing

    Correct

    Intelligent prompt routing provides a single serverless endpoint that routes requests between models in the same family, predicting the response quality of each to optimize for quality and cost.

  • DProvisioned Throughput for the large model

    Incorrect

    Provisioned Throughput adds capacity for one model at a fixed cost; every request still goes to the large model.

Routing simple requests to smaller models in the family cuts cost while keeping quality where it is needed.

Question 9 · choose 1

An ML platform team runs SageMaker AI and Amazon Bedrock workloads in many accounts. Last month a forgotten tuning job ran up costs that were noticed only on the invoice. The team wants alerts on unusual spend, found by machine learning rather than fixed thresholds, with root causes by service, account and Region. What should the team use?

  1. AAWS Budgets with a fixed monthly amount
  2. BCloudWatch alarms on each job metric
  3. CCost allocation tags on all resources
  4. DAWS Cost Anomaly Detection
Show the answer and why
  • AAWS Budgets with a fixed monthly amount

    Incorrect

    Budgets alert against amounts you set, which is the fixed-threshold approach the team wants to avoid.

  • BCloudWatch alarms on each job metric

    Incorrect

    Job metrics show resource use, not spend, and per-job alarms would miss new resources.

  • CCost allocation tags on all resources

    Incorrect

    Tags organize costs in reports; they do not raise alerts on their own.

  • DAWS Cost Anomaly Detection

    Correct

    Cost Anomaly Detection uses machine learning models to detect and alert on anomalous spend, and ranks root causes by service, account, Region and usage type.

Combine budgets for planned limits with anomaly detection for surprises such as runaway jobs.

Question 10 · choose 1

A knowledge base stores 300 million chunks embedded with Amazon Titan Text Embeddings V2 at the default 1,024 dimensions, and vector storage is now a major cost. Retrieval tests show that 256-dimension vectors keep accuracy acceptable for this use case. What should the ML engineer do?

  1. ARebuild the index with 256-dimension embeddings
  2. BCut the stored vectors down to their first 256 values
  3. CReturn more chunks for each query
  4. DStore the vectors as text in the metadata field
Show the answer and why
  • ARebuild the index with 256-dimension embeddings

    Correct

    Titan Text Embeddings V2 can output 1,024, 512 or 256 dimensions, and the vector field must match the model's output, so the data must be embedded again at 256 into a matching index.

  • BCut the stored vectors down to their first 256 values

    Incorrect

    The index is set up for the configured number of dimensions; the documented way to get smaller vectors is to request a smaller output size from the model.

  • CReturn more chunks for each query

    Incorrect

    The number of results changes retrieval, not the size of the stored vectors.

  • DStore the vectors as text in the metadata field

    Incorrect

    Vectors belong in the vector field; storing them as text would break similarity search.

Smaller embedding dimensions cut vector storage and search cost; test retrieval quality before you switch.

Practise domain 4 →Practise all domains →