Skip to content
BytePatterns

MLA-C02 · Domain 3: Deployment and Orchestration of ML and AI Workflows · 24% of the exam

Task 3.2: Provision and configure resources for ML and AI workloads based on existing architecture and requirements.

On-demand versus provisioned capacity, infrastructure as code, containers for training and inference, endpoints inside a VPC, deploying with the SageMaker Python SDK and Boto3, auto scaling metrics, Bedrock knowledge bases, and the infrastructure that agents run on.

Study it

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

A team hosts a large language model on a SageMaker AI real-time endpoint with GPU instances for internal testing. The endpoint is idle every night and weekend, and the team wants it to drop to zero instances then while keeping the same real-time invocation interface during working hours. Cold starts of several minutes are acceptable. How should the ML engineer configure this?

  1. AMove the model to serverless inference so it scales to zero automatically
  2. BSet the variant minimum capacity to 0 in a target tracking policy on InvocationsPerInstance
  3. CDeploy it as an inference component with min 0 instances and step scaling
  4. DConvert the endpoint to asynchronous inference
Show the answer and why
  • AMove the model to serverless inference so it scales to zero automatically

    Incorrect

    Serverless inference does scale to zero, but it does not support GPUs, which this model needs.

  • BSet the variant minimum capacity to 0 in a target tracking policy on InvocationsPerInstance

    Incorrect

    Scaling to zero is tied to inference components; a variant without them cannot scale in to zero instances. Once at zero, a scale-out also needs a step scaling policy and alarm.

  • CDeploy it as an inference component with min 0 instances and step scaling

    Correct

    An endpoint can scale in to and out from zero instances only if it hosts inference components, with MinInstanceCount set to 0. A step scaling policy tied to an alarm provisions an instance again when requests arrive.

  • DConvert the endpoint to asynchronous inference

    Incorrect

    Asynchronous inference can scale to zero, but callers must place payloads in S3 and use asynchronous invocation, which changes the real-time interface the team wants to keep.

Real-time endpoints can scale to zero when models are deployed as inference components. Expect errors during the several minutes it takes to provision the first instance again.

Question 2 · choose 1

A SageMaker AI endpoint serves a text generation model that streams tokens back to callers. Requests vary from one second to two minutes. Target tracking on InvocationsPerInstance reacts slowly, and during spikes requests queue up inside the containers before more instances arrive. Which metric should the ML engineer use for the scaling policy?

  1. AInvocationsPerInstance with a lower target value
  2. BCPUUtilization as a custom metric
  3. CModelLatency, with a scheduled scaling action each morning
  4. DConcurrentRequestsPerModel
Show the answer and why
  • AInvocationsPerInstance with a lower target value

    Incorrect

    A lower target scales out earlier, but the metric still counts invocations per minute, not how many long-running requests are in flight, and it is reported less often.

  • BCPUUtilization as a custom metric

    Incorrect

    On a GPU-bound generation model, CPU usage says little about how many requests are waiting.

  • CModelLatency, with a scheduled scaling action each morning

    Incorrect

    A fixed schedule cannot follow unpredictable spikes, and latency rises only after requests are already queuing.

  • DConcurrentRequestsPerModel

    Correct

    This high-resolution metric counts the simultaneous requests a model container is handling, including queued ones, and for streaming responses it tracks each request until the last token. It emits data more often than standard metrics, so scaling reacts faster.

For long or streaming requests, concurrency is a better measure of load than invocation counts. The high-resolution concurrency metrics also let auto scaling react sooner.

Question 3 · choose 2

A model hosted on a SageMaker AI real-time endpoint must read reference files from Amazon S3 at runtime and query an Amazon RDS database that sits in private subnets. Security rules say the model container must run inside the company VPC and have no route to the internet. Which steps should the ML engineer take? (Choose TWO.)

  1. ASet VpcConfig with private subnets in two Availability Zones and security groups
  2. BEnable network isolation on the model
  3. CAdd a route from the private subnets to an internet gateway
  4. DCreate an S3 VPC endpoint so the container reaches the bucket privately
  5. ECreate an interface VPC endpoint for the SageMaker AI API and Runtime
Show the answer and why
  • ASet VpcConfig with private subnets in two Availability Zones and security groups

    Correct

    VpcConfig makes SageMaker AI attach network interfaces in your subnets, giving the container a network connection inside the VPC that is not connected to the internet. At least two subnets in different Availability Zones are required.

  • BEnable network isolation on the model

    Incorrect

    With network isolation the container cannot make any outbound network calls, including to S3, so it could not read the files or query the database.

  • CAdd a route from the private subnets to an internet gateway

    Incorrect

    An internet route breaks the security rule. Private connectivity to S3 and RDS does not need the internet.

  • DCreate an S3 VPC endpoint so the container reaches the bucket privately

    Correct

    With no internet route, the container needs a VPC endpoint to reach S3. The hosting guide lists creating an S3 VPC endpoint, optionally with a restrictive endpoint policy.

  • ECreate an interface VPC endpoint for the SageMaker AI API and Runtime

    Incorrect

    An interface endpoint for the SageMaker AI API keeps API calls such as invoking the endpoint private. It does not give the container access to S3 objects.

VPC-attached endpoints reach private resources through the subnets and security groups you choose, and reach AWS services such as S3 through VPC endpoints. Network isolation is a different control that blocks outbound calls entirely.

Question 4 · choose 1

A team is deploying a research agent on Amazon Bedrock AgentCore Runtime. Some research sessions must keep working for several days, and the agent runs a small open-weight model that needs a GPU. Which configuration should the ML engineer choose?

  1. AThe Instances compute type for AgentCore Runtime
  2. BThe microVM compute type, with a longer session timeout
  3. CAgentCore Memory to keep the session running
  4. DAn AWS Lambda function with the maximum timeout
Show the answer and why
  • AThe Instances compute type for AgentCore Runtime

    Correct

    Instances run agents on AWS managed EC2 infrastructure in your account and support persistent, multi-day sessions of up to 14 days and GPU-accelerated workloads.

  • BThe microVM compute type, with a longer session timeout

    Incorrect

    MicroVM sessions are serverless and start instantly, but they run for up to 8 hours, not several days.

  • CAgentCore Memory to keep the session running

    Incorrect

    Memory stores short-term and long-term context for agents. It does not change how long a runtime session can run or add a GPU.

  • DAn AWS Lambda function with the maximum timeout

    Incorrect

    A Lambda function can run for at most 900 seconds (15 minutes), far short of a session that lasts several days.

AgentCore Runtime offers two compute types: microVMs for serverless, isolated sessions of up to 8 hours, and Instances for long sessions of up to 14 days and GPU workloads.

Question 5 · choose 1

A travel-booking agent runs on Amazon Bedrock AgentCore Runtime. Within a conversation it must remember what the user said a few turns earlier, and across conversations weeks apart it should recall preferences such as "always book aisle seats". The team wants the least custom infrastructure. What should the ML engineer use?

  1. AFiles written to the runtime session's local filesystem
  2. BA larger context window on the foundation model
  3. CAgentCore Observability traces of past sessions
  4. DAgentCore Memory, with short-term and long-term memory
Show the answer and why
  • AFiles written to the runtime session's local filesystem

    Incorrect

    Each session runs in its own microVM that is terminated, with memory sanitized, when the session ends, so nothing would remain for the next conversation.

  • BA larger context window on the foundation model

    Incorrect

    A larger window holds more of the current conversation, but it does not carry anything over to a conversation weeks later.

  • CAgentCore Observability traces of past sessions

    Incorrect

    Observability records traces, metrics and logs for debugging and monitoring. It is not a store the agent reads preferences from.

  • DAgentCore Memory, with short-term and long-term memory

    Correct

    AgentCore Memory is a managed service with short-term memory for turn-by-turn context in a session and long-term memory that extracts and keeps insights such as user preferences across sessions.

Agents are stateless by default. AgentCore Memory adds managed state at two levels: within a session and across sessions.

Question 6 · choose 1

A research team needs 16 instances of a high-demand GPU type for a two-week training run that starts on a fixed date next month. On-demand requests for that type often fail for lack of capacity. Which option gives the team predictable access to the capacity for that window?

  1. AA SageMaker training plan for those dates
  2. BManaged spot training with checkpoints enabled
  3. CA SageMaker AI Savings Plan for one year
  4. DWarm pools that keep instances after each job
Show the answer and why
  • AA SageMaker training plan for those dates

    Correct

    Training plans reserve GPU capacity as Reserved Capacity blocks defined by instance type, quantity, Availability Zone, duration and start and end times, giving predictable access within your timelines.

  • BManaged spot training with checkpoints enabled

    Incorrect

    Spot instances can be interrupted and jobs can wait for Spot capacity, so spot training does not secure capacity for set dates.

  • CA SageMaker AI Savings Plan for one year

    Incorrect

    SageMaker AI Savings Plans lower the price of SageMaker AI usage; they are a pricing commitment and do not secure instances for set dates.

  • DWarm pools that keep instances after each job

    Incorrect

    Warm pools retain infrastructure after a training job to speed up later jobs; they do not reserve capacity in advance.

For scarce accelerators on a known schedule, reserve the capacity ahead of time rather than relying on on-demand or spot availability.

Question 7 · choose 1

A distributed training job with 16 EFA-enabled instances must run in a small private subnet in the team's VPC. The job fails to start, and the subnet turns out to have very few free IP addresses. What should the ML engineer check against the subnet's free addresses?

  1. AOne public IP address per training instance
  2. BAt least 5 private IPs per EFA instance
  3. CA NAT gateway per training instance
  4. DAn Elastic IP for the training job
Show the answer and why
  • AOne public IP address per training instance

    Incorrect

    Training in private subnets does not need public IP addresses.

  • BAt least 5 private IPs per EFA instance

    Correct

    AWS guidance says training instances that use an Elastic Fabric Adapter should have at least 5 private IP addresses, and instances without EFA at least 2, so the subnet must have enough free addresses.

  • CA NAT gateway per training instance

    Incorrect

    NAT gateways provide outbound internet access, not private IP addresses for instances.

  • DAn Elastic IP for the training job

    Incorrect

    Elastic IPs are public addresses and do not solve a shortage of private addresses in the subnet.

Plan subnet sizes for ML jobs: multi-instance and EFA-enabled training uses several private IP addresses per instance.

Question 8 · choose 2

A team creates its own Amazon OpenSearch Serverless vector index to back an Amazon Bedrock knowledge base. Which settings must the index and the knowledge base field mapping include? (Choose TWO.)

  1. AThe faiss engine for the vector field
  2. BA text field and a Bedrock-managed metadata field
  3. CThe nmslib engine for the vector field
  4. DA partition key named after each source document
  5. EA DynamoDB table that holds the chunk text
Show the answer and why
  • AThe faiss engine for the vector field

    Correct

    The documentation says the vector engine used for search must be faiss.

  • BA text field and a Bedrock-managed metadata field

    Correct

    The field mapping includes a text field for the raw chunk text and a Bedrock-managed metadata field, alongside the vector field.

  • CThe nmslib engine for the vector field

    Incorrect

    The documentation states that nmslib is not supported.

  • DA partition key named after each source document

    Incorrect

    A partition key is not part of the vector store field mapping for a knowledge base.

  • EA DynamoDB table that holds the chunk text

    Incorrect

    Chunk text is stored in the vector store's text field, not in a separate DynamoDB table.

Bring-your-own vector stores need the right engine and a field mapping: a vector field, a text field and a metadata field.

Question 9 · choose 1

An agent on Amazon Bedrock AgentCore must create calendar events in a third-party SaaS application on behalf of each signed-in user, without the team storing user credentials in its own code. Which AgentCore service is designed for this?

  1. AAgentCore Memory
  2. BAgentCore Observability
  3. CAgentCore Identity
  4. DAgentCore Code Interpreter
Show the answer and why
  • AAgentCore Memory

    Incorrect

    Memory stores conversation context; it does not manage user credentials for third-party services.

  • BAgentCore Observability

    Incorrect

    Observability traces and monitors agents; it does not grant access to external services.

  • CAgentCore Identity

    Correct

    AgentCore Identity provides identity and credential management for agents, enabling them to access AWS resources and third-party services on behalf of users while keeping security controls and audit trails.

  • DAgentCore Code Interpreter

    Incorrect

    Code Interpreter runs code in a sandbox; it is not a credential management service.

Agents acting for users need delegated, auditable credentials; AgentCore Identity provides them so secrets stay out of agent code.

Question 10 · choose 1

A newly created account runs its first SageMaker AI training job on two large GPU instances. The job fails at once because the account's applied quota for that instance type for training job usage is lower than the requested instance count. What should the ML engineer do?

  1. AAttach a broader IAM policy to the execution role
  2. BAdd more private subnets to the VPC settings
  3. CRetry the job until enough capacity is free
  4. DRequest a quota increase in Service Quotas
Show the answer and why
  • AAttach a broader IAM policy to the execution role

    Incorrect

    The failure is a quota limit, not missing permissions; a wider role policy does not raise account quotas.

  • BAdd more private subnets to the VPC settings

    Incorrect

    Subnets supply IP addresses for the instances; they do not change the account's instance quota.

  • CRetry the job until enough capacity is free

    Incorrect

    Retrying does not help, because the quota stays the same until an increase is approved.

  • DRequest a quota increase in Service Quotas

    Correct

    You can view SageMaker AI quotas and request increases for adjustable quotas through Service Quotas; smaller increases are often approved automatically.

Plan quotas before large ML jobs: instance quotas for training, hosting and other uses are separate and may need increases in new accounts.

Practise domain 3 →Practise all domains →