Skip to content
BytePatterns

MLA-C02 · Domain 3: Deployment and Orchestration of ML and AI Workflows · 24% of the exam

Task 3.1: Manage deployment infrastructure for ML and AI model types.

Choosing between real-time, serverless, asynchronous and batch inference, multi-model and multi-container endpoints, hosting foundation models in Bedrock or SageMaker AI, importing models built elsewhere, and deploying agents and RAG configurations.

Study it

  • Inference options on SageMaker AI: real-time, serverless, asynchronous and batch transform

    Partly covered by: Training vs Inference

  • Multi-model, multi-container and inference component endpoints

    Lesson coming

  • Hosting foundation models: Bedrock on-demand, Provisioned Throughput, Custom Model Import and SageMaker AI

    Partly covered by: Quantization, The KV Cache

  • Deploying agents: Amazon Bedrock AgentCore Runtime, Gateway, Memory and agent protocols

    Partly covered by: Agents and Tools, The Tool-Use Loop

  • Amazon Bedrock Knowledge Bases and retrieval configuration

    Partly covered by: Retrieval-Augmented Generation, Vector Databases

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

A media company runs a SageMaker AI model that analyzes uploaded video clips of up to 400 MB. Each analysis takes 10 to 20 minutes. Clips arrive at random times during the day, results are needed within the hour, and there are long periods with no uploads, during which the company does not want to pay for instances. Which inference option should the ML engineer use?

  1. AA real-time endpoint on a large GPU instance
  2. BAsynchronous inference that scales to zero when idle
  3. CServerless inference with the maximum memory setting
  4. DA nightly batch transform job over all clips uploaded that day
Show the answer and why
  • AA real-time endpoint on a large GPU instance

    Incorrect

    Real-time inference supports payloads up to 25 MB and 60-second processing for regular responses, and the instances keep running during idle periods.

  • BAsynchronous inference that scales to zero when idle

    Correct

    Asynchronous inference queues requests, accepts payloads up to 1 GB with processing times up to one hour, and can scale the instance count to zero when there are no requests to process.

  • CServerless inference with the maximum memory setting

    Incorrect

    Serverless inference supports payloads up to 4 MB and processing times up to 60 seconds, far below a 400 MB clip that takes 20 minutes.

  • DA nightly batch transform job over all clips uploaded that day

    Incorrect

    Batch transform suits offline processing of data that is available up front. Waiting for a nightly run would miss the one-hour target for clips uploaded in the morning.

Choose the inference option by payload size, processing time, latency target and traffic pattern. Large payloads with long processing and near real-time needs point to asynchronous inference, which can also scale to zero.

Question 2 · choose 1

A SaaS company trains one XGBoost model per customer, about 2,000 models in total, all using the same framework and of similar size. A few customers call their model often, but most call it only a few times a day, and an occasional extra delay on a rarely used model is acceptable. Which hosting approach minimizes cost and operational effort?

  1. AA separate real-time endpoint for each customer model
  2. BA multi-container endpoint with one container for each customer model
  3. CA batch transform job for every customer request
  4. DA multi-model endpoint that loads models from S3 as needed
Show the answer and why
  • AA separate real-time endpoint for each customer model

    Incorrect

    Two thousand single-model endpoints would each keep their own instances running, most of them idle, which multiplies cost and operational work.

  • BA multi-container endpoint with one container for each customer model

    Incorrect

    Multi-container endpoints run several different containers, such as different frameworks, on one endpoint. They are not built to hold thousands of models that share one framework.

  • CA batch transform job for every customer request

    Incorrect

    Batch transform is for offline processing without a persistent endpoint. Starting a job per request would add far more delay than an occasional cold start.

  • DA multi-model endpoint that loads models from S3 as needed

    Correct

    Multi-model endpoints host large numbers of models that use the same framework on a shared serving container and fleet, loading models into memory as needed. They fit a mix of frequently and rarely used models when occasional cold starts are acceptable.

Many similar models with uneven traffic is the textbook case for a multi-model endpoint. Models with much higher traffic or strict latency can still get dedicated endpoints.

Question 3 · choose 2

An ML engineer is reviewing several models that could be hosted with SageMaker AI Serverless Inference. Which workloads are good fits for serverless inference? (Choose TWO.)

  1. AA computer vision model that needs a GPU to meet its latency target
  2. BA model that must reach a private database through the company VPC
  3. CA document classification model that receives 50 MB payloads
  4. DAn internal tool called a few times an hour that tolerates cold starts
  5. EA predictable morning burst, served with provisioned concurrency
Show the answer and why
  • AA computer vision model that needs a GPU to meet its latency target

    Incorrect

    GPUs are among the real-time inference features that serverless inference does not support.

  • BA model that must reach a private database through the company VPC

    Incorrect

    VPC configuration is not supported for serverless inference, so the endpoint could not reach resources inside the VPC.

  • CA document classification model that receives 50 MB payloads

    Incorrect

    Serverless inference supports payloads up to 4 MB, so 50 MB requests would need another option such as asynchronous inference.

  • DAn internal tool called a few times an hour that tolerates cold starts

    Correct

    On-demand serverless inference suits workloads with idle periods between bursts that can tolerate cold starts. It scales the endpoint down to zero when there are no requests, so the tool costs nothing while idle.

  • EA predictable morning burst, served with provisioned concurrency

    Correct

    Serverless inference with provisioned concurrency keeps capacity initialized for predictable bursts, so requests are served with predictable performance.

Serverless inference trades a few capabilities (GPUs, VPC configuration, large payloads) for zero idle cost and no instance management. Provisioned concurrency removes cold starts for predictable peaks.

Question 4 · choose 1

A company fine-tuned an open-source large language model in SageMaker AI and stored the resulting weights in Amazon S3. Its application already calls other models through the Amazon Bedrock runtime API. The team wants to serve the new model through the same API, paying on demand, without managing inference instances. What should the ML engineer use?

  1. AAmazon Bedrock Custom Model Import
  2. BA SageMaker AI real-time endpoint on GPU instances
  3. CAmazon Bedrock Model Distillation
  4. DAn Amazon Bedrock prompt router
Show the answer and why
  • AAmazon Bedrock Custom Model Import

    Correct

    Custom Model Import brings models customized in other environments, such as SageMaker AI, into Bedrock. The imported model is used with on-demand throughput through InvokeModel and InvokeModelWithResponseStream.

  • BA SageMaker AI real-time endpoint on GPU instances

    Incorrect

    A real-time endpoint can host the model, but the team would choose and run the instances and call a different API, which the requirements rule out.

  • CAmazon Bedrock Model Distillation

    Incorrect

    Distillation trains a smaller student model from a teacher's responses. It does not bring an existing set of weights into Bedrock.

  • DAn Amazon Bedrock prompt router

    Incorrect

    A prompt router sends requests between foundation models in the same model family. It cannot serve a model whose weights the team trained elsewhere.

To use the Bedrock API with a model trained outside Bedrock, import it. To manage the serving infrastructure yourself instead, host it on SageMaker AI.

Question 5 · choose 1

A RAG application retrieves 20 chunks from an Amazon Bedrock knowledge base for each question. Many of them are only loosely related, which makes answers vague and raises the token cost of each generation call. The team wants fewer, more relevant chunks in the prompt without rebuilding the index. What should the ML engineer configure?

  1. AA higher numberOfResults value for retrieval
  2. BHierarchical chunking on the data source
  3. CA global cross-Region inference profile for generation
  4. DA reranker model in the retrieval request
Show the answer and why
  • AA higher numberOfResults value for retrieval

    Incorrect

    Retrieving more chunks adds more loosely related text and more input tokens, which works against both goals.

  • BHierarchical chunking on the data source

    Incorrect

    Changing the chunking strategy means re-ingesting the data, and returning parent chunks makes each result larger, not more selective.

  • CA global cross-Region inference profile for generation

    Incorrect

    Global cross-Region inference changes where requests are processed and can lower price, but it does not change which chunks are retrieved.

  • DA reranker model in the retrieval request

    Correct

    A reranker scores each retrieved chunk's relevance to the query and reorders the results, so the application can pass fewer but more relevant chunks to the model, which also lowers cost and latency.

Vector similarity gets candidates; a reranker orders them by relevance to the actual question. Retrieving a wider set and reranking to a small top set is a common RAG configuration.

Question 6 · choose 1

A payments service must score each card transaction synchronously while the customer waits, with a response in well under a second, around the clock and at steady high volume. Payloads are a few kilobytes. Which SageMaker AI inference option fits?

  1. AAsynchronous inference with an SNS notification
  2. BBatch transform run every few minutes
  3. CServerless inference with on-demand capacity
  4. DA real-time endpoint on chosen instances
Show the answer and why
  • AAsynchronous inference with an SNS notification

    Incorrect

    Asynchronous inference queues requests and writes results to S3 later, with an optional notification, so the caller cannot get the score while waiting.

  • BBatch transform run every few minutes

    Incorrect

    Batch transform is for offline processing of data available up front; running it every few minutes still cannot answer a waiting customer.

  • CServerless inference with on-demand capacity

    Incorrect

    Serverless inference suits intermittent or unpredictable traffic, not steady high-volume traffic with strict latency.

  • DA real-time endpoint on chosen instances

    Correct

    Real-time inference suits online inferences with low latency or high throughput needs, using a persistent, fully managed endpoint backed by the instance type you choose.

Steady, latency-critical, synchronous traffic is the home of real-time endpoints.

Question 7 · choose 1

A team built an agent with LangGraph that calls a model from another provider as well as an Amazon Bedrock model. It wants to deploy the agent on AWS without rewriting it for a specific framework or model, with serverless scaling. What should the ML engineer use?

  1. AAmazon Bedrock AgentCore Runtime
  2. BBedrock Agents Classic, rebuilt in its console
  3. CA SageMaker AI batch transform job
  4. DAmazon Bedrock Prompt management
Show the answer and why
  • AAmazon Bedrock AgentCore Runtime

    Correct

    AgentCore Runtime is framework agnostic, working with LangGraph, Strands, CrewAI and custom agents, and works with any large language model, providing a secure, serverless hosting environment.

  • BBedrock Agents Classic, rebuilt in its console

    Incorrect

    Rebuilding in another service is the rewrite the team wants to avoid, and Agents Classic is no longer open to new customers.

  • CA SageMaker AI batch transform job

    Incorrect

    Batch transform scores datasets offline; it does not host interactive agents.

  • DAmazon Bedrock Prompt management

    Incorrect

    Prompt management stores and versions prompts; it does not host agent code.

AgentCore services work with any framework and any model, so existing agents can move to managed infrastructure without a rewrite.

Question 8 · choose 1

An internal HR assistant must answer policy questions from an Amazon Bedrock knowledge base and show the source documents behind each answer. The team wants the simplest possible integration, with retrieval and generation in one call. Which API should the application use?

  1. ARetrieve only, then display the returned chunks
  2. BRetrieveAndGenerate on the knowledge base
  3. CInvokeModel without the knowledge base
  4. DStartIngestionJob for each question
Show the answer and why
  • ARetrieve only, then display the returned chunks

    Incorrect

    Retrieve returns chunks without generating an answer, so the application would still need its own model call to answer the question.

  • BRetrieveAndGenerate on the knowledge base

    Correct

    RetrieveAndGenerate queries the knowledge base and generates a response based on the retrieved chunks, returning citations to the original source data.

  • CInvokeModel without the knowledge base

    Incorrect

    Calling the model alone gives no retrieval and no citations to policy documents.

  • DStartIngestionJob for each question

    Incorrect

    Ingestion jobs sync data into the knowledge base; they do not answer queries.

For quick RAG integration with citations, RetrieveAndGenerate does retrieval and generation together; use Retrieve when you need full control of the prompt.

Question 9 · choose 1

A team imported its own fine-tuned open-source model into Amazon Bedrock with Custom Model Import and uses it on demand. It now also needs to score 2 million documents overnight and tries to submit a Bedrock batch inference job with the imported model. What should the ML engineer know?

  1. ABatch inference needs Provisioned Throughput first
  2. BBatch jobs are limited to 1,000 documents
  3. CImported models cannot use Bedrock batch inference
  4. DBatch inference needs the model in another Region
Show the answer and why
  • ABatch inference needs Provisioned Throughput first

    Incorrect

    The documented limitation is that imported models are not supported for batch inference at all.

  • BBatch jobs are limited to 1,000 documents

    Incorrect

    There is no such limit stated; the issue is that imported models are not supported for batch inference.

  • CImported models cannot use Bedrock batch inference

    Correct

    Custom Model Import cannot be used with Bedrock batch inference, so the overnight run needs another approach, such as on-demand calls or hosting the model on SageMaker AI for batch transform.

  • DBatch inference needs the model in another Region

    Incorrect

    Region is not the issue; imported models are excluded from batch inference.

Check feature compatibility for imported models: they work with on-demand inference but not with every Bedrock feature, such as batch inference.

Question 10 · choose 1

A team must host a 70-billion-parameter open-weight language model on a SageMaker AI real-time endpoint. The weights do not fit in the memory of a single GPU, but they fit across the 8 GPUs of one large instance. Which deployment approach fits?

  1. AA CPU instance with the default inference container
  2. BA batch transform job on a single GPU instance
  3. CThe training container image reused for hosting
  4. DAn LMI container with tensor parallelism
Show the answer and why
  • AA CPU instance with the default inference container

    Incorrect

    A CPU instance gives up the accelerators the model needs, and the default setup does not shard the model across devices.

  • BA batch transform job on a single GPU instance

    Incorrect

    Batch transform is for offline processing without a persistent endpoint, and one GPU still cannot hold the weights.

  • CThe training container image reused for hosting

    Incorrect

    A training image is built to run training jobs; serving a large LLM needs an inference stack such as an LMI container.

  • DAn LMI container with tensor parallelism

    Correct

    Large model inference (LMI) containers are specialized containers for LLM inference that support features such as tensor parallelism, which spreads the model across the instance's GPUs, and continuous batching.

When a model is too large for one accelerator, use a serving container that shards it across the accelerators of the instance.

Practise domain 3 →Practise all domains →