Skip to content
BytePatterns

AIP-C01 · Domain 2: Implementation and Integration · 26% of the exam

Task 2.2: Implement model deployment strategies.

Serving models the right way: on-demand, reserved and provisioned capacity in Bedrock, SageMaker AI endpoints and inference components for large and adapted models, and smaller models where they suffice.

Study it

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

A payments company runs a fraud-explanation service on one Amazon Bedrock model around the clock. Its steady load is about 300,000 input tokens and 40,000 output tokens per minute, the service cannot tolerate throttling at that level, and the company wants input and output capacity sized separately. Short spikes above the steady level must still be served automatically at on-demand rates. The model supports every Bedrock service tier. Which option should the company use?

  1. AThe Reserved tier, sized for the steady input and output tokens per minute and arranged with the AWS account team
  2. BThe Flex tier for all requests, with retries that use exponential backoff to absorb throttling
  3. CThe Standard tier with a request for a higher tokens-per-minute quota through Service Quotas
  4. DThe Priority tier, set with the service_tier parameter on every request the service sends
Show the answer and why
  • AThe Reserved tier, sized for the steady input and output tokens per minute and arranged with the AWS account team

    Correct

    The Reserved tier reserves prioritized capacity with separately sized input and output tokens per minute and automatically overflows to the Standard tier when traffic exceeds the reservation.

  • BThe Flex tier for all requests, with retries that use exponential backoff to absorb throttling

    Incorrect

    Flex trades longer processing times for a discount and suits workloads that tolerate delay, not a service that cannot be throttled.

  • CThe Standard tier with a request for a higher tokens-per-minute quota through Service Quotas

    Incorrect

    A higher quota raises the upper bound, but a quota is not a guarantee that every on-demand request is served immediately at peak demand.

  • DThe Priority tier, set with the service_tier parameter on every request the service sends

    Incorrect

    Priority requests are served ahead of Standard and Flex without a reservation, but they still draw on the shared on-demand quota, so capacity at the steady level is not reserved.

Service tiers trade price for priority. Mission-critical, steady traffic that must not be throttled calls for reserved capacity; the Reserved tier sizes input and output separately and overflows to Standard so spikes are still served.

Question 2 · choose 1

A team deploys a 140 GB open-weight model on an Amazon SageMaker AI real-time endpoint with a large model inference (LMI) container. The instance type has enough GPU memory, but creation fails: CloudWatch Logs show that the model download from Amazon S3 is still running when the container health check times out, and the instance has no local NVMe disk. Which change should the team make?

  1. ASwitch the endpoint to SageMaker AI Serverless Inference so that compute is allocated after the model is downloaded
  2. BRaise the model download timeout, the container startup health check timeout and the EBS volume size
  3. CMove the model to an asynchronous inference endpoint because it allows payloads of up to 1 GB
  4. DIncrease the endpoint's InitialInstanceCount so that the download is spread across more instances
Show the answer and why
  • ASwitch the endpoint to SageMaker AI Serverless Inference so that compute is allocated after the model is downloaded

    Incorrect

    Serverless Inference does not support GPUs, so it cannot host a model of this size.

  • BRaise the model download timeout, the container startup health check timeout and the EBS volume size

    Correct

    For large models, SageMaker AI recommends a longer model download timeout, a longer container startup health check timeout and an EBS volume larger than the model when the instance has no local disk.

  • CMove the model to an asynchronous inference endpoint because it allows payloads of up to 1 GB

    Incorrect

    Asynchronous inference handles large request payloads and long processing times, but the endpoint still has to download the model and pass its health check at startup.

  • DIncrease the endpoint's InitialInstanceCount so that the download is spread across more instances

    Incorrect

    Every instance downloads and loads the full model, so more instances do not make the download or health check any faster.

LLM deployments fail in ways classic ML deployments rarely do: model artifacts of tens or hundreds of gigabytes must be downloaded and loaded before the container answers health checks. SageMaker AI exposes the download timeout, the startup health check timeout and the EBS volume size for exactly this.

Question 3 · choose 1

A company serves six fine-tuned 7B-parameter models, one per business unit, on SageMaker AI. Traffic per model is uneven: two models are busy all day, the others get a few requests per hour and none at night. The company wants the models to share GPU instances, to reserve accelerator and memory per model, to scale each model independently, and to let idle models scale down to no copies. Which deployment approach meets these requirements?

  1. ADeploy each model as an inference component on one endpoint with its own compute requirements and copy count
  2. BDeploy each model to a SageMaker AI serverless endpoint so that idle models cost nothing at night
  3. CDeploy each model to its own real-time endpoint with a minimum of one GPU instance per endpoint
  4. DDeploy all models to one multi-model endpoint so that SageMaker AI loads them from Amazon S3 on demand
Show the answer and why
  • ADeploy each model as an inference component on one endpoint with its own compute requirements and copy count

    Correct

    Inference components share an endpoint's instances while each model declares its own CPU, accelerator and memory needs, scales its copies independently and can scale down to zero copies.

  • BDeploy each model to a SageMaker AI serverless endpoint so that idle models cost nothing at night

    Incorrect

    Serverless endpoints scale to zero, but they do not support GPUs, which 7B-parameter models need for acceptable latency.

  • CDeploy each model to its own real-time endpoint with a minimum of one GPU instance per endpoint

    Incorrect

    Separate endpoints scale independently, but six always-on GPU instances waste money on the quiet models and do not share hardware.

  • DDeploy all models to one multi-model endpoint so that SageMaker AI loads them from Amazon S3 on demand

    Incorrect

    Multi-model endpoints load models on demand from S3, but you cannot reserve accelerator and memory per model or set copy counts for each model.

Inference components decouple models from the endpoint: several models share the instances, each with its own resource allocation and copy count, and copies can scale to zero to free room for busier models.

Question 4 · choose 2

A help desk sends 200,000 tickets a day to one large reasoning model on Amazon Bedrock. Analysis shows that 85% of tickets only need an intent label and a priority, which a small model handles with the same accuracy in tests, while 15% need multi-step troubleshooting. The company wants to cut inference cost substantially without lowering the quality of the complex answers. Which actions should the developer take? (Choose TWO.)

  1. ASend every ticket to both models in parallel and keep the answer with the higher self-reported confidence
  2. BEscalate a ticket to the large reasoning model only when the small model flags it as complex or its output fails validation
  3. CPurchase Provisioned Throughput for the large model and keep sending every ticket to it
  4. DSend every ticket first to a smaller, cheaper model that returns the intent label and priority
  5. EMove all requests to the Priority service tier so that the large model answers each ticket faster
Show the answer and why
  • ASend every ticket to both models in parallel and keep the answer with the higher self-reported confidence

    Incorrect

    Running both models on every ticket increases cost rather than reducing it.

  • BEscalate a ticket to the large reasoning model only when the small model flags it as complex or its output fails validation

    Correct

    Cascading sends only the hard cases to the expensive model, so the complex answers keep their quality while most tickets never reach it.

  • CPurchase Provisioned Throughput for the large model and keep sending every ticket to it

    Incorrect

    Fixed capacity for the large model keeps paying large-model rates for the 85% of tickets a small model could handle.

  • DSend every ticket first to a smaller, cheaper model that returns the intent label and priority

    Correct

    Using a smaller model for the routine classification task cuts the per-token price for 85% of the traffic while tests show equal accuracy.

  • EMove all requests to the Priority service tier so that the large model answers each ticket faster

    Incorrect

    The Priority tier costs more than Standard. It improves response time, not cost.

Right-size the model for the task. A small model for routine work plus a cascade that escalates only complex or invalid results to the large model is the standard way to cut cost without hurting the hard cases.

Question 5 · choose 1

A SaaS company gives each of its 3,000 customers a small fine-tuned text classifier, all built with the same framework and each under 500 MB, that tags incoming messages before a generative model drafts a reply. The classifiers need GPUs to meet the reply-time target. Most customers send only a few messages a day, new customers are added every week without a redeployment, and some extra latency on a model's first request after a quiet period is acceptable. Hosting cost must be as low as possible. Which deployment fits?

  1. AOne real-time endpoint for each customer's classifier
  2. BOne serverless endpoint for each customer's classifier
  3. CA multi-model endpoint on CPU instances for all classifiers
  4. DA multi-model endpoint on GPU instances for all classifiers
Show the answer and why
  • AOne real-time endpoint for each customer's classifier

    Incorrect

    Dedicated endpoints fit models with high traffic or strict latency needs. Three thousand mostly idle endpoints would cost far more than shared hosting.

  • BOne serverless endpoint for each customer's classifier

    Incorrect

    Serverless Inference suits intermittent traffic, but it does not support GPUs.

  • CA multi-model endpoint on CPU instances for all classifiers

    Incorrect

    A CPU-backed multi-model endpoint fits models that are fast enough on CPU. These classifiers need GPUs to meet the target.

  • DA multi-model endpoint on GPU instances for all classifiers

    Correct

    Multi-model endpoints host large numbers of models on a shared fleet and container, support GPU-backed models, and load a newly copied model from S3 on its first invocation, accepting cold-start latency for rarely used models.

Many similar, rarely used models belong on a multi-model endpoint. Check each option against the hard constraints first: GPU support rules out Serverless Inference and CPU fleets here.

Question 6 · choose 1

A legal drafting assistant self-hosts a 70B open-weight model on a SageMaker AI endpoint. Users say long clauses appear too slowly. Legal reviewers validated the current outputs and will not accept any change in output quality or in the model itself, and the endpoint already runs on the largest instance type the budget allows. Which optimization should the developer apply?

  1. AQuantize the model weights to INT4-AWQ with an optimization job
  2. BApply speculative decoding with a draft model through an optimization job
  3. CMove the endpoint to an instance type with more GPU memory
  4. DAdd instances and a target tracking auto scaling policy
Show the answer and why
  • AQuantize the model weights to INT4-AWQ with an optimization job

    Incorrect

    Quantization reduces hardware requirements so the model can run on cheaper GPUs, but the quantized model might be less accurate than the source model.

  • BApply speculative decoding with a draft model through an optimization job

    Correct

    Speculative decoding speeds up decoding without compromising the quality of the generated text: a smaller draft model proposes tokens and the original model verifies them.

  • CMove the endpoint to an instance type with more GPU memory

    Incorrect

    A larger instance fits when the model or its batches do not fit in GPU memory. The endpoint already uses the largest instance the budget allows.

  • DAdd instances and a target tracking auto scaling policy

    Incorrect

    Auto scaling adds instances as the workload rises, which serves more concurrent requests. It does not make a single response generate faster.

Pick the optimization by what may change: quantization trades accuracy for cheaper hardware, while speculative decoding cuts latency and leaves the output quality of the original model intact.

Question 7 · choose 1

A startup serves a fine-tuned 70B open-weight model on a SageMaker AI endpoint that needs a large multi-GPU instance, which costs more than its budget allows. Offline tests show it can accept a small loss of accuracy as long as its regression test set still passes. It wants to keep its own fine-tuned model without retraining, cannot commit to one or three years of spend, and wants to run on less expensive GPU instances. What should the team do?

  1. AQuantize the model, rerun the tests, and host it on a smaller GPU instance
  2. BDistill the model into a smaller student with Amazon Bedrock Model Distillation
  3. CCompile the model for the current instance type with an optimization job
  4. DBuy a SageMaker AI Savings Plan that covers the current instance
Show the answer and why
  • AQuantize the model, rerun the tests, and host it on a smaller GPU instance

    Correct

    Quantization uses a less precise data type for weights and activations, which reduces hardware requirements so the model can run on less expensive GPUs. It might lose some accuracy, which the regression tests check.

  • BDistill the model into a smaller student with Amazon Bedrock Model Distillation

    Incorrect

    Distillation fine-tunes a smaller student model on a teacher model's responses. That is the retraining the startup wants to avoid.

  • CCompile the model for the current instance type with an optimization job

    Incorrect

    Compilation optimizes the model for the chosen hardware type without loss of accuracy and shortens deployment and scaling time. It does not move the model to cheaper hardware.

  • DBuy a SageMaker AI Savings Plan that covers the current instance

    Incorrect

    Savings Plans lower prices in exchange for a one- or three-year usage commitment, which the startup cannot make.

When the constraint is hardware cost and some accuracy loss is acceptable, quantization is the lever that changes the instance a model needs.

Question 8 · choose 1

A legal team's drafting tool streams responses from a fine-tuned open-weight model on a SageMaker AI real-time endpoint with one GPU instance. The tool is used only on weekdays during office hours. Nights and weekends have no traffic, yet the instance is billed around the clock. The team wants to pay for no instances while the tool is idle and to keep streaming responses on the same GPU instance type. It accepts that the first requests after an idle period fail for several minutes while capacity starts, because the tool retries them. Which approach meets these requirements?

  1. AA serverless endpoint at the largest memory size for the model
  2. BAn asynchronous endpoint whose instance count scales in to zero
  3. CAn inference component for the model on an endpoint that can scale to zero
  4. DScheduled scaling actions that set the variant's instance count to zero after hours
Show the answer and why
  • AA serverless endpoint at the largest memory size for the model

    Incorrect

    Serverless Inference scales to zero when there are no requests, but it does not support GPUs, and its largest memory size is 6 GB.

  • BAn asynchronous endpoint whose instance count scales in to zero

    Incorrect

    Asynchronous Inference can scale to zero, but it queues requests and writes results to Amazon S3, and such an endpoint accepts only asynchronous invocations, so the tool could not stream responses.

  • CAn inference component for the model on an endpoint that can scale to zero

    Correct

    An endpoint that hosts inference components can scale in to zero instances when MinInstanceCount is 0 and the component's minimum copy count is 0. A step scaling policy on a NoCapacityInvocationFailures alarm brings capacity back, which takes several minutes.

  • DScheduled scaling actions that set the variant's instance count to zero after hours

    Incorrect

    Scheduled actions change capacity at set times for predictable load. A real-time endpoint can scale to zero instances only if it hosts inference components, so this variant cannot reach zero.

For an idle real-time GPU endpoint, scale-to-zero needs inference components plus a policy that adds capacity when the next request arrives.

Question 9 · choose 1

A document-analysis model on SageMaker AI receives requests with 400 MB payloads that take up to 30 minutes each. Traffic arrives in bursts, and the company wants requests queued and capacity scaled down when idle. Which inference option fits?

  1. ASageMaker Asynchronous Inference
  2. BA Lambda function that runs the model
  3. CA real-time endpoint with a 30-minute client timeout
  4. DSageMaker Serverless Inference
Show the answer and why
  • ASageMaker Asynchronous Inference

    Correct

    Asynchronous inference queues requests, supports large payloads and long processing times, and can scale to zero when idle.

  • BA Lambda function that runs the model

    Incorrect

    Lambda functions stop after 15 minutes.

  • CA real-time endpoint with a 30-minute client timeout

    Incorrect

    Real-time endpoints are for short synchronous requests, not 30-minute jobs with large payloads.

  • DSageMaker Serverless Inference

    Incorrect

    Serverless endpoints have payload and duration limits far below these requests.

Large payloads and long processing call for asynchronous inference, which queues work and scales with demand.

Practise domain 2 →Practise all domains →