Skip to content
BytePatterns

AIP-C01 · Domain 1: Foundation Model Integration, Data Management, and Compliance · 31% of the exam

Task 1.2: Select and configure FMs.

Picking foundation models by capability, cost and limits, switching models without code changes, staying up during disruptions with cross-Region inference and fallbacks, and deploying and retiring customized models.

Study it

  • Choosing a foundation model: capability, context window, cost, Regional availability and lifecycle

    Partly covered by: Context Windows, What Is an LLM

  • Switching models without code changes and staying up: Converse, AppConfig, cross-Region inference and circuit breakers

    Partly covered by: Rate Limiting

  • Customized models: fine-tuning, distillation, LoRA adapters, Model Registry and safe rollouts

    Partly covered by: Fine-Tuning vs Prompting, Adapters and LoRA

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

A company runs a summarization service on AWS Lambda that calls Amazon Bedrock through the Converse API. The team wants to move traffic to a newer model, and later possibly to another provider's model, without deploying new function code or a new function version. The switch must roll out to a percentage of invocations at a time, must be checked against a schema before it takes effect, and must roll back automatically if an error alarm fires. Which solution meets these requirements?

  1. AReplace the Converse API with InvokeModel calls so that each provider's native request body can be built in code
  2. BKeep the model ID and settings in AWS AppConfig with a validator and a gradual deployment strategy, read through its Lambda extension
  3. CKeep the model ID in a Lambda environment variable and update the variable when the team wants to switch models
  4. DStore the model ID in an AWS Systems Manager Parameter Store parameter that the function reads at the start of every invocation
Show the answer and why
  • AReplace the Converse API with InvokeModel calls so that each provider's native request body can be built in code

    Incorrect

    Provider-specific request bodies make switching harder and need code changes. The Converse API is what lets the same code work across models that support messages.

  • BKeep the model ID and settings in AWS AppConfig with a validator and a gradual deployment strategy, read through its Lambda extension

    Correct

    AppConfig changes application behavior without redeploying code, validates configuration before deployment, rolls changes out gradually with a deployment strategy, and rolls back automatically when a monitored CloudWatch alarm fires.

  • CKeep the model ID in a Lambda environment variable and update the variable when the team wants to switch models

    Incorrect

    Changing an environment variable updates the function configuration for every invocation at once. There is no schema validation, percentage rollout or alarm-based rollback.

  • DStore the model ID in an AWS Systems Manager Parameter Store parameter that the function reads at the start of every invocation

    Incorrect

    Parameter Store holds the value, but a parameter update applies to every reader immediately, with no deployment strategy, validator or automatic rollback.

Two things make model switching code-free: a model-agnostic API (Converse) and an external, governed configuration source. AppConfig adds what plain parameters lack: validators, gradual deployment strategies and automatic rollback on CloudWatch alarms.

Question 2 · choose 1

A healthcare company in Frankfurt (eu-central-1) has chosen a model that Amazon Bedrock offers in that Region only through cross-Region inference. Patient data must be processed only in AWS Regions inside the EU. During peak hours the application also needs more throughput than a single Region can provide. The company wants to pay on-demand prices. What should the developer configure?

  1. AInvoke the model through the global cross-Region inference profile to get the lowest token price and the most capacity
  2. BMove the application to us-east-1, where the model is offered in single-Region mode with higher quotas
  3. CInvoke the model through the EU geographic cross-Region inference profile from the application in eu-central-1
  4. DPurchase Provisioned Throughput for the model in eu-central-1 and attach it to the cross-Region inference profile
Show the answer and why
  • AInvoke the model through the global cross-Region inference profile to get the lowest token price and the most capacity

    Incorrect

    A global profile can route requests to commercial Regions worldwide, which breaks the EU-only processing requirement.

  • BMove the application to us-east-1, where the model is offered in single-Region mode with higher quotas

    Incorrect

    Processing in a US Region violates the requirement that patient data stay in EU Regions.

  • CInvoke the model through the EU geographic cross-Region inference profile from the application in eu-central-1

    Correct

    A geographic profile keeps processing inside the named geography, spreads requests across its Regions for throughput, and uses on-demand pricing.

  • DPurchase Provisioned Throughput for the model in eu-central-1 and attach it to the cross-Region inference profile

    Incorrect

    Inference profiles don't support Provisioned Throughput, and Provisioned Throughput is a fixed hourly commitment rather than on-demand pricing.

Cross-Region inference solves both limited Regional availability and peak throughput. The choice between profile types is a data residency decision: geographic profiles keep requests within a geography, while global profiles trade that boundary for lower cost.

Question 3 · choose 1

An order-enrichment workflow calls a foundation model for every order. During a recent provider incident, each call retried with backoff for several seconds before failing, and thousands of concurrent workflows kept retrying, which delayed orders for an hour. The team wants calls to the model to stop immediately for a cooling-off period once failures cross a threshold, the state to be shared by all running workflows, and requests to go to a fallback model while the primary is unavailable. Which design meets these requirements?

  1. ASend failed calls to an Amazon SQS dead-letter queue and reprocess them when the provider incident is over
  2. BIncrease the SDK retry attempts and the maximum backoff delay so that each call waits much longer before it finally gives up
  3. CSet reserved concurrency on the Lambda function that calls the model so that fewer of the failing calls can run at the same time
  4. DBuild a Step Functions circuit breaker that keeps the circuit status in DynamoDB and routes to the fallback model while it is open
Show the answer and why
  • ASend failed calls to an Amazon SQS dead-letter queue and reprocess them when the provider incident is over

    Incorrect

    A dead-letter queue captures failures for later processing. It neither stops new calls quickly nor serves orders from a fallback model during the incident.

  • BIncrease the SDK retry attempts and the maximum backoff delay so that each call waits much longer before it finally gives up

    Incorrect

    More retries keep every workflow busy for longer and add load to a dependency that is already failing, which is the behavior that caused the delay.

  • CSet reserved concurrency on the Lambda function that calls the model so that fewer of the failing calls can run at the same time

    Incorrect

    Reserved concurrency caps parallel executions, but calls that do run still fail slowly, and nothing routes requests to a fallback model.

  • DBuild a Step Functions circuit breaker that keeps the circuit status in DynamoDB and routes to the fallback model while it is open

    Correct

    A circuit breaker stops callers from retrying a failing dependency and detects recovery. Keeping the status in DynamoDB shares it across all executions, and a Choice state can route to the fallback model while the circuit is open.

Retries with backoff handle transient errors. When a dependency is down for minutes, a circuit breaker is the pattern: trip after a threshold, fail fast or fall back while open, and probe again after the expiry. AWS prescriptive guidance implements it with Step Functions and a DynamoDB status table.

Question 4 · choose 2

A company fine-tunes an open-weight model with LoRA and serves it on an Amazon SageMaker AI real-time endpoint. Every retraining run must produce a traceable model version, only versions that a reviewer approves may reach production, and a new version must receive a small share of live traffic first and be removed automatically if latency or error alarms fire. Which actions meet these requirements? (Choose TWO.)

  1. AEnable SageMaker Model Monitor on the endpoint and let it replace the model when it detects drift
  2. BUpdate the endpoint with a blue/green deployment that uses canary traffic shifting and alarm-based automatic rollback
  3. CStore each trained model in a versioned Amazon S3 bucket and deploy the newest object version after each training run
  4. DRun a SageMaker AI shadow test that mirrors production traffic to the new version for a fixed period
  5. ERegister each trained model as a version in a SageMaker Model Registry model group and deploy only Approved versions
Show the answer and why
  • AEnable SageMaker Model Monitor on the endpoint and let it replace the model when it detects drift

    Incorrect

    Model Monitor is closed to new customers, and monitoring never deploys or removes a model version by itself.

  • BUpdate the endpoint with a blue/green deployment that uses canary traffic shifting and alarm-based automatic rollback

    Correct

    Deployment guardrails shift a canary share of traffic to the new fleet, watch the configured alarms during the baking period and roll back automatically if one fires.

  • CStore each trained model in a versioned Amazon S3 bucket and deploy the newest object version after each training run

    Incorrect

    S3 versioning keeps old artifacts, but it has no approval status and would deploy every run whether or not it was reviewed.

  • DRun a SageMaker AI shadow test that mirrors production traffic to the new version for a fixed period

    Incorrect

    A shadow test is useful to compare a new variant on live traffic, but its responses are never returned to users and it does not shift traffic or roll back a deployment.

  • ERegister each trained model as a version in a SageMaker Model Registry model group and deploy only Approved versions

    Correct

    The Model Registry catalogs model versions in a model group and tracks their approval status. Moving a version to Approved can start the CI/CD deployment, so only approved versions reach production.

Lifecycle management for customized models has two halves: governance of which version may ship (Model Registry versions and approval status) and a safe rollout with automatic rollback (deployment guardrails with canary traffic shifting and CloudWatch alarms).

Question 5 · choose 1

A team is comparing four foundation models from different providers on Amazon Bedrock for a support assistant that sends product photos and lets the model call two tools. It wants one code path for building requests, reading responses and handling tool calls, so that the winning model, and later replacements, can be adopted with minimal code change. Which approach should the developer use?

  1. ACall each provider through its own SDK and API outside Amazon Bedrock
  2. BCall InvokeModel with each provider's native request body
  3. CCall the Converse API with image content blocks and a toolConfig
  4. DDeploy each model to its own SageMaker AI endpoint behind one API
Show the answer and why
  • ACall each provider through its own SDK and API outside Amazon Bedrock

    Incorrect

    A provider's own API is the route for a model that Bedrock does not offer. With four providers it means four request formats, four credential setups and four ways to handle tool calls.

  • BCall InvokeModel with each provider's native request body

    Incorrect

    InvokeModel fits when a feature exists only in a model's native format. The request body is model-specific, so each switch changes request and tool-call code.

  • CCall the Converse API with image content blocks and a toolConfig

    Correct

    Converse provides a consistent API for all Bedrock models that support messages, including images and tool use, so the same code works with different models.

  • DDeploy each model to its own SageMaker AI endpoint behind one API

    Incorrect

    SageMaker AI endpoints fit self-hosted models. Here they add four endpoints to deploy, scale and pay for while the team only wants to compare models.

A unified inference API keeps the model choice a configuration value. Reserve native request bodies for features that only exist there.

Question 6 · choose 1

A warehouse operator stores 2-minute safety camera clips in Amazon S3 as MP4 files of about 60 MB each, with no audio track. Supervisors want to ask free-form questions about a clip, such as "Did the forklift stop before the crossing, and did anyone step back to let it pass?", and get an answer within seconds. The team has no labeled clips, does not want to train or host models, and wants to pay per request. Which approach meets these requirements?

  1. ASend the clip's S3 location and the question to a Nova model that accepts video
  2. BTranscribe each clip with Amazon Transcribe and send the transcript to a text model
  3. CRun Amazon Rekognition Video label detection and send the labels to a text model
  4. DEmbed frames with Titan Multimodal Embeddings and search for the most similar frames
Show the answer and why
  • ASend the clip's S3 location and the question to a Nova model that accepts video

    Correct

    Amazon Nova 2 Lite can analyze video and answer questions about it. An Amazon S3 URI supports videos up to 1 GB, while video sent as bytes in the request is limited to 25 MB.

  • BTranscribe each clip with Amazon Transcribe and send the transcript to a text model

    Incorrect

    Transcription fits when the answer is in what people say. These clips have no audio track, so there is nothing to transcribe.

  • CRun Amazon Rekognition Video label detection and send the labels to a text model

    Incorrect

    Label detection returns objects, scenes and actions it recognizes, which suits tagging and search. Stored-video label detection is an asynchronous job, and a label list does not let a model reason about the sequence of events supervisors ask about.

  • DEmbed frames with Titan Multimodal Embeddings and search for the most similar frames

    Incorrect

    Multimodal embeddings support image search and recommendations. Similarity search returns matching frames, not an answer to a question.

Assess models by capability first: the input modality, size limits and how the input is delivered. A video-capable model with S3 input answers questions about large clips without training or hosting.

Question 7 · choose 1

A customer-facing claims assistant needs the fastest response times the company can get on demand during business hours, and the company accepts a price premium. It does not want a 24x7 capacity reservation. Which service tier should requests use?

  1. AThe default Standard tier with a higher client timeout
  2. BThe Reserved tier
  3. CThe Flex tier
  4. DThe Priority tier
Show the answer and why
  • AThe default Standard tier with a higher client timeout

    Incorrect

    A longer timeout does not make responses faster.

  • BThe Reserved tier

    Incorrect

    Reserved capacity is a reservation, which the company does not want.

  • CThe Flex tier

    Incorrect

    Flex trades speed for a lower price, which is the opposite of the goal.

  • DThe Priority tier

    Correct

    Priority delivers the fastest response times for a price premium and fits mission-critical customer-facing workflows that do not need a 24x7 reservation.

Service tiers let you trade cost against speed per request. Priority is for latency-sensitive traffic without a standing reservation.

Question 8 · choose 1

A research team needs an open-weight foundation model that is not offered as a serverless model in Amazon Bedrock. It wants a catalog of pre-trained models that it can deploy to its own SageMaker AI endpoint with a few steps and fine-tune later. Where should it start?

  1. AAmazon Bedrock Provisioned Throughput for the model
  2. BAmazon SageMaker JumpStart foundation models
  3. CAmazon SageMaker Ground Truth
  4. DAmazon Comprehend custom classification
Show the answer and why
  • AAmazon Bedrock Provisioned Throughput for the model

    Incorrect

    Provisioned Throughput applies to models offered in Amazon Bedrock, not to an arbitrary open-weight model.

  • BAmazon SageMaker JumpStart foundation models

    Correct

    JumpStart offers pre-trained foundation models that you can deploy and fine-tune with SageMaker AI features.

  • CAmazon SageMaker Ground Truth

    Incorrect

    Ground Truth is a labeling service and is closed to new customers.

  • DAmazon Comprehend custom classification

    Incorrect

    Comprehend custom models are for NLP tasks like classification, not for hosting foundation models.

When a model is not available serverless, JumpStart is a fast path to host and adapt open foundation models on SageMaker AI.

Question 9 · choose 1

A bank serves a LoRA fine-tuned open-weight model on a SageMaker AI real-time endpoint with four instances in eu-west-1. The recovery plan must survive the loss of the whole Region with traffic moving within 15 minutes, regulators allow processing only in EU Regions, and client applications call one hostname that cannot be changed during an incident. Which design meets these requirements?

  1. AAdd instances so that the endpoint spreads across more Availability Zones in eu-west-1
  2. BAdd a second production variant to the endpoint as a standby copy of the model
  3. CReplicate the model artifacts to another EU Region and create an endpoint there after an outage
  4. DRun a standby endpoint in another EU Region and use Route 53 failover routing with health checks
Show the answer and why
  • AAdd instances so that the endpoint spreads across more Availability Zones in eu-west-1

    Incorrect

    Multiple instances across Availability Zones protect an endpoint from Availability Zone outages and instance failures. They do not help when the whole Region is impaired.

  • BAdd a second production variant to the endpoint as a standby copy of the model

    Incorrect

    Production variants split or direct traffic between models behind one endpoint, which suits A/B testing. They live in the same endpoint and Region as the primary.

  • CReplicate the model artifacts to another EU Region and create an endpoint there after an outage

    Incorrect

    S3 replication copies objects to another Region automatically. Creating the endpoint only after the outage starts suits a plan with a longer recovery target, and traffic still has to be moved.

  • DRun a standby endpoint in another EU Region and use Route 53 failover routing with health checks

    Correct

    With failover routing, Route 53 sends traffic to the primary while it is healthy and to the secondary when it is not, so the same hostname moves to an endpoint that is already running in another EU Region.

Match the resilience mechanism to the failure: Availability Zones cover zone outages, a warm standby in a second Region covers a Regional outage, and DNS failover moves clients without code changes.

Practise domain 1 →Practise all domains →