Skip to content
BytePatterns

MLA-C02 · Domain 3: Deployment and Orchestration of ML and AI Workflows · 24% of the exam

Task 3.3: Implement automated orchestration and continuous integration and continuous delivery (CI/CD) pipelines for MLOps and AI workloads.

SageMaker Pipelines and the AWS developer tools, safe deployment with rollback, retraining triggers, model versions in the Model Registry and MLflow, prompt and agent versions, testing models and prompts in a pipeline, and keeping a knowledge base in sync with its sources.

Study it

  • SageMaker Pipelines, Step Functions, MWAA and event-driven retraining

    Partly covered by: SQS vs SNS vs EventBridge

  • CI/CD for models: CodePipeline, CodeBuild, deployment guardrails and rollback

    Lesson coming

  • Versioning models, prompts and agents: Model Registry, MLflow, Prompt Management

    Lesson coming

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

A SageMaker AI pipeline processes data, trains a model and evaluates it. The model should be registered in the SageMaker Model Registry only when its evaluation RMSE is below a threshold; otherwise the run should end without registering anything. Which pipeline step should the ML engineer add before the registration step?

  1. AA callback step that waits for a message on an Amazon SQS queue
  2. BA condition step that checks the RMSE against the threshold
  3. CStep caching on the training step
  4. DA transform step that scores the test set again
Show the answer and why
  • AA callback step that waits for a message on an Amazon SQS queue

    Incorrect

    A callback step adds processes outside SageMaker Pipelines to the workflow. It does not compare a metric with a threshold by itself.

  • BA condition step that checks the RMSE against the threshold

    Correct

    A condition step evaluates step properties, such as a metric produced by the evaluation step, and decides which steps run next, so registration happens only when the condition is met.

  • CStep caching on the training step

    Incorrect

    Caching reuses the output of an earlier successful run with the same inputs. It saves time but does not decide whether to register.

  • DA transform step that scores the test set again

    Incorrect

    A transform step runs batch transform. Scoring the data again does not stop a weak model from being registered.

Quality gates in SageMaker Pipelines are condition steps: evaluate, compare with a threshold, then register (and later approve and deploy) only models that pass.

Question 2 · choose 1

An ML engineer is updating the model behind a busy SageMaker AI real-time endpoint. The new fleet should first take 10% of traffic for 30 minutes while CloudWatch alarms watch latency and errors. If an alarm fires, all traffic must return to the old fleet automatically; otherwise the rest of the traffic should move over in one step. Which deployment configuration meets these requirements?

  1. AA blue/green deployment with all-at-once traffic shifting
  2. BA blue/green deployment with linear traffic shifting in ten equal steps
  3. CA shadow test that copies 10% of requests to the new model
  4. DBlue/green with canary traffic shifting and auto-rollback alarms
Show the answer and why
  • AA blue/green deployment with all-at-once traffic shifting

    Incorrect

    All-at-once moves 100% of traffic to the new fleet in one step, so there is no 10% trial period first.

  • BA blue/green deployment with linear traffic shifting in ten equal steps

    Incorrect

    Linear shifting moves traffic in several equal steps. The requirement is one small trial step followed by a single full shift.

  • CA shadow test that copies 10% of requests to the new model

    Incorrect

    A shadow test never returns the new model's responses to callers, so it is a validation step, not a way to move production traffic.

  • DBlue/green with canary traffic shifting and auto-rollback alarms

    Correct

    Canary shifting sends a portion of traffic to the new fleet for a baking period. If no alarm trips, the rest of the traffic shifts; if one trips, SageMaker AI rolls all traffic back to the old fleet.

Deployment guardrails give blue/green updates with canary, linear or all-at-once traffic shifting, plus rolling deployments, all with baking periods and alarm-based auto-rollback.

Question 3 · choose 2

A team wants its SageMaker AI model-building pipeline to start automatically, with no one clicking "Start". It should run every Sunday night, and also whenever a new labeled data file lands in a specific Amazon S3 bucket. Which mechanisms can start the pipeline in these cases? (Choose TWO.)

  1. ASetting the model version's status to Approved in the Model Registry
  2. BAn EventBridge rule on a weekly schedule that targets the pipeline
  3. CTurning on step caching for every step of the pipeline
  4. DA callback step placed at the beginning of the pipeline
  5. EAn EventBridge rule on S3 object-created events that targets the pipeline
Show the answer and why
  • ASetting the model version's status to Approved in the Model Registry

    Incorrect

    In the SageMaker project templates, approving a model version starts CI/CD deployment of that version. It does not start the model-building pipeline.

  • BAn EventBridge rule on a weekly schedule that targets the pipeline

    Correct

    SageMaker Pipelines is supported as an EventBridge target, and a scheduled rule can start a pipeline execution on a fixed schedule.

  • CTurning on step caching for every step of the pipeline

    Incorrect

    Caching reuses the outputs of earlier successful steps when a pipeline runs. It does not start a run.

  • DA callback step placed at the beginning of the pipeline

    Incorrect

    A callback step only runs after an execution has started, so it cannot be what starts the pipeline.

  • EAn EventBridge rule on S3 object-created events that targets the pipeline

    Correct

    EventBridge can start a pipeline execution from events on the bus, including a new file being uploaded to an S3 bucket.

Retraining triggers are usually events: a schedule, new data, or an alert. EventBridge rules turn those events into StartPipelineExecution calls.

Question 4 · choose 1

A team keeps its customer-support prompt in Amazon Bedrock Prompt management and keeps editing it. The production application must keep using the exact prompt that passed testing until a newer one is approved. How should the application reference the prompt?

  1. AReference the working draft so the newest edits are used
  2. BCreate a version of the tested prompt and reference that version
  3. CCreate a new guardrail version for each prompt change
  4. DCopy the prompt text into an environment variable of the application
Show the answer and why
  • AReference the working draft so the newest edits are used

    Incorrect

    The draft changes every time the team saves edits, so production would pick up untested changes.

  • BCreate a version of the tested prompt and reference that version

    Correct

    A prompt version is a snapshot taken when you are satisfied with a configuration. The application keeps using it while the team continues to edit the draft.

  • CCreate a new guardrail version for each prompt change

    Incorrect

    Guardrail versions capture guardrail policies. They do not snapshot the prompt text, model or inference settings.

  • DCopy the prompt text into an environment variable of the application

    Incorrect

    This freezes the text, but it loses the managed versions, variants and model settings that Prompt management keeps.

Treat prompts like code: iterate on a draft, then release immutable versions that applications reference. Prompt management provides drafts, variants and versions.

Question 5 · choose 1

A company's ETL job updates product manuals in an Amazon S3 bucket every night. The bucket is the data source of an Amazon Bedrock knowledge base (customer-managed, with an OpenSearch Serverless vector store), and answers must reflect the night's changes by 7 AM. What should the ML engineer automate after the ETL job completes?

  1. AStart an ingestion job (sync) for the data source
  2. BDelete and re-create the knowledge base
  3. CEnable model invocation logging on the generation model
  4. DIncrease the number of results returned by each query
Show the answer and why
  • AStart an ingestion job (sync) for the data source

    Correct

    After files are added, modified or removed, the data source must be synced to re-index them. Syncing is incremental, so only the documents changed since the last sync are processed.

  • BDelete and re-create the knowledge base

    Incorrect

    Recreating everything re-embeds every document and changes resource IDs that the application uses, when an incremental sync handles the changes.

  • CEnable model invocation logging on the generation model

    Incorrect

    Invocation logging records requests and responses for auditing. It does not update the index with the new manuals.

  • DIncrease the number of results returned by each query

    Incorrect

    More results come from the same stale index, so the night's changes would still be missing.

A knowledge base does not see data source changes until it is synced. Trigger a StartIngestionJob after each data refresh, for example from the ETL job's completion event.

Question 6 · choose 1

A team deploys its agent to Amazon Bedrock AgentCore Runtime from a CI/CD pipeline that updates the runtime on every merge to main. Production callers must stay on the version that passed acceptance tests until the team promotes a newer one. How should the team route production traffic?

  1. APoint production at the DEFAULT endpoint
  2. BA named endpoint that points to a specific runtime version
  3. CCreate a new AgentCore Runtime for every merge
  4. DStore the agent code in Amazon S3 with S3 Versioning turned on
Show the answer and why
  • APoint production at the DEFAULT endpoint

    Incorrect

    The DEFAULT endpoint automatically moves to the latest version whenever the runtime is updated, so production would receive every merge.

  • BA named endpoint that points to a specific runtime version

    Correct

    Each update creates a new immutable runtime version. A named endpoint can point to a specific version and moves only when explicitly updated, which separates production from new builds.

  • CCreate a new AgentCore Runtime for every merge

    Incorrect

    Versions already capture each update. New runtimes per merge add resources and require callers to change identifiers each time.

  • DStore the agent code in Amazon S3 with S3 Versioning turned on

    Incorrect

    Versioned objects keep old code, but they do not control which version of the runtime production callers reach.

AgentCore Runtime versions are immutable snapshots; endpoints are movable pointers. Pin production to an endpoint and update it only when promoting a tested version.

Question 7 · choose 1

A SageMaker AI pipeline runs a six-hour data processing step, then training and evaluation. In last night's run, training failed because of a wrong instance type. The engineer fixed it and wants to run only training and evaluation, reusing the processing outputs from last night's run. What should the ML engineer use?

  1. ASelective execution of the two steps
  2. BA new full execution with the fixed instance type
  3. CA Condition step placed before the training step
  4. DA manual training job outside the pipeline
Show the answer and why
  • ASelective execution of the two steps

    Correct

    Selective execution runs a connected subset of steps and reuses outputs from a reference pipeline execution for the steps that are not selected, so expensive data preparation is not repeated.

  • BA new full execution with the fixed instance type

    Incorrect

    A full execution runs the six-hour processing step again.

  • CA Condition step placed before the training step

    Incorrect

    A Condition step chooses a path based on values; it does not reuse outputs from an earlier run.

  • DA manual training job outside the pipeline

    Incorrect

    A manual job leaves the pipeline's repeatable workflow and still leaves evaluation to be run by hand.

When an intermediate step fails after expensive upstream work, rerun only the affected connected steps against a reference execution.

Question 8 · choose 1

A nightly SageMaker AI pipeline sometimes fails because a training step hits a temporary capacity error, and an engineer restarts it by hand the next morning. How should the ML engineer make the pipeline handle these transient failures automatically?

  1. AAdd a Fail step right after the training step
  2. BAdd a retry policy to the training step
  3. CTurn on step caching for training
  4. DAdd a condition step before training
Show the answer and why
  • AAdd a Fail step right after the training step

    Incorrect

    A Fail step ends the execution with an error message; it does not retry the step.

  • BAdd a retry policy to the training step

    Correct

    Retry policies automatically retry supported pipeline steps, including training, processing and transform steps, after an error, for the exception types you choose.

  • CTurn on step caching for training

    Incorrect

    Caching reuses earlier successful outputs; it does not retry a failed step.

  • DAdd a condition step before training

    Incorrect

    A condition step chooses paths based on values; it does not retry after errors.

Make pipelines resilient to transient errors with step retry policies, rather than relying on manual restarts.

Question 9 · choose 1

A team uses a SageMaker project template in which approving a model version starts CI/CD deployment. The newest approved version, now in production, turns out to have a data problem. What is the quickest way to return production to the previous good version through the same automation?

  1. ADelete the model package group
  2. BRegister the faulty model again as a new version
  3. CSet the faulty version to Rejected
  4. DStop the endpoint until a new model is trained
Show the answer and why
  • ADelete the model package group

    Incorrect

    Deleting the group removes all versions, including the good one, and does not redeploy anything.

  • BRegister the faulty model again as a new version

    Incorrect

    A new registration of the same faulty model does not restore the previous version.

  • CSet the faulty version to Rejected

    Correct

    In the project templates, changing a version from Approved to Rejected starts CI/CD to deploy the latest model version that still has an Approved status.

  • DStop the endpoint until a new model is trained

    Incorrect

    Stopping the endpoint causes an outage and waits for a retrain that is not needed.

Registry statuses drive deployment in project templates: rejecting the current version rolls production back to the latest approved one.

Question 10 · choose 1

A Bedrock flow is called by a production application. The team keeps changing the flow's draft and wants to release tested changes to production in a controlled way and switch back quickly if needed, without changing the application's code. What should the ML engineer use?

  1. AThe working draft, updated after each test
  2. BA copy of the flow for each release
  3. CA guardrail version for each flow change
  4. DFlow versions with an alias the app calls
Show the answer and why
  • AThe working draft, updated after each test

    Incorrect

    The draft changes with every edit, so production would receive untested changes.

  • BA copy of the flow for each release

    Incorrect

    New copies have new identifiers, which forces application changes for each release.

  • CA guardrail version for each flow change

    Incorrect

    Guardrail versions capture guardrail policies, not flow definitions.

  • DFlow versions with an alias the app calls

    Correct

    Flows are deployed to applications using versions and aliases; the application calls an alias, which can be pointed at a new version or back to an earlier one.

Versions freeze what was tested; aliases let applications follow a moving pointer, which makes releases and rollbacks configuration changes.

Question 11 · choose 2

A forecasting team must retrain 50 regional models every week with AWS Step Functions, all with the same steps, using the createTrainingJob.sync integration. The account's quota allows only 10 of these training jobs to run at the same time. Which Step Functions features should the ML engineer use? (Choose TWO.)

  1. AA Map state that iterates over the list of regions
  2. BThe Wait for Callback pattern for SageMaker jobs
  3. CA MaxConcurrency value of 10 on that Map state
  4. DA Choice state that trains one region at a time
  5. EFifty state machines, each started by hand
Show the answer and why
  • AA Map state that iterates over the list of regions

    Correct

    The Map state runs a set of workflow steps for each item in a dataset, and its iterations run in parallel.

  • BThe Wait for Callback pattern for SageMaker jobs

    Incorrect

    The SageMaker AI integration does not support the Wait for a Callback with Task Token pattern.

  • CA MaxConcurrency value of 10 on that Map state

    Correct

    MaxConcurrency sets the upper bound on Map iterations that run in parallel, so a value of 10 keeps concurrent training jobs within the quota.

  • DA Choice state that trains one region at a time

    Incorrect

    A Choice state branches on input values; it does not iterate over the 50 regions.

  • EFifty state machines, each started by hand

    Incorrect

    Separate manual runs lose the single repeatable workflow and do not enforce the concurrency limit.

Fan out identical work with Map, and cap its concurrency to respect service quotas.

Practise domain 3 →Practise all domains →