Skip to content
BytePatterns

MLA-C02 · Domain 4: Operating, Monitoring, and Securing ML and AI Solutions · 24% of the exam

Task 4.1: Monitor ML and AI model inference and performance.

Watching models in production: drift in data and predictions, A/B tests between variants, errors in data and inference workflows, CloudWatch generative AI observability, Bedrock evaluations, and tool and coordination failures in agents.

Study it

  • Monitoring models: data drift, prediction drift and A/B tests between variants

    Lesson coming

  • Observability for generative AI and agents: CloudWatch, AgentCore Observability and X-Ray

    Partly covered by: CloudWatch, Alarms & X-Ray

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

A new AWS account hosts a credit-risk model on a SageMaker AI real-time endpoint. The team must detect when the distribution of incoming features drifts away from the training data, store drift results next to the training metrics, and email the on-call engineer when more than a set share of features drift. Which approach should the ML engineer implement?

  1. ACloudWatch alarms on the endpoint Invocations and ModelLatency metrics
  2. BData capture plus a scheduled Evidently drift job with SNS alerts
  3. CA weekly retraining schedule that runs whether or not the data has changed
  4. DMore instances behind the endpoint, chosen with auto scaling
Show the answer and why
  • ACloudWatch alarms on the endpoint Invocations and ModelLatency metrics

    Incorrect

    These metrics show traffic and speed. A feature distribution can shift completely while invocation counts and latency stay normal.

  • BData capture plus a scheduled Evidently drift job with SNS alerts

    Correct

    This is the pattern of the open-source SageMaker AI monitoring solutions that AWS recommends: data capture from the endpoint, Evidently drift presets that compare current data with the training reference, results logged to a SageMaker AI MLflow App, and SNS alerts above a drift threshold.

  • CA weekly retraining schedule that runs whether or not the data has changed

    Incorrect

    Retraining on a timer may help, but it does not detect or report drift, store drift results, or alert anyone.

  • DMore instances behind the endpoint, chosen with auto scaling

    Incorrect

    Scaling changes capacity, not the data. It has no effect on whether the inputs still look like the training data.

Drift detection compares the statistics of live inputs with a training baseline. AWS documents an open-source path built on endpoint data capture, Evidently AI, MLflow, EventBridge, Lambda and SNS, with CloudWatch for system-level metrics.

Question 2 · choose 1

A retailer has two recommendation models and wants to know which one leads to more purchases. Purchases can only happen if customers actually see a model's recommendations. Both models are deployed on one SageMaker AI endpoint. How should the ML engineer compare them in production?

  1. ARun the second model as a shadow variant and compare its responses offline
  2. BTwo production variants with 90/10 traffic weights
  3. CScore last month's sessions with both models in batch transform jobs
  4. DCompare the models' validation AUC from training
Show the answer and why
  • ARun the second model as a shadow variant and compare its responses offline

    Incorrect

    Only the production variant's responses go back to callers in a shadow test, so customers never see the shadow model's recommendations and its effect on purchases cannot be measured.

  • BTwo production variants with 90/10 traffic weights

    Correct

    Production variants on a multi-variant endpoint split live invocations by the weights you set, so each model serves real customers and their purchases can be attributed to the variant that served them.

  • CScore last month's sessions with both models in batch transform jobs

    Incorrect

    Offline scoring shows what each model would have recommended, not whether customers would have bought anything.

  • DCompare the models' validation AUC from training

    Incorrect

    Offline metrics are useful for selection, but they do not measure the business outcome the retailer cares about.

A/B testing exposes real users to each model and measures business outcomes. Shadow testing is safer but can only compare technical behavior, because users never see the shadow responses.

Question 3 · choose 2

A chatbot calls Amazon Bedrock with the Converse API. The team needs two things: the full text of each prompt and response kept for weekly quality reviews, and an alert within minutes when Bedrock starts throttling the application's requests. Which actions meet these needs? (Choose TWO.)

  1. AEnable CloudTrail data events for Bedrock to record the prompts
  2. BRun Amazon Inspector on the application
  3. CTurn on Bedrock model invocation logging to CloudWatch Logs or S3
  4. DTurn on AWS Config recording for Bedrock resources
  5. ECreate a CloudWatch alarm on the InvocationThrottles metric
Show the answer and why
  • AEnable CloudTrail data events for Bedrock to record the prompts

    Incorrect

    CloudTrail records API activity, such as who called which operation and when. The full prompt and response text is what model invocation logging collects.

  • BRun Amazon Inspector on the application

    Incorrect

    Inspector scans workloads for software vulnerabilities. It neither keeps prompts nor watches throttling.

  • CTurn on Bedrock model invocation logging to CloudWatch Logs or S3

    Correct

    Model invocation logging collects the request data, response data and metadata of model invocations and delivers them to CloudWatch Logs, Amazon S3, or both. It is off by default.

  • DTurn on AWS Config recording for Bedrock resources

    Incorrect

    AWS Config records resource configurations and evaluates them against rules. It does not capture inference content or request-level throttling.

  • ECreate a CloudWatch alarm on the InvocationThrottles metric

    Correct

    Bedrock publishes runtime metrics such as InvocationThrottles to CloudWatch, so an alarm on that metric can notify the team soon after throttling begins.

For foundation model applications, keep three signals apart: runtime metrics in CloudWatch for health and throttling, invocation logs for content, and CloudTrail for who called what.

Question 4 · choose 1

An agent hosted on Amazon Bedrock AgentCore Runtime sometimes takes 40 seconds to answer, and users report occasional failed answers. The team needs to see each step of a slow or failed session, such as which tool call or model call took long or returned an error. What should the ML engineer use?

  1. AAWS CloudTrail management events for the account
  2. BAn Amazon Bedrock model evaluation job on sample prompts
  3. CAgentCore Memory with long-term memory turned on
  4. DAgentCore Observability traces in Amazon CloudWatch
Show the answer and why
  • AAWS CloudTrail management events for the account

    Incorrect

    CloudTrail records API calls made in the account. It does not show the internal steps, tool calls and latencies of an agent session.

  • BAn Amazon Bedrock model evaluation job on sample prompts

    Incorrect

    An evaluation job scores model responses offline. It does not trace live production sessions step by step.

  • CAgentCore Memory with long-term memory turned on

    Incorrect

    Memory stores conversation context and extracted insights. It is not a tracing or monitoring tool.

  • DAgentCore Observability traces in Amazon CloudWatch

    Correct

    AgentCore Observability visualizes each step of the agent workflow, so you can inspect the execution path, audit intermediate outputs and find bottlenecks and failures. Its metrics, spans and logs are stored in CloudWatch.

Agents fail in the middle: a slow tool, a failed call, a bad intermediate output. Step-level traces from AgentCore Observability, viewed in CloudWatch, are how those problems are found.

Question 5 · choose 1

A nightly SageMaker AI pipeline retrains and evaluates a model. Last week a processing step failed and nobody noticed for three days. The team wants an email whenever any execution of this pipeline fails, with the least custom code. What should the ML engineer set up?

  1. AStep caching turned on for every step of the nightly pipeline
  2. BAn EventBridge rule on failed pipeline executions, sending to SNS
  3. CAn AWS Config rule that evaluates the pipeline resource
  4. DA CloudTrail trail that records the StartPipelineExecution call
Show the answer and why
  • AStep caching turned on for every step of the nightly pipeline

    Incorrect

    Caching reuses earlier successful outputs. It does not notify anyone about a failure.

  • BAn EventBridge rule on failed pipeline executions, sending to SNS

    Correct

    SageMaker AI sends pipeline execution state change events to EventBridge, and a rule can send matching events to an SNS topic that emails the team.

  • CAn AWS Config rule that evaluates the pipeline resource

    Incorrect

    Config evaluates resource configurations against rules. It does not react to an individual pipeline execution failing.

  • DA CloudTrail trail that records the StartPipelineExecution call

    Incorrect

    The trail would show that a run started, not that a step inside it failed later.

Pipeline and job state changes arrive in EventBridge in near real time. Matching the failure states and routing them to SNS is the simplest failure alert.

Question 6 · choose 1

A model on a SageMaker AI real-time endpoint started returning errors after a new container version was deployed. The container prints a stack trace to stderr whenever a request fails. Where should the ML engineer look for these stack traces?

  1. AThe endpoint's log group in CloudWatch Logs
  2. BThe AWS CloudTrail event history for the endpoint
  3. CThe endpoint's Invocations metric in CloudWatch
  4. DThe S3 bucket that stores the model artifacts
Show the answer and why
  • AThe endpoint's log group in CloudWatch Logs

    Correct

    Anything a model container sends to stdout or stderr is also sent to CloudWatch Logs, in the /aws/sagemaker/Endpoints/[EndpointName] log group with a stream per variant and instance.

  • BThe AWS CloudTrail event history for the endpoint

    Incorrect

    CloudTrail records API activity in SageMaker AI, not what a container writes to stderr.

  • CThe endpoint's Invocations metric in CloudWatch

    Incorrect

    Metrics count requests and errors over time, but they do not contain the container's output.

  • DThe S3 bucket that stores the model artifacts

    Incorrect

    The artifact bucket holds model files; container output goes to CloudWatch Logs, not to that bucket.

Container stdout and stderr land in CloudWatch Logs per endpoint, variant and instance, which is the first place to debug inference errors.

Question 7 · choose 1

A serverless inference endpoint sometimes answers slowly after quiet periods. The ML engineer wants to chart, in CloudWatch, how long it takes to launch new compute resources for the endpoint, which depends on model size, model download and container start-up. Which metric shows this?

  1. AModelLatency
  2. BModelSetupTime
  3. COverheadLatency
  4. DInvocationsPerInstance
Show the answer and why
  • AModelLatency

    Incorrect

    ModelLatency is the time the model takes to respond as seen from SageMaker AI, including the inference in the container, not the time to launch compute.

  • BModelSetupTime

    Correct

    ModelSetupTime is the time it takes to launch new compute resources for a serverless endpoint, and it varies with model size, model download time and container start-up time.

  • COverheadLatency

    Incorrect

    OverheadLatency is the time SageMaker AI adds to a request beyond ModelLatency, not the time to launch compute resources.

  • DInvocationsPerInstance

    Incorrect

    This metric counts invocations per instance; it does not measure launch time.

For serverless endpoints, ModelSetupTime shows the cold-start cost of new compute, separately from model and overhead latency.

Question 8 · choose 1

A chatbot streams answers from an Amazon Bedrock model with ConverseStream. InvocationLatency has risen, but the team recently lengthened its system prompt and answers have grown longer. The team wants an alarm that fires only when the model itself generates tokens more slowly. What should the alarm be based on?

  1. AThe InvocationLatency p50 statistic alone
  2. BThe OutputTokenCount metric alone
  3. COutput tokens per second from metric math
  4. DThe InvocationThrottles metric
Show the answer and why
  • AThe InvocationLatency p50 statistic alone

    Incorrect

    InvocationLatency alone cannot tell slower generation apart from more tokens per request, which is exactly what changed here.

  • BThe OutputTokenCount metric alone

    Incorrect

    Output token count reflects how long answers are, not how fast the model generates them.

  • COutput tokens per second from metric math

    Correct

    Output tokens per second (OTPS) is OutputTokenCount divided by InvocationLatency minus TimeToFirstToken, times 1,000; it isolates throughput, so longer prompts or answers do not cause false alarms.

  • DThe InvocationThrottles metric

    Incorrect

    Throttles count requests the system rejected; they do not measure generation speed.

Separate workload changes from service changes: OTPS stays stable when only prompt or answer length grows. TimeToFirstToken is published for the streaming operations.

Question 9 · choose 2

An inference container writes the line LOW_CONFIDENCE to its endpoint log each time a prediction's top score is below 0.5. The team wants an email when more than 200 of these lines appear in an hour, with the least custom code. Which steps meet the need? (Choose TWO.)

  1. ASubscribe a Lambda function that counts lines
  2. BCreate a CloudTrail trail for the endpoint
  3. CCreate a metric filter on the endpoint's log group
  4. DAdd an AWS Config rule for the endpoint
  5. ECreate an alarm on that metric with an SNS action
Show the answer and why
  • ASubscribe a Lambda function that counts lines

    Incorrect

    Subscriptions can deliver log events to Lambda, but counting and alerting would then be custom code.

  • BCreate a CloudTrail trail for the endpoint

    Incorrect

    CloudTrail records API calls, not lines that a container writes to its log.

  • CCreate a metric filter on the endpoint's log group

    Correct

    Metric filters define terms and patterns to look for in log data and turn matching log events into CloudWatch metrics that you can graph or alarm on.

  • DAdd an AWS Config rule for the endpoint

    Incorrect

    Config evaluates resource configurations, not log contents.

  • ECreate an alarm on that metric with an SNS action

    Correct

    A CloudWatch alarm watches a metric against a threshold over a period and can notify an Amazon SNS topic, which can send email.

Turn log patterns into metrics with metric filters, then alarm on them, so model behavior signals need no custom code.

Question 10 · choose 1

After a client release, many Amazon Bedrock calls from an application fail. The team needs a quick signal of whether the failures come from requests the new client sends or from the service side. Which pair of CloudWatch metrics should the ML engineer compare?

  1. AInvocationThrottles and Invocations
  2. BInputTokenCount and OutputTokenCount per model ID
  3. CInvocationLatency and TimeToFirstToken by model
  4. DInvocationClientErrors and InvocationServerErrors
Show the answer and why
  • AInvocationThrottles and Invocations

    Incorrect

    Throttled requests are not counted in Invocations or in the error metrics, so this pair cannot separate client from server errors.

  • BInputTokenCount and OutputTokenCount per model ID

    Incorrect

    Token counts measure request and response volume, not where errors come from.

  • CInvocationLatency and TimeToFirstToken by model

    Incorrect

    Latency metrics measure speed, not the source of failures.

  • DInvocationClientErrors and InvocationServerErrors

    Correct

    InvocationClientErrors counts invocations that end in client-side errors and InvocationServerErrors counts those that end in AWS server-side errors, so comparing them shows where failures start.

Split error metrics by origin first: a jump in client errors after a client release points at the requests, not the service.

Practise domain 4 →Practise all domains →