Skip to content
BytePatterns

AIP-C01 · Domain 4: Operational Efficiency and Optimization for GenAI Applications · 12% of the exam

Task 4.3: Implement monitoring systems for GenAI applications.

Seeing what the application does: Bedrock metrics and invocation logs, CloudWatch generative AI observability, agent and tool tracing, vector store health and GenAI-specific failure signals.

Study it

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

After a prompt change, the InvocationLatency of a streaming chat feature on Amazon Bedrock rose by 40%. The team needs an alarm that fires when the model itself generates tokens more slowly, but not when responses simply get longer because of prompt or product changes. The application uses ConverseStream. What should the team alarm on?

  1. AOutputTokenCount at the average statistic, with a CloudWatch anomaly detection band
  2. BEstimatedTPMQuotaUsage with a threshold at 80% of the tokens-per-minute quota
  3. CInvocationLatency at p99 with a static threshold 40% above the old baseline
  4. DOutput tokens per second, computed with metric math from three runtime metrics
Show the answer and why
  • AOutputTokenCount at the average statistic, with a CloudWatch anomaly detection band

    Incorrect

    Output token count shows that responses got longer, but it says nothing about how fast tokens are generated.

  • BEstimatedTPMQuotaUsage with a threshold at 80% of the tokens-per-minute quota

    Incorrect

    This metric approximates quota consumption and is not meant as a latency or throughput signal.

  • CInvocationLatency at p99 with a static threshold 40% above the old baseline

    Incorrect

    InvocationLatency rises for both reasons, so it cannot tell slower generation apart from longer responses.

  • DOutput tokens per second, computed with metric math from three runtime metrics

    Correct

    Output tokens per second isolates generation speed: it stays stable when responses get longer and drops when the model slows down. It uses TimeToFirstToken, which streaming operations publish.

InvocationLatency is TimeToFirstToken plus output tokens divided by output tokens per second. Computing OTPS with metric math separates service-side throughput changes from workload changes, which makes a far better alarm.

Question 2 · choose 1

An operations dashboard graphs InvocationClientErrors for each model ID of an Amazon Bedrock application. During an incident, the application logs showed thousands of failed requests, but the dashboard showed only a few hundred errors. The failing requests used a malformed model identifier. What should the team change so that the dashboard counts every failed request?

  1. AEnable model invocation logging so that failed requests are counted in CloudWatch metrics
  2. BGraph the error metrics without the ModelId dimension in addition to the per-model graphs
  3. CSwitch the dashboard to the InvocationThrottles metric, which includes all rejected requests
  4. DIncrease the metric period from one minute to one hour so that late data points are included
Show the answer and why
  • AEnable model invocation logging so that failed requests are counted in CloudWatch metrics

    Incorrect

    Invocation logging records requests and responses in logs. It does not change how the runtime metrics are published.

  • BGraph the error metrics without the ModelId dimension in addition to the per-model graphs

    Correct

    When a request fails before Bedrock determines its model, the metric is published without the ModelId dimension, so per-model graphs undercount errors. The dimensionless series includes them.

  • CSwitch the dashboard to the InvocationThrottles metric, which includes all rejected requests

    Incorrect

    Throttles count requests rejected for rate limits. Malformed requests are client errors, not throttles.

  • DIncrease the metric period from one minute to one hour so that late data points are included

    Incorrect

    A longer period aggregates the same data points. The missing errors were never published with the ModelId dimension.

Bedrock publishes its runtime metrics both with and without the ModelId dimension. Errors that occur before the model is resolved only appear in the series without a dimension, so error dashboards should include it.

Question 3 · choose 1

A team runs a multi-agent application built with an open-source framework on Amazon EKS, not on AgentCore Runtime. It wants end-to-end traces of each session, including model calls, tool calls and their latencies, in the CloudWatch generative AI observability dashboards alongside its AgentCore-hosted agents. What must the team do?

  1. ATurn on model invocation logging so that the dashboards can rebuild the session traces from the logs
  2. BMove the agents to AgentCore Runtime, because the dashboards cannot show agents that run elsewhere
  3. CEnable CloudWatch Transaction Search and instrument the agents with the ADOT SDK
  4. DCreate a CloudTrail trail with data events for every AWS service and tool that the agents call
Show the answer and why
  • ATurn on model invocation logging so that the dashboards can rebuild the session traces from the logs

    Incorrect

    Invocation logs capture individual model calls, not the spans of tool calls and agent steps that make up a session trace.

  • BMove the agents to AgentCore Runtime, because the dashboards cannot show agents that run elsewhere

    Incorrect

    AgentCore observability also supports agents hosted outside AgentCore when they are instrumented with ADOT.

  • CEnable CloudWatch Transaction Search and instrument the agents with the ADOT SDK

    Correct

    AgentCore observability needs Transaction Search to view spans and traces, and agents hosted outside AgentCore emit telemetry through the ADOT SDK with the framework's instrumentation.

  • DCreate a CloudTrail trail with data events for every AWS service and tool that the agents call

    Incorrect

    CloudTrail records API activity for auditing, not OpenTelemetry spans with latencies for each step of an agent session.

AgentCore observability is OpenTelemetry-based: turn on Transaction Search once per account, then instrument agents with ADOT. Agents on EKS, ECS, Lambda or EC2 appear in the same GenAI observability views as AgentCore-hosted agents.

Question 4 · choose 1

Users of a web chat assistant in 40 countries say that nothing happens for several seconds after they send a message, yet the backend's model latency metrics and API latency look normal. The team suspects delays inside the browser, such as slow page scripts, network distance and JavaScript errors on some devices. It needs measurements from real user sessions, broken down by country, browser and device, and it wants to sample only part of the sessions to control cost. What should the team use?

  1. AAlarms on the Amazon Bedrock runtime latency metrics for each model the assistant calls
  2. BAWS X-Ray active tracing on the backend Lambda functions and their downstream calls
  3. CStandard logs for the CloudFront distribution that serves the web application
  4. DA CloudWatch RUM app monitor added to the web client, sampling a share of user sessions
Show the answer and why
  • AAlarms on the Amazon Bedrock runtime latency metrics for each model the assistant calls

    Incorrect

    Runtime metrics measure the service side of each model call. They cannot see time spent in the user's browser or network.

  • BAWS X-Ray active tracing on the backend Lambda functions and their downstream calls

    Incorrect

    Backend traces show where server-side time goes, not page load times or JavaScript errors on the user's device.

  • CStandard logs for the CloudFront distribution that serves the web application

    Incorrect

    Distribution logs record requests that reach CloudFront. They do not capture rendering time or client-side errors in the browser.

  • DA CloudWatch RUM app monitor added to the web client, sampling a share of user sessions

    Correct

    RUM collects client-side performance data, errors and user behavior from real sessions, breaks it down by characteristics such as browser and device, shows user locations, and lets you choose the percentage of sessions to sample.

User interaction tracking complements service metrics: a model can respond quickly while users still wait. Real user monitoring measures the experience where it happens and can be correlated with backend traces.

Question 5 · choose 1

A GenAI chat API runs on Amazon EKS, and CloudWatch Application Signals already collects its latency and availability. Leadership has agreed that 99% of chat requests in any rolling 28 days must complete within 8 seconds. It wants a dashboard of how much of that error budget remains, and an early warning when the budget is being used up quickly, rather than an alert only after the target is missed. The team does not want to write code that computes the indicator. What should the team configure?

  1. AA static CloudWatch alarm that fires when the service's p99 latency exceeds 8 seconds in a 5-minute period
  2. BA request-based Application Signals SLO with an 8-second latency threshold, a 99% goal and burn rate alarms
  3. CA CloudWatch dashboard that shows the service's average latency for each day of the month
  4. DContainer Insights for the EKS cluster, with alarms on pod CPU and memory utilization
Show the answer and why
  • AA static CloudWatch alarm that fires when the service's p99 latency exceeds 8 seconds in a 5-minute period

    Incorrect

    A threshold alarm reacts to individual periods. It does not track the 28-day goal, the remaining error budget or how fast it is being used.

  • BA request-based Application Signals SLO with an 8-second latency threshold, a 99% goal and burn rate alarms

    Correct

    An SLO tracks the indicator against a threshold and an attainment goal over a rolling interval, reports the remaining error budget, and can raise burn rate alarms before the budget is exhausted.

  • CA CloudWatch dashboard that shows the service's average latency for each day of the month

    Incorrect

    A daily average hides slow requests and gives no error budget or early warning.

  • DContainer Insights for the EKS cluster, with alarms on pod CPU and memory utilization

    Incorrect

    Resource metrics show cluster health, not whether requests meet the latency objective that leadership agreed.

Service level objectives turn an operational target into an error budget that can be watched and spent deliberately. Burn rate alarms warn early enough to act before the objective is missed.

Question 6 · choose 2

A RAG service stores 60 million vectors in an approximate k-NN (HNSW) index on an Amazon OpenSearch Service domain, and the index grows by about one million vectors a week. Twice last quarter, query latency spiked and new vectors were rejected before anyone noticed. The team wants one alarm that warns before the vector graphs outgrow the memory set aside for them, and one that fires when graphs start being pushed out of memory and reloaded. Which CloudWatch metrics should these alarms use? (Choose TWO.)

  1. AJVMMemoryPressure on the data nodes
  2. BKNNGraphMemoryUsagePercentage on each data node
  3. CSysMemoryUtilization on the data nodes
  4. DCPUUtilization of the data nodes
  5. EKNNEvictionCount on each data node
Show the answer and why
  • AJVMMemoryPressure on the data nodes

    Incorrect

    JVMMemoryPressure measures the Java heap. k-NN graphs live in native memory outside the heap, so this metric does not track them.

  • BKNNGraphMemoryUsagePercentage on each data node

    Correct

    This metric is the share of native memory used by k-NN graphs relative to the circuit breaker limit. At 100% the circuit breaker triggers and new vector indexing is rejected, so an alarm well below that warns in time.

  • CSysMemoryUtilization on the data nodes

    Incorrect

    High values of this metric are normal and usually do not indicate a problem, so it cannot warn about graph memory specifically.

  • DCPUUtilization of the data nodes

    Incorrect

    CPU load can rise for many reasons and says nothing about whether the graphs fit in their memory budget.

  • EKNNEvictionCount on each data node

    Correct

    This metric counts graphs evicted from the cache because of memory constraints or idle time. A rising count means graphs are being reloaded, which slows queries.

Vector stores have failure modes that general infrastructure metrics miss. Watching graph memory against the circuit breaker limit and the eviction rate gives early warning, so the team can scale the domain or reduce vector memory before users notice.

Question 7 · choose 1

During a promotion, some customers saw errors from an assistant that calls two models on Amazon Bedrock. The application's SDK retries failed calls, so its own logs show only the final failures. The team must confirm whether Bedrock throttled requests, as opposed to server-side errors or a client bug, see the trend for each model on a dashboard, and be alarmed early during the next promotion. Which runtime metric should the dashboard and alarm use?

  1. AInvocationThrottles for each model ID
  2. BInvocationServerErrors for each model ID
  3. CInvocations for each model ID, watching for a drop during the peak
  4. DEstimatedTPMQuotaUsage for each model ID
Show the answer and why
  • AInvocationThrottles for each model ID

    Correct

    This metric counts invocations that the service throttled, and throttled requests are not counted in the other invocation and error metrics, so it isolates this failure mode.

  • BInvocationServerErrors for each model ID

    Incorrect

    Server errors count AWS server-side failures. Throttled requests are not included, so this metric cannot confirm throttling.

  • CInvocations for each model ID, watching for a drop during the peak

    Incorrect

    Invocations does not include throttled requests, so a drop cannot be told apart from lower demand.

  • DEstimatedTPMQuotaUsage for each model ID

    Incorrect

    This metric estimates tokens-per-minute quota consumption. It does not count the requests that were actually throttled.

Choose the metric that isolates the failure mode. Throttles are counted separately from accepted invocations and from client and server errors, and the count you see also depends on the SDK's retry settings.

Question 8 · choose 1

A team self-hosts an open-weight LLM on Amazon EKS GPU nodes that use the EKS-optimized accelerated AMI with the NVIDIA device plugin installed. To tune its batching and decide when to add nodes, the team needs GPU utilization and GPU memory metrics for each node and each GPU device in CloudWatch, next to its pod and container metrics. It does not want to run and maintain its own metrics collection stack. What should the team do?

  1. ATurn on detailed monitoring for the GPU instances so that EC2 publishes metrics every minute
  2. BUse the Amazon Bedrock runtime metrics for the model in the AWS/Bedrock namespace
  3. CInstall the CloudWatch Observability EKS add-on with Container Insights enhanced observability
  4. DInstrument the model server with AWS X-Ray and analyze the time spent in each request segment
Show the answer and why
  • ATurn on detailed monitoring for the GPU instances so that EC2 publishes metrics every minute

    Incorrect

    Detailed monitoring publishes the standard EC2 metrics more often. It does not add GPU utilization or GPU memory metrics.

  • BUse the Amazon Bedrock runtime metrics for the model in the AWS/Bedrock namespace

    Incorrect

    Bedrock runtime metrics describe calls to models served by Amazon Bedrock, not a model that the team hosts on its own EKS nodes.

  • CInstall the CloudWatch Observability EKS add-on with Container Insights enhanced observability

    Correct

    Container Insights with enhanced observability for EKS collects NVIDIA GPU metrics, such as node and pod GPU utilization and memory, with dimensions down to the GPU device, alongside the cluster, pod and container metrics.

  • DInstrument the model server with AWS X-Ray and analyze the time spent in each request segment

    Incorrect

    Traces show where a request spends time, but they do not report how busy each GPU is or how much GPU memory is in use.

Self-hosted models need accelerator-level observability that managed model services provide for you. Collecting GPU metrics with the cluster's other metrics shows whether capacity is saturated before latency suffers.

Practise domain 4 →Practise all domains →