Question 1 · choose 1
After a prompt change, the InvocationLatency of a streaming chat feature on Amazon Bedrock rose by 40%. The team needs an alarm that fires when the model itself generates tokens more slowly, but not when responses simply get longer because of prompt or product changes. The application uses ConverseStream. What should the team alarm on?
- AOutputTokenCount at the average statistic, with a CloudWatch anomaly detection band
- BEstimatedTPMQuotaUsage with a threshold at 80% of the tokens-per-minute quota
- CInvocationLatency at p99 with a static threshold 40% above the old baseline
- DOutput tokens per second, computed with metric math from three runtime metrics
Show the answer and why
AOutputTokenCount at the average statistic, with a CloudWatch anomaly detection band
Incorrect
Output token count shows that responses got longer, but it says nothing about how fast tokens are generated.
BEstimatedTPMQuotaUsage with a threshold at 80% of the tokens-per-minute quota
Incorrect
This metric approximates quota consumption and is not meant as a latency or throughput signal.
CInvocationLatency at p99 with a static threshold 40% above the old baseline
Incorrect
InvocationLatency rises for both reasons, so it cannot tell slower generation apart from longer responses.
DOutput tokens per second, computed with metric math from three runtime metrics
Correct
Output tokens per second isolates generation speed: it stays stable when responses get longer and drops when the model slows down. It uses TimeToFirstToken, which streaming operations publish.
InvocationLatency is TimeToFirstToken plus output tokens divided by output tokens per second. Computing OTPS with metric math separates service-side throughput changes from workload changes, which makes a far better alarm.
AWS documentation