Question 1 · choose 1
A summarization service on the bedrock-runtime endpoint receives ThrottlingException errors at only a third of its tokens-per-minute quota. Each request has about 3,000 input tokens, the summaries average 700 output tokens and never exceed 1,200, and the code sets max_tokens to 32,000 "to be safe". The model has a 5x output token burndown rate. What should the developer change first to serve more requests within the same quota?
- ARequest a higher tokens-per-minute quota and keep max_tokens at 32,000 for safety
- BMove the requests to the Priority service tier so that they are admitted ahead of other traffic
- CEnable prompt caching on the 3,000-token input so that cached input tokens no longer count toward the quota
- DLower max_tokens to slightly above the longest expected summary, such as 1,500 tokens
Show the answer and why
ARequest a higher tokens-per-minute quota and keep max_tokens at 32,000 for safety
Incorrect
A higher quota helps, but the oversized reservation keeps wasting most of it, so fixing max_tokens is the first step.
BMove the requests to the Priority service tier so that they are admitted ahead of other traffic
Incorrect
Priority affects ordering and price, and on-demand quota is shared across Priority, Standard and Flex, so the reservation problem remains.
CEnable prompt caching on the 3,000-token input so that cached input tokens no longer count toward the quota
Incorrect
Cache reads are not counted toward the quota, but each input here is a different document, and the 32,000-token output reservation would remain the main problem.
DLower max_tokens to slightly above the longest expected summary, such as 1,500 tokens
Correct
At the start of each request, the input tokens plus max_tokens are deducted from the quota, and unused tokens are returned only at the end. A 32,000 reservation limits concurrency even though actual use is far lower.
Quota is reserved up front as input tokens plus max_tokens and corrected at the end to input plus cache writes plus output times the burndown rate. Setting max_tokens close to real output size is the cheapest way to raise concurrency within the same quota.
AWS documentation