Question 1 · choose 1
A reporting feature builds a quarterly review by calling a model six times, once for each section. The sections do not depend on one another, each call takes about 10 seconds, and the whole report takes a minute because the calls run one after another. The model's quota easily covers six concurrent requests. Which change reduces the report's total latency the most?
- ARun the six section calls concurrently, for example in an AWS Step Functions Parallel state
- BSend all six calls with the Flex service tier to get a lower price per token
- CSet a higher max_tokens value on each call so that the model never has to stop early in a section
- DCombine the six section prompts into one long request so that the model writes the whole report in one response
Show the answer and why
ARun the six section calls concurrently, for example in an AWS Step Functions Parallel state
Correct
Independent calls can run at the same time, so the total time approaches the slowest single call instead of the sum of all six.
BSend all six calls with the Flex service tier to get a lower price per token
Incorrect
Flex offers a discount for workloads that can accept longer processing times, so it can increase latency.
CSet a higher max_tokens value on each call so that the model never has to stop early in a section
Incorrect
max_tokens caps the output length. Raising it does not make generation faster.
DCombine the six section prompts into one long request so that the model writes the whole report in one response
Incorrect
Output tokens are generated sequentially, so one long response still takes about as long as the six responses combined.
For independent subtasks, parallel requests are the main latency lever. Step Functions Parallel (or concurrent SDK calls) runs them together, as long as the quota allows the concurrency.
AWS documentation