Skip to content
BytePatterns

AIP-C01 · Domain 4: Operational Efficiency and Optimization for GenAI Applications · 12% of the exam

Task 4.2: Optimize application performance.

Faster and higher-throughput GenAI: streaming, parallel calls, retrieval tuning, inference parameters and capacity planning for token-heavy traffic.

Study it

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

A reporting feature builds a quarterly review by calling a model six times, once for each section. The sections do not depend on one another, each call takes about 10 seconds, and the whole report takes a minute because the calls run one after another. The model's quota easily covers six concurrent requests. Which change reduces the report's total latency the most?

  1. ARun the six section calls concurrently, for example in an AWS Step Functions Parallel state
  2. BSend all six calls with the Flex service tier to get a lower price per token
  3. CSet a higher max_tokens value on each call so that the model never has to stop early in a section
  4. DCombine the six section prompts into one long request so that the model writes the whole report in one response
Show the answer and why
  • ARun the six section calls concurrently, for example in an AWS Step Functions Parallel state

    Correct

    Independent calls can run at the same time, so the total time approaches the slowest single call instead of the sum of all six.

  • BSend all six calls with the Flex service tier to get a lower price per token

    Incorrect

    Flex offers a discount for workloads that can accept longer processing times, so it can increase latency.

  • CSet a higher max_tokens value on each call so that the model never has to stop early in a section

    Incorrect

    max_tokens caps the output length. Raising it does not make generation faster.

  • DCombine the six section prompts into one long request so that the model writes the whole report in one response

    Incorrect

    Output tokens are generated sequentially, so one long response still takes about as long as the six responses combined.

For independent subtasks, parallel requests are the main latency lever. Step Functions Parallel (or concurrent SDK calls) runs them together, as long as the quota allows the concurrency.

Question 2 · choose 1

One application uses the same model for two tasks. Ticket classification returns different labels when the same ticket is submitted twice, which breaks reporting. Marketing tagline generation, in contrast, produces options that the marketing team finds too similar to each other. How should the developer configure inference parameters?

  1. AAdd the same stop sequence to both tasks so that responses end at a consistent point
  2. BUse a high temperature for classification and a low temperature for tagline generation
  3. CUse a low temperature for classification and a higher temperature for tagline generation
  4. DIncrease max_tokens for both tasks so that the model has room to settle on an answer
Show the answer and why
  • AAdd the same stop sequence to both tasks so that responses end at a consistent point

    Incorrect

    Stop sequences end generation at a marker. They do not affect which tokens are chosen before that point.

  • BUse a high temperature for classification and a low temperature for tagline generation

    Incorrect

    This is reversed: it makes classification more random and the taglines even more alike.

  • CUse a low temperature for classification and a higher temperature for tagline generation

    Correct

    A lower temperature favors high-probability tokens, giving more consistent output, while a higher temperature allows lower-probability tokens and more varied taglines.

  • DIncrease max_tokens for both tasks so that the model has room to settle on an answer

    Incorrect

    max_tokens limits length. It does not control how deterministic or varied the output is.

Temperature, top K and top P shape how the model samples tokens. Tasks that need repeatable answers use low randomness; creative tasks benefit from more. Length controls such as max_tokens and stop sequences are a separate concern.

Question 3 · choose 2

A knowledge base on an Amazon OpenSearch Serverless vector collection has p95 retrieval latency of 1.8 seconds at peak. CloudWatch shows that search OCUs sit at the collection group's configured maximum during peaks, while indexing OCUs are mostly idle. Each query retrieves 50 chunks that a reranker then reduces to 5, and evaluations show answers are just as good when 20 chunks are retrieved. Which actions reduce retrieval latency? (Choose TWO.)

  1. ARaise the maximum indexing OCUs of the collection group
  2. BSwitch the data source to hierarchical chunking so that each result contains more text
  3. CRaise the maximum search OCUs of the collection group so that search can scale at peak
  4. DRetrieve 20 chunks instead of 50 before reranking
  5. EMove the vectors to an Amazon S3 Vectors index to lower the storage cost
Show the answer and why
  • ARaise the maximum indexing OCUs of the collection group

    Incorrect

    Indexing capacity is idle, so more indexing OCUs do not help search latency.

  • BSwitch the data source to hierarchical chunking so that each result contains more text

    Incorrect

    Hierarchical chunking returns larger parent chunks, which adds tokens and does not address the capacity limit.

  • CRaise the maximum search OCUs of the collection group so that search can scale at peak

    Correct

    OpenSearch Serverless scales search capacity within the configured minimum and maximum. A ceiling that is reached at peak limits how far search can scale.

  • DRetrieve 20 chunks instead of 50 before reranking

    Correct

    Fewer candidates mean less work for the vector search and the reranker, and evaluations show no quality loss.

  • EMove the vectors to an Amazon S3 Vectors index to lower the storage cost

    Incorrect

    S3 Vectors is cost-effective for infrequently queried workloads, but it does not lower query latency for a busy interactive application.

Retrieval latency has two common causes: capacity limits in the vector store and too much work per query. Here search OCUs are capped at the maximum, and the candidate count is larger than quality requires.

Question 4 · choose 1

A support assistant's responses take a long time because the model writes long paragraphs, while users only need a short direct answer. Time to first token is fine. Which change reduces total response time most?

  1. AAsk for concise answers and cap output length
  2. BAdd more retrieved passages to the prompt
  3. CIncrease the temperature for every request
  4. DSwitch to a model with a larger context window
Show the answer and why
  • AAsk for concise answers and cap output length

    Correct

    Output tokens are generated one after another, so shorter outputs finish sooner, and maxTokens caps the length.

  • BAdd more retrieved passages to the prompt

    Incorrect

    More passages give the model more to write about, not less.

  • CIncrease the temperature for every request

    Incorrect

    Temperature affects variety, not length or speed.

  • DSwitch to a model with a larger context window

    Incorrect

    Context size does not make generation faster.

Generation time grows with output length. Ask for what users need and limit the rest.

Question 5 · choose 1

A SageMaker AI endpoint for a summarization model runs on a fixed number of instances. Daytime traffic overloads it, and nights are almost idle. The team wants capacity to follow the request rate automatically. What should it configure?

  1. AA second endpoint that the team switches to by hand
  2. BA Lambda function that queues requests in front of the endpoint
  3. CA larger instance type running all day
  4. DTarget tracking on InvocationsPerInstance
Show the answer and why
  • AA second endpoint that the team switches to by hand

    Incorrect

    Manual switching does not respond automatically to traffic.

  • BA Lambda function that queues requests in front of the endpoint

    Incorrect

    Queuing delays requests and does not add capacity to the endpoint.

  • CA larger instance type running all day

    Incorrect

    Bigger instances still sit idle at night and do not follow the load.

  • DTarget tracking on InvocationsPerInstance

    Correct

    A target tracking policy on InvocationsPerInstance adds and removes instances to keep the metric near the target value.

Auto scaling matches capacity to demand. Target tracking keeps a chosen metric near its target.

Question 6 · choose 1

A mobile app's backend runs a short orchestration for every user action: a Lambda function validates the input, a model on Amazon Bedrock drafts a reply, and the result is written to Amazon DynamoDB with a PUT. At peak there are about 4,000 runs per second, each run finishes in two to four seconds while the user waits for the result, and every step is safe to repeat. Orchestration cost must stay low, and logs are needed only to investigate failed runs. Which AWS Step Functions design fits best?

  1. AA Standard workflow for each user action, using its exactly-once execution model
  2. BA Synchronous Express workflow for each user action, with logging to CloudWatch Logs
  3. CA Standard workflow with a Distributed Map state that processes user actions in large batches
  4. DAn Amazon SQS queue in front of a Standard workflow to smooth out the peak load
Show the answer and why
  • AA Standard workflow for each user action, using its exactly-once execution model

    Incorrect

    Standard workflows are billed per state transition and are built for durable, auditable, non-idempotent work. This workload is short, idempotent and very high volume, which is the profile AWS describes for Express workflows instead.

  • BA Synchronous Express workflow for each user action, with logging to CloudWatch Logs

    Correct

    Express workflows are designed for high-volume event processing such as mobile backends and are billed by executions, duration and memory. The synchronous type returns the result to the waiting caller, and CloudWatch Logs captures the execution details.

  • CA Standard workflow with a Distributed Map state that processes user actions in large batches

    Incorrect

    Distributed Map is for large-scale parallel batch processing. Batching would delay each user's reply, and it does not lower the per-transition cost of Standard workflows.

  • DAn Amazon SQS queue in front of a Standard workflow to smooth out the peak load

    Incorrect

    A queue decouples work that can wait, but users are waiting for each result, and every run would still be billed as a Standard workflow.

Pick the workflow type by execution semantics, volume and how the caller receives results. Short, idempotent, high-volume request paths fit Synchronous Express workflows; long-running or non-idempotent processes fit Standard workflows.

Question 7 · choose 1

A research portal's Amazon Bedrock knowledge base holds 300,000 analyst reports, each with company, ticker, fiscal_year and region metadata. Users type free-text questions such as "Spanish telecom outlook for fiscal 2025", but retrieval returns chunks from other years and countries, so answers mix periods. The UI team will not add filter controls, the metadata schema is stable, one extra model call of about 300 ms per query is acceptable, and the reports carry no access restrictions. What should the developer configure?

  1. ARaise numberOfResults to 50 so that chunks from the requested year are more likely to be included
  2. BRe-ingest the reports with semantic chunking so that each chunk keeps one topic together
  3. CImplicit metadata filtering, with a description of each metadata attribute for the model
  4. DAdd the fiscal year and region as words at the top of every report before ingestion
Show the answer and why
  • ARaise numberOfResults to 50 so that chunks from the requested year are more likely to be included

    Incorrect

    More results add more chunks from the wrong years and countries; they do not restrict retrieval to the period in the question.

  • BRe-ingest the reports with semantic chunking so that each chunk keeps one topic together

    Incorrect

    Chunking changes how text is split, not which documents a query may match, so reports from other years still compete.

  • CImplicit metadata filtering, with a description of each metadata attribute for the model

    Correct

    With implicit filtering, the knowledge base uses a model and the attribute descriptions to generate a filter from each query, so "fiscal 2025" and "Spanish" become metadata conditions without UI changes.

  • DAdd the fiscal year and region as words at the top of every report before ingestion

    Incorrect

    The extra words may nudge similarity scores, but they do not limit results to the requested year or country.

Query preprocessing can turn constraints hidden in free text into structured filters, so retrieval is narrowed to the documents the user actually asked about instead of relying on similarity alone.

Question 8 · choose 1

A retailer shows an AI-written summary of customer reviews on each of 200,000 product pages, which get 2 million views a day. Generating one summary with a model on Amazon Bedrock takes about four seconds. Every page view must render within 300 ms, including the first view after a product's reviews change, while a summary may lag behind a new review by up to an hour. Reviews for a product change only a few times a week. Which design meets these requirements?

  1. AGenerate the summary on each page view with ConverseStream so that text appears while it is written
  2. BRegenerate a product's summary when its reviews change and store it in DynamoDB for the page to read
  3. CCache summaries in ElastiCache with a 24-hour TTL and generate one when a page view misses the cache
  4. DSend every summary request with the Priority service tier to get the fastest responses
Show the answer and why
  • AGenerate the summary on each page view with ConverseStream so that text appears while it is written

    Incorrect

    Streaming shows the first words sooner, but the full summary still takes seconds, and every one of the 2 million daily views pays for a model call.

  • BRegenerate a product's summary when its reviews change and store it in DynamoDB for the page to read

    Correct

    The summaries are predictable, so they can be computed ahead of time. Page views then read a stored item with single-digit millisecond latency, and model calls happen only when reviews change.

  • CCache summaries in ElastiCache with a 24-hour TTL and generate one when a page view misses the cache

    Incorrect

    With lazy loading, the first view after a miss or expiry waits for the model, which breaks the 300 ms target, and a summary could be up to a day out of date.

  • DSend every summary request with the Priority service tier to get the fastest responses

    Incorrect

    Priority processing speeds up inference but cannot bring a four-second generation under 300 ms, and it charges a premium on every page view.

When the questions are predictable, pre-computation moves model latency out of the user's request path entirely. Event-driven regeneration keeps the stored answers fresh while calling the model only when the inputs change.

Practise domain 4 →Practise all domains →