Skip to content
BytePatterns

AIP-C01 · Domain 4: Operational Efficiency and Optimization for GenAI Applications · 12% of the exam

Task 4.1: Implement cost optimization and resource efficiency strategies.

Spending fewer tokens and less capacity: token budgets, model tiering, batch inference and service tiers, prompt caching and semantic caching, and cost attribution per application.

Study it

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

A summarization service on the bedrock-runtime endpoint receives ThrottlingException errors at only a third of its tokens-per-minute quota. Each request has about 3,000 input tokens, the summaries average 700 output tokens and never exceed 1,200, and the code sets max_tokens to 32,000 "to be safe". The model has a 5x output token burndown rate. What should the developer change first to serve more requests within the same quota?

  1. ARequest a higher tokens-per-minute quota and keep max_tokens at 32,000 for safety
  2. BMove the requests to the Priority service tier so that they are admitted ahead of other traffic
  3. CEnable prompt caching on the 3,000-token input so that cached input tokens no longer count toward the quota
  4. DLower max_tokens to slightly above the longest expected summary, such as 1,500 tokens
Show the answer and why
  • ARequest a higher tokens-per-minute quota and keep max_tokens at 32,000 for safety

    Incorrect

    A higher quota helps, but the oversized reservation keeps wasting most of it, so fixing max_tokens is the first step.

  • BMove the requests to the Priority service tier so that they are admitted ahead of other traffic

    Incorrect

    Priority affects ordering and price, and on-demand quota is shared across Priority, Standard and Flex, so the reservation problem remains.

  • CEnable prompt caching on the 3,000-token input so that cached input tokens no longer count toward the quota

    Incorrect

    Cache reads are not counted toward the quota, but each input here is a different document, and the 32,000-token output reservation would remain the main problem.

  • DLower max_tokens to slightly above the longest expected summary, such as 1,500 tokens

    Correct

    At the start of each request, the input tokens plus max_tokens are deducted from the quota, and unused tokens are returned only at the end. A 32,000 reservation limits concurrency even though actual use is far lower.

Quota is reserved up front as input tokens plus max_tokens and corrected at the end to input plus cache writes plus output times the burndown rate. Setting max_tokens close to real output size is the cheapest way to raise concurrency within the same quota.

Question 2 · choose 2

A compliance assistant sends the same 30,000-token rulebook and instructions with every question, followed by the user's question. Traffic is steady, with a request every few seconds, and the model supports explicit prompt caching through the Converse API. The team wants to cut input cost and latency and to verify that the change works. Which actions should the developer take? (Choose TWO.)

  1. ACheck cacheReadInputTokens and cacheWriteInputTokens in the response usage to confirm cache hits
  2. BSubmit the questions through batch inference so that the cached rulebook is shared across records
  3. CPut the user's question first and the rulebook after it so that the newest content is read first
  4. DSet the cache time to live to 24 hours so that the rulebook stays cached overnight
  5. EPut the rulebook and instructions at the start of the prompt and add a cachePoint block right after them
Show the answer and why
  • ACheck cacheReadInputTokens and cacheWriteInputTokens in the response usage to confirm cache hits

    Correct

    The usage fields report how many tokens were read from and written to the cache, which shows whether requests hit the cache.

  • BSubmit the questions through batch inference so that the cached rulebook is shared across records

    Incorrect

    Prompt caching is supported only for on-demand inference, not with the batch inference API, and batch does not suit interactive questions.

  • CPut the user's question first and the rulebook after it so that the newest content is read first

    Incorrect

    A prefix that changes with every request causes cache misses, because caching reuses an identical prompt prefix.

  • DSet the cache time to live to 24 hours so that the rulebook stays cached overnight

    Incorrect

    Cache TTLs are short, such as 5 minutes or 1 hour depending on the model, and reset with each hit. A steady request rate keeps the cache warm anyway.

  • EPut the rulebook and instructions at the start of the prompt and add a cachePoint block right after them

    Correct

    Explicit caching caches the prefix up to the checkpoint. Static content must come first and stay identical, with the changing question after the checkpoint.

Prompt caching rewards stable prefixes. Order the prompt static-first, place the checkpoint after the static block, and confirm hits in the usage fields or the CacheReadInputTokenCount metric. Cache reads are billed at a lower rate and do not count toward the tokens-per-minute quota.

Question 3 · choose 1

An internal IT help chatbot receives 40,000 questions a day, and analysis shows that most of them are rephrasings of about 300 common questions, such as several ways of asking how to set up the VPN. Answers change only when IT updates a procedure, roughly weekly. The company wants to stop paying for model calls on repeated questions while keeping answers accurate. Which solution meets these requirements?

  1. AA semantic cache of question embeddings and answers in Amazon ElastiCache for Valkey, with a TTL
  2. BPrompt caching of the system prompt so that each question costs fewer input tokens
  3. CAmazon CloudFront in front of the chat API so that identical requests are served from edge caches
  4. DAn exact-match cache in Amazon DynamoDB keyed by a hash of each question's text, with a one-week TTL
Show the answer and why
  • AA semantic cache of question embeddings and answers in Amazon ElastiCache for Valkey, with a TTL

    Correct

    A semantic cache matches new questions to earlier ones by meaning and returns the stored answer above a similarity threshold, avoiding the model call. A TTL limits how long answers are reused.

  • BPrompt caching of the system prompt so that each question costs fewer input tokens

    Incorrect

    Prompt caching lowers the cost of a repeated prefix, but every question still triggers a full model invocation and output generation.

  • CAmazon CloudFront in front of the chat API so that identical requests are served from edge caches

    Incorrect

    CloudFront caches responses by cache key, such as the URL and selected headers or query strings, so paraphrased questions do not match.

  • DAn exact-match cache in Amazon DynamoDB keyed by a hash of each question's text, with a one-week TTL

    Incorrect

    Hash keys only match identical strings, so rephrased questions, the majority here, would still invoke the model.

When users ask the same thing in different words, only a semantic cache avoids the model call. Start with a strict similarity threshold, lower it while watching accuracy, and set TTLs that match how often the answers change.

Question 4 · choose 1

A company runs nightly agentic research jobs and weekly content summaries on Amazon Bedrock. The work can take longer to complete without harm, and the team wants a discount compared with standard on-demand pricing without changing to batch files. Which option fits?

  1. ABuy Provisioned Throughput for the nightly window
  2. BSend these requests with the Priority service tier
  3. CRaise maxTokens so that each job finishes in fewer calls
  4. DSend these requests with the Flex service tier
Show the answer and why
  • ABuy Provisioned Throughput for the nightly window

    Incorrect

    Provisioned Throughput reserves capacity and is billed while it exists, not a discount for flexible work.

  • BSend these requests with the Priority service tier

    Incorrect

    Priority costs more than standard pricing in exchange for speed.

  • CRaise maxTokens so that each job finishes in fewer calls

    Incorrect

    A higher output limit does not lower the price per token.

  • DSend these requests with the Flex service tier

    Correct

    Flex offers a pricing discount for workloads that can handle longer processing times, such as summarization and agentic workflows.

Match urgency to price. Work that can wait fits the Flex tier's discount.

Question 5 · choose 1

An invoice-extraction service sends 3 million requests a month to a large model on Amazon Bedrock, and the large model is accurate enough. A smaller model from the same family costs far less per token but scores 12 points lower on the team's test set. The team has no human-labeled outputs, but model invocation logging has stored months of production prompts and the large model's responses, tagged with request metadata for this use case. The team has no capacity to run its own training infrastructure. Which approach best lowers the cost per request while keeping accuracy close to the large model?

  1. APurchase Provisioned Throughput for the large model and route all extraction traffic through it
  2. BA Bedrock Model Distillation job with the large model as teacher and the small model as student
  3. CSwitch to the small model and add twenty worked examples from the logs to every prompt
  4. DFine-tune the small model on a few hundred invoices that analysts label by hand first
Show the answer and why
  • APurchase Provisioned Throughput for the large model and route all extraction traffic through it

    Incorrect

    Provisioned Throughput reserves dedicated capacity, but every request would still run on the large model, so the cost gap to the small model remains.

  • BA Bedrock Model Distillation job with the large model as teacher and the small model as student

    Correct

    Distillation fine-tunes the smaller student on the teacher's responses for this use case. It can use prompt-response pairs from invocation logs, filtered by request metadata, so no labeled data or training infrastructure is needed.

  • CSwitch to the small model and add twenty worked examples from the logs to every prompt

    Incorrect

    Twenty examples add input tokens to every one of 3 million calls, which eats into the saving, and they do not train the small model on the large model's behavior for this task.

  • DFine-tune the small model on a few hundred invoices that analysts label by hand first

    Incorrect

    The team has no labeled data and no time to create it, and a few hundred examples is far less than the logged production traffic.

When a large model is good enough but too expensive, transfer its behavior for the specific task to a smaller model. Distillation automates that transfer from teacher responses, which production logs may already hold.

Question 6 · choose 1

Every night a partner system re-exports all 500,000 product manuals to an Amazon S3 bucket and rewrites every object, although only about 2% of the manuals actually change. A workflow then summarizes each manual with a model on Amazon Bedrock and stores the summaries for a support assistant. Model costs are dominated by summarizing manuals that did not change. The partner's export process cannot be changed, and the summary of an unchanged manual must stay exactly the same. What should the developer do?

  1. AStart summarization from S3 event notifications so that only new or updated objects are processed
  2. BSubmit the nightly summaries as an Amazon Bedrock batch inference job instead of individual calls
  3. CKeep a content checksum with each summary and call the model only when the checksum changes
  4. DEnable prompt caching for the summarization instructions that precede each manual's text
Show the answer and why
  • AStart summarization from S3 event notifications so that only new or updated objects are processed

    Incorrect

    Because the export rewrites every object each night, a notification fires for all 500,000 manuals and nothing is skipped.

  • BSubmit the nightly summaries as an Amazon Bedrock batch inference job instead of individual calls

    Incorrect

    Batch inference lowers the price per token, but it still summarizes every manual, including the 98% that did not change.

  • CKeep a content checksum with each summary and call the model only when the checksum changes

    Correct

    A checksum of the content stays the same when an unchanged manual is rewritten. Amazon S3 stores one for objects uploaded through AWS clients, or the workflow can compute it, and a match means the model call is skipped and the stored summary is kept.

  • DEnable prompt caching for the summarization instructions that precede each manual's text

    Incorrect

    Caching can only reuse the shared instructions. Each manual is still sent and summarized, and its text is different in every request.

The cheapest model call is the one you do not make. Fingerprinting the input detects real changes even when the upstream system rewrites everything, and it keeps existing outputs stable.

Question 7 · choose 1

A ticket-triage service makes 2 million model calls a day. The model writes the category on the first line and then a paragraph that starts with "Rationale:" and runs 150 to 400 tokens, which the service discards. Category names range from one to twelve words, and output tokens cost several times more than input tokens. The prompt template belongs to a compliance team and cannot be changed this quarter, and categories must never be cut off. Which change reduces output cost the most?

  1. AAdd "Rationale:" as a stop sequence in the inference configuration of each request
  2. BSet the maximum response length to 15 tokens on every request to cut off the rationale
  3. CLower the temperature to 0 so that the model writes shorter responses
  4. DEnable prompt caching for the fixed template that precedes each ticket
Show the answer and why
  • AAdd "Rationale:" as a stop sequence in the inference configuration of each request

    Correct

    Generation stops as soon as the model produces the stop sequence, so the discarded paragraph is never generated or billed, whatever the length of the category line before it.

  • BSet the maximum response length to 15 tokens on every request to cut off the rationale

    Incorrect

    A fixed cap would also cut off the longest category names, which the service must never do.

  • CLower the temperature to 0 so that the model writes shorter responses

    Incorrect

    Temperature changes how likely tokens are chosen, not how long the response is, so the rationale is still generated.

  • DEnable prompt caching for the fixed template that precedes each ticket

    Incorrect

    Prompt caching lowers the cost of repeated input tokens. The expensive part here is the output, which caching does not reduce.

Response size controls are a direct lever on token cost. A stop sequence ends generation at a known marker, which suits outputs whose useful part varies in length; a fixed token cap suits outputs of predictable size.

Question 8 · choose 1

A legal research assistant sends a 40,000-token case file with every question in an attorney's session. Explicit prompt caching is already configured for the case file, and questions asked within a few minutes of each other read it from the cache. Attorneys usually pause 10 to 30 minutes between questions while they read, and billing shows that most follow-up questions write the case file to the cache again. The model's card lists both 5-minute and 1-hour cache TTLs, and the team must not add model calls. What should the developer change?

  1. ASend a short keep-alive request in each open session every four minutes to keep the cache warm
  2. BAdd more cache checkpoints inside the case file so that smaller parts of it are cached
  3. CSet the ttl field of the existing cache checkpoint to 1h for these sessions
  4. DMove the follow-up questions to Amazon Bedrock batch inference jobs
Show the answer and why
  • ASend a short keep-alive request in each open session every four minutes to keep the cache warm

    Incorrect

    A cache hit does reset the TTL, but this adds a model call to every session every four minutes, which the team ruled out.

  • BAdd more cache checkpoints inside the case file so that smaller parts of it are cached

    Incorrect

    Extra checkpoints change which prefixes can be cached, not how long a cache entry lives, so the entries still expire during the pauses.

  • CSet the ttl field of the existing cache checkpoint to 1h for these sessions

    Correct

    For models that support it, a 1-hour TTL keeps the cached prefix across pauses longer than 5 minutes, which AWS describes as the case for prompts reused less often than every 5 minutes but within the hour.

  • DMove the follow-up questions to Amazon Bedrock batch inference jobs

    Incorrect

    Prompt caching is not supported with the batch inference API, and attorneys need their answers interactively.

Match the cache TTL to the traffic pattern. The 5-minute TTL keeps refreshing under steady traffic, while sessions with longer pauses need the 1-hour TTL to avoid writing the same prefix to the cache again and again.

Practise domain 4 →Practise all domains →