Skip to content
BytePatterns

AIP-C01 · Domain 5: Testing, Validation, and Troubleshooting · 11% of the exam

Task 5.2: Troubleshoot GenAI applications.

Finding the cause when GenAI misbehaves: context and truncation problems, API errors and throttling, prompt regressions, retrieval and embedding issues, and format drift.

Study it

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

A contract summarizer that uses the Converse API sometimes returns summaries that stop mid-sentence. The responses have no error, the input is well within the context window, and the stopReason field of the affected responses is max_tokens. What is the most likely cause and fix?

  1. AA guardrail intervened in the output, so review the guardrail's content filter strengths
  2. BThe output reached the maxTokens limit, so raise it or ask for a shorter summary
  3. CThe input exceeded the context window, so split the contract into smaller chunks
  4. DA stop sequence in the prompt ended generation, so remove all stop sequences
Show the answer and why
  • AA guardrail intervened in the output, so review the guardrail's content filter strengths

    Incorrect

    A guardrail intervention reports a different stop reason and replaces the output with the blocked message.

  • BThe output reached the maxTokens limit, so raise it or ask for a shorter summary

    Correct

    A max_tokens stop reason means the model hit the output limit set in inferenceConfig, so the limit or the requested length must change.

  • CThe input exceeded the context window, so split the contract into smaller chunks

    Incorrect

    The input is within the context window, and an oversized input is rejected with a validation error rather than cut off silently.

  • DA stop sequence in the prompt ended generation, so remove all stop sequences

    Incorrect

    A stop sequence ends generation with a stop_sequence reason, not max_tokens.

Check stopReason first when output looks cut off. max_tokens means the output budget ran out; end_turn means the model finished; other values point to stop sequences, tool use or guardrail interventions.

Question 2 · choose 1

An application in us-east-1 switched from a single-Region model ID to the US geographic cross-Region inference profile for the same model. Its IAM role allows bedrock:InvokeModel on the inference profile ARN and on the foundation model ARN in us-east-1. Many requests now succeed, but others fail intermittently with AccessDeniedException. What should the developer do?

  1. ARequest a higher cross-Region requests-per-minute quota for the model in us-east-1
  2. BEnable the destination Regions in the account settings so that requests can be routed to them
  3. CReplace the inference profile with the global profile, which needs only the profile permission
  4. DAllow bedrock:InvokeModel on the model ARN in every destination Region of the profile
Show the answer and why
  • ARequest a higher cross-Region requests-per-minute quota for the model in us-east-1

    Incorrect

    Quota problems return throttling errors, not access denied errors.

  • BEnable the destination Regions in the account settings so that requests can be routed to them

    Incorrect

    Cross-Region inference can route to Regions that are not manually enabled, so Region opt-in is not the cause.

  • CReplace the inference profile with the global profile, which needs only the profile permission

    Incorrect

    The global profile also needs model permissions, and it would route requests worldwide, which changes the data residency posture.

  • DAllow bedrock:InvokeModel on the model ARN in every destination Region of the profile

    Correct

    Requests through a geographic profile can be routed to any destination Region in it, so the role needs permission for the foundation model in each of those Regions as well as for the profile.

Intermittent AccessDeniedException after moving to cross-Region inference usually means the policy covers the source Region only. Grant the inference profile plus the foundation model in all of the profile's destination Regions, optionally conditioned on the profile ARN.

Question 3 · choose 1

A service on Amazon EKS calls Amazon Bedrock through an interface VPC endpoint. After quiet periods of several minutes, the first request often fails with a connection reset or hangs for more than a minute before it succeeds, while requests during busy periods are fine. Bedrock metrics show no errors or throttles. What should the developer do?

  1. AIncrease the SDK's maximum retry attempts so that a reset connection is retried more often
  2. BPurchase Provisioned Throughput so that capacity stays warm during quiet periods
  3. CReplace the interface VPC endpoint with a NAT gateway so that connections use the public endpoint
  4. DEnable TCP keep-alive in the SDK and set the OS keep-alive time below 350 seconds
Show the answer and why
  • AIncrease the SDK's maximum retry attempts so that a reset connection is retried more often

    Incorrect

    Retries eventually succeed, but each first call still waits for the dead connection to time out, so the delay remains.

  • BPurchase Provisioned Throughput so that capacity stays warm during quiet periods

    Incorrect

    Bedrock reports no errors or throttles. The dropped connection is on the network path, not a capacity problem.

  • CReplace the interface VPC endpoint with a NAT gateway so that connections use the public endpoint

    Incorrect

    NAT gateways have the same 350-second idle timeout, so the problem would remain.

  • DEnable TCP keep-alive in the SDK and set the OS keep-alive time below 350 seconds

    Correct

    Interface endpoints, NAT gateways and Network Load Balancers drop idle TCP connections after 350 seconds. Keep-alive needs both the SDK setting and an OS interval shorter than that timeout.

Long-idle pooled connections are silently dropped by NAT gateways, VPC interface endpoints and NLBs after 350 seconds. Turning on SDK keep-alive alone is not enough, because the Linux default keep-alive interval is two hours; lower the kernel setting as well.

Question 4 · choose 1

A team runs its own RAG pipeline: documents were embedded last year with one embedding model and stored in an Amazon OpenSearch Service index. This month the team switched the query path to a newer embedding model with the same vector dimension, and retrieval relevance collapsed even for simple questions. Nothing else changed. What should the team do?

  1. AIncrease the number of results per query so that relevant documents are more likely to appear
  2. BSwitch the index to hybrid search so that keyword matching compensates for the new model
  3. CRe-embed the documents with the new model into a new index and query that index
  4. DAdd a reranker model after retrieval so that the best documents are moved to the top
Show the answer and why
  • AIncrease the number of results per query so that relevant documents are more likely to appear

    Incorrect

    More results from incompatible vectors are still poorly ranked, so this treats the symptom without fixing the mismatch.

  • BSwitch the index to hybrid search so that keyword matching compensates for the new model

    Incorrect

    Keyword matching can mask part of the problem, but the vector half of the search would stay broken.

  • CRe-embed the documents with the new model into a new index and query that index

    Correct

    Query and document vectors must come from the same embedding model to be comparable. Vectors from different models live in different spaces, even when the dimensions match.

  • DAdd a reranker model after retrieval so that the best documents are moved to the top

    Incorrect

    A reranker can only reorder what retrieval returned, and mismatched vectors return the wrong candidates.

Embedding spaces are model-specific. Changing the query model without re-embedding the corpus breaks similarity search. In Bedrock knowledge bases, queries are converted with the same embedding model used during ingestion, and changing the embedding approach means creating a new knowledge base.

Question 5 · choose 1

A field-inspection app sends technicians' photos to a model through the Converse API as image content blocks. Photos from Android phones work, but every photo taken on iPhones fails at once with a ValidationException, including small ones well under the size limits. Logging shows that the failing files are HEIC images and that the code passes the file extension as the image format. The model must keep receiving the photos as images. What should the developer change?

  1. ASend the iPhone photos as document content blocks instead of image blocks
  2. BDownscale the HEIC photos before sending them, keeping their original format
  3. CConvert HEIC photos to JPEG or PNG and set the matching image format
  4. DRetry the failed calls with exponential backoff and jitter in the SDK client
Show the answer and why
  • ASend the iPhone photos as document content blocks instead of image blocks

    Incorrect

    Document blocks accept formats such as PDF, CSV, Word, Excel, HTML, text and Markdown, not photos, and the model must receive images.

  • BDownscale the HEIC photos before sending them, keeping their original format

    Incorrect

    The small photos fail too, so size is not the cause. The format value itself is rejected.

  • CConvert HEIC photos to JPEG or PNG and set the matching image format

    Correct

    The image block's format accepts only png, jpeg, gif and webp. Converting the photos and declaring a supported format fixes the validation error.

  • DRetry the failed calls with exponential backoff and jitter in the SDK client

    Incorrect

    A validation error means the input breaks a constraint of the API. Retrying the same request fails the same way.

Validate request inputs against the API's allowed values before calling the model. A ValidationException that appears for one input type points to a request constraint, not to capacity or transient faults.

Question 6 · choose 1

A team created its own vector index in an Amazon OpenSearch Serverless collection and connected it to an Amazon Bedrock knowledge base, choosing the nmslib engine for the vector field. Unfiltered retrieval works, and every source document has a valid .metadata.json file with a department attribute. However, metadata filtering on Retrieve requests does not work for this knowledge base. The team must keep OpenSearch Serverless and must filter results by department. What should the developer do?

  1. AStart a new sync of the data source so that the metadata attributes are indexed again
  2. BCreate a faiss-engine vector index, attach it to a new knowledge base and ingest again
  3. CCopy the department name from each metadata file into the document text before ingestion
  4. DChange the space type of the vector field to cosine similarity and rebuild the index
Show the answer and why
  • AStart a new sync of the data source so that the metadata attributes are indexed again

    Incorrect

    The metadata is already ingested. Filtering requires a supported index engine, which a resync does not change.

  • BCreate a faiss-engine vector index, attach it to a new knowledge base and ingest again

    Correct

    Metadata filtering on an OpenSearch Serverless vector store needs the index to use the faiss engine. With nmslib, AWS directs you to create a new index with faiss, or let Bedrock create one, and a new knowledge base.

  • CCopy the department name from each metadata file into the document text before ingestion

    Incorrect

    Words in the text can influence similarity, but they do not enforce a filter, so results from other departments can still appear.

  • DChange the space type of the vector field to cosine similarity and rebuild the index

    Incorrect

    The space type sets how vector distance is computed. It does not make the nmslib engine support metadata filtering.

Some retrieval features depend on how the vector store was configured, not on the query. When a feature silently does nothing, check the knowledge base's documented requirements for that vector store before tuning queries.

Question 7 · choose 1

A ticket classifier on Amazon Nova assigns one of nine categories. Its prompt contains six fixed examples, five of which are billing tickets. An evaluation on 2,000 labeled tickets shows high accuracy for billing, but many technical and account tickets are labeled as billing too. The team must fix this in the prompt rather than by fine-tuning, and the output format must stay a single category name. What should the prompt engineer change?

  1. AAdd five more billing examples so that the model learns the billing category's boundaries better
  2. BMove the six examples from the user message into the system prompt in an examples section
  3. CReplace the examples with ones that cover all nine categories, including common edge cases
  4. DRaise the maximum output tokens so that the model can reason before it names the category
Show the answer and why
  • AAdd five more billing examples so that the model learns the billing category's boundaries better

    Incorrect

    More billing examples make the example set even more one-sided, which pushes outputs further toward billing.

  • BMove the six examples from the user message into the system prompt in an examples section

    Incorrect

    Examples can sit in either place, so moving them changes their position, not the imbalance that biases the outputs.

  • CReplace the examples with ones that cover all nine categories, including common edge cases

    Correct

    Examples should represent the expected distribution of inputs and outputs, from common cases to edge cases, because biased examples lead to biased outputs.

  • DRaise the maximum output tokens so that the model can reason before it names the category

    Incorrect

    More output room does not correct the skew that the examples create, and the output must remain a single category name.

Few-shot examples teach the model what the label distribution looks like. When one class dominates the examples, the model over-predicts it, so diagnose by per-class results and rebalance the examples before reaching for heavier fixes.

Question 8 · choose 1

A request with structured outputs fails at once with a 400 error. The JSON schema requires a minLength of 3 for a product code. The developer wants to keep structured outputs. What should the developer do?

  1. ADrop minLength and check the length in code
  2. BRetry the same request with exponential backoff
  3. CSwitch to a higher service tier
  4. DRaise maxTokens for the request
Show the answer and why
  • ADrop minLength and check the length in code

    Correct

    String constraints such as minLength are not supported, and schemas with unsupported features return a 400 error.

  • BRetry the same request with exponential backoff

    Incorrect

    The schema will be rejected again on every retry.

  • CSwitch to a higher service tier

    Incorrect

    Service tiers do not change which schema features are supported.

  • DRaise maxTokens for the request

    Incorrect

    Output length is unrelated to schema validation.

Structured outputs support a subset of JSON Schema. Move unsupported constraints into application-side validation.

Question 9 · choose 1

An Amazon Bedrock flow with six nodes, including two prompt nodes, a knowledge base node, a condition node and a Lambda function node, sends about 3% of refund requests down the wrong branch. The team has collected several inputs that reproduce the problem. It must find which node produced the wrong intermediate result and which conditions were satisfied, without changing the published flow definition or adding logging code to the Lambda function. How should the developer investigate?

  1. ATurn on model invocation logging and read the prompts and responses of the two prompt nodes
  2. BRecord CloudTrail data events for the flow alias and review each InvokeFlow event
  3. CInsert a Lambda function node after every node that writes its output to CloudWatch Logs
  4. DInvoke the flow with the failing inputs and enableTrace set to true, then read each trace
Show the answer and why
  • ATurn on model invocation logging and read the prompts and responses of the two prompt nodes

    Incorrect

    Invocation logs cover model calls only. They do not show the knowledge base node's output or which conditions the condition node evaluated as satisfied.

  • BRecord CloudTrail data events for the flow alias and review each InvokeFlow event

    Incorrect

    Data events show who invoked the flow and when, not the inputs and outputs of the nodes inside it.

  • CInsert a Lambda function node after every node that writes its output to CloudWatch Logs

    Incorrect

    This would change the flow definition, which is not allowed, and it adds code to maintain for a diagnosis the service already supports.

  • DInvoke the flow with the failing inputs and enableTrace set to true, then read each trace

    Correct

    Flow traces return the input and output of each node and which conditions a condition node satisfied, so the node that produced the wrong intermediate result can be found without changing the flow.

Reasoning-path tracing turns a wrong final answer into a specific failing step. Reproduce with known failing inputs and follow the trace node by node before changing prompts or logic.

Practise domain 5 →Practise all domains →