Skip to content
BytePatterns

AIP-C01 · Domain 2: Implementation and Integration · 26% of the exam

Task 2.4: Implement FM API integrations.

Calling models reliably: synchronous, streaming and asynchronous patterns, retries with backoff, rate limits and fallbacks, and routing each request to the right model.

Study it

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

A chat feature calls ConverseStream from a Lambda function behind an existing Amazon API Gateway REST API that uses a Lambda authorizer. Users wait up to 90 seconds before they see any text, and some long answers fail at 29 seconds. The team wants tokens shown as they are generated, wants to keep the REST API and its authorizer, and does not want to manage persistent client connections. What should the team do?

  1. AMove the chat to an API Gateway WebSocket API and push each token to the client with the connections endpoint
  2. BRequest a higher integration timeout for the REST API and keep the default buffered transfer mode
  3. CSet the proxy integration's response transfer mode to STREAM and stream the model output
  4. DReplace the REST API with an HTTP API so that responses are sent with chunked transfer encoding
Show the answer and why
  • AMove the chat to an API Gateway WebSocket API and push each token to the client with the connections endpoint

    Incorrect

    WebSocket APIs can deliver tokens in real time, but they require the team to manage connections and client changes, which the team wants to avoid.

  • BRequest a higher integration timeout for the REST API and keep the default buffered transfer mode

    Incorrect

    A longer timeout stops the failures, but buffered mode still waits for the complete response, so users see nothing until the end.

  • CSet the proxy integration's response transfer mode to STREAM and stream the model output

    Correct

    Response streaming on REST API proxy integrations sends output as it is produced, lowers time to first byte, and can run up to 15 minutes without an integration timeout increase.

  • DReplace the REST API with an HTTP API so that responses are sent with chunked transfer encoding

    Incorrect

    Response streaming is supported only for REST APIs, and changing the API type would also drop the existing setup the team wants to keep.

For token-by-token delivery over an existing REST API, set the proxy integration's response transfer mode to STREAM. It lowers time to first byte, can exceed the 29-second integration timeout and the 10 MB payload limit, and keeps the authorizer in place. Endpoint caching is not available in streaming mode.

Question 2 · choose 1

A batch summarizer running on Amazon ECS calls InvokeModel with boto3. During busy periods it receives ThrottlingException errors (HTTP 429), and the code currently retries immediately in a loop until a call succeeds, which makes the throttling worse. The team wants a bounded number of attempts that fits a 60-second latency budget and that avoids synchronized retries across tasks. What should the developer do?

  1. ATreat HTTP 429 as a permanent error and drop the request so that the queue keeps moving
  2. BSwitch the client to the bedrock-mantle endpoint because it does not enforce requests-per-minute quotas
  3. CKeep the immediate retries but add a fixed one-second sleep between attempts in every ECS task
  4. DUse the SDK's standard retry mode with a bounded total_max_attempts value so that retries back off with jitter
Show the answer and why
  • ATreat HTTP 429 as a permanent error and drop the request so that the queue keeps moving

    Incorrect

    A throttle is transient and should be retried with backoff. Dropping requests loses work that would succeed moments later.

  • BSwitch the client to the bedrock-mantle endpoint because it does not enforce requests-per-minute quotas

    Incorrect

    bedrock-mantle does not enforce RPM quotas, but it still applies token quotas and capacity limits, so the client still needs retry handling.

  • CKeep the immediate retries but add a fixed one-second sleep between attempts in every ECS task

    Incorrect

    A fixed delay without jitter keeps the tasks retrying in step, which recreates bursts of synchronized retries.

  • DUse the SDK's standard retry mode with a bounded total_max_attempts value so that retries back off with jitter

    Correct

    The SDK's standard retry mode backs off exponentially with random jitter, and total_max_attempts caps the attempts, including the first one, to fit the latency budget.

Throttles and transient capacity errors are retried with exponential backoff and jitter within a bounded budget. Most AWS SDKs implement this, and in botocore total_max_attempts includes the initial request.

Question 3 · choose 2

A legal team's web portal asks a model on Amazon Bedrock to draft contracts. Each draft takes 5 to 12 minutes to generate. Browsers must not hold an open request while the draft is produced, no draft request may be lost if a worker fails, and the portal must show the finished draft when the user returns. Which actions should the developer take? (Choose TWO.)

  1. ACall the model from the browser with temporary credentials and keep the tab open until the draft completes
  2. BSubmit each draft request as an Amazon Bedrock batch inference job so that drafts are produced asynchronously
  3. CIncrease the API Gateway integration timeout so that the browser receives the draft in the same request
  4. DPut each request on an Amazon SQS queue with a dead-letter queue and process it with a Lambda worker
  5. EReturn a job ID at once and have the worker save the draft and its status in Amazon DynamoDB
Show the answer and why
  • ACall the model from the browser with temporary credentials and keep the tab open until the draft completes

    Incorrect

    Holding the tab open is the same long-running connection the portal must avoid, and a closed tab loses the draft.

  • BSubmit each draft request as an Amazon Bedrock batch inference job so that drafts are produced asynchronously

    Incorrect

    Batch inference processes large sets of records with a minimum record count and a processing window measured in hours, which does not fit single interactive requests.

  • CIncrease the API Gateway integration timeout so that the browser receives the draft in the same request

    Incorrect

    This keeps the browser connection open for the whole generation, which the portal must avoid.

  • DPut each request on an Amazon SQS queue with a dead-letter queue and process it with a Lambda worker

    Correct

    The queue decouples the browser from generation, keeps the request durably until a worker succeeds, and the dead-letter queue holds requests that keep failing.

  • EReturn a job ID at once and have the worker save the draft and its status in Amazon DynamoDB

    Correct

    An immediate job ID frees the browser, and a status record lets the portal show the draft whenever the user returns.

Long generations follow the asynchronous request pattern: accept, enqueue durably, return an identifier, let workers process at their own pace, and store the result for later retrieval or notification.

Question 4 · choose 1

An assistant answers legal, coding and general questions. The company wants legal questions handled by a model it customized on legal text, coding questions by a code-specialized model from a different provider, and everything else by a low-cost general model. Routing must follow a classifier the company controls, each decision must be visible for audits, and new categories must be added without changing application code. Which approach meets these requirements?

  1. AHardcoded if-else statements in the application code that pick a model ID from keywords found in the user's question
  2. BAmazon Bedrock intelligent prompt routing with a router that includes the three models
  3. CA Lambda alias with weighted routing that splits traffic across three function versions, one per model
  4. DAn AWS Step Functions workflow with a classifier step and a Choice state that invokes the model for each category
Show the answer and why
  • AHardcoded if-else statements in the application code that pick a model ID from keywords found in the user's question

    Incorrect

    Static routing in application code means each new category needs a code change and deployment.

  • BAmazon Bedrock intelligent prompt routing with a router that includes the three models

    Incorrect

    Intelligent prompt routing predicts quality among models of the same family and cannot use the company's own classifier or route to a customized model from another provider.

  • CA Lambda alias with weighted routing that splits traffic across three function versions, one per model

    Incorrect

    Weighted aliases split traffic randomly by percentage. They do not look at the content of the question.

  • DAn AWS Step Functions workflow with a classifier step and a Choice state that invokes the model for each category

    Correct

    Step Functions implements content-based routing to specialized models, records every state transition in the execution history for audits, and new branches are added in the workflow definition rather than the application.

Static routing in code is simple but rigid; intelligent prompt routing optimizes cost and quality within one model family. When routing must follow your own classification across providers and be auditable, an orchestrated, content-based router such as a Step Functions Choice state is the fit.

Question 5 · choose 1

An internal Amazon API Gateway REST API has two routes, POST /chat and POST /bulk-summarize, both integrated with the same Lambda function that calls a model on Amazon Bedrock. A nightly job from another team floods /bulk-summarize, and the resulting model throttling makes interactive chat fail. All callers use IAM authorization through one shared role, so they cannot be told apart, and the team wants /bulk-summarize held to 20 requests per second while /chat keeps its current limits, without code changes. What should the developer configure?

  1. AA method-level throttling limit for POST /bulk-summarize on the stage
  2. BReserved concurrency on the shared Lambda function behind both routes
  3. CA lower account-level throttling limit for the API Gateway Region
  4. DAPI caching on the stage for the POST /bulk-summarize route
Show the answer and why
  • AA method-level throttling limit for POST /bulk-summarize on the stage

    Correct

    Stage throttling can be set differently for each method, so one route can be held to a lower rate while the other keeps its limits.

  • BReserved concurrency on the shared Lambda function behind both routes

    Incorrect

    Reserved concurrency caps one function's concurrent requests. Both routes use this function, so chat would be capped as well.

  • CA lower account-level throttling limit for the API Gateway Region

    Incorrect

    Account-level limits apply to all APIs in the account and Region, including /chat.

  • DAPI caching on the stage for the POST /bulk-summarize route

    Incorrect

    Caching returns stored responses for repeated requests, and only GET methods are cached by default. POST requests with different documents are not repeats.

Throttle at the narrowest level that matches the problem: per method when routes share a backend and callers cannot be told apart.

Question 6 · choose 1

An internal assistant runs as a Python service on Amazon ECS and calls the Converse API with a model and prompt that the legal team validated. Answers average 1,500 tokens, and users stare at an empty screen for about 20 seconds before the whole answer appears, although the total generation time is acceptable. The model, prompt and answer length must not change, and the service must keep the Converse message format, including its tool-use blocks. Which change lets users see the answer sooner?

  1. ALower maxTokens to 500 tokens for every request
  2. BPurchase Provisioned Throughput for the validated model
  3. CCall InvokeModel with the model's native request body
  4. DCall ConverseStream and render text deltas as they arrive
Show the answer and why
  • ALower maxTokens to 500 tokens for every request

    Incorrect

    A response length limit caps the tokens generated. Answers would be cut off, and the answer length must not change.

  • BPurchase Provisioned Throughput for the validated model

    Incorrect

    Provisioned Throughput provides a higher level of throughput for a model. The response would still arrive only when it is complete.

  • CCall InvokeModel with the model's native request body

    Incorrect

    InvokeModel fits features that exist only in a model's native format. It still returns the complete response, and it would replace the Converse tool-use code.

  • DCall ConverseStream and render text deltas as they arrive

    Correct

    ConverseStream returns the response as a stream of events, and contentBlockDelta events carry partial text, and partial tool input, as the model generates it.

Streaming improves perceived latency without changing the model, the prompt or the answer: the first tokens appear while the rest is still being generated.

Question 7 · choose 2

A travel company's English-language help assistant sends every request through the Converse API to the larger of two models from one model family on Amazon Bedrock. Most questions are simple, but some need the larger model's quality. The company wants to cut inference cost without training any model or building a classifier, wants the larger model used only when it is predicted to answer meaningfully better, and finance wants to know which model served each request. Which actions should the developer take? (Choose TWO.)

  1. ASend prompts shorter than 200 tokens to the smaller model from application code
  2. BCreate a prompt router with both models and the smaller one as fallback
  3. CRun an Amazon Bedrock Model Distillation job with the larger model as the teacher
  4. DAdd a Bedrock Flows condition node that routes on a list of keywords
  5. ERead the invoked model ID from the prompt router trace in each response and log it
Show the answer and why
  • ASend prompts shorter than 200 tokens to the smaller model from application code

    Incorrect

    A static rule in code is a simple form of routing. Prompt length says little about how hard a question is, and the rule needs manual tuning.

  • BCreate a prompt router with both models and the smaller one as fallback

    Correct

    A prompt router predicts response quality per request and switches from the fallback model to the other model only when the predicted response quality difference meets the routing criteria.

  • CRun an Amazon Bedrock Model Distillation job with the larger model as the teacher

    Incorrect

    Distillation fine-tunes a smaller student model on the teacher's responses. That is the training the company wants to avoid.

  • DAdd a Bedrock Flows condition node that routes on a list of keywords

    Incorrect

    Condition nodes branch a flow on defined conditions, which suits deterministic logic. A keyword list does not predict answer quality and must be maintained.

  • ERead the invoked model ID from the prompt router trace in each response and log it

    Correct

    Responses from a prompt router include information about the model that was used; the router trace carries the invoked model ID.

Intelligent prompt routing lowers cost within a model family without a classifier of your own, and its trace tells you which model answered each request.

Question 8 · choose 1

A support assistant calls ConverseStream and offers the model a create_ticket tool. Text answers stream to users correctly, but whenever the model decides to use the tool, the application's JSON parser throws an error on the first toolUse delta and the turn fails. The team must keep streaming for text, and the tool must run exactly once with the complete arguments. What should the developer change?

  1. ASwitch every turn from ConverseStream to the Converse API
  2. BRaise maxTokens so that the tool input arrives in a single delta
  3. CJoin the block's toolUse fragments and parse them at contentBlockStop
  4. DRead the complete tool arguments from the messageStop event
Show the answer and why
  • ASwitch every turn from ConverseStream to the Converse API

    Incorrect

    Converse returns the whole message, including tool input, in one response, which suits callers that do not need streaming. The team must keep streaming.

  • BRaise maxTokens so that the tool input arrives in a single delta

    Incorrect

    maxTokens limits the length of the response, which matters when output is cut off. It does not control how the input is split across delta events.

  • CJoin the block's toolUse fragments and parse them at contentBlockStop

    Correct

    In a stream, toolUse deltas carry partial input JSON as strings. The input is complete only when the content block stops, so the application joins the fragments, parses once and runs the tool once.

  • DRead the complete tool arguments from the messageStop event

    Incorrect

    messageStop carries the reason the model stopped generating, such as a tool use request. It does not carry the tool input.

Streaming changes how tool calls arrive: buffer partial tool input per content block and act only when the block is complete.

Question 9 · choose 1

A company exposes a document question-answering API to 15 partner developers through an Amazon API Gateway REST API whose Lambda integration builds a prompt and calls a model. Logs show that 12 percent of requests have a missing question field, an unsupported language code or a malformed document ID, and each still runs the Lambda function and a model call before failing. The team wants such requests rejected before the integration runs, with an error that tells the caller what was wrong, and without adding code. What should the developer configure?

  1. AChecks in the Lambda function that return HTTP 400 for bad input
  2. BA request validator on the method with a JSON schema model for the body
  3. CA Lambda authorizer that inspects each incoming request
  4. DA custom gateway response for the BAD_REQUEST_BODY type
Show the answer and why
  • AChecks in the Lambda function that return HTTP 400 for bad input

    Incorrect

    Checks in the function fit rules that need application data, such as whether a document exists. Here the integration would still run, and it means writing code.

  • BA request validator on the method with a JSON schema model for the body

    Correct

    API Gateway can validate the request body against a JSON schema model before the integration request. A failed validation returns 400 at once and avoids unnecessary backend calls.

  • CA Lambda authorizer that inspects each incoming request

    Incorrect

    A Lambda authorizer implements a custom authorization scheme based on a token or request parameters. It decides who may call the API, not whether the body is well formed.

  • DA custom gateway response for the BAD_REQUEST_BODY type

    Incorrect

    Gateway responses customize the error that API Gateway returns. By themselves they do not validate anything.

Reject malformed requests at the API layer, before they cost a function invocation and a model call.

Practise domain 2 →Practise all domains →