Skip to content
BytePatterns

AIP-C01 · Domain 1: Foundation Model Integration, Data Management, and Compliance · 31% of the exam

Task 1.5: Design retrieval mechanisms for FM augmentation.

Making retrieval return the right context: chunking strategies, embedding model choice, hybrid search, reranking, query decomposition and agentic retrieval, and exposing retrieval as a tool.

Study it

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

An aircraft maintenance team searches 3,000-page technical manuals. Technicians type precise questions such as "torque value for the left actuator bolt", and evaluations show that small chunks find the right passage but the model then lacks the surrounding procedure, while large chunks return the procedure but rank poorly. The manuals have a clear section structure. Which chunking strategy should the developer use for the knowledge base?

  1. ANo chunking, after splitting each manual into one separate file per chapter beforehand
  2. BSemantic chunking with a high breakpoint percentile threshold and a larger buffer size
  3. CFixed-size chunking with 100-token chunks and a 50% overlap between consecutive chunks
  4. DHierarchical chunking with small child chunks and larger parent chunks
Show the answer and why
  • ANo chunking, after splitting each manual into one separate file per chapter beforehand

    Incorrect

    With no chunking each file becomes one chunk, so whole chapters are embedded as single vectors, which hurts precision and inflates the context sent to the model.

  • BSemantic chunking with a high breakpoint percentile threshold and a larger buffer size

    Incorrect

    Semantic chunking groups sentences by meaning, which helps unstructured text, but a high threshold just makes chunks larger and reintroduces the ranking problem. It has no parent expansion.

  • CFixed-size chunking with 100-token chunks and a 50% overlap between consecutive chunks

    Incorrect

    Very small fixed chunks with heavy overlap improve matching but still hand the model only a fragment of the procedure.

  • DHierarchical chunking with small child chunks and larger parent chunks

    Correct

    Hierarchical chunking searches on small child chunks for precise matching and then replaces them with their parent chunks, so the model receives the broader procedure around the match.

The symptoms describe the precision versus context trade-off that hierarchical chunking was built for: match on small children, answer with the larger parent. Expect fewer results than numberOfResults, because children that share a parent collapse into one.

Question 2 · choose 1

An industrial parts distributor's knowledge base uses an Amazon OpenSearch Serverless index created by Amazon Bedrock. Customers ask questions such as "is the HX-4471-B valve rated for steam?". Semantic search returns passages about similar valves but often misses the exact part number, and adding a reranker did not help because the right passage was never retrieved. What should the developer change?

  1. ARe-create the data source with semantic chunking so that each part's description stays in one chunk
  2. BAdd a reranker model with a higher number of reranked results to reorder the retrieved passages
  3. CIncrease numberOfResults to 100 so that the passage with the part number is more likely to be included
  4. DSet overrideSearchType to HYBRID so that keyword matching on the raw text is combined with vector search
Show the answer and why
  • ARe-create the data source with semantic chunking so that each part's description stays in one chunk

    Incorrect

    Chunk boundaries are not the cause. The issue is that dense vectors do not reliably match exact identifiers, and chunking changes cannot fix that.

  • BAdd a reranker model with a higher number of reranked results to reorder the retrieved passages

    Incorrect

    A reranker reorders what retrieval returned. It cannot surface a passage that was never retrieved, which the team already observed.

  • CIncrease numberOfResults to 100 so that the passage with the part number is more likely to be included

    Incorrect

    More semantic results add noise and tokens but do not make an embedding model better at matching an arbitrary code, so the passage can still be missing.

  • DSet overrideSearchType to HYBRID so that keyword matching on the raw text is combined with vector search

    Correct

    Hybrid search combines vector search with a search of the raw text, so exact tokens such as part numbers are matched. OpenSearch Serverless with a filterable text field supports it.

Embeddings capture meaning but are weak at exact identifiers such as SKUs, error codes and part numbers. Hybrid search adds keyword matching on the raw text. It is available for Aurora (RDS), OpenSearch Serverless and MongoDB vector stores that have a filterable text field.

Question 3 · choose 1

A wealth management firm uses a customer-managed knowledge base on Amazon Aurora PostgreSQL. Advisors ask compound questions such as "Compare the fee schedules of the Growth and Income portfolios and list which one allows quarterly withdrawals." A single retrieval returns chunks about only one portfolio. The firm wants the knowledge base to split such questions into smaller searches without building its own orchestration, and it calls RetrieveAndGenerate. What should the developer configure?

  1. ACall AgenticRetrieveStream so that a foundation model plans sub-queries across iterations
  2. BEnable query decomposition in the orchestration configuration of RetrieveAndGenerate
  3. CEnable implicit metadata filtering so that the model generates a filter on the portfolio name
  4. DAdd a reranker model so that the chunks about the second portfolio are moved to the top
Show the answer and why
  • ACall AgenticRetrieveStream so that a foundation model plans sub-queries across iterations

    Incorrect

    Agentic retrieval also decomposes questions, but it currently supports only managed knowledge bases, and this one is customer-managed.

  • BEnable query decomposition in the orchestration configuration of RetrieveAndGenerate

    Correct

    Query decomposition breaks a complex query into smaller sub-queries and runs them, which retrieves chunks for each portfolio. It is configured with the QUERY_DECOMPOSITION transformation type.

  • CEnable implicit metadata filtering so that the model generates a filter on the portfolio name

    Incorrect

    Implicit filtering derives one filter from the query. A filter for one portfolio would exclude the other, and the question needs both.

  • DAdd a reranker model so that the chunks about the second portfolio are moved to the top

    Incorrect

    A reranker can only reorder retrieved chunks. If no chunks about the second portfolio were retrieved, reranking cannot add them.

Multi-part questions need multiple retrievals. On a customer-managed knowledge base, the built-in way is query decomposition in RetrieveAndGenerate. Agentic retrieval goes further, with iterative planning across knowledge bases, but requires a managed knowledge base.

Question 4 · choose 2

Five product teams build agents with different open-source frameworks, and all of them need to search the company's Amazon Bedrock managed knowledge base. The platform team wants every agent to discover and call the search the same way through the Model Context Protocol (MCP), without each team writing integration code, and wants callers authenticated against the corporate identity provider. Which actions should the platform team take? (Choose TWO.)

  1. APublish an Amazon API Gateway REST API with an OpenAPI definition that wraps the Retrieve API
  2. BGive each team an IAM role that allows the Retrieve API and let each team write its own retrieval tool
  3. CConfigure the gateway's inbound authorizer to validate JSON Web Tokens issued by the corporate identity provider
  4. DCopy the knowledge base content into a separate vector store for each team and let each team host its own MCP server
  5. EAdd the managed knowledge base to an Amazon Bedrock AgentCore Gateway as a target through the built-in connector
Show the answer and why
  • APublish an Amazon API Gateway REST API with an OpenAPI definition that wraps the Retrieve API

    Incorrect

    A REST API with an OpenAPI definition is a fine API-first design, but it is not an MCP interface, so each framework would still need custom integration code.

  • BGive each team an IAM role that allows the Retrieve API and let each team write its own retrieval tool

    Incorrect

    Direct Retrieve calls work, but every team would write and maintain its own integration, which is what the platform team wants to avoid.

  • CConfigure the gateway's inbound authorizer to validate JSON Web Tokens issued by the corporate identity provider

    Correct

    A gateway must have inbound authorization, and the OAuth (JWT) option validates tokens from the company's identity provider before any tool call reaches the knowledge base.

  • DCopy the knowledge base content into a separate vector store for each team and let each team host its own MCP server

    Incorrect

    Duplicating the index multiplies storage and synchronization work and spreads security controls across five teams.

  • EAdd the managed knowledge base to an Amazon Bedrock AgentCore Gateway as a target through the built-in connector

    Correct

    A managed knowledge base can be a target of an AgentCore Gateway, which exposes it as an MCP tool that any MCP-compatible framework can discover and call.

Consistent access to retrieval means one standard interface. An AgentCore Gateway turns the managed knowledge base into an MCP tool, and its inbound JWT authorizer ties every call to the corporate identity provider.

Question 5 · choose 1

A knowledge base uses fixed-size chunking at 300 tokens with no overlap. Evaluations show that answers miss facts whose key sentence is split across two chunks. The team must keep fixed-size chunking because its citation viewer depends on it, the documents have no consistent structure, and the index can grow by about 20 percent. Which change helps most?

  1. ASwitch to no chunking so that each file becomes one chunk
  2. BReduce the fixed chunk size from 300 to 100 tokens
  3. CSet an overlap percentage between consecutive chunks
  4. DRaise numberOfResults from 5 to 10 for each query
Show the answer and why
  • ASwitch to no chunking so that each file becomes one chunk

    Incorrect

    No chunking fits documents already split into small files. Whole files make retrieval less precise, and the team must keep fixed-size chunking.

  • BReduce the fixed chunk size from 300 to 100 tokens

    Incorrect

    Smaller chunks fit when chunks mix too much unrelated text. They add boundaries, so more sentences get split.

  • CSet an overlap percentage between consecutive chunks

    Correct

    Fixed-size chunking lets you set an overlap percentage between consecutive chunks, so text at a boundary also appears in the neighboring chunk, at the cost of some extra index size.

  • DRaise numberOfResults from 5 to 10 for each query

    Incorrect

    More results help when relevant chunks rank just below the cutoff. A split fact stays split, and both halves still have to be retrieved.

Boundaries are the weak point of fixed-size chunking. Overlap repeats text across neighbors so a sentence near a boundary survives intact in at least one chunk.

Question 6 · choose 1

A knowledge base often returns the most useful passage in fifth place, behind loosely related passages, and the generation prompt's token budget allows only the top three. The team cannot re-embed or re-index its 20 million chunks this quarter, and it can accept a few hundred milliseconds of extra latency per question. Which change should the team make?

  1. ARe-embed the corpus with a higher-dimension embedding model
  2. BSend the top ten passages to the model instead of three
  3. CRe-create the data source with larger chunks and re-sync it
  4. DAdd a reranker model before the top three passages are kept
Show the answer and why
  • ARe-embed the corpus with a higher-dimension embedding model

    Incorrect

    Vectors with more dimensions need a new index filled with re-embedded chunks, which the team cannot build this quarter.

  • BSend the top ten passages to the model instead of three

    Incorrect

    More passages help when the prompt budget allows them. Here the budget allows only three.

  • CRe-create the data source with larger chunks and re-sync it

    Incorrect

    Larger chunks help when chunks lack context, but changing chunking requires re-ingesting the data.

  • DAdd a reranker model before the top three passages are kept

    Correct

    A reranker scores the relevance of retrieved results to the query and reorders them, which works on the existing index and adds only a reranking call.

Retrieval finds candidates and reranking orders them. When the right passage is found but ranked too low, rerank before truncating.

Question 7 · choose 1

A knowledge base holds manuals for five product lines. Customers of one product line keep getting passages from manuals of other lines. Each document already has a product_line metadata attribute. What should the application do?

  1. ACreate a separate embedding model per product line
  2. BApply a metadata filter on product_line in each request
  3. CAsk the model to ignore other product lines
  4. DLower the chunk size of every manual
Show the answer and why
  • ACreate a separate embedding model per product line

    Incorrect

    Different embedding models do not separate the documents, and queries must use the same model as the index.

  • BApply a metadata filter on product_line in each request

    Correct

    Metadata filters restrict retrieval to documents whose attributes match, such as the customer's product line.

  • CAsk the model to ignore other product lines

    Incorrect

    Instructions do not stop other passages from taking result slots.

  • DLower the chunk size of every manual

    Incorrect

    Chunk size does not limit results to one product line.

Use metadata you already have. Filters narrow retrieval to relevant documents before generation.

Question 8 · choose 1

An IT help center searches 9,000 articles with keyword search on Amazon OpenSearch Service. Employees describe problems in their own words, such as "can't get into my laptop after the update", while the articles use terms such as "BitLocker recovery key". Analysis shows thousands of different phrasings, typos are rare, and the right article is usually not retrieved at all. Which change addresses the root cause?

  1. AAdd semantic search with vector embeddings of the articles and the queries
  2. BUpload a synonym file as a custom package for the index analyzer
  3. CAdd a reranker model on top of the keyword results
  4. DExtract key phrases from each query with Amazon Comprehend before searching
Show the answer and why
  • AAdd semantic search with vector embeddings of the articles and the queries

    Correct

    Embeddings represent meaning, so an article can match a question that shares no keywords with it.

  • BUpload a synonym file as a custom package for the index analyzer

    Incorrect

    Synonym files fit a known, bounded list of equivalent terms. Thousands of free-form phrasings cannot be kept in synonym lists.

  • CAdd a reranker model on top of the keyword results

    Incorrect

    A reranker reorders retrieved results. It cannot surface an article that keyword search never retrieved.

  • DExtract key phrases from each query with Amazon Comprehend before searching

    Incorrect

    Key phrases are noun phrases taken from the user's own words. Searching on them is still keyword matching.

When users and authors use different words, the fix is a different retrieval method, not more keyword tuning.

Question 9 · choose 1

A media research firm builds a knowledge base over 8,000 interview transcripts. The transcripts have no headings, and each interview moves between several topics without any markers. With default chunking, retrieved chunks often start in the middle of a topic and blend two subjects, so answers mix unrelated quotes. The firm wants chunk boundaries to follow topic shifts, wants Amazon Bedrock to keep parsing and embedding the data without any preprocessing code of its own, and accepts extra ingestion cost. Which data source setting should the developer choose?

  1. ANo chunking, so that each transcript is kept as a single chunk
  2. BFixed-size chunking with a larger maximum token size per chunk
  3. CSemantic chunking with a tuned breakpoint percentile threshold
  4. DDefault chunking with a higher numberOfResults on each query
Show the answer and why
  • ANo chunking, so that each transcript is kept as a single chunk

    Incorrect

    No chunking treats each document as one chunk, so every transcript would become one long chunk that spans many topics.

  • BFixed-size chunking with a larger maximum token size per chunk

    Incorrect

    Larger fixed-size chunks fit when answers need longer passages. Boundaries still follow token counts, so each chunk would mix even more topics.

  • CSemantic chunking with a tuned breakpoint percentile threshold

    Correct

    Semantic chunking divides text by meaning, splitting where sentences become dissimilar beyond a breakpoint percentile threshold. It uses a foundation model, which is the extra cost the firm accepts.

  • DDefault chunking with a higher numberOfResults on each query

    Incorrect

    More results help when relevant chunks rank just below the cutoff. Here they would add more blended chunks to each answer.

Choose the chunking strategy from the shape of the content: structure-based splits need structure, and unmarked topic shifts need boundaries chosen by meaning.

Practise domain 1 →Practise all domains →