Skip to content
BytePatterns

AIP-C01 · Domain 1: Foundation Model Integration, Data Management, and Compliance · 31% of the exam

Task 1.4: Design and implement vector store solutions.

Where embeddings live and how they stay current: Bedrock Knowledge Bases, OpenSearch, Aurora with pgvector and S3 Vectors, metadata for precise search, indexing at scale and keeping the index in sync.

Study it

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

A legal research firm will index 900 million passages for semantic search in an Amazon Bedrock knowledge base. To cut storage cost, the team chose an embedding model setting that outputs binary vectors instead of float32 vectors and accepts the small loss of precision. The firm wants a fully managed AWS vector store. Which vector store can the knowledge base use?

  1. AAn Amazon Aurora PostgreSQL cluster with the pgvector extension
  2. BAn Amazon Neptune Analytics graph with GraphRAG enabled
  3. CAn Amazon OpenSearch Serverless vector search collection
  4. DAn Amazon S3 Vectors vector bucket and index
Show the answer and why
  • AAn Amazon Aurora PostgreSQL cluster with the pgvector extension

    Incorrect

    Aurora with pgvector is a supported vector store for float32 embeddings, but binary vectors are supported only on OpenSearch Serverless and OpenSearch managed clusters.

  • BAn Amazon Neptune Analytics graph with GraphRAG enabled

    Incorrect

    Neptune Analytics adds graph relationships for multi-hop retrieval, which is a different need. It is not a binary vector store option.

  • CAn Amazon OpenSearch Serverless vector search collection

    Correct

    OpenSearch Serverless and OpenSearch managed clusters are the vector stores that support binary vector embeddings for knowledge bases.

  • DAn Amazon S3 Vectors vector bucket and index

    Incorrect

    S3 Vectors is a low-cost store for vectors, but it supports only floating-point embeddings, not binary ones.

Binary vectors use 1 bit per dimension instead of 32, which saves storage at some cost in precision. Choosing them narrows the vector store choice: only OpenSearch Serverless and OpenSearch managed clusters support binary embeddings for knowledge bases.

Question 2 · choose 1

A developer creates the table for an Amazon Bedrock knowledge base in an existing Amazon Aurora PostgreSQL cluster with pgvector. The table has an embedding column, a chunks text column, a metadata column and a custom_metadata jsonb column. The application will use hybrid search and will filter on custom metadata. Which set of indexes should the developer create?

  1. AA GIN index on custom_metadata only, because the knowledge base creates the vector index itself
  2. BAn HNSW index on embedding, a GIN index on to_tsvector of chunks, and a GIN index on custom_metadata
  3. CA B-tree index on embedding, a GIN index on to_tsvector of chunks, and a B-tree index on the metadata column
  4. DAn HNSW index on embedding and a B-tree index on chunks, with no index on custom_metadata
Show the answer and why
  • AA GIN index on custom_metadata only, because the knowledge base creates the vector index itself

    Incorrect

    When you bring your own Aurora table, you create the indexes on the vector and text columns yourself. Only the console quick-create flow sets the database up for you.

  • BAn HNSW index on embedding, a GIN index on to_tsvector of chunks, and a GIN index on custom_metadata

    Correct

    The vector column needs an HNSW index, hybrid search needs a GIN full-text index on the chunks column, and metadata filtering on the jsonb column needs a GIN index.

  • CA B-tree index on embedding, a GIN index on to_tsvector of chunks, and a B-tree index on the metadata column

    Incorrect

    B-tree indexes cannot accelerate vector similarity search. The embedding column needs an HNSW index with a vector operator class.

  • DAn HNSW index on embedding and a B-tree index on chunks, with no index on custom_metadata

    Incorrect

    A B-tree index does not serve full-text search, so the keyword half of hybrid search would have no suitable index, and jsonb filtering would scan the table.

For a self-managed Aurora vector store, the knowledge base expects you to index the vector column (HNSW), the text column for full-text search (GIN on to_tsvector) and, when you use a custom metadata column, that column (GIN). Without the text index, hybrid search has nothing to use for keywords.

Question 3 · choose 1

A news company stores articles in an Amazon S3 bucket that is the data source of an Amazon Bedrock knowledge base. Editors publish and correct articles all day, and corrections must be searchable within two minutes. The bucket holds 2 million articles, and the team wants to avoid processing that only scans for changes. Which approach meets these requirements?

  1. ADelete and re-create the knowledge base every night so that the index always matches the bucket
  2. BSchedule StartIngestionJob every hour with Amazon EventBridge Scheduler so that only changed articles are processed
  3. CCall IngestKnowledgeBaseDocuments from a Lambda function triggered by S3 event notifications for new or changed articles
  4. DRun an AWS Glue crawler on the bucket every five minutes so that the knowledge base sees the new objects
Show the answer and why
  • ADelete and re-create the knowledge base every night so that the index always matches the bucket

    Incorrect

    Rebuilding re-embeds all 2 million articles, which is slow and costly, and corrections would still wait until the next night.

  • BSchedule StartIngestionJob every hour with Amazon EventBridge Scheduler so that only changed articles are processed

    Incorrect

    Syncs are incremental, which is good for periodic refreshes, but an hourly schedule cannot meet a two-minute freshness target and each sync still checks the whole data source for changes.

  • CCall IngestKnowledgeBaseDocuments from a Lambda function triggered by S3 event notifications for new or changed articles

    Correct

    Direct ingestion indexes the submitted documents into the vector store in one action, without waiting for a sync that scans the data source, so a change is searchable shortly after it is written.

  • DRun an AWS Glue crawler on the bucket every five minutes so that the knowledge base sees the new objects

    Incorrect

    A Glue crawler updates the Data Catalog with table metadata. It does not create embeddings or update a knowledge base index.

Syncing a data source is right for scheduled, incremental refreshes. When individual documents must be searchable within minutes, direct ingestion through the KnowledgeBaseDocuments operations, driven by S3 event notifications, skips the scan and indexes just the changed objects.

Question 4 · choose 3

A pharmaceutical company keeps study reports as PDF files in an Amazon S3 bucket that feeds an Amazon Bedrock knowledge base on Amazon OpenSearch Serverless. Researchers ask for results "from the last two years" and "from the oncology or immunology teams", and the current answers mix in old reports from unrelated teams. The developer must make these constraints enforceable at retrieval time. Which steps should the developer take? (Choose THREE.)

  1. AAdd a metadata file named <file>.pdf.metadata.json next to each report with the attributes
  2. BName the custom attributes with the x-amz-bedrock prefix so that the knowledge base indexes them automatically
  3. CStore the report date as a NUMBER attribute, such as epoch seconds, for greaterThan filters
  4. DAdd the attributes as S3 object tags on each PDF so that the next sync copies them into the index
  5. EPass a Retrieve filter that joins a date condition and an in condition on team with andAll
Show the answer and why
  • AAdd a metadata file named <file>.pdf.metadata.json next to each report with the attributes

    Correct

    For an S3 data source, metadata is read from a sidecar file that shares the source file's name and extension plus .metadata.json, and the attributes are stored with each chunk for filtering.

  • BName the custom attributes with the x-amz-bedrock prefix so that the knowledge base indexes them automatically

    Incorrect

    Fields with the x-amz-bedrock prefix are reserved by the service for custom knowledge bases and cannot be used for your own attributes.

  • CStore the report date as a NUMBER attribute, such as epoch seconds, for greaterThan filters

    Correct

    Range operators such as greaterThan apply to number attributes, so a numeric timestamp makes "last two years" a precise filter.

  • DAdd the attributes as S3 object tags on each PDF so that the next sync copies them into the index

    Incorrect

    S3 object tags are a storage feature. The knowledge base reads metadata from .metadata.json files, not from object tags.

  • EPass a Retrieve filter that joins a date condition and an in condition on team with andAll

    Correct

    Retrieval filters accept operators such as greaterThan and in, and andAll requires every condition to match, so only recent reports from the listed teams are returned.

Metadata turns soft wishes in a prompt into hard retrieval constraints. The attributes must arrive through sidecar metadata files with the right data types, and the query must pass a filter expression. OpenSearch Serverless indexes need the faiss engine for filtering, which knowledge-base-created indexes use.

Question 5 · choose 1

A three-person team is building its first Amazon Bedrock knowledge base for an internal help assistant. It does not want to design index mappings or manage servers, expects a few hundred queries a minute during working hours, and needs retrieval to stay well under 100 milliseconds so that answers start quickly. Which vector store approach fits?

  1. ALet the console quick-create an Amazon OpenSearch Serverless vector store
  2. BLet the console quick-create an Amazon S3 Vectors bucket and index
  3. CCreate an OpenSearch Serverless collection and index by hand, then map the fields
  4. DInstall a vector database on Amazon EC2 and connect the knowledge base to it
Show the answer and why
  • ALet the console quick-create an Amazon OpenSearch Serverless vector store

    Correct

    Quick create builds an OpenSearch Serverless vector search collection and index with the required fields, and vector search collections provide scalable, high-performing similarity search without clusters to manage.

  • BLet the console quick-create an Amazon S3 Vectors bucket and index

    Incorrect

    S3 Vectors suits cost-effective storage of large vector sets. It delivers subsecond latency for infrequent queries and as low as 100 milliseconds for frequent ones, which does not meet this target.

  • CCreate an OpenSearch Serverless collection and index by hand, then map the fields

    Incorrect

    Bringing your own index fits when you need custom index settings. It is the mapping work the team wants to avoid.

  • DInstall a vector database on Amazon EC2 and connect the knowledge base to it

    Incorrect

    A self-installed database fits when a specific engine is required. It means managing servers, and a knowledge base connects only to its supported vector stores.

For a first knowledge base, let Amazon Bedrock create the store, and pick the store type by latency and query rate: S3 Vectors trades latency for cost, while OpenSearch Serverless serves frequent, fast queries.

Question 6 · choose 1

A retailer's orders, inventory and store data live in an Amazon Aurora PostgreSQL cluster that its engineers operate and query in SQL every day. It wants a knowledge base over product manuals whose vectors sit next to that data, so that analysts can join similarity results with the inventory tables in SQL, and the security team will not approve a new database engine. Which vector store should the team use?

  1. AA new Amazon OpenSearch Serverless vector search collection
  2. BAn Amazon S3 Vectors vector bucket and vector index for the manuals
  3. CThe existing Aurora PostgreSQL cluster with the pgvector extension
  4. DA new Amazon DocumentDB cluster with vector search enabled
Show the answer and why
  • AA new Amazon OpenSearch Serverless vector search collection

    Incorrect

    OpenSearch Serverless suits high-performing similarity search without clusters to manage. It is a new engine, and its results cannot be joined with Aurora tables in SQL.

  • BAn Amazon S3 Vectors vector bucket and vector index for the manuals

    Incorrect

    S3 Vectors suits low-cost storage of large vector sets. It is a separate store with its own query API, not a SQL table.

  • CThe existing Aurora PostgreSQL cluster with the pgvector extension

    Correct

    An Aurora PostgreSQL cluster can serve as a knowledge base vector store, so the vectors live in PostgreSQL tables that SQL can join with existing data.

  • DA new Amazon DocumentDB cluster with vector search enabled

    Incorrect

    DocumentDB vector search suits data that already lives in DocumentDB. It would be a new engine for this team.

The best vector store is often the database the team already runs. pgvector keeps vectors next to relational data and inside existing approval and operations processes.

Question 7 · choose 2

A team created an index on an Amazon OpenSearch Service domain for 1,024-dimension vectors from Amazon Titan Text Embeddings V2 and bulk-loaded the documents without defining a mapping first. Every knn query now fails with errors about the vector field. The team must keep approximate k-NN search for latency, query volume is low, and the data nodes have plenty of free memory. Which actions fix the errors? (Choose TWO.)

  1. AMap the embedding field as knn_vector with a dimension of 1,024
  2. BAdd a post_filter clause to each knn query to narrow results
  3. CCreate the index with the index.knn setting set to true
  4. DMove the data nodes to an instance type with more memory
  5. ESend the searches in batches through the _msearch API
Show the answer and why
  • AMap the embedding field as knn_vector with a dimension of 1,024

    Correct

    The knn_vector type takes a required dimension that sets the number of floats per vector, and it must match the model's 1,024-dimension output.

  • BAdd a post_filter clause to each knn query to narrow results

    Incorrect

    A post_filter narrows the results after the k-NN search and can return fewer than k results. It does not change the field type.

  • CCreate the index with the index.knn setting set to true

    Correct

    To use k-NN, the index must be created with the index.knn setting.

  • DMove the data nodes to an instance type with more memory

    Incorrect

    A larger instance helps when k-NN graphs outgrow the memory available for them. Memory is not the problem here.

  • ESend the searches in batches through the _msearch API

    Incorrect

    The _msearch API batches a large volume of queries into one request. Volume is low, and batching does not fix the mapping.

A vector index needs two things before data arrives: an index created with k-NN enabled and a knn_vector field sized to the embedding model.

Question 8 · choose 1

An archive assistant stores 2 billion vectors that are queried only a few times a day, and sub-second latency is acceptable. The company wants the lowest storage cost for the vectors in a knowledge base. Which vector store fits?

  1. AAn OpenSearch Service domain sized to hold every vector in memory
  2. BAmazon MemoryDB with vector search
  3. CAurora PostgreSQL with an HNSW index on large instances
  4. DAmazon S3 Vectors
Show the answer and why
  • AAn OpenSearch Service domain sized to hold every vector in memory

    Incorrect

    Keeping billions of vectors in memory is costly for a few queries a day.

  • BAmazon MemoryDB with vector search

    Incorrect

    MemoryDB keeps vectors in memory on always-on nodes, which costs far more than a few queries a day justify.

  • CAurora PostgreSQL with an HNSW index on large instances

    Incorrect

    Large always-on database instances cost more than a storage-based option for rare queries.

  • DAmazon S3 Vectors

    Correct

    S3 Vectors is designed for durable, low-cost storage of large vector datasets with sub-second query performance for infrequent queries.

Query frequency drives vector store cost. For huge, rarely queried datasets, storage-based vector buckets are the economical choice.

Question 9 · choose 1

An editor updates 40 articles and deletes 3 in the S3 bucket behind a knowledge base. The assistant still answers from the old versions. What must happen for the knowledge base to reflect the changes?

  1. ADelete and re-create the knowledge base
  2. BRun a sync of the knowledge base data source
  3. CIncrease the number of retrieved results
  4. DNothing, because the knowledge base reads S3 at query time
Show the answer and why
  • ADelete and re-create the knowledge base

    Incorrect

    Re-creating is unnecessary because syncs update the index incrementally.

  • BRun a sync of the knowledge base data source

    Correct

    Each change to the data source requires a sync, and syncing is incremental over changes since the last sync.

  • CIncrease the number of retrieved results

    Incorrect

    Retrieving more results still returns the stale indexed content.

  • DNothing, because the knowledge base reads S3 at query time

    Incorrect

    Content is indexed at ingestion, so changes appear only after a sync.

Knowledge bases index content at sync time. Sync after changes so that new, updated and deleted documents are reflected.

Practise domain 1 →Practise all domains →