Skip to content
BytePatterns

AIF-C01 · Domain 3: Applications of Foundation Models · 28% of the exam

Task 3.1: Describe design considerations for applications that use foundation models (FMs).

How to choose a foundation model, what inference parameters such as temperature change, how Retrieval Augmented Generation and vector stores work on AWS, what each customization approach costs, and what agents add.

Study it

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

A legal team uses a foundation model in Amazon Bedrock to summarize contracts. Reviewers complain that the summaries use unexpected, creative wording and vary too much from run to run. Which inference parameter change addresses this most directly?

  1. ARaise the temperature so the model explores more word choices
  2. BRaise Top P so that a larger share of candidate tokens is considered
  3. CLower the temperature so the model favors higher-probability tokens
  4. DIncrease the maximum response length so the model has more room
Show the answer and why
  • ARaise the temperature so the model explores more word choices

    Incorrect

    A higher temperature flattens the probability distribution and makes lower-probability tokens more likely, which leads to more random output.

  • BRaise Top P so that a larger share of candidate tokens is considered

    Incorrect

    A higher Top P widens the pool of candidate tokens and lets the model consider less likely outputs, the opposite of what the reviewers want.

  • CLower the temperature so the model favors higher-probability tokens

    Correct

    A lower temperature steepens the probability distribution, makes the model select higher-probability tokens, and leads to more deterministic responses.

  • DIncrease the maximum response length so the model has more room

    Incorrect

    Response length only limits how many tokens are returned. It does not change how the model chooses its words.

Temperature, Top K and Top P control randomness; lowering them makes output more focused and repeatable. Length parameters only cap or penalize length.

Question 2 · choose 1

An HR team wants a chatbot that answers employee questions from the company's policy documents, which change every month. Answers must reflect the latest version and show which document they came from, and the team does not want to retrain a model each month. Which approach fits best?

  1. AFine-tune the foundation model on the updated policy documents at the start of every month
  2. BRaise the temperature so the model generalizes to the new policies
  3. CUse a larger model whose training data is more recent
  4. DUse Retrieval Augmented Generation with Amazon Bedrock Knowledge Bases
Show the answer and why
  • AFine-tune the foundation model on the updated policy documents at the start of every month

    Incorrect

    Retraining foundation models for organization-specific information has high computational and financial costs, and the team wants to avoid monthly training.

  • BRaise the temperature so the model generalizes to the new policies

    Incorrect

    Temperature only changes how random the output is. It cannot give the model information it has never seen.

  • CUse a larger model whose training data is more recent

    Incorrect

    LLM training data is static and has a cut-off date, so a newer model would still miss the company's internal policies and next month's changes.

  • DUse Retrieval Augmented Generation with Amazon Bedrock Knowledge Bases

    Correct

    RAG has the model consult an authoritative knowledge base before it answers, without retraining. Knowledge Bases can answer with direct quotations from the sources or with responses generated from them.

Changing facts plus a need for sources is the classic case for RAG: update the documents, not the model.

Question 3 · choose 2

A company is building a Retrieval Augmented Generation application and needs somewhere to store document embeddings and run similarity searches on them. Which AWS services can serve as the vector store? (Choose TWO.)

  1. AAmazon OpenSearch Service
  2. BAWS Glue DataBrew
  3. CAmazon Polly
  4. DAmazon RDS for PostgreSQL with the pgvector extension
  5. EAmazon Translate
Show the answer and why
  • AAmazon OpenSearch Service

    Correct

    OpenSearch Service offers k-nearest neighbor (k-NN) search over vectors and is a vector database option for generative AI applications.

  • BAWS Glue DataBrew

    Incorrect

    DataBrew is a visual data preparation tool for cleaning and normalizing data. It does not store or search embeddings.

  • CAmazon Polly

    Incorrect

    Amazon Polly converts text into lifelike speech. It has nothing to do with storing vectors.

  • DAmazon RDS for PostgreSQL with the pgvector extension

    Correct

    Amazon RDS for PostgreSQL and Aurora PostgreSQL support the pgvector extension to store embeddings and perform similarity searches.

  • EAmazon Translate

    Incorrect

    Amazon Translate translates text between languages. It does not index embeddings.

On AWS, vectors can live in a dedicated search engine (OpenSearch Service) or next to existing relational data (Aurora or RDS for PostgreSQL with pgvector), among other options.

Question 4 · choose 1

A company uses a large, highly capable model in Amazon Bedrock to classify support emails. The accuracy is good, but the model is slow and expensive at the company's volume. The company wants similar accuracy on this one task from a smaller, faster and cheaper model. Which customization approach fits?

  1. AModel distillation from the large model into a smaller model
  2. BContinued pre-training of the large model on a large set of unlabeled email text
  3. CProvisioned Throughput for the large model
  4. DA higher temperature for the large model's requests
Show the answer and why
  • AModel distillation from the large model into a smaller model

    Correct

    Distillation transfers knowledge from a larger model (teacher) to a smaller, faster, cost-efficient model (student), improving the student's performance for a specific use case.

  • BContinued pre-training of the large model on a large set of unlabeled email text

    Incorrect

    Continued pre-training adapts a model to a domain with unlabeled data. The large model would stay just as slow and costly to run.

  • CProvisioned Throughput for the large model

    Incorrect

    Provisioned Throughput buys a higher level of throughput for a model at a fixed hourly cost. It does not make the model smaller or cheaper per request.

  • DA higher temperature for the large model's requests

    Incorrect

    Temperature changes the randomness of output, not the model's size, speed or cost.

When a big model's quality is right but its cost and latency are not, distillation captures that quality in a smaller model for the specific task.

Question 5 · choose 1

A company is deciding where an AI agent would add more value than a single prompt to a foundation model. Which use case is the best fit for an agent?

  1. ATranslating one product description from English into Spanish for the online catalog
  2. BResolving a customer query by asking questions, checking records and deciding on a handoff
  3. CClassifying the sentiment of a single customer review
  4. DSummarizing one long internal policy document that an employee pastes directly into the prompt
Show the answer and why
  • ATranslating one product description from English into Spanish for the online catalog

    Incorrect

    Translation is a single-step task a model completes from one prompt; it needs no planning or actions in other systems.

  • BResolving a customer query by asking questions, checking records and deciding on a handoff

    Correct

    AWS describes this contact center case as an AI agent: it asks questions, looks up information, responds with a solution and decides whether to pass the query to a human.

  • CClassifying the sentiment of a single customer review

    Incorrect

    Sentiment classification is a single prompt-and-answer task that a model handles without tools or planning.

  • DSummarizing one long internal policy document that an employee pastes directly into the prompt

    Incorrect

    Summarizing text supplied in the prompt is a one-step generation task; there is nothing to look up or act on.

Agents earn their extra complexity on multi-step goals that need decisions, tool calls and data lookups, not on single-step generation tasks.

Question 6 · choose 1

In a company's chatbot, each user uploads a long product manual and then asks many questions about it in the same session. Every question resends the whole manual, so responses are slow and input token costs are high. Which design choice addresses this most directly?

  1. AFine-tune the model on each uploaded manual before the user's chat session starts
  2. BSend the questions through Amazon Bedrock batch inference
  3. CChoose a model that supports prompt caching and cache the manual
  4. DRaise the temperature so the model reads the manual faster
Show the answer and why
  • AFine-tune the model on each uploaded manual before the user's chat session starts

    Incorrect

    Model customization is a training job that adjusts the model's parameters; running one per uploaded manual would add cost and delay, not remove it.

  • BSend the questions through Amazon Bedrock batch inference

    Incorrect

    Batch inference processes prompts asynchronously from files in S3. It does not suit an interactive chat session.

  • CChoose a model that supports prompt caching and cache the manual

    Correct

    Prompt caching reduces inference latency and input token costs for long contexts that are reused across queries, such as a document users ask several questions about.

  • DRaise the temperature so the model reads the manual faster

    Incorrect

    Temperature only influences how random the output is. It does not change how much input the model processes.

Prompt caching support is a model selection criterion: for long, repeated context it cuts both latency and the input token bill.

Question 7 · choose 1

A product team finds that a model's answers are often far longer than the two or three sentences users need, which also raises output token costs. Which inference parameter should the team set?

  1. AA maximum response length
  2. BA higher temperature
  3. CA higher Top K value
  4. DA larger context window for the model
Show the answer and why
  • AA maximum response length

    Correct

    Response length parameters set the minimum or maximum number of tokens returned in the generated response.

  • BA higher temperature

    Incorrect

    Temperature changes how random the token choices are. It does not cap the length of the answer.

  • CA higher Top K value

    Incorrect

    Top K sets how many of the most likely candidates are considered for each next token; a higher value widens choice, not length limits.

  • DA larger context window for the model

    Incorrect

    The context window is how much input the model can consider. It does not limit how long the response is.

Length parameters cap or penalize output length; randomness parameters (temperature, Top K, Top P) shape word choice.

Question 8 · choose 1

A retailer's product catalog already lives in Amazon Aurora PostgreSQL. It wants to add semantic search over product descriptions by storing embeddings next to the existing rows, without adding another database service. What should it use?

  1. AA new Amazon OpenSearch Service domain for the vectors
  2. BAmazon Translate to normalize the product descriptions
  3. CThe AWS Glue Data Catalog to index the product rows
  4. DThe pgvector extension in Aurora PostgreSQL
Show the answer and why
  • AA new Amazon OpenSearch Service domain for the vectors

    Incorrect

    OpenSearch Service is a strong vector store, but it is a separate service, which the retailer wants to avoid.

  • BAmazon Translate to normalize the product descriptions

    Incorrect

    Amazon Translate translates text between languages. It does not store embeddings or run similarity searches.

  • CThe AWS Glue Data Catalog to index the product rows

    Incorrect

    The Data Catalog stores metadata about data sources for discovery and ETL. It does not run vector similarity searches.

  • DThe pgvector extension in Aurora PostgreSQL

    Correct

    Aurora PostgreSQL supports the pgvector extension to store embeddings in the database and perform efficient similarity searches.

When the data already lives in PostgreSQL, pgvector adds vector search in place; a dedicated store such as OpenSearch Service suits large search workloads.

Question 9 · choose 2

A team wants to improve a foundation model's answers for its use case while keeping cost low. Which approaches change the model's behavior without any training job that changes the model's weights? (Choose TWO.)

  1. ASupervised fine-tuning on labeled examples
  2. BIn-context learning with a few examples in the prompt
  3. CContinued pre-training on unlabeled domain text
  4. DDistillation into a smaller student model
  5. ERetrieval Augmented Generation over the company's documents
Show the answer and why
  • ASupervised fine-tuning on labeled examples

    Incorrect

    Fine-tuning adjusts the model's parameters with a training dataset.

  • BIn-context learning with a few examples in the prompt

    Correct

    Few-shot prompting, also called in-context learning, puts examples in the prompt to calibrate the output; nothing in the model is trained.

  • CContinued pre-training on unlabeled domain text

    Incorrect

    Continued pre-training is a training job on large amounts of unlabeled data that adapts the model to a domain.

  • DDistillation into a smaller student model

    Incorrect

    Distillation fine-tunes a student model on responses from a teacher, so it is a training job.

  • ERetrieval Augmented Generation over the company's documents

    Correct

    RAG adds relevant information to the prompt at query time and extends the model to the company's knowledge without retraining it.

The cost ladder runs from prompt changes and RAG (no training) through fine-tuning and distillation to continued and full pre-training (most expensive).

Question 10 · choose 1

In a team's knowledge base, small chunks give precise retrieval matches, but the model then lacks the surrounding context to answer well. Which Amazon Bedrock Knowledge Bases chunking strategy addresses both needs?

  1. ANo chunking, so that each whole document is a single chunk
  2. BHierarchical chunking with parent and child chunks
  3. CFixed-size chunking with the smallest possible chunk size
  4. DDefault chunking of about 300 tokens per chunk
Show the answer and why
  • ANo chunking, so that each whole document is a single chunk

    Incorrect

    With no chunking each document is one chunk, which gives up the precision of small chunks, the opposite of the goal.

  • BHierarchical chunking with parent and child chunks

    Correct

    Hierarchical chunking retrieves small, precise child chunks and then replaces them with broader parent chunks, so the model gets more comprehensive context.

  • CFixed-size chunking with the smallest possible chunk size

    Incorrect

    Smaller fixed chunks make matches more precise but give the model even less surrounding context.

  • DDefault chunking of about 300 tokens per chunk

    Incorrect

    Default chunking splits text into chunks of about 300 tokens; it does not combine precise matching with broader context.

Hierarchical chunking separates what you search with (small child chunks) from what you give the model (larger parent chunks).

Question 11 · choose 1

A company already has a RAG chatbot that answers questions about its travel policy. Managers now want the assistant to also book approved trips by calling the booking system's API. What does this new capability require?

  1. AA larger embeddings model for the knowledge base
  2. BA lower temperature for the existing chatbot
  3. CAn AI agent that can plan steps and call tools
  4. DA second knowledge base for the booking records
Show the answer and why
  • AA larger embeddings model for the knowledge base

    Incorrect

    Embeddings improve retrieval of relevant text; they do not let the assistant call an API.

  • BA lower temperature for the existing chatbot

    Incorrect

    Temperature affects how varied the generated text is, not whether the model can take actions.

  • CAn AI agent that can plan steps and call tools

    Correct

    AI agents perform self-directed tasks toward a goal, and they can invoke APIs and company systems to execute tasks, which goes beyond answering from retrieved text.

  • DA second knowledge base for the booking records

    Incorrect

    A knowledge base retrieves information to ground answers. Retrieval alone does not perform a booking.

RAG lets a model know things; agents let it do things, by planning and calling tools such as APIs.

Practise domain 3 →Practise all domains →