Skip to content
BytePatterns

AIF-C01 · Domain 2: Fundamentals of GenAI · 24% of the exam

Task 2.1: Explain the basic concepts of generative AI (GenAI).

Tokens, embeddings, transformers, diffusion and multimodal models, the foundation model lifecycle, how token pricing shapes cost, and the building blocks of agents: tools, memory, orchestration and the Model Context Protocol.

Study it

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

A company has two million support articles. It wants a search feature that finds articles with the same meaning as a customer's question even when the article uses different words. What should the company create from the articles to make this possible?

  1. AEmbeddings, numeric vectors that can be compared for similarity
  2. BToken lists that are matched against the question word for word
  3. CA higher temperature setting so that the model considers related words
  4. DA model with a context window large enough to hold every article
Show the answer and why
  • AEmbeddings, numeric vectors that can be compared for similarity

    Correct

    An embedding turns input into a vector of numbers so that objects can be compared; sentences can be compared to determine how similar their meaning is, even when the wording differs.

  • BToken lists that are matched against the question word for word

    Incorrect

    A token is a unit of text a model reads, such as a word or part of a word. Matching tokens exactly still needs the same words, which is the limitation the company wants to remove.

  • CA higher temperature setting so that the model considers related words

    Incorrect

    Temperature changes how random a model's generated output is. It has no role in finding which stored articles are similar to a question.

  • DA model with a context window large enough to hold every article

    Incorrect

    The context window is the number of tokens a model can take into account at one time. Two million articles will not fit, and a bigger window does not provide a similarity search.

Semantic search compares meaning, and meaning is compared through embedding vectors: similar texts get vectors that are close to each other.

Question 2 · choose 1

A team uses on-demand inference for a text model in Amazon Bedrock. It plans to add a 2,000-token block of instructions to the start of every request. How will this change affect the application?

  1. ACosts stay the same, because on-demand inference is billed per hour of model runtime
  2. BEvery request costs more, because input tokens are billed, and latency may also rise
  3. CCosts stay the same, because only output tokens are billed and the instructions are input
  4. DCosts stay the same, because each request is billed at a flat price whatever its length
Show the answer and why
  • ACosts stay the same, because on-demand inference is billed per hour of model runtime

    Incorrect

    Hourly billing applies to Provisioned Throughput, which is bought in model units. On-demand text models are priced per input and output token.

  • BEvery request costs more, because input tokens are billed, and latency may also rise

    Correct

    On-demand text models are priced per million input tokens and per million output tokens, so a longer prompt costs more on every call. More tokens in the input can also increase inference latency.

  • CCosts stay the same, because only output tokens are billed and the instructions are input

    Incorrect

    Input tokens are billed too; the pricing tables list a price per million input tokens next to the price per million output tokens.

  • DCosts stay the same, because each request is billed at a flat price whatever its length

    Incorrect

    On-demand text inference is not priced per request. The price depends on how many input and output tokens are processed.

With token-based pricing, prompt length is a cost and performance lever: every token you add to a reused prompt is paid for on every request.

Question 3 · choose 1

A design team is reading about image generation models. One model type adds noise to data step by step during training and then learns to reverse the process, removing noise gradually to produce a new image. Which model type is this?

  1. AA transformer-based large language model
  2. BAn embeddings model
  3. CA diffusion model
  4. DA recommendation model trained on user clicks
Show the answer and why
  • AA transformer-based large language model

    Incorrect

    Large language models are transformer networks trained to understand and generate text. Gradual noising and denoising is not how they are described.

  • BAn embeddings model

    Incorrect

    An embeddings model turns input into numeric vectors for comparing similarity. It does not generate new images.

  • CA diffusion model

    Correct

    Diffusion models add controlled noise to data over several iterations and then reverse the process, denoising step by step to produce a new data sample.

  • DA recommendation model trained on user clicks

    Incorrect

    A recommendation model suggests items to users from their past interactions. It does not create images.

Noise in, noise out, step by step: that iterative denoising is the defining mechanism of diffusion models, which many image generators use.

Question 4 · choose 1

A company connects its policy documents to Amazon Bedrock Knowledge Bases. What happens to the documents during ingestion, before any question is asked?

  1. AThey are used to fine-tune the foundation model so that it memorizes the policies
  2. BEach data source is turned into one embedding that represents all of its documents at once
  3. CThey are stored as token IDs, and questions are answered by exact token matching
  4. DThey are split into chunks, and each chunk is turned into an embedding in a vector index
Show the answer and why
  • AThey are used to fine-tune the foundation model so that it memorizes the policies

    Incorrect

    A knowledge base removes the need to keep training the model on your private data. The model stays as it is; the documents are retrieved at query time.

  • BEach data source is turned into one embedding that represents all of its documents at once

    Incorrect

    Documents are split into many chunks first, and each chunk gets its own embedding, so retrieval can return the relevant passage.

  • CThey are stored as token IDs, and questions are answered by exact token matching

    Incorrect

    The stored chunks are converted to vector embeddings so that queries and text can be compared by semantic similarity, not by exact match.

  • DThey are split into chunks, and each chunk is turned into an embedding in a vector index

    Correct

    During ingestion, documents are split into manageable chunks, the chunks are converted to embeddings and written to a vector index, and a mapping to the original document is kept.

Chunking decides what a retrieval can return: the knowledge base embeds and searches chunks, then hands the best ones to the model as context.

Question 5 · choose 1

An architect explains that the company's AI agents will reach internal APIs and third-party services through the Model Context Protocol (MCP). What is the role of MCP in this design?

  1. AIt gives agents a standard way to connect to external tools and data sources
  2. BIt compresses prompts so that the agents use fewer input tokens per request
  3. CIt fine-tunes the agent's foundation model on the APIs' documentation
  4. DIt encrypts all traffic between the agents and the company's on-premises data center
Show the answer and why
  • AIt gives agents a standard way to connect to external tools and data sources

    Correct

    MCP is a standard for exposing tools and context to models and agents. AgentCore Gateway, for example, converts APIs and Lambda functions into MCP-compatible tools that agents can discover and call.

  • BIt compresses prompts so that the agents use fewer input tokens per request

    Incorrect

    MCP is about connecting agents to tools and context. Reusing repeated prompt prefixes to cut input token cost is what prompt caching does.

  • CIt fine-tunes the agent's foundation model on the APIs' documentation

    Incorrect

    MCP makes tools available for an agent to call at run time. Adapting a model to your data by training it is model customization, such as fine-tuning, which is a separate process.

  • DIt encrypts all traffic between the agents and the company's on-premises data center

    Incorrect

    MCP is not an encryption protocol. Private connectivity to AWS services is provided by features such as AWS PrivateLink.

MCP standardizes how an agent discovers and calls tools, so one integration works with any MCP-compatible agent instead of custom code for each pair.

Question 6 · choose 1

An insurer designs an agentic system for claims. One agent receives each claim, splits the work into document checks, fraud screening and payout calculation, hands each step to a specialist agent in a set order, and tracks every step until the claim is done. Which agentic pattern does this describe?

  1. AMulti-agent collaboration among peer agents that negotiate roles
  2. BA workflow orchestration agent that delegates to subagents
  3. CA single retrieval-augmented agent answering from a knowledge base
  4. DA single reasoning agent that works through the claim on its own
Show the answer and why
  • AMulti-agent collaboration among peer agents that negotiate roles

    Incorrect

    Multi-agent collaboration emphasizes peer-to-peer, emergent coordination without a central coordinator. This design has one agent in charge and a predefined sequence.

  • BA workflow orchestration agent that delegates to subagents

    Correct

    Workflow orchestration agents coordinate multistep tasks: a central agent delegates work to subagents, tracks execution state and passes results to the next agent in the sequence.

  • CA single retrieval-augmented agent answering from a knowledge base

    Incorrect

    A retrieval-augmented agent is a single-agent pattern that answers from retrieved data. It does not split work across other agents.

  • DA single reasoning agent that works through the claim on its own

    Incorrect

    Workflow orchestration agents differ from agents that reason and act in isolation; here the work is delegated to several specialists.

A central coordinator with a predefined flow is workflow orchestration; peers that negotiate and divide work dynamically are multi-agent collaboration.

Question 7 · choose 1

Amazon Bedrock bills text models by the token. Which statement best describes a token?

  1. AA unit of text, such as a word, part of a word or a punctuation mark
  2. BExactly one complete word, so the token count always equals the word count of a prompt
  3. CA numeric vector that represents the meaning of a whole document for similarity search
  4. DA credential that grants an application permission to call the model
Show the answer and why
  • AA unit of text, such as a word, part of a word or a punctuation mark

    Correct

    A token is a sequence of characters a model interprets or predicts as a single unit, which can be a word, part of a word such as "-ed", a punctuation mark or a common phrase.

  • BExactly one complete word, so the token count always equals the word count of a prompt

    Incorrect

    Tokens can be parts of words, punctuation marks or common phrases, so the token count does not simply equal the word count.

  • CA numeric vector that represents the meaning of a whole document for similarity search

    Incorrect

    A vector of numbers that represents meaning for comparison is an embedding, not a token.

  • DA credential that grants an application permission to call the model

    Incorrect

    In billing, a token is a unit of text the model processes. Permission to call a model is a separate access control matter, managed with IAM.

Models read and write tokens, not words or characters, which is why prompts and responses are measured and priced in tokens.

Question 8 · choose 1

Which mechanism lets a transformer-based model look at all parts of an input sequence at once and work out which parts matter most for each word?

  1. ATemperature sampling
  2. BSelf-attention
  3. CRecurrent processing, one token at a time
  4. DChunking the input into fixed-size pieces
Show the answer and why
  • ATemperature sampling

    Incorrect

    Temperature shapes the probability distribution used to pick the next output token; it does not relate parts of the input to each other.

  • BSelf-attention

    Correct

    The self-attention mechanism lets the model look at different parts of the sequence all at once and determine which parts are most important, instead of processing the data in order.

  • CRecurrent processing, one token at a time

    Incorrect

    Earlier recurrent neural networks processed inputs sequentially; transformers process entire sequences in parallel.

  • DChunking the input into fixed-size pieces

    Incorrect

    Chunking splits documents for retrieval in a knowledge base. It is not how a transformer relates words within a sequence.

Self-attention is the core of the transformer: every position weighs every other position, which also lets training run in parallel.

Question 9 · choose 1

An insurer wants customers to upload a photo of a damaged car together with a typed description, and have one model answer questions about both. What kind of model does this require?

  1. AA unimodal text model
  2. BAn embeddings model
  3. CA multimodal model
  4. DA text-to-speech model
Show the answer and why
  • AA unimodal text model

    Incorrect

    A unimodal model processes a single data type; a text-only model cannot read the photo.

  • BAn embeddings model

    Incorrect

    An embeddings model turns input into vectors for comparing similarity. It does not answer questions.

  • CA multimodal model

    Correct

    Multimodal models can integrate multiple data types, such as images and text, in the same request.

  • DA text-to-speech model

    Incorrect

    Text-to-speech models such as those in Amazon Polly turn text into speech; they do not interpret photos.

Modality is a model selection criterion: if the input mixes images and text, the model must be multimodal.

Question 10 · choose 1

A team has selected a pre-trained foundation model and fine-tuned it for its claims-processing use case. In the foundation model lifecycle, which step should come next, before the model is deployed?

  1. APre-train the model again from scratch on the same claims data to be safe
  2. BSkip testing, because fine-tuning always improves a model
  3. CCollect user feedback from production before any testing
  4. DEvaluate the model on use-case data against the success criteria
Show the answer and why
  • APre-train the model again from scratch on the same claims data to be safe

    Incorrect

    Pre-training from scratch is the most expensive step in the lifecycle and would discard the pre-trained model the team just adapted.

  • BSkip testing, because fine-tuning always improves a model

    Incorrect

    Fine-tuning changes the model, and results have to be verified; AWS advises evaluating models on your own data before choosing one.

  • CCollect user feedback from production before any testing

    Incorrect

    Feedback from production comes after deployment. A model should be evaluated before it reaches users.

  • DEvaluate the model on use-case data against the success criteria

    Correct

    Evaluation, with programmatic metrics, human review or a judge model, shows whether the customized model meets the use case before it is deployed.

The lifecycle runs data selection, model selection, pre-training, fine-tuning, evaluation, deployment and feedback. Evaluation is the gate before deployment.

Question 11 · choose 1

A travel agent built on a foundation model treats every conversation as new: it forgets a returning customer's seating and meal preferences from last week. What does the agent need?

  1. ALong-term memory that persists across sessions
  2. BA lower temperature setting for every request it sends
  3. CA larger model with more parameters
  4. DA guardrail with a denied topic
Show the answer and why
  • ALong-term memory that persists across sessions

    Correct

    Without memory, agents treat each interaction as new. Long-term memory, such as AgentCore Memory provides, persists knowledge across sessions, while short-term memory covers a single multi-turn conversation.

  • BA lower temperature setting for every request it sends

    Incorrect

    Temperature controls randomness in the output. It does not make the model remember anything.

  • CA larger model with more parameters

    Incorrect

    Models do not recall prior prompts and previous requests unless the earlier interaction is included in the current prompt, whatever their size.

  • DA guardrail with a denied topic

    Incorrect

    Denied topics block undesirable subjects. They do not store customer preferences.

Foundation models are stateless between requests; agent memory is the component that keeps context within and across sessions.

Practise domain 2 →Practise all domains →