Skip to content
BytePatterns

AIB-C01 · Domain 1: AI Fundamentals and Literacy · 24% of the exam

Task 1.3: Apply GenAI concepts and techniques.

Getting useful output from generative AI: prompt engineering basics, where tokens and context windows limit a system, and when retrieval augmented generation or fine-tuning is the right way to adapt a model to the business.

Study it

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 2

A support manager uses a foundation model to label incoming emails as billing, technical or account. The prompt currently says "Classify this email." The answers come back as long free-form summaries with labels in many different words, which breaks the routing spreadsheet. Which changes to the prompt will most likely fix this? (Choose TWO.)

  1. ARaise the temperature to vary the wording
  2. BName the three allowed categories and ask for exactly one of them
  3. CAdd the company's full product history as background before the email
  4. DAsk for only the category word in the response, with no explanation
  5. EAsk the model to explain its reasoning in detail before giving a label
Show the answer and why
  • ARaise the temperature to vary the wording

    Incorrect

    A higher temperature makes responses more random. That helps with creative variety, but it works against the consistent labels needed here.

  • BName the three allowed categories and ask for exactly one of them

    Correct

    Models work best with simple, clear and complete instructions. Naming the choices explicitly turns a vague request into a defined classification task.

  • CAdd the company's full product history as background before the email

    Incorrect

    Extra context that is unrelated to the task makes the prompt longer and costlier without clarifying what output is expected.

  • DAsk for only the category word in the response, with no explanation

    Correct

    Output indicators that specify the form and length of the response keep the model from adding free-form text the spreadsheet cannot use.

  • EAsk the model to explain its reasoning in detail before giving a label

    Incorrect

    Step-by-step reasoning can help with complex problems, but it makes the response longer and adds exactly the free-form text that breaks the routing.

Inconsistent output usually means an ambiguous prompt. Telling the model exactly which choices exist and exactly what the answer should look like is the first and cheapest fix, before any change of model or customization.

Question 2 · choose 1

A legal operations team wants a single summary of a 900-page master services agreement. When they submit the whole agreement in one request, the request fails because the input is larger than the model can accept at once. What is the most direct way to get the summary?

  1. AIncrease the maximum output length setting for the request
  2. BSummarize it in sections that each fit the context window, then combine them
  3. CFine-tune the model on the agreement so that it can summarize the text from memory
  4. DLower the temperature so the model processes the text more carefully
Show the answer and why
  • AIncrease the maximum output length setting for the request

    Incorrect

    The maximum output setting limits how many tokens the model may generate. It does not let a larger input fit into the context window.

  • BSummarize it in sections that each fit the context window, then combine them

    Correct

    The context window limits how many tokens a model can take in at once. Chunking breaks large documents into pieces small enough to fit, so the work can be done in parts.

  • CFine-tune the model on the agreement so that it can summarize the text from memory

    Incorrect

    Fine-tuning adapts a model's behavior with training examples; it is slow and costly, and it is not a reliable way to make a model hold one specific long document for a summary.

  • DLower the temperature so the model processes the text more carefully

    Incorrect

    Temperature changes how random the generated tokens are. It has no effect on how much input the model can accept.

Every model has a context window, a limit on the tokens it can consider at once. When a document exceeds it, the input must be reduced, typically by chunking the document and working section by section, or by retrieving only the relevant parts. Context window size is also a criterion when choosing a model for long-document work.

Question 3 · choose 1

An insurer wants a generative AI assistant that answers employees' questions about internal underwriting guidelines. The guidelines run to thousands of pages, are revised every week, and every answer must reflect the latest version. Which approach should the business sponsor support?

  1. ARetrieval augmented generation that retrieves current guideline passages per question
  2. BFine-tune a foundation model on the guidelines every time they are revised
  3. CWrite a detailed system prompt that summarizes the guidelines
  4. DContinue pretraining a foundation model on the full guideline archive
Show the answer and why
  • ARetrieval augmented generation that retrieves current guideline passages per question

    Correct

    RAG references an authoritative knowledge base outside the model's training data before generating a response, so updated guidelines are used as soon as they are indexed, without retraining the model.

  • BFine-tune a foundation model on the guidelines every time they are revised

    Incorrect

    Fine-tuning changes the model's weights through a training job. Doing that every week is slow and costly, and for injecting knowledge AWS recommends RAG as the better starting point.

  • CWrite a detailed system prompt that summarizes the guidelines

    Incorrect

    A prompt cannot hold thousands of pages, and a summary loses the details employees need. Prompting alone fits tasks that rely on the model's general knowledge.

  • DContinue pretraining a foundation model on the full guideline archive

    Incorrect

    Continued pretraining adapts a model to domain language using large amounts of data. It is the costliest option and still leaves the model out of date after the next weekly revision.

When answers must rest on specific, proprietary and frequently changing documents, RAG is the standard choice: it grounds the response in the retrieved source and stays current without touching the model. It also lets the system cite the passage it used.

Question 4 · choose 1

A cosmetics brand generates thousands of product descriptions a month. The facts come from the product database, but even after weeks of prompt work, the descriptions do not reliably follow the brand's distinctive voice and fixed seven-part structure. The brand has 5,000 approved descriptions written by its copy team. Which next step is most appropriate?

  1. AFine-tune a model on the 5,000 descriptions the copy team approved
  2. BAdd a retrieval layer that pulls facts from the product database at request time
  3. CPretrain a new foundation model from scratch on the brand's content
  4. DRaise the temperature so the descriptions sound more creative
Show the answer and why
  • AFine-tune a model on the 5,000 descriptions the copy team approved

    Correct

    Fine-tuning suits teaching a model a specific style, format or niche terminology that is hard to reproduce through prompting alone, and the approved descriptions are the high-quality examples it needs.

  • BAdd a retrieval layer that pulls facts from the product database at request time

    Incorrect

    Retrieval brings in facts the model lacks, but the facts are already available. The gap is style and structure, which retrieval does not teach.

  • CPretrain a new foundation model from scratch on the brand's content

    Incorrect

    Building a foundation model is the most expensive option and is reserved for cases where no available model meets the need. 5,000 examples are far too few for pretraining.

  • DRaise the temperature so the descriptions sound more creative

    Incorrect

    Higher temperature increases randomness. It makes the voice and the seven-part structure less consistent, not more.

Start with prompt engineering and add retrieval when the model lacks facts. When the model has the facts but cannot reliably produce a particular style or format, fine-tuning on high-quality examples is the right next step, after weighing its cost against the expected gain.

Question 5 · choose 1

A finance analyst is building a cost estimate for a generative AI summarization service on Amazon Bedrock. The draft assumes the price is charged per word of input and output, because "a token is just a word". What should the analyst change?

  1. AKeep the per-word estimate, because Bedrock converts words to tokens one for one
  2. BEstimate per request, because Bedrock charges a flat fee for each API call
  3. CEstimate only the input, because output tokens are free on Bedrock
  4. DEstimate in tokens, which can be words, word fragments or punctuation
Show the answer and why
  • AKeep the per-word estimate, because Bedrock converts words to tokens one for one

    Incorrect

    Tokens do not map one to one to words; a single word can be split into several tokens, and punctuation counts too.

  • BEstimate per request, because Bedrock charges a flat fee for each API call

    Incorrect

    On-demand model inference is priced by input and output tokens, so long requests cost more than short ones.

  • CEstimate only the input, because output tokens are free on Bedrock

    Incorrect

    Bedrock prices output tokens separately from input tokens, so the length of the generated summaries matters to cost.

  • DEstimate in tokens, which can be words, word fragments or punctuation

    Correct

    A token is a sequence of characters the model treats as one unit, such as a word, part of a word or a punctuation mark. Bedrock's on-demand prices are stated per million input and output tokens.

Tokens are the unit in which models read, generate, are limited and are billed. Because tokens do not map one to one to words, cost estimates and capacity plans should be built in tokens for both the input and the output.

Question 6 · choose 2

A legal assistant built on a foundation model sometimes ignores its original instructions after users paste very long documents into a single conversation. Security reviewers also worry that attackers could flood it with text on purpose to push those instructions out. Lawyers still need to work with long contracts. Which mitigations address both concerns? (Choose TWO.)

  1. AMove to a model with a larger context window and accept any input size
  2. BRaise the maximum output tokens so the model can answer in full
  3. CLimit input size and keep the essential information first
  4. DRepeat the full original instructions inside every user message
  5. EAlert when a conversation nears the context window's capacity
Show the answer and why
  • AMove to a model with a larger context window and accept any input size

    Incorrect

    A larger window only moves the limit; accepting any input size still lets attackers flood the model, and AWS's mitigation is to limit the size of the input.

  • BRaise the maximum output tokens so the model can answer in full

    Incorrect

    Output length parameters limit the response, not how much input the model must hold, so they do nothing about overflow from long inputs.

  • CLimit input size and keep the essential information first

    Correct

    AWS's security guidance on context window overflow recommends input management: limit the size of input, prioritize essential information and sanitize excessive content.

  • DRepeat the full original instructions inside every user message

    Incorrect

    Repeating long instructions adds text to every turn, so the window fills faster, and it does nothing to stop deliberate flooding.

  • EAlert when a conversation nears the context window's capacity

    Correct

    AWS also recommends real-time monitoring that triggers alerts when the context window is nearing capacity, so overflow can be managed before instructions are lost.

Models can only consider a limited amount of information at once. When inputs exceed that limit, earlier instructions can be lost, so input size must be managed and monitored rather than simply allowed to grow.

Question 7 · choose 1

A healthcare provider wants an assistant that answers staff questions using patient-related procedures and records. Its privacy officer worries that putting sensitive data into a model could expose it to the wrong users. Which approach does AWS security guidance favor?

  1. AFine-tuning a model on all patient records so it answers from memory
  2. BPasting full patient records into every prompt for context
  3. CRetrieval with strong authorization, not fine-tuning
  4. DRemoving all access controls
Show the answer and why
  • AFine-tuning a model on all patient records so it answers from memory

    Incorrect

    Training sensitive data into a model makes it hard to control who can draw it out; retrieval keeps access controls in force.

  • BPasting full patient records into every prompt for context

    Incorrect

    AWS advises not to include sensitive data in prompts beyond what the task needs.

  • CRetrieval with strong authorization, not fine-tuning

    Correct

    AWS's guidance on isolating sensitive data recommends using RAG with strong authentication and authorization to augment model data over fine-tuning, alongside data minimization.

  • DRemoving all access controls

    Incorrect

    Removing controls exposes records to people without a right to them.

How a model is adapted affects data security. Keeping sensitive data in governed sources and retrieving it with authorization limits exposure better than embedding it in a model.

Question 8 · choose 1

A marketing team uses a foundation model to brainstorm slogans for a new product. Every run returns nearly the same few slogans. The team wants more varied ideas, while each slogan stays short and follows the same product brief. Which inference setting change fits?

  1. ALower the temperature so the model favors its most likely wording
  2. BRaise the temperature to allow less likely word choices
  3. CLower Top P so the model considers fewer candidate words
  4. DRaise the maximum response length so each run writes more
Show the answer and why
  • ALower the temperature so the model favors its most likely wording

    Incorrect

    A lower temperature steepens the probability distribution and makes responses more deterministic, which narrows variety further.

  • BRaise the temperature to allow less likely word choices

    Correct

    A higher temperature increases the likelihood of lower-probability outputs and flattens the distribution, which leads to more varied responses.

  • CLower Top P so the model considers fewer candidate words

    Incorrect

    A lower Top P considers a smaller share of the probability distribution, cutting less likely words and so reducing variety.

  • DRaise the maximum response length so each run writes more

    Incorrect

    Response length only sets how many tokens may be generated; longer output breaks the short format without making the word choices more varied.

Inference parameters let the same model serve different needs. For brainstorming, more randomness yields more diverse options to choose from, while length settings and the prompt keep the output on format.

Question 9 · choose 1

A finance team is budgeting a fine-tuning project on Amazon Bedrock and assumes the only cost is the model's normal per-request price. What else should the budget include?

  1. ANothing else, because fine-tuning is included in inference pricing
  2. BTraining charges per token processed, plus monthly storage
  3. COnly a one-time license fee for the base model that is being tuned
  4. DOnly the cost of labeling, because training itself is free
Show the answer and why
  • ANothing else, because fine-tuning is included in inference pricing

    Incorrect

    Training and storage are charged separately from inference.

  • BTraining charges per token processed, plus monthly storage

    Correct

    Bedrock charges for model training based on the number of tokens processed (tokens in the training data times epochs) and for model storage per month per model.

  • COnly a one-time license fee for the base model that is being tuned

    Incorrect

    Bedrock's documented customization charges are based on training tokens and monthly storage, not a license fee.

  • DOnly the cost of labeling, because training itself is free

    Incorrect

    Training is not free; tokens processed during training are charged.

Model adaptation choices have different cost profiles. Fine-tuning adds training and storage charges on top of usage, which belongs in the business case alongside the expected improvement.

Question 10 · choose 1

A team prefixes every request to summarize a short news item with ten long example summaries. Output quality is fine, but costs are high. The task relies only on general language skills. What should the team test?

  1. AA zero-shot prompt without the ten long examples
  2. BAdding ten more examples to improve quality further
  3. CFine-tuning a model first
  4. DDoubling the maximum output length for each summary
Show the answer and why
  • AA zero-shot prompt without the ten long examples

    Correct

    AWS's cost guidance suggests experimenting with zero-shot prompting for common knowledge tasks and keeping prompts as short as possible while meeting performance requirements.

  • BAdding ten more examples to improve quality further

    Incorrect

    More examples add tokens and cost when quality is already fine.

  • CFine-tuning a model first

    Incorrect

    Customization costs more than first testing a leaner prompt.

  • DDoubling the maximum output length for each summary

    Incorrect

    Longer outputs raise cost without addressing the prompt overhead.

Prompt design affects cost as well as quality. When a task needs only general skills, a lean prompt can deliver the same results for fewer tokens.

Practise domain 1 →Practise all domains →