Skip to content
BytePatterns

AIF-C01 · Domain 3: Applications of Foundation Models · 28% of the exam

Task 3.4: Describe methods to evaluate FM performance.

Human, benchmark and model-as-judge evaluation, metrics such as ROUGE, BLEU and BERTScore, Amazon Bedrock evaluations, and the business measures that show whether an application meets its goal.

Study it

  • Evaluating foundation models: human review, benchmarks, ROUGE, BLEU, BERTScore and model-as-judge

    Partly covered by: Evaluating LLMs, LLM as a Judge

  • Evaluating RAG, agents and business outcomes

    Lesson coming

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

A team evaluates a model that writes summaries of news articles. For each article it has a reference summary written by an editor, and it wants a metric that counts the overlapping word sequences (n-grams) between the model's summary and the reference. Which metric fits?

  1. ABERTScore
  2. BROUGE
  3. CToxicity score
  4. DClassification accuracy
Show the answer and why
  • ABERTScore

    Incorrect

    BERTScore compares embeddings of the generated and reference text, not overlapping word sequences.

  • BROUGE

    Correct

    ROUGE-N computes the n-gram overlap between the reference summary and the model's summary, and it is used to score summarization accuracy.

  • CToxicity score

    Incorrect

    A toxicity score comes from a toxicity detector model and measures harmful content, not agreement with a reference.

  • DClassification accuracy

    Incorrect

    Accuracy is the ratio of correctly classified items to all items. A summary is generated text, not a class label.

ROUGE counts shared n-grams with a reference; BERTScore compares meaning through embeddings. Both are used to judge summaries against a gold standard.

Question 2 · choose 1

A team needs to rate thousands of chatbot responses for helpfulness and wants an explanation for each score. It does not want to recruit or manage human reviewers for this round. Which Amazon Bedrock capability fits best?

  1. AA model evaluation job that uses human workers
  2. BA programmatic evaluation that computes BERTScore
  3. CA model evaluation job that uses an LLM as a judge
  4. DThe contextual grounding check in Amazon Bedrock Guardrails
Show the answer and why
  • AA model evaluation job that uses human workers

    Incorrect

    This job type brings in employees or subject-matter experts to rate responses, which is exactly what the team wants to avoid.

  • BA programmatic evaluation that computes BERTScore

    Incorrect

    BERTScore is a computed similarity metric against reference text. It does not judge helpfulness or explain each score.

  • CA model evaluation job that uses an LLM as a judge

    Correct

    In a judge-model evaluation, a second LLM scores the generator model's responses on the metrics you select, such as helpfulness, and explains each score.

  • DThe contextual grounding check in Amazon Bedrock Guardrails

    Incorrect

    Guardrails filter content at run time; the grounding check flags responses not supported by the source. It is not an evaluation job that rates helpfulness.

LLM-as-a-judge scales subjective evaluation that once needed people, and it explains its scores. Spot-check it with humans for high-stakes uses.

Question 3 · choose 2

A team runs a RAG chatbot on Amazon Bedrock Knowledge Bases. Two problems appear in testing: some retrieved passages have nothing to do with the question, and some answers state facts that are not in the retrieved passages. Which built-in Amazon Bedrock RAG evaluation metrics measure these two problems? (Choose TWO.)

  1. AStereotyping
  2. BRefusal
  3. CLogical coherence
  4. DContext relevance
  5. EFaithfulness
Show the answer and why
  • AStereotyping

    Incorrect

    Stereotyping measures generalized statements about individuals or groups in responses, not retrieval quality or grounding.

  • BRefusal

    Incorrect

    Refusal measures how evasive the responses are in answering questions.

  • CLogical coherence

    Incorrect

    Logical coherence checks responses for logical gaps, inconsistencies or contradictions, not whether they stick to the retrieved texts.

  • DContext relevance

    Correct

    Context relevance measures how contextually relevant the retrieved texts are to the questions, the first problem.

  • EFaithfulness

    Correct

    Faithfulness measures how well responses avoid hallucination with respect to the retrieved texts, the second problem.

Evaluate RAG in two halves: retrieval (context relevance, context coverage) and generation (faithfulness, correctness, completeness).

Question 4 · choose 1

A company deploys an AI agent to handle password-reset requests from start to finish. Leadership wants one metric that shows whether the agent meets that business objective. Which metric fits best?

  1. AThe average number of words in the agent's replies
  2. BThe BERTScore of the agent's replies measured against a set of sample replies
  3. CThe number of tokens the agent processes per day
  4. DThe task completion rate, the share of requests resolved without a person
Show the answer and why
  • AThe average number of words in the agent's replies

    Incorrect

    Reply length says nothing about whether a reset was completed; a short or long reply can both fail the task.

  • BThe BERTScore of the agent's replies measured against a set of sample replies

    Incorrect

    BERTScore compares generated text with reference text through embeddings. It measures wording similarity, not task success.

  • CThe number of tokens the agent processes per day

    Incorrect

    Token volume drives cost on token-based pricing; it measures usage, not whether users got their passwords reset.

  • DThe task completion rate, the share of requests resolved without a person

    Correct

    Task completion rate is one of the business objective alignment metrics the exam guide lists for AI applications, alongside user satisfaction and cost per interaction.

For agents, measure outcomes: did the task get done, how satisfied was the user, and what did each interaction cost.

Question 5 · choose 1

A luxury brand must judge whether candidate models write in its distinctive brand voice. Only its own copywriters can make that call reliably. Which Amazon Bedrock evaluation approach fits?

  1. AA human-based evaluation job staffed by the brand's copywriters
  2. BAn automatic evaluation that computes toxicity
  3. CA retrieve-only RAG evaluation job
  4. DAn automatic evaluation job that runs on a built-in open-source prompt dataset
Show the answer and why
  • AA human-based evaluation job staffed by the brand's copywriters

    Correct

    Human-based evaluations bring human input to the process; the workers can be employees of your company or subject-matter experts.

  • BAn automatic evaluation that computes toxicity

    Incorrect

    Toxicity scores measure harmful content, not whether text matches a brand voice.

  • CA retrieve-only RAG evaluation job

    Incorrect

    Retrieve-only RAG evaluations measure retrieval from a knowledge base, not writing style.

  • DAn automatic evaluation job that runs on a built-in open-source prompt dataset

    Incorrect

    Built-in datasets are generic prompt sets; they cannot capture the brand's own voice.

When the quality bar is subjective and only experts can judge it, human-in-the-loop evaluation is the right tool.

Question 6 · choose 1

A team wants a quick first comparison of three foundation models on general text generation, without writing its own test prompts. Which Amazon Bedrock option fits?

  1. AA human-based evaluation job with a newly recruited team of reviewers
  2. BA Guardrails contextual grounding check
  3. CAn automatic evaluation with built-in prompt datasets
  4. DA Provisioned Throughput purchase for each model
Show the answer and why
  • AA human-based evaluation job with a newly recruited team of reviewers

    Incorrect

    Human evaluation is slower to set up and suits subjective criteria; the team wants a quick automated first pass.

  • BA Guardrails contextual grounding check

    Incorrect

    The grounding check filters ungrounded responses at run time; it is not a model comparison.

  • CAn automatic evaluation with built-in prompt datasets

    Correct

    Amazon Bedrock provides built-in prompt datasets, each based on an open-source dataset, that can be used in automatic model evaluation jobs.

  • DA Provisioned Throughput purchase for each model

    Incorrect

    Provisioned Throughput buys capacity for a model; it does not evaluate or compare models.

Benchmark-style built-in datasets give a fast, repeatable first comparison; follow up with your own data for the final choice.

Question 7 · choose 1

A team's reference summaries and the model's summaries often say the same thing in different words, so word-overlap scores look unfairly low. Which metric compares meaning rather than exact word overlap?

  1. AROUGE-N
  2. BClassification accuracy
  3. CWord error rate
  4. DBERTScore
Show the answer and why
  • AROUGE-N

    Incorrect

    ROUGE-N counts n-gram overlaps between the reference and the model's text, so different wording lowers it.

  • BClassification accuracy

    Incorrect

    Accuracy is the ratio of correctly classified items, used for class labels rather than generated text.

  • CWord error rate

    Incorrect

    Word error rate is a robustness metric in Bedrock evaluations; it does not compare the meaning of two texts.

  • DBERTScore

    Correct

    BERTScore compares embeddings of the summary and the reference, so it reflects similarity of meaning.

ROUGE rewards matching words; BERTScore rewards matching meaning. Using both gives a fuller picture of summary quality.

Question 8 · choose 1

A company runs AI agents in production and wants automated, ongoing scoring of how well they complete tasks and handle edge cases, using built-in or custom evaluators. Which AWS capability fits?

  1. AAmazon Bedrock AgentCore Memory
  2. BAmazon Bedrock AgentCore Evaluations
  3. CAmazon Bedrock AgentCore Identity
  4. DAmazon Bedrock AgentCore Observability dashboards
Show the answer and why
  • AAmazon Bedrock AgentCore Memory

    Incorrect

    Memory lets agents remember past interactions; it does not score them.

  • BAmazon Bedrock AgentCore Evaluations

    Correct

    AgentCore Evaluations measures how well agents and tools perform tasks, handle edge cases and stay consistent, using built-in and custom evaluators, before and after deployment.

  • CAmazon Bedrock AgentCore Identity

    Incorrect

    Identity manages agent identities and credentials for accessing resources; it does not assess quality.

  • DAmazon Bedrock AgentCore Observability dashboards

    Incorrect

    Observability traces, debugs and monitors agent workflows step by step; scoring task success with evaluators is what Evaluations adds.

Evaluating an agent means scoring task outcomes over whole sessions, which AgentCore Evaluations does with LLM-as-a-judge evaluators.

Question 9 · choose 2

A team's RAG application on Amazon Bedrock Knowledge Bases gives weak answers, and the team suspects the problem starts before the model writes anything: the search may not be finding the right passages. The team has written ground-truth texts for a set of test questions and wants Amazon Bedrock evaluations to assess only the search step against them. Which choices fit? (Choose TWO.)

  1. ARun a retrieve-only RAG evaluation job
  2. BRun an automatic model evaluation job on a built-in prompt dataset
  3. CScore the job with the citation precision metric
  4. DScore the job with the context coverage metric
  5. ETurn on the contextual grounding check in Amazon Bedrock Guardrails
Show the answer and why
  • ARun a retrieve-only RAG evaluation job

    Correct

    A retrieve-only job bases its report on the data retrieved from the RAG source, so it isolates the search step from response generation.

  • BRun an automatic model evaluation job on a built-in prompt dataset

    Incorrect

    A model evaluation job scores a model's own outputs for a task type such as text generation or summarization; it does not look at what a knowledge base retrieves.

  • CScore the job with the citation precision metric

    Incorrect

    Citation precision is a retrieve-and-generate metric about the passages a generated response cites, so it does not apply to a retrieve-only job.

  • DScore the job with the context coverage metric

    Correct

    Context coverage is a built-in retrieve-only metric that measures how much of the information in the ground-truth texts the retrieved texts cover, and it needs exactly the ground truth the team has prepared.

  • ETurn on the contextual grounding check in Amazon Bedrock Guardrails

    Incorrect

    The contextual grounding check detects and filters hallucinations in model responses; it is a safeguard on responses, not an evaluation of what the retrieval step finds.

Amazon Bedrock RAG evaluations come in two job types. Retrieve only scores what the knowledge base returns (context relevance, context coverage); retrieve and generate also scores the generated answers (correctness, faithfulness, citation metrics and more). To find out whether retrieval is the weak link, evaluate retrieval on its own first.

Practise domain 3 →Practise all domains →