Skip to content
BytePatterns

AIP-C01 · Domain 5: Testing, Validation, and Troubleshooting · 11% of the exam

Task 5.1: Implement evaluation systems for GenAI.

Measuring quality before and after release: Bedrock model and RAG evaluations, LLM-as-a-judge and human review, agent evaluations, regression gates and validation of model updates.

Study it

  • Evaluating models and RAG: automatic metrics, LLM-as-a-judge, human review and RAG evaluations

    Partly covered by: Evaluating LLMs, LLM as a Judge

  • Evaluating agents and gating releases: AgentCore Evaluations, regression tests and deployment validation

    Partly covered by: Evaluating LLMs

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

A RAG assistant built on an Amazon Bedrock knowledge base gives incomplete answers to about a quarter of test questions. The team has 400 test questions with the passages that should be retrieved for each and wants to find out whether the problem lies in retrieval or in generation before it changes either one. Which evaluation should the team run first?

  1. AA retrieve-only RAG evaluation job with context relevance and context coverage against the expected passages
  2. BAn automatic model evaluation job on a built-in question answering dataset
  3. CA human-based evaluation job in which reviewers rate the final answers on a five-point scale
  4. DA retrieve-and-generate RAG evaluation job that scores only the helpfulness of the final answers from the assistant
Show the answer and why
  • AA retrieve-only RAG evaluation job with context relevance and context coverage against the expected passages

    Correct

    Retrieve-only jobs score the retrieved texts alone. Context coverage compares them with the ground truth passages, so low scores point to retrieval rather than generation.

  • BAn automatic model evaluation job on a built-in question answering dataset

    Incorrect

    Built-in datasets test the model's general ability, not this knowledge base's retrieval of the company's passages.

  • CA human-based evaluation job in which reviewers rate the final answers on a five-point scale

    Incorrect

    Human ratings of final answers are slow and still do not separate retrieval quality from generation quality.

  • DA retrieve-and-generate RAG evaluation job that scores only the helpfulness of the final answers from the assistant

    Incorrect

    Scoring only final answers mixes retrieval and generation, so it cannot show which stage causes the incomplete answers.

Separate the stages when diagnosing RAG. Retrieve-only evaluations measure context relevance and, with ground truth, context coverage; retrieve-and-generate evaluations add answer metrics such as correctness and faithfulness once retrieval is known to be sound.

Question 2 · choose 1

A company wants to replace a support chatbot that runs on a model hosted outside AWS with a model on Amazon Bedrock. It has 1,000 real questions and the current chatbot's stored answers, and it wants both systems scored with the same judge model and the same built-in quality metrics, without calling the external system again. How should the team set up the evaluation?

  1. AImport the external model into Amazon Bedrock with custom model import so that both models can be invoked by the job
  2. BRun a programmatic evaluation job with a built-in dataset for both models and compare their accuracy scores
  3. CRun LLM-as-a-judge jobs that take the stored answers as the external system's own inference responses
  4. DAsk support agents to compare answers from both systems in a spreadsheet and count which one they prefer
Show the answer and why
  • AImport the external model into Amazon Bedrock with custom model import so that both models can be invoked by the job

    Incorrect

    Custom model import works only for supported open-weight model architectures and is unnecessary when the answers already exist.

  • BRun a programmatic evaluation job with a built-in dataset for both models and compare their accuracy scores

    Incorrect

    A built-in dataset does not reflect the company's real questions, and the external system cannot be invoked by the job.

  • CRun LLM-as-a-judge jobs that take the stored answers as the external system's own inference responses

    Correct

    Judge-based jobs accept your own inference responses in the prompt dataset, in which case Bedrock skips the invocation and scores the supplied answers with the same metrics as the Bedrock model.

  • DAsk support agents to compare answers from both systems in a spreadsheet and count which one they prefer

    Incorrect

    Manual comparison is slow and inconsistent, and it does not use the same judge model and metrics that the company requires.

Bedrock evaluations can score models outside Bedrock by taking their responses as input data. That makes like-for-like comparisons on real traffic possible without re-running the external system.

Question 3 · choose 3

A team is about to ship a new version of an order-support agent that runs on Amazon Bedrock AgentCore. For release, it needs a pre-release regression test on 300 curated past sessions with expected tool call sequences, a deterministic check that refund amounts in tool calls never exceed the order value, and continuous quality scoring of a sample of live sessions after release. Which actions meet these requirements? (Choose THREE.)

  1. AEnable SageMaker Model Monitor on the agent so that it flags sessions with wrong tool calls
  2. BRun a batch evaluation over the curated sessions with expected trajectories as ground truth and trajectory match evaluators
  3. CConfigure online evaluation that samples a percentage of production sessions with the chosen evaluators
  4. DRun an Amazon Bedrock automatic model evaluation job with the built-in robustness metric on the agent's model
  5. ECreate a custom code-based evaluator in Lambda that compares each refund amount with the order value
Show the answer and why
  • AEnable SageMaker Model Monitor on the agent so that it flags sessions with wrong tool calls

    Incorrect

    Model Monitor is closed to new customers and does not evaluate agent sessions or tool calls.

  • BRun a batch evaluation over the curated sessions with expected trajectories as ground truth and trajectory match evaluators

    Correct

    Batch evaluation scores many sessions in one job and supports ground truth such as expected tool trajectories, which fits pre-release regression testing.

  • CConfigure online evaluation that samples a percentage of production sessions with the chosen evaluators

    Correct

    Online evaluation continuously scores live traffic with sampling rules, which provides quality monitoring after release.

  • DRun an Amazon Bedrock automatic model evaluation job with the built-in robustness metric on the agent's model

    Incorrect

    Model evaluations score a model on prompts, not an agent's tool use across multi-step sessions.

  • ECreate a custom code-based evaluator in Lambda that compares each refund amount with the order value

    Correct

    Code-based evaluators run your own Lambda logic for deterministic checks instead of an LLM judge.

Agent evaluation needs session-aware tools. AgentCore Evaluations offers batch jobs with ground truth for regression testing, online sampling of production, and evaluators that can be built-in, LLM-based or code-based.

Question 4 · choose 1

A cosmetics brand must choose between two candidate models for product descriptions. The deciding criteria are subjective, such as whether the copy matches the brand voice, and the brand wants its own copywriters to rate the outputs side by side on custom rating scales inside a managed workflow. Which approach should the team use?

  1. AA Bedrock evaluation job that uses the brand's own work team to compare the two models
  2. BAn automatic model evaluation job on the built-in text generation dataset with the accuracy metric
  3. CA guardrail with a denied topic that blocks descriptions that do not match the brand voice
  4. DAn LLM-as-a-judge evaluation job with the professional style and tone metric for both models
Show the answer and why
  • AA Bedrock evaluation job that uses the brand's own work team to compare the two models

    Correct

    Human-based evaluation jobs let a work team of your own employees rate and compare the responses of up to two models with custom metrics.

  • BAn automatic model evaluation job on the built-in text generation dataset with the accuracy metric

    Incorrect

    Built-in accuracy metrics on public datasets do not measure brand voice.

  • CA guardrail with a denied topic that blocks descriptions that do not match the brand voice

    Incorrect

    Guardrails filter content in production. They do not compare models for a selection decision.

  • DAn LLM-as-a-judge evaluation job with the professional style and tone metric for both models

    Incorrect

    A judge model is fast and repeatable, but the brand explicitly wants its copywriters' judgment on a subjective criterion.

Use human evaluation when the criterion is subjective and the business wants its own experts to decide. Bedrock manages the work team, task instructions and rating methods, and supports comparing two models.

Question 5 · choose 1

A telecom company's support assistant renders every customer question into a prompt template whose directions say: answer in no more than 80 words, use numbered steps for any procedure, never quote a price, and end with a link to the self-service portal. The company wants to move to a cheaper model on Amazon Bedrock, but spot checks suggest the cheaper model ignores some of these directions. Before deciding, the team needs a managed, repeatable score of how closely each model obeys the directions in its prompts, with an explanation for every score, on 800 logged questions rendered through the template. There are no reference answers, and nobody is available to grade responses by hand. Which approach meets these requirements?

  1. AAn LLM-as-a-judge job for each model with the Following instructions metric on the 800 rendered prompts
  2. BAn LLM-as-a-judge job for each model with the Completeness metric on the 800 rendered prompts
  3. CAn LLM-as-a-judge job for each model with the Faithfulness metric on the 800 rendered prompts
  4. DA Logs Insights query over the model invocation logs that counts responses longer than 80 words
Show the answer and why
  • AAn LLM-as-a-judge job for each model with the Following instructions metric on the 800 rendered prompts

    Correct

    Following instructions measures how well a response respects the exact directions in its prompt, which is the failure the spot checks found. It needs no reference answers, the judge explains each score, and 800 prompts fit within the 1,000-prompt limit of a job.

  • BAn LLM-as-a-judge job for each model with the Completeness metric on the 800 rendered prompts

    Incorrect

    Completeness measures whether a response answers every question in the prompt. An answer can be complete and still run past 80 words, quote a price or leave out the portal link.

  • CAn LLM-as-a-judge job for each model with the Faithfulness metric on the 800 rendered prompts

    Incorrect

    Faithfulness flags information that is not found in the prompt. It scores whether answers stay within the given context, not whether they follow the template's directions.

  • DA Logs Insights query over the model invocation logs that counts responses longer than 80 words

    Incorrect

    Invocation logs can be queried with CloudWatch Logs Insights, but a length count checks only one of the four directions and gives no explanation for each response.

Match the metric to the failure. When the risk is a model that drifts from the template's rules, evaluate the fully rendered prompts with an instruction-adherence metric, so the judge sees the same directions that the model saw.

Question 6 · choose 1

A team self-hosts an LLM on a SageMaker AI real-time endpoint and has built a quantized version that passed offline quality tests. Before promoting it, the team must compare its latency and error rate with the current model on real production traffic for a week, and keep the new version's responses for an offline quality review. No customer may receive a response from the new version during this period, and the calling application cannot be changed. What should the team use?

  1. AA second production variant with 5% of the traffic weight on the same endpoint
  2. BA second inference component for the new version, with the application calling both components
  3. CA SageMaker AI shadow test with the new version deployed as a shadow variant
  4. DA batch transform job that replays last week's captured requests against the new version
Show the answer and why
  • AA second production variant with 5% of the traffic weight on the same endpoint

    Incorrect

    Weighted variants split live traffic, so about 5% of customers would receive responses from the new version.

  • BA second inference component for the new version, with the application calling both components

    Incorrect

    This needs changes to the calling application, which must duplicate every request and discard one response, work that the endpoint can do without code.

  • CA SageMaker AI shadow test with the new version deployed as a shadow variant

    Correct

    A shadow variant receives a copy of the inference requests in real time, only the production variant's responses go back to callers, and the shadow responses can be logged for offline comparison while a live dashboard shows the metrics.

  • DA batch transform job that replays last week's captured requests against the new version

    Incorrect

    Replaying captured requests in batch does not measure latency and errors under live production load on the endpoint.

Validating a new model version on real traffic does not have to expose any user to it. The endpoint itself can mirror requests and keep the new version's responses for comparison.

Question 7 · choose 2

A billing agent on Amazon Bedrock AgentCore Runtime sometimes calls get_invoice when it should call refund_payment, and sometimes passes arguments that do not match what the customer said, such as the wrong invoice number. The team already scores whole sessions with the built-in Goal success rate evaluator, but those scores do not show which step went wrong. It wants built-in AgentCore Evaluations evaluators that score each individual tool call, with no ground truth and no evaluator code to write. Which evaluators should the team add? (Choose TWO.)

  1. ATool selection accuracy
  2. BFaithfulness
  3. CCoherence
  4. DTool parameter accuracy
  5. EGoal success rate with ground truth
Show the answer and why
  • ATool selection accuracy

    Correct

    This tool-level evaluator judges whether the agent chose the appropriate tool at that point in the conversation, which targets the get_invoice and refund_payment confusion.

  • BFaithfulness

    Incorrect

    Faithfulness is a trace-level evaluator of the assistant's response, not a score for each tool call.

  • CCoherence

    Incorrect

    Coherence is a trace-level evaluator of how logically consistent the assistant's response is. It does not judge tool calls.

  • DTool parameter accuracy

    Correct

    This tool-level evaluator checks whether the parameters of a tool call are accurately derived from the conversation, which catches the wrong invoice numbers.

  • EGoal success rate with ground truth

    Incorrect

    This is a session-level evaluator that needs a set of success assertions, so it neither scores individual calls nor works without ground truth.

Agent evaluation works at several levels. Session-level scores show whether the user's goal was met, while tool-level evaluators pinpoint the step where the agent chose the wrong tool or filled in the wrong arguments.

Question 8 · choose 1

A support organization must choose between two candidate models on Amazon Bedrock. It has 300 real support questions, each with a reference answer approved by the legal team. The decision must rest on whether each answer is factually right and covers every part of its question as judged against the references, with a written explanation for each score so that legal can audit disagreements. The evaluation must be repeatable and managed, and the team cannot grade 600 answers by hand. Which evaluation should the team run?

  1. AAn automatic evaluation job for question answering with the accuracy metric on the 300 questions and references
  2. BAn LLM-as-a-judge job with the Logical coherence and Professional style and tone metrics
  3. CA CloudWatch comparison of output token counts and latency for the two models on the same questions
  4. DAn LLM-as-a-judge job with the Correctness and Completeness metrics and the references in the dataset
Show the answer and why
  • AAn automatic evaluation job for question answering with the accuracy metric on the 300 questions and references

    Incorrect

    For question answering, the automatic accuracy metric is computed as an NLP-F1 score against the reference. It gives no written explanation per answer and does not judge whether every part of a question was answered.

  • BAn LLM-as-a-judge job with the Logical coherence and Professional style and tone metrics

    Incorrect

    These metrics judge consistency and tone. Neither compares an answer's facts or coverage with the approved reference.

  • CA CloudWatch comparison of output token counts and latency for the two models on the same questions

    Incorrect

    Token counts and latency describe cost and speed, not whether the answers are right or complete.

  • DAn LLM-as-a-judge job with the Correctness and Completeness metrics and the references in the dataset

    Correct

    Correctness and Completeness are judge metrics that take reference responses into account when they are in the prompt dataset, and the judge explains how it scored each response.

Approved references let a judge model score correctness and completeness repeatably, while its explanations keep the scores auditable. Computed overlap scores are useful for quick checks but cannot explain a judgment.

Question 9 · choose 1

A media company will add a headline-summary feature and must shortlist one of three candidate models on Amazon Bedrock by tomorrow. It has no reference summaries or prompt dataset of its own yet, and it does not want to design judge prompts or rubrics for this first pass. The shortlist must rest on a repeatable, managed score of summary quality, plus a toxicity check, for each model. Which approach meets these requirements?

  1. AAn LLM-as-a-judge evaluation job with the Completeness metric on prompts that the team writes first
  2. BAn automatic evaluation job for text summarization with its built-in dataset, accuracy and toxicity
  3. CA side-by-side comparison of ten summaries per model in the Amazon Bedrock playground
  4. DAn automatic evaluation job for question answering with the built-in BoolQ dataset
Show the answer and why
  • AAn LLM-as-a-judge evaluation job with the Completeness metric on prompts that the team writes first

    Incorrect

    Judge jobs score responses to prompts you supply in your own dataset, which the team does not have and cannot write by tomorrow.

  • BAn automatic evaluation job for text summarization with its built-in dataset, accuracy and toxicity

    Correct

    For text summarization, automatic evaluation offers the built-in Gigaword dataset with an accuracy metric computed as BERTScore and a toxicity metric, so the three models get comparable scores without any team-built data.

  • CA side-by-side comparison of ten summaries per model in the Amazon Bedrock playground

    Incorrect

    Reading a handful of outputs by hand is neither a managed score nor repeatable across models.

  • DAn automatic evaluation job for question answering with the built-in BoolQ dataset

    Incorrect

    BoolQ holds yes/no questions about passages for the question answering task type, so its scores say little about summary quality.

Built-in automatic evaluations give a fast, repeatable first pass for standard tasks. Judge or human evaluations with the team's own data can follow for criteria that the built-in metrics do not capture.

Practise domain 5 →Practise all domains →