Question 1 · choose 1
A RAG assistant built on an Amazon Bedrock knowledge base gives incomplete answers to about a quarter of test questions. The team has 400 test questions with the passages that should be retrieved for each and wants to find out whether the problem lies in retrieval or in generation before it changes either one. Which evaluation should the team run first?
- AA retrieve-only RAG evaluation job with context relevance and context coverage against the expected passages
- BAn automatic model evaluation job on a built-in question answering dataset
- CA human-based evaluation job in which reviewers rate the final answers on a five-point scale
- DA retrieve-and-generate RAG evaluation job that scores only the helpfulness of the final answers from the assistant
Show the answer and why
AA retrieve-only RAG evaluation job with context relevance and context coverage against the expected passages
Correct
Retrieve-only jobs score the retrieved texts alone. Context coverage compares them with the ground truth passages, so low scores point to retrieval rather than generation.
BAn automatic model evaluation job on a built-in question answering dataset
Incorrect
Built-in datasets test the model's general ability, not this knowledge base's retrieval of the company's passages.
CA human-based evaluation job in which reviewers rate the final answers on a five-point scale
Incorrect
Human ratings of final answers are slow and still do not separate retrieval quality from generation quality.
DA retrieve-and-generate RAG evaluation job that scores only the helpfulness of the final answers from the assistant
Incorrect
Scoring only final answers mixes retrieval and generation, so it cannot show which stage causes the incomplete answers.
Separate the stages when diagnosing RAG. Retrieve-only evaluations measure context relevance and, with ground truth, context coverage; retrieve-and-generate evaluations add answer metrics such as correctness and faithfulness once retrieval is known to be sound.
AWS documentation
- Use metrics to understand RAG system performance (opens in a new tab)
- Evaluate the performance of RAG sources using Amazon Bedrock evaluations (opens in a new tab)
- Model evaluation task types in Amazon Bedrock (opens in a new tab)
- Creating a model evaluation job that uses human workers in Amazon Bedrock (opens in a new tab)