Skip to content
BytePatterns

MLA-C02 · Domain 2: ML Model and Foundation Model (FM) Development · 24% of the exam

Task 2.3: Analyze and evaluate the performance of ML and AI systems.

Reproducible experiments with MLflow on SageMaker AI, baselines and drift, shadow variants, explaining predictions, debugging training that does not converge, the metrics for classifiers, regressors and text generation, human and model-based evaluation, and judging RAG retrieval.

Study it

  • Experiments and baselines: MLflow on SageMaker AI, shadow variants and explaining predictions

    Lesson coming

  • Metrics for classifiers and regressors: confusion matrix, precision, recall, F1, AUC, RMSE

    Lesson coming

  • Evaluating generative AI: BLEU, ROUGE, BERTScore, human review, LLM-as-a-judge and Amazon Bedrock evaluations

    Partly covered by: Evaluating LLMs, LLM as a Judge

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

A data science team on SageMaker AI wants every training run to record its parameters, metrics and model artifacts automatically, so runs can be compared side by side and the best model can be registered and reproduced later. The team uses the current SageMaker Studio experience. Which capability should the team adopt?

  1. AManaged MLflow on SageMaker AI for experiment tracking
  2. BSageMaker Experiments Classic through the Experiments Python SDK
  3. CAmazon CloudWatch Logs for each training job
  4. DS3 Versioning on the bucket that stores model artifacts
Show the answer and why
  • AManaged MLflow on SageMaker AI for experiment tracking

    Correct

    Managed MLflow tracks and compares experiment runs, keeps the best models in the MLflow Model Registry, and can register them as SageMaker AI models for deployment.

  • BSageMaker Experiments Classic through the Experiments Python SDK

    Incorrect

    Experiment tracking with the SageMaker Experiments Python SDK is only available in Studio Classic, and AWS recommends the MLflow integration with the new Studio experience instead.

  • CAmazon CloudWatch Logs for each training job

    Incorrect

    Training logs and metrics go to CloudWatch, but logs do not group runs with their parameters and artifacts for side-by-side comparison and registration.

  • DS3 Versioning on the bucket that stores model artifacts

    Incorrect

    Versioning keeps every version of an object. It does not link an artifact to the parameters, data and metrics of the run that produced it.

Reproducible experiments need parameters, metrics and artifacts tied together per run. Managed MLflow is the current SageMaker AI tool for this; Experiments Classic remains only in Studio Classic.

Question 2 · choose 1

A new fraud model has passed offline evaluation. Before it replaces the current model on a SageMaker AI real-time endpoint, the team wants to measure its latency and error rate on live production requests, without any customer ever receiving one of its predictions. What should the ML engineer set up?

  1. AAn A/B test that sends 10% of traffic to a second production variant
  2. BA shadow test that deploys the new model as a shadow variant
  3. CA batch transform job on last month's requests
  4. DA serverless endpoint that hosts the new model
Show the answer and why
  • AAn A/B test that sends 10% of traffic to a second production variant

    Incorrect

    With traffic distribution across production variants, the requests routed to the new variant are answered by it, so some customers would receive its predictions.

  • BA shadow test that deploys the new model as a shadow variant

    Correct

    In a shadow test SageMaker AI sends a copy of live requests to the shadow variant, and only the production variant's responses go back to callers. The shadow responses can be logged for comparison.

  • CA batch transform job on last month's requests

    Incorrect

    Batch transform scores stored data offline. It shows prediction quality on old data but not latency and errors on live traffic.

  • DA serverless endpoint that hosts the new model

    Incorrect

    Serverless inference is a hosting option, not a comparison method, and endpoints that use serverless inference cannot run shadow tests.

Shadow testing compares a candidate with production on real traffic at zero user impact. A/B testing compares variants by actually serving users, which comes after the shadow phase.

Question 3 · choose 2

A team wants to evaluate only the retrieval step of its Amazon Bedrock knowledge base, before any answer is generated. Its prompt dataset includes each query and the ground-truth texts that should be retrieved. Which built-in metrics should the team select for a retrieve-only RAG evaluation job? (Choose TWO.)

  1. AFaithfulness
  2. BContext relevance
  3. CContext coverage
  4. DCitation precision
  5. EBERTScore
Show the answer and why
  • AFaithfulness

    Incorrect

    Faithfulness measures how well generated responses avoid hallucination relative to the retrieved texts, so it belongs to retrieve-and-generate jobs.

  • BContext relevance

    Correct

    Context relevance is a built-in retrieve-only metric that measures how relevant the retrieved texts are to the questions.

  • CContext coverage

    Correct

    Context coverage is a built-in retrieve-only metric that measures how much of the ground-truth information the retrieved texts cover. It requires ground truth in the dataset, which the team has.

  • DCitation precision

    Incorrect

    Citation precision judges how correctly a generated response cites passages. It needs a generated answer, so it is a retrieve-and-generate metric.

  • EBERTScore

    Incorrect

    BERTScore is a computed metric for accuracy and robustness in automatic model evaluation of summarization. It is not one of the RAG retrieval metrics.

Bedrock RAG evaluations have two job types. Retrieve-only jobs score the retrieved context (relevance and coverage); retrieve-and-generate jobs score the answer (correctness, faithfulness, citations and more).

Question 4 · choose 1

A team must compare two candidate models on 5,000 customer questions for qualities such as helpfulness and correctness within a day. It has no group of reviewers available, and it wants a score plus a short written reason for each response. Which Amazon Bedrock evaluation type should the team use?

  1. AA model evaluation job that uses human workers
  2. BAn automatic job with the built-in summarization task and BERTScore
  3. CAmazon Comprehend sentiment analysis of each response
  4. DA model evaluation job that uses a judge model
Show the answer and why
  • AA model evaluation job that uses human workers

    Incorrect

    Human-based evaluation brings in a team of reviewers to rate responses. The team has no reviewers, and 5,000 items in a day is a heavy load for people.

  • BAn automatic job with the built-in summarization task and BERTScore

    Incorrect

    Built-in automatic tasks compute fixed metrics such as BERTScore against reference data. They do not score helpfulness or write a reason for each score.

  • CAmazon Comprehend sentiment analysis of each response

    Incorrect

    Sentiment shows whether a response sounds positive or negative, not whether it is helpful or correct.

  • DA model evaluation job that uses a judge model

    Correct

    With a judge model, Bedrock uses a second LLM to score each response on the selected metrics and to explain each score, so no human workers are needed.

Bedrock evaluations come in three kinds: computed metrics on built-in tasks, human reviews, and LLM-as-a-judge. The judge option scales qualitative criteria without a workforce and explains each score.

Question 5 · choose 1

A team evaluates a model that summarizes news articles. Its reference summaries are often worded differently from the model's summaries even when the meaning is the same, and the team wants a metric that does not penalize good paraphrases. Which metric fits best?

  1. AROUGE-2
  2. BWord error rate
  3. CBERTScore
  4. DToxicity
Show the answer and why
  • AROUGE-2

    Incorrect

    ROUGE counts overlapping word units (here bigrams) between the generated and reference summaries, so a correct paraphrase with different words scores low.

  • BWord error rate

    Incorrect

    In Bedrock's built-in tasks, word error rate is the computed robustness metric for general text generation. It does not measure whether a summary means the same as its reference.

  • CBERTScore

    Correct

    BERTScore compares sentence embeddings from a BERT-family model by cosine similarity, which gives more linguistic flexibility, because sentences with similar meaning are embedded close to each other.

  • DToxicity

    Incorrect

    Toxicity measures harmful content in the output, not whether a summary matches the reference.

N-gram metrics such as ROUGE and BLEU reward shared words; embedding-based metrics such as BERTScore reward shared meaning. Choose the latter when valid outputs can be phrased in many ways.

Question 6 · choose 1

Users of a question-answering assistant often type with small typos and inconsistent capitalization. Before launch, the team wants to measure how much the model's answers change when inputs are slightly perturbed, using an automated Amazon Bedrock evaluation. Which metric category should it select?

  1. AToxicity
  2. BContext coverage
  3. CCitation coverage
  4. DRobustness
Show the answer and why
  • AToxicity

    Incorrect

    Toxicity measures harmful content in outputs, not sensitivity to small input changes.

  • BContext coverage

    Incorrect

    Context coverage is a RAG retrieval metric.

  • CCitation coverage

    Incorrect

    Citation coverage checks whether cited passages support a RAG answer.

  • DRobustness

    Correct

    Bedrock automatic evaluations compute robustness metrics, such as word error rate, deltaBERTScore or deltaF1, which reflect a model's semantic robustness to small changes in input.

Robustness metrics compare outputs on original and perturbed inputs; large changes indicate fragile behavior.

Question 7 · choose 2

A team is comparing two configurations of its RAG application, with different chunking and reranking, using Amazon Bedrock retrieve-and-generate evaluation jobs on the same question set. It mainly cares whether answers are right and whether they address every part of each question. Which built-in metrics should it compare? (Choose TWO.)

  1. AContext relevance
  2. BRefusal
  3. CCorrectness
  4. DCompleteness
  5. EWord error rate
Show the answer and why
  • AContext relevance

    Incorrect

    Context relevance is a retrieve-only metric about retrieved texts, not the final answers.

  • BRefusal

    Incorrect

    Refusal measures how evasive responses are, not whether they are right and complete.

  • CCorrectness

    Correct

    Correctness measures how accurate the responses are in answering the questions.

  • DCompleteness

    Correct

    Completeness measures how well responses answer and resolve all aspects of the questions.

  • EWord error rate

    Incorrect

    Word error rate is a robustness metric for model evaluations, not a RAG answer-quality metric.

Compare RAG configurations on the same dataset with the metrics that match the goal: correctness and completeness for answer quality.

Question 8 · choose 1

After 40 training runs with different learning rates and tree depths, all tracked in managed MLflow on SageMaker AI, a data scientist wants to see the runs side by side, sort them by validation AUC and pick the best configuration. Where should this be done?

  1. AIn the MLflow UI, comparing the runs
  2. BIn AWS Cost Explorer
  3. CIn AWS CloudTrail event history
  4. DIn the S3 console, by opening each artifact
Show the answer and why
  • AIn the MLflow UI, comparing the runs

    Correct

    With managed MLflow you can compare model performance, parameters and metrics across experiments in the MLflow UI and keep track of the best models.

  • BIn AWS Cost Explorer

    Incorrect

    Cost Explorer shows spending, not model metrics.

  • CIn AWS CloudTrail event history

    Incorrect

    CloudTrail lists API calls, not training metrics.

  • DIn the S3 console, by opening each artifact

    Incorrect

    Opening artifacts one by one is slow and does not compare metrics across runs.

Experiment tracking pays off when choosing among runs: compare parameters and metrics in one view, then register the winner.

Question 9 · choose 1

A regulator asks a lender to describe which features drive its credit model's decisions overall, across all applicants, rather than for one applicant. The account cannot use SageMaker Clarify. What should the ML engineer produce?

  1. AOne applicant's per-prediction SHAP values
  2. BGlobal feature importance from SHAP values
  3. CThe model's training loss curve
  4. DThe endpoint's ModelLatency metric
Show the answer and why
  • AOne applicant's per-prediction SHAP values

    Incorrect

    A single applicant's attributions explain one decision, not the model's behavior overall.

  • BGlobal feature importance from SHAP values

    Correct

    AWS's guidance for replacing Clarify points to the SHAP library, and its reference notebook generates global feature importance as well as individual prediction explanations.

  • CThe model's training loss curve

    Incorrect

    The loss curve shows how training progressed, not which features drive decisions.

  • DThe endpoint's ModelLatency metric

    Incorrect

    Latency is an operational metric and says nothing about feature influence.

Global explanations summarize what drives a model overall; local explanations justify individual decisions. Regulators often ask for both.

Question 10 · choose 1

After a fine-tuning job in Amazon Bedrock, the ML engineer opens the training and validation metric files. Training loss falls steadily across all 5 epochs, while validation loss is lowest after epoch 2 and rises afterward. What should the engineer do in the next customization run?

  1. AIncrease the epochs to the maximum
  2. BDelete the validation dataset
  3. CReduce the epoch count to about 2
  4. DRaise the learning rate to its maximum
Show the answer and why
  • AIncrease the epochs to the maximum

    Incorrect

    More epochs would continue past the point where validation loss started rising.

  • BDelete the validation dataset

    Incorrect

    Without validation data, the team would lose the signal that revealed the problem.

  • CReduce the epoch count to about 2

    Correct

    Bedrock writes step-wise training metrics and validation metrics, including validation loss, for analysis. Rising validation loss with falling training loss signals overfitting after epoch 2, and each extra epoch also adds cost.

  • DRaise the learning rate to its maximum

    Incorrect

    A higher learning rate risks instability and does not address training too long.

Use validation curves from customization jobs to pick epochs: stop where validation loss bottoms out.

Practise domain 2 →Practise all domains →