Skip to content
BytePatterns

MLA-C02 · Domain 2: ML Model and Foundation Model (FM) Development · 24% of the exam

Task 2.2: Train, fine-tune, and customize models for ML and AI solutions.

Training with SageMaker AI built-in algorithms and script mode, automatic model tuning, early stopping and distributed training, fighting overfitting and catastrophic forgetting, ensembles, the core hyperparameters, and customizing foundation models and retrieval.

Study it

  • Selecting a foundation model in Amazon Bedrock and choosing between prompting, RAG and fine-tuning

    Partly covered by: Fine-Tuning vs Prompting, Retrieval-Augmented Generation, Adapters and LoRA

  • Training on SageMaker AI: built-in algorithms, script mode and training jobs

    Lesson coming

  • Automatic model tuning, early stopping and distributed training

    Lesson coming

  • Overfitting, underfitting, regularization, catastrophic forgetting and ensembles

    Lesson coming

  • Tuning retrieval: embedding models, chunk size, hybrid search and reranking

    Partly covered by: Chunking and Reranking, Embeddings

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

An ML engineer configures SageMaker AI automatic model tuning for an XGBoost model with five continuous hyperparameters. The budget is 200 training jobs, and to finish overnight the engineer wants to run 50 jobs at the same time without the tuning results getting worse because of that parallelism. Which tuning strategy should the engineer choose?

  1. ABayesian optimization
  2. BGrid search
  3. CRandom search
  4. DA warm start job with IDENTICAL_DATA_AND_ALGORITHM
Show the answer and why
  • ABayesian optimization

    Incorrect

    Bayesian optimization uses the results of completed jobs to choose the next combinations. Running many jobs at once means each choice is made with less information.

  • BGrid search

    Incorrect

    Grid search supports only categorical hyperparameters, so it cannot search five continuous ranges.

  • CRandom search

    Correct

    Random search picks each combination independently of earlier results, so the maximum number of concurrent jobs can run without changing the performance of the tuning.

  • DA warm start job with IDENTICAL_DATA_AND_ALGORITHM

    Incorrect

    Warm start reuses earlier tuning jobs as a starting point. There is no parent job here, and warm start is not a way to run more jobs in parallel.

Bayesian optimization learns from finished jobs, which makes it efficient but sequential in spirit. Random search does not learn between jobs, so it parallelizes freely.

Question 2 · choose 1

A PyTorch training job on SageMaker AI runs for about 30 hours on GPU instances, and the team wants to cut its compute cost substantially. The job can tolerate interruptions as long as it does not lose most of its progress. What should the ML engineer do?

  1. AUse managed Spot training and restart interrupted jobs from scratch
  2. BTurn on SageMaker AI managed warm pools for the job
  3. CUse managed Spot training with checkpoints in Amazon S3
  4. DMove the job to a larger On-Demand instance type
Show the answer and why
  • AUse managed Spot training and restart interrupted jobs from scratch

    Incorrect

    Without checkpoints an interrupted job starts over, and AWS recommends checkpointing for any Spot job that does not finish quickly. A 30-hour job would risk repeating hours of work.

  • BTurn on SageMaker AI managed warm pools for the job

    Incorrect

    Warm pools keep infrastructure after a job ends so the next similar job starts faster. They do not lower the price of the compute, and the retained instances are billable.

  • CUse managed Spot training with checkpoints in Amazon S3

    Correct

    Managed Spot training can cut training cost by up to 90% compared with On-Demand. With checkpointing, SageMaker AI copies checkpoints to S3 and restores them after an interruption, so the job resumes instead of restarting.

  • DMove the job to a larger On-Demand instance type

    Incorrect

    A larger instance may finish sooner but is billed at a higher On-Demand rate, so it does not give the large saving the team wants.

Spot capacity is the main lever for training cost; checkpoints are what make long jobs safe on it. Set EnableManagedSpotTraining and a MaxWaitTimeInSeconds larger than MaxRuntimeInSeconds.

Question 3 · choose 2

While training a neural network, an ML engineer sees training loss keep falling every epoch, while validation loss stops improving after epoch 5 and then rises steadily through epoch 30. Which changes address this behavior? (Choose TWO.)

  1. AStop training early, when validation loss stops improving
  2. BTrain for more epochs so the model can converge
  3. CIncrease regularization, such as a stronger weight penalty
  4. DAdd more layers and engineered features to the model
  5. ERaise the learning rate so training converges faster
Show the answer and why
  • AStop training early, when validation loss stops improving

    Correct

    The rising validation loss shows the model is overfitting. Early stopping ends training before the model learns noise in the training data.

  • BTrain for more epochs so the model can converge

    Incorrect

    More epochs help an underfit model. Here the model already fits the training data too closely, so more passes widen the gap.

  • CIncrease regularization, such as a stronger weight penalty

    Correct

    Regularization reduces model flexibility by penalizing parameters that add little to predictions, which counters overfitting.

  • DAdd more layers and engineered features to the model

    Incorrect

    A more complex model is more likely to learn the noise in the training data, which makes overfitting worse.

  • ERaise the learning rate so training converges faster

    Incorrect

    A higher learning rate can speed convergence but risks instability. It does not address the gap between training and validation loss.

Falling training loss with rising validation loss is the classic sign of overfitting (high variance). Reduce flexibility or stop earlier; do the opposite only for underfitting.

Question 4 · choose 1

A team fine-tuned an Amazon Nova model on SageMaker HyperPod using only its own insurance-claims data. The model now handles claims well, but its general instruction following and reasoning dropped sharply on benchmark tests. The team wants to keep the domain gains and recover the general skills in the next run. What should the ML engineer change?

  1. AIncrease the number of epochs on the claims data
  2. BRaise the learning rate so the next run finishes sooner
  3. CRemove the validation set so all claims data is used for training
  4. DTurn on data mixing so Nova training data is blended in
Show the answer and why
  • AIncrease the number of epochs on the claims data

    Incorrect

    More passes over the same narrow data push the model further toward the domain and away from what it knew before.

  • BRaise the learning rate so the next run finishes sooner

    Incorrect

    A higher learning rate makes larger updates to the weights and can make training unstable. It does nothing to preserve general skills.

  • CRemove the validation set so all claims data is used for training

    Incorrect

    Dropping the validation set adds a little data but removes the main way to see the problem during training.

  • DTurn on data mixing so Nova training data is blended in

    Correct

    Data mixing combines the custom dataset with samples from Nova's own training data. AWS documents it as a way to prevent overfitting on custom data and catastrophic forgetting of the model's existing capabilities.

Catastrophic forgetting happens when training on new, narrow data overwrites what a model already knew. Mixing general data back into fine-tuning, and watching general benchmarks as well as domain metrics, are the usual defenses.

Question 5 · choose 1

A parts catalog assistant uses an Amazon Bedrock knowledge base with an Amazon OpenSearch Serverless vector store that has a filterable text field. Users often type exact part numbers such as "XJ-4471-B", and semantic retrieval frequently returns chunks about similar but different parts. What should the ML engineer change first?

  1. AIncrease numberOfResults so more chunks come back
  2. BSet the knowledge base retrieval search type to hybrid
  3. CSwitch to 256-dimension Titan Text Embeddings V2 vectors
  4. DTurn off chunking so each document is one chunk
Show the answer and why
  • AIncrease numberOfResults so more chunks come back

    Incorrect

    Returning more chunks adds more near-miss results to the context. It does not make the exact part number rank higher.

  • BSet the knowledge base retrieval search type to hybrid

    Correct

    Hybrid search combines vector (semantic) search with search through the raw text, which helps with exact tokens such as part numbers. It is supported for OpenSearch Serverless stores that contain a filterable text field.

  • CSwitch to 256-dimension Titan Text Embeddings V2 vectors

    Incorrect

    Smaller vectors reduce storage and cost, but they are still semantic embeddings and will still blur similar part numbers together.

  • DTurn off chunking so each document is one chunk

    Incorrect

    With no chunking each whole document is one embedding, which makes matching on one specific part even less precise.

Semantic search is good at meaning and weak at exact identifiers. Hybrid search adds keyword matching on the raw text, which is the usual fix for codes, IDs and rare terms.

Question 6 · choose 1

An ML engineer runs automatic model tuning on a custom PyTorch training script. The tuning job starts, but it cannot find the objective metric, val_f1, that the script prints to its logs as "val_f1=0.83;". What must the engineer provide in the tuning job configuration?

  1. AA larger instance for the training jobs
  2. BThe Hyperband strategy instead of Bayesian
  3. CA warm start from a previous tuning job
  4. DA metric definition with a name and a regex
Show the answer and why
  • AA larger instance for the training jobs

    Incorrect

    Instance size does not tell the tuning job how to read the metric from the logs.

  • BThe Hyperband strategy instead of Bayesian

    Incorrect

    The strategy chooses hyperparameters; it still needs a metric definition to read the objective.

  • CA warm start from a previous tuning job

    Incorrect

    Warm start reuses earlier results; it does not define how to parse the metric.

  • DA metric definition with a name and a regex

    Correct

    For custom algorithms, you define metrics by specifying a name and a regular expression; tuning parses the training job's stdout and stderr with it to find the values.

Built-in algorithms emit known metrics; custom scripts need metric definitions so tuning can read the objective from the logs.

Question 7 · choose 2

An ML engineer enables checkpointing for a long SageMaker AI training job, with the default local checkpoint path and an S3 URI for checkpoint storage. Which statements about how checkpoints behave are correct? (Choose TWO.)

  1. ACheckpoints added to S3 mid-job are copied in
  2. BFiles written to /opt/ml/checkpoints sync to S3
  3. CExisting S3 checkpoints are copied in at job start
  4. DDeleting a local checkpoint keeps the S3 copy
  5. ECheckpoints cannot use S3 Express One Zone
Show the answer and why
  • ACheckpoints added to S3 mid-job are copied in

    Incorrect

    Checkpoints added to the S3 folder after the job has started are not copied to the training container.

  • BFiles written to /opt/ml/checkpoints sync to S3

    Correct

    Checkpoint files are saved under a local directory, by default /opt/ml/checkpoints, and SageMaker AI copies them to S3 and keeps the directory in sync during training.

  • CExisting S3 checkpoints are copied in at job start

    Correct

    Existing checkpoints in S3 are written to the container at the start of the job, which lets a job resume from a checkpoint.

  • DDeleting a local checkpoint keeps the S3 copy

    Incorrect

    If a checkpoint is deleted in the container, it is also deleted in the S3 folder.

  • ECheckpoints cannot use S3 Express One Zone

    Incorrect

    Checkpoints can use S3 Express One Zone for faster access.

Checkpointing syncs a local folder with S3 both ways at defined moments: restore at start, write during training, mirror deletions.

Question 8 · choose 1

A data scientist launches dozens of short, similar SageMaker AI training jobs one after another while experimenting, and each job waits several minutes for instances to be provisioned. How can the ML engineer reduce this wait between consecutive jobs?

  1. AManaged warm pools with a keep-alive period
  2. BUse managed Spot training for the jobs
  3. CUse a larger instance for each job
  4. DUse pipe input mode for the data
Show the answer and why
  • AManaged warm pools with a keep-alive period

    Correct

    Managed warm pools retain provisioned infrastructure after a job completes so matching subsequent jobs reuse it, reducing start-up latency; you set KeepAlivePeriodInSeconds, and warm pools are billable.

  • BUse managed Spot training for the jobs

    Incorrect

    Spot capacity lowers cost but can take longer to start and may be interrupted.

  • CUse a larger instance for each job

    Incorrect

    A larger instance can run faster but still has to be provisioned for each job.

  • DUse pipe input mode for the data

    Incorrect

    Input modes change how data is read, not how long instances take to be provisioned.

For rapid iteration with many similar jobs, warm pools trade some idle cost for much faster starts.

Question 9 · choose 1

A retailer uses one prompt for product descriptions across 12 product categories. Only the category name, tone and a few required attributes change between them. The team wants one reusable, task-specific prompt that the application fills in at run time, managed in Amazon Bedrock. What should the ML engineer use?

  1. ATwelve fine-tuned models, one per category
  2. BA Prompt management prompt with variables
  3. CA guardrail with a denied topic per category
  4. DA knowledge base with one document per category
Show the answer and why
  • ATwelve fine-tuned models, one per category

    Incorrect

    Fine-tuning twelve models adds cost and maintenance for differences a prompt variable can express.

  • BA Prompt management prompt with variables

    Correct

    Prompt management lets you include variables in a prompt and supply their values when you test or invoke the model, so one prompt serves different use cases.

  • CA guardrail with a denied topic per category

    Incorrect

    Guardrails filter content; they do not template prompts.

  • DA knowledge base with one document per category

    Incorrect

    Retrieval adds context from documents; it does not parameterize the prompt itself.

Prompt templates with variables are a lightweight customization technique; reach for fine-tuning only when prompting cannot get the behavior.

Question 10 · choose 1

A support classifier built on a large foundation model in Amazon Bedrock is accurate but too expensive and slow at production volume. A smaller model in the same family is fast and cheap but noticeably less accurate. The team wants close to the large model's quality at the small model's cost. Which customization approach fits?

  1. AContinued use of the large model with Priority tier
  2. BA larger context window on the small model
  3. CDistillation from the large model to the small one
  4. DProvisioned Throughput for the large model
Show the answer and why
  • AContinued use of the large model with Priority tier

    Incorrect

    Priority makes responses faster for a premium, which increases cost.

  • BA larger context window on the small model

    Incorrect

    More context does not transfer the larger model's capability.

  • CDistillation from the large model to the small one

    Correct

    Distillation transfers knowledge from a larger, more capable teacher model to a smaller, faster and cost-efficient student by fine-tuning the student on responses the teacher generates.

  • DProvisioned Throughput for the large model

    Incorrect

    Reserved capacity can add throughput but does not make the large model cheaper per request.

Combining models through distillation is a standard way to cut inference cost while keeping much of a larger model's quality.

Practise domain 2 →Practise all domains →