Skip to content
BytePatterns

MLA-C02 · Domain 1: Data Preparation for ML and AI · 28% of the exam

Task 1.3: Validate data quality and manage bias.

Checking data with DataBrew and AWS Glue Data Quality, labeling, finding and reducing bias, fixing class imbalance, cleaning outliers, gaps and duplicates, and screening training pairs before they reach a model.

Study it

  • Data quality with DataBrew and AWS Glue Data Quality; cleaning outliers, gaps and duplicates

    Lesson coming

  • Labeling, bias in data, class imbalance and validating training pairs

    Lesson coming

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

An AWS Glue ETL job builds the daily training table for a churn model. The team wants the job itself to check the data in transit, for example that customer_id is always present and unique and that tenure is within a valid range, and to stop the pipeline before a bad table reaches training. What should the ML engineer add?

  1. AAn AWS Glue crawler that runs on the output prefix after each job
  2. BAWS Glue Data Quality rules in DQDL, evaluated in the ETL job
  3. CAmazon CloudWatch alarms on the Glue job metrics
  4. DAn Amazon Macie job on the output bucket
Show the answer and why
  • AAn AWS Glue crawler that runs on the output prefix after each job

    Incorrect

    A crawler infers schemas and adds or updates tables in the Data Catalog. It does not test the values in the data against rules.

  • BAWS Glue Data Quality rules in DQDL, evaluated in the ETL job

    Correct

    Glue Data Quality evaluates rules such as IsComplete, IsUnique and range checks written in the Data Quality Definition Language, both on Data Catalog tables and inside Glue ETL jobs, and the job can act on the results.

  • CAmazon CloudWatch alarms on the Glue job metrics

    Incorrect

    Glue job metrics describe the job's execution, such as memory and data moved, not whether each value in the table is valid.

  • DAn Amazon Macie job on the output bucket

    Incorrect

    Macie discovers sensitive data such as PII in S3 objects. It does not check completeness, uniqueness or value ranges.

Glue Data Quality, built on the open-source Deequ framework, brings declarative data checks into the pipeline so bad data can be stopped or quarantined before it is used for training.

Question 2 · choose 2

A binary fraud classifier is trained with the SageMaker AI built-in XGBoost algorithm on transactions where only 1% are fraud. The model reaches 99% accuracy but catches almost no fraud. Which actions should the ML engineer take to improve how the model learns the minority class? (Choose TWO.)

  1. AKeep accuracy as the objective metric for model selection
  2. BUndersample the test set too, so it has the same class ratio as training
  3. COversample the fraud class in the training split, for example with SMOTE
  4. DSet the scale_pos_weight hyperparameter to weight the positive class more
  5. EShuffle the rows before splitting the data
Show the answer and why
  • AKeep accuracy as the objective metric for model selection

    Incorrect

    On a dataset this imbalanced, predicting "not fraud" every time already scores 99%. A metric such as balanced accuracy, F1 or recall shows whether fraud is caught.

  • BUndersample the test set too, so it has the same class ratio as training

    Incorrect

    Changing the test set's class mix makes the evaluation unlike production traffic, so the reported numbers stop predicting real performance.

  • COversample the fraud class in the training split, for example with SMOTE

    Correct

    Balancing operations such as random oversampling or SMOTE add minority-class samples, which helps a binary classifier learn the rare class. Only the training data should be balanced, so evaluation still reflects real traffic.

  • DSet the scale_pos_weight hyperparameter to weight the positive class more

    Correct

    scale_pos_weight controls the balance of positive and negative weights in XGBoost and is meant for unbalanced classes.

  • EShuffle the rows before splitting the data

    Incorrect

    Shuffling changes which rows land in each split, but every split still has about 1% fraud, so the imbalance is untouched.

Fix imbalance where the model learns (resampling the training data or reweighting the classes) and judge the result with metrics that do not reward always predicting the majority class.

Question 3 · choose 1

A retailer trains a model to forecast next week's demand from three years of daily sales. An engineer splits the rows 80/20 at random and gets excellent validation scores, but the model performs poorly after launch. Which split should the team use in SageMaker Canvas data preparation to get a realistic validation score?

  1. AA stratified split on the demand column
  2. BA randomized split with a different random seed
  3. CA split by key on store_id, so each store is in one split
  4. DAn ordered split that holds out the last 20% of days
Show the answer and why
  • AA stratified split on the demand column

    Incorrect

    A stratified split keeps the proportion of each target value the same in every split. Rows are still drawn from the whole time range, so future days still inform training.

  • BA randomized split with a different random seed

    Incorrect

    A different seed gives a different random sample, but every random split still mixes future days into the training set.

  • CA split by key on store_id, so each store is in one split

    Incorrect

    Split by key keeps each store in only one split. Every store's history still spans all dates, so the time leakage remains.

  • DAn ordered split that holds out the last 20% of days

    Correct

    An ordered split keeps the sequence of the observations: the first 80% train and the last 20% test. Validating on later days than training mirrors how the forecast is used.

For time-ordered data a random split leaks the future into training and inflates validation scores. Hold out the most recent period instead.

Question 4 · choose 1

A company is merging customer records from two CRM systems to build a training set. The same person often appears in both with small differences, such as "Jon Smith, 12 Main St." and "John Smith, 12 Main Street", and there is no shared customer ID. Which approach finds these duplicates with the least custom code?

  1. AA dropDuplicates() step on all columns in a Spark job
  2. BA DQDL IsUnique rule on the name column
  3. CAn AWS Glue crawler on both sources
  4. DA FindMatches transform in an AWS Glue job
Show the answer and why
  • AA dropDuplicates() step on all columns in a Spark job

    Incorrect

    Dropping duplicate rows removes rows whose values are identical. Records that differ by a spelling or an abbreviation are kept as two customers.

  • BA DQDL IsUnique rule on the name column

    Incorrect

    IsUnique checks whether values in a column are unique. It reports a problem but neither decides which differently spelled records refer to the same person nor merges them.

  • CAn AWS Glue crawler on both sources

    Incorrect

    A crawler catalogs schemas. It does not compare records with each other.

  • DA FindMatches transform in an AWS Glue job

    Correct

    FindMatches is a machine learning transform that finds duplicate or matching records even when they have no common unique identifier and no fields match exactly. You teach it with labeled examples and call it from a Spark-based Glue job.

Exact deduplication only catches identical rows. Fuzzy record matching across sources without a common key is what FindMatches is for.

Question 5 · choose 1

In a loan dataset, 12% of the values in the numeric "debt_ratio" column are missing, and the column has a long right tail with some extreme values. The team wants to keep every row and let the model learn whether a value was originally missing. Which preparation steps should the ML engineer apply?

  1. AImpute the mean of the column, then drop the original column
  2. BImpute the median and add a missing-value indicator column
  3. CDrop every row that has a missing debt_ratio value
  4. DFill every missing value with zero
Show the answer and why
  • AImpute the mean of the column, then drop the original column

    Incorrect

    The mean of a long-tailed column is pulled toward its extreme values, and dropping the original column throws the feature away.

  • BImpute the median and add a missing-value indicator column

    Correct

    Median imputation fills gaps with a value that extreme values barely move, and the "add indicator for missing" transform creates a Boolean column so the model can still use the fact that a value was missing.

  • CDrop every row that has a missing debt_ratio value

    Incorrect

    Drop missing removes 12% of the rows, which the team wants to avoid, and loses the signal in the missingness itself.

  • DFill every missing value with zero

    Incorrect

    Zero is a real, valid-looking debt ratio, so the model cannot tell filled rows from customers who truly have no debt.

Imputation keeps rows; the median resists extreme values; an indicator column keeps the information that the value was missing. All three are standard cleaning steps in Canvas data preparation.

Question 6 · choose 1

A company is preparing 50,000 prompt-completion pairs, collected from employees, to fine-tune a model in Amazon Bedrock. Before training, every pair must be screened for insults, hate speech and a list of topics the company does not want the model to learn about. Which approach should the ML engineer use?

  1. AScreen every pair with ApplyGuardrail and a configured guardrail
  2. BRun AWS Glue Data Quality IsComplete and IsUnique rules on both columns
  3. CRun an Amazon Macie sensitive data discovery job on the dataset bucket
  4. DRun Amazon Comprehend sentiment analysis and drop every negative pair
Show the answer and why
  • AScreen every pair with ApplyGuardrail and a configured guardrail

    Correct

    ApplyGuardrail assesses any text against a configured guardrail's content filters, denied topics, PII detectors and word lists without invoking a foundation model, so it can screen a dataset before training.

  • BRun AWS Glue Data Quality IsComplete and IsUnique rules on both columns

    Incorrect

    These rules check that values are present and unique. They say nothing about whether the text is insulting or touches a denied topic.

  • CRun an Amazon Macie sensitive data discovery job on the dataset bucket

    Incorrect

    Macie looks for sensitive data such as PII in S3 objects. It does not detect hate speech or company-specific denied topics.

  • DRun Amazon Comprehend sentiment analysis and drop every negative pair

    Incorrect

    Sentiment shows whether text is positive or negative. Polite text can still cover a denied topic, and negative text is not necessarily harmful.

Training data integrity for foundation models includes content safety. Because ApplyGuardrail is decoupled from model invocation, the same guardrail used at runtime can screen fine-tuning data in advance.

Question 7 · choose 1

Before using a new dataset of 3 million rows, a data preparation team wants a quick overview: the distribution of each column, how many values are missing, and which columns appear to contain personal information. It works in AWS Glue DataBrew. What should the team run?

  1. AA DataBrew recipe job with no steps
  2. BA DataBrew profile job on the dataset
  3. CAn Amazon Athena CTAS query on the data
  4. DAn S3 Inventory report on the bucket
Show the answer and why
  • AA DataBrew recipe job with no steps

    Incorrect

    A recipe job applies transformations; without steps it produces a copy, not a profile.

  • BA DataBrew profile job on the dataset

    Correct

    Profile jobs run a series of evaluations on a dataset and output the results to S3, helping you understand the data before preparing it; profile jobs can also be configured to detect PII.

  • CAn Amazon Athena CTAS query on the data

    Incorrect

    CTAS converts query results into new tables; it does not profile column distributions or detect PII.

  • DAn S3 Inventory report on the bucket

    Incorrect

    An inventory lists objects and their properties, not the contents of the data.

Profile before you prepare: statistics, missing values and sensitive fields guide which cleaning steps are needed.

Question 8 · choose 1

An AWS Glue job's data quality score dropped sharply after a supplier changed its export, but the team cannot tell which rows caused it. It wants to find the exact failing records and keep them out of the training table while still loading the good rows. What should the ML engineer use?

  1. AFail the whole job on any rule failure
  2. BAn AWS Glue crawler run after each job
  3. CGlue DQ record-level results to quarantine rows
  4. DAn S3 Lifecycle rule on the output prefix
Show the answer and why
  • AFail the whole job on any rule failure

    Incorrect

    Stopping the job also blocks all the good rows, which the team wants to keep loading.

  • BAn AWS Glue crawler run after each job

    Incorrect

    A crawler updates schemas; it does not identify failing records.

  • CGlue DQ record-level results to quarantine rows

    Correct

    AWS Glue Data Quality helps identify the exact records that caused quality scores to go down, so they can be quarantined and fixed while good data continues through the pipeline.

  • DAn S3 Lifecycle rule on the output prefix

    Incorrect

    Lifecycle rules move or expire objects by age; they do not check data quality.

Quarantining bad records keeps pipelines flowing while problems are investigated, instead of choosing between bad data and no data.

Question 9 · choose 1

The categorical "region" column has missing values. Instead of guessing a region, the team wants every missing value replaced by the literal string "Unknown", so the model can treat it as its own category. Which SageMaker Canvas transform does this?

  1. AImpute missing for the categorical column
  2. BDrop missing on the region column
  3. CRobust Scaler on the region column
  4. DFill missing with the value "Unknown"
Show the answer and why
  • AImpute missing for the categorical column

    Incorrect

    For categorical data, Impute missing uses the most frequent value, which guesses a region instead of marking the value as unknown.

  • BDrop missing on the region column

    Incorrect

    Drop missing removes the rows, which loses data instead of labeling it.

  • CRobust Scaler on the region column

    Incorrect

    Scalers transform numeric data; region is a text category.

  • DFill missing with the value "Unknown"

    Correct

    Impute missing fills categorical gaps with the most frequent value; to use a custom string instead, the documentation points to the Fill missing transform.

"Missing" can itself be informative; filling with an explicit placeholder category keeps that information.

Question 10 · choose 1

Two exports of the same customer table, from January and from February, must be combined into one dataset for training. Many customers appear in both files with identical rows. Which SageMaker Canvas data preparation step combines the files and handles this in one operation?

  1. AConcatenate, removing duplicates afterward
  2. BAn inner join of the files on customer_id
  3. CA randomized split of the two files
  4. DSimilarity encode on the customer_id column
Show the answer and why
  • AConcatenate, removing duplicates afterward

    Correct

    Concatenate appends the rows of one dataset to another, and you can select an option to remove duplicates after concatenation.

  • BAn inner join of the files on customer_id

    Incorrect

    A join combines columns from two tables side by side; it does not stack the monthly rows into one dataset.

  • CA randomized split of the two files

    Incorrect

    Splitting divides data into subsets; it does not combine files.

  • DSimilarity encode on the customer_id column

    Incorrect

    Similarity encoding creates embeddings for categories; it does not merge or deduplicate datasets.

Stack datasets with concatenation, combine columns with joins, and remove exact duplicates as part of the merge when sources overlap.

Question 11 · choose 1

A fine-tuning job in Amazon Bedrock fails immediately with the message "Unable to parse Amazon S3 file: train.jsonl. Data files must conform to JSONL format." The file was produced by a script that writes all records as one JSON array. What should the ML engineer fix?

  1. AIncrease the number of training epochs
  2. BOne JSON object per line, not an array
  3. CCompress the training file with gzip
  4. DGrant the role kms:Encrypt on the bucket key
Show the answer and why
  • AIncrease the number of training epochs

    Incorrect

    The job fails while parsing the file, before training settings matter.

  • BOne JSON object per line, not an array

    Correct

    The error indicates the file is not valid JSONL. Customization datasets are .jsonl files in which each line is a JSON object for one record.

  • CCompress the training file with gzip

    Incorrect

    Compression does not turn a JSON array into JSON Lines.

  • DGrant the role kms:Encrypt on the bucket key

    Incorrect

    A parsing error points to the file format, not to key permissions.

Validate training files before submitting jobs: correct JSONL structure first, then field names and size limits.

Practise domain 1 →Practise all domains →