MLA-C02 · Domain 1: Data Preparation for ML and AI · 28% of the exam
Task 1.2: Perform data transformation, feature engineering, and pre-processing.
Transforming batch and streaming data with Glue, DataBrew, EMR and Lambda, the feature engineering toolbox (scaling, encoding, binning, log transforms), embeddings and text pre-processing, chunking documents for RAG, masking sensitive data, and shaping datasets for fine-tuning and distillation.
Study it
Merging sources with AWS Glue and Spark on Amazon EMR
Lesson coming
SageMaker Feature Store: online and offline stores, ingestion and feature groups
Lesson coming
Transforming data with Glue, DataBrew, SageMaker Canvas data preparation and streaming Lambda functions
Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.
Question 1 · choose 1
A k-nearest neighbors model uses a numeric "monthly spend" feature. A few enterprise customers spend hundreds of times more than everyone else, and the team must keep those rows. When the feature is scaled, the extreme values should not squash all ordinary customers into a tiny part of the range. Which transform in SageMaker Canvas data preparation fits best?
ARobust Scaler
BMin Max Scaler
COrdinal encode
DDrop missing
Show the answer and why
ARobust Scaler
Correct
The Robust Scaler scales the column with statistics that are robust to outliers, so a handful of extreme values do not dominate how the ordinary values are spread.
BMin Max Scaler
Incorrect
Min-max scaling maps the column to a fixed range from its minimum and maximum. A few huge values set the maximum, so nearly every other customer ends up packed near the bottom of the range.
COrdinal encode
Incorrect
Ordinal encoding turns categories into integers. Monthly spend is already numeric, and the problem is its scale, not its type.
DDrop missing
Incorrect
Drop missing removes rows that have missing values in the input column. It neither scales the feature nor addresses the outliers.
Distance-based models such as k-NN need features on comparable scales. When a feature has extreme values that must stay in the data, scale it with statistics such as the median and quantiles that are not pulled by those values.
A dataset for a linear model has a "country" column with 40 values such as Brazil, Japan and Kenya. The values have no natural order. How should the ML engineer encode this column?
AOrdinal encode it, so each country becomes an integer from 0 to 39
BApply the Standard Scaler to the column
COne-hot encode it, one indicator column per country
DApply the Quantile Numeric Outliers transform to the column
Show the answer and why
AOrdinal encode it, so each country becomes an integer from 0 to 39
Incorrect
Ordinal encoding suits categories with a natural order, such as education levels. For a linear model, integers would imply that country 39 is "larger" than country 1.
BApply the Standard Scaler to the column
Incorrect
The Standard Scaler centers and scales numeric columns. A string category has to be encoded before any scaling makes sense.
COne-hot encode it, one indicator column per country
Correct
Country is a nominal category with no inherent order. One-hot encoding gives every category its own indicator, so the model does not read any order or distance into the codes.
DApply the Quantile Numeric Outliers transform to the column
Incorrect
This transform detects and fixes outliers in numeric features. The country column holds text categories, so there are no numeric outliers to fix and it still needs encoding.
Nominal categories (no order) are usually one-hot encoded; ordinal categories (an inherent order) can use ordered integer codes. The distinction matters most for linear and distance-based models.
A team is building an Amazon Bedrock knowledge base over long equipment manuals, using an Amazon OpenSearch Serverless vector store. Retrieval should match small, precise passages, but the model should then receive the broader section around each match so its answers have full context. Which chunking strategy should the team choose for the data source?
AHierarchical chunking with a parent and a child chunk size
BDefault chunking of about 300 tokens per chunk
CNo chunking, so each manual is one chunk
DSemantic chunking with a higher breakpoint percentile threshold
Show the answer and why
AHierarchical chunking with a parent and a child chunk size
Correct
Hierarchical chunking embeds small child chunks for precise matching and, at retrieval time, replaces them with their larger parent chunks so the model gets more complete context.
BDefault chunking of about 300 tokens per chunk
Incorrect
Default chunking produces chunks of roughly 300 tokens that honor sentence boundaries. What is retrieved is what is returned, so there is no wider section around the match.
CNo chunking, so each manual is one chunk
Incorrect
With no chunking each document becomes a single chunk. A whole manual as one embedding cannot match small passages precisely.
DSemantic chunking with a higher breakpoint percentile threshold
Incorrect
Semantic chunking splits text where meaning changes, and a higher threshold gives fewer, larger chunks. It does not retrieve on small chunks and then return a larger parent.
Small chunks embed precisely; large chunks give context. Hierarchical chunking gets both by searching on child chunks and returning their parents. The documentation does not recommend it with an S3 vector bucket as the store, which is not the case here.
Before fine-tuning a model on 2 million support tickets stored as text files in Amazon S3, a company must remove personal data. Each name, address and phone number must be replaced with a label of its type, such as [NAME], so the model still learns where such details appear in a ticket. What should the ML engineer use?
AAn Amazon Macie sensitive data discovery job on the bucket
BAn Amazon Comprehend sentiment analysis job on the tickets
CA Comprehend PII redaction job that replaces each entity with its type
DA Comprehend PII redaction job that masks each character with an asterisk
Show the answer and why
AAn Amazon Macie sensitive data discovery job on the bucket
Incorrect
Macie discovers sensitive data in S3 objects and reports it as findings. It does not write a redacted copy of the text.
BAn Amazon Comprehend sentiment analysis job on the tickets
Incorrect
Sentiment analysis returns whether text is positive, negative, neutral or mixed. It does not find or replace personal data.
CA Comprehend PII redaction job that replaces each entity with its type
Correct
Comprehend redaction jobs read text from S3 and write a redacted copy. With MaskMode set to REPLACE_WITH_PII_ENTITY_TYPE, an entity such as "Jane Doe" becomes "[NAME]".
DA Comprehend PII redaction job that masks each character with an asterisk
Incorrect
MASK mode replaces every character of an entity with the mask character, such as "**** ***". The text is protected, but the type label the team wants to keep is lost.
Comprehend can redact PII in batch from S3 in two ways: mask the characters or replace each entity with its type. Keeping the type preserves useful structure for training while removing the personal value.
A company wants to use Amazon Bedrock Model Distillation to create a smaller, cheaper model that answers questions about its insurance policies as well as a large teacher model does. Which inputs can the ML engineer use to prepare the training data for the distillation job? (Choose TWO.)
AA JSONL file of representative prompts, from which the teacher model generates the responses
BModel invocation logs from production traffic to the teacher model, stored in Amazon S3
CA CSV file with one numeric feature per column and a label column
DThousands of human-written answers, because distillation cannot use model-generated responses
EThe teacher model weights, exported from Amazon Bedrock to Amazon S3
Show the answer and why
AA JSONL file of representative prompts, from which the teacher model generates the responses
Correct
Distillation input data is a collection of prompts. Bedrock sends them to the teacher model and uses the generated responses to fine-tune the student.
BModel invocation logs from production traffic to the teacher model, stored in Amazon S3
Correct
If invocation logging is enabled, the job can use the logged prompts, or the logged prompt-response pairs when the teacher in the job matches the logged model.
CA CSV file with one numeric feature per column and a label column
Incorrect
Distillation datasets are JSONL records in a conversation format for text-to-text models, not tabular feature files.
DThousands of human-written answers, because distillation cannot use model-generated responses
Incorrect
The opposite is true: distillation is built on responses generated by the teacher model, so human-written answers are not required.
EThe teacher model weights, exported from Amazon Bedrock to Amazon S3
Incorrect
Distillation needs prompts (and optionally logged responses), not the teacher's weights. The teacher is invoked by Bedrock during the job.
In distillation the teacher writes the training answers. You supply prompts, either directly as JSONL or from invocation logs, and Bedrock creates the synthetic data and fine-tunes the student on it.
A data engineering team prepares customer data in AWS Glue DataBrew for several ML teams. Customer email addresses must not be readable in the output, but the ML teams still need to join three datasets on the email column, and an authorized auditor must be able to recover the original value. Which DataBrew technique should the team apply to the email column?
AProbabilistic encryption
BShuffling the values within the column
CDeterministic encryption
DNulling out the column
Show the answer and why
AProbabilistic encryption
Incorrect
Probabilistic encryption produces different ciphertext each time it is applied, so the same email would not match across the three datasets.
BShuffling the values within the column
Incorrect
Shuffling moves values between rows, which breaks the link between each email and its own customer record, so joins return wrong matches.
CDeterministic encryption
Correct
Deterministic encryption always produces the same ciphertext for the same value, so the encrypted column still joins across datasets, and holders of the key can decrypt it.
DNulling out the column
Incorrect
Replacing the values with nulls hides them but leaves nothing to join on and nothing for an auditor to recover.
Pick the masking method from the requirements: joinable and reversible points to deterministic encryption; irreversible but joinable points to hashing; not needed at all points to nulling out or deleting the column.
A knowledge base uses default chunking, which suits the team's documents. The team now wants each chunk tagged with the section heading it came from, so queries can be filtered by section. What should the ML engineer configure?
ANo chunking with the whole document as one chunk
BDefault chunking plus a metadata Lambda function
CSemantic chunking with a larger buffer size
DA reranker model applied at query time
Show the answer and why
ANo chunking with the whole document as one chunk
Incorrect
No chunking makes each document one chunk, which loses the section-level granularity and cannot tag sections separately.
BDefault chunking plus a metadata Lambda function
Correct
You can keep a built-in chunking strategy and provide a custom transformation Lambda function that receives the pre-chunked files and adds chunk-level metadata before the knowledge base continues processing.
CSemantic chunking with a larger buffer size
Incorrect
Buffer size changes how semantic boundaries are found; it does not add section metadata to chunks.
DA reranker model applied at query time
Incorrect
Reranking reorders results; it does not attach metadata for filtering.
Custom transformation Lambda functions can replace chunking entirely or just enrich built-in chunks with metadata.
An analyst built a set of cleaning and transformation steps for January's sales file in an AWS Glue DataBrew project. The same steps must now run on every new monthly file without being rebuilt by hand. What should the analyst do?
ARecreate the steps manually each month
BRun a profile job on each new monthly file
CSave the steps as a recipe for a recipe job
DUse a Glue crawler to transform the file
Show the answer and why
ARecreate the steps manually each month
Incorrect
Rebuilding the steps by hand is slow and error-prone when a saved recipe can be reused.
BRun a profile job on each new monthly file
Incorrect
Profile jobs evaluate and summarize data; they do not apply transformation steps.
CSave the steps as a recipe for a recipe job
Correct
DataBrew saves transformations as steps in a recipe, which you can update or reuse later with other datasets and deploy on a continuing basis.
DUse a Glue crawler to transform the file
Incorrect
Crawlers catalog schemas; they do not transform data.
Make data preparation repeatable: capture steps once as a recipe and apply it to each new batch.
A team consumes events continuously from an Apache Kafka cluster on Amazon MSK and must clean and transform them with Spark code into an S3 data lake for later training. The job should run continuously, and the team prefers a serverless Spark environment. What should the ML engineer use?
AAn AWS Glue crawler on the topic
BAn Athena CTAS query on the topic
CA DataBrew profile job on the stream
DAn AWS Glue streaming ETL job
Show the answer and why
AAn AWS Glue crawler on the topic
Incorrect
Crawlers catalog stored data; they do not consume or transform streams.
BAn Athena CTAS query on the topic
Incorrect
Athena queries data at rest in S3; it does not read Kafka topics continuously.
CA DataBrew profile job on the stream
Incorrect
Profile jobs summarize datasets; they do not run continuous stream transformations.
DAn AWS Glue streaming ETL job
Correct
Glue streaming ETL jobs run continuously, consume sources such as Kinesis Data Streams, Apache Kafka and Amazon MSK, cleanse and transform the data, and load it into S3 data lakes.
For continuous Spark-based stream transformation without managing clusters, Glue streaming ETL is a managed option alongside Flink-based processing.
A dataset has an "education" column with values such as High school, Bachelors, Masters and Doctorate. For a linear model, the team wants these mapped to the ordered numbers 1 to 4, with any unexpected value replaced by 0, in an AWS Glue DataBrew recipe. Which recipe step fits?
ACATEGORICAL_MAPPING with an other value
BONE_HOT_ENCODING of the education column
CBUCKETIZATION of the education column
DTOKENIZATION of the education column
Show the answer and why
ACATEGORICAL_MAPPING with an other value
Correct
CATEGORICAL_MAPPING maps categorical values to numeric or other values through a category map, and its "other" parameter replaces all non-mapped values.
BONE_HOT_ENCODING of the education column
Incorrect
One-hot encoding creates a column per category and discards the order the team wants to keep.
CBUCKETIZATION of the education column
Incorrect
Bucketization groups numeric values into ranges; the column holds text categories.
DTOKENIZATION of the education column
Incorrect
Tokenization splits text into words; it does not assign ordered numbers.
Ordinal categories can be mapped to ordered numbers explicitly; nominal categories are usually one-hot encoded instead.
Chat transcripts flow into a pipeline that will create training data for an intent classifier. Personal details in each message must be detected and masked as messages arrive, one message at a time, and a nightly batch of older transcripts in S3 must also be redacted. Which Amazon Comprehend options fit these two needs? (Choose TWO.)
AComprehend sentiment analysis on each message
BReal-time PII detection on each message, then masking
CComprehend topic modeling on the transcripts
DComprehend language detection on each message
EAn asynchronous PII redaction job for the S3 batch
Show the answer and why
AComprehend sentiment analysis on each message
Incorrect
Sentiment classifies opinion; it does not find personal data.
BReal-time PII detection on each message, then masking
Correct
Comprehend can detect PII entities in a single document in real time, returning their types and offsets so the application can mask them as messages arrive.
CComprehend topic modeling on the transcripts
Incorrect
Topic modeling groups documents by topics, does not redact PII, and is not available to new customers.
DComprehend language detection on each message
Incorrect
Language detection identifies the language; it does not find or mask personal data.
EAn asynchronous PII redaction job for the S3 batch
Correct
Comprehend asynchronous jobs read text from S3 and write a redacted copy, which fits the nightly batch.
Use synchronous PII detection for streaming, per-message masking, and asynchronous redaction jobs for large batches in S3.