Skip to content
BytePatterns

MLA-C02 · Domain 1: Data Preparation for ML and AI · 28% of the exam

Task 1.2: Perform data transformation, feature engineering, and pre-processing.

Transforming batch and streaming data with Glue, DataBrew, EMR and Lambda, the feature engineering toolbox (scaling, encoding, binning, log transforms), embeddings and text pre-processing, chunking documents for RAG, masking sensitive data, and shaping datasets for fine-tuning and distillation.

Study it

  • Merging sources with AWS Glue and Spark on Amazon EMR

    Lesson coming

  • SageMaker Feature Store: online and offline stores, ingestion and feature groups

    Lesson coming

  • Transforming data with Glue, DataBrew, SageMaker Canvas data preparation and streaming Lambda functions

    Partly covered by: Lambda & Event-Driven Design

  • Feature engineering: scaling, standardization, encoding, binning, log transforms and feature splitting

    Lesson coming

  • Text for models: tokenization, embeddings and domain-specific augmentation

    Partly covered by: Tokenization, BPE vs WordPiece, Embeddings, Cosine Similarity

  • Preparing documents for RAG: chunking strategies and metadata

    Partly covered by: Chunking and Reranking, Retrieval-Augmented Generation

  • Masking and redacting sensitive data; preparing data for fine-tuning, continued pre-training and distillation

    Partly covered by: Fine-Tuning vs Prompting

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

A k-nearest neighbors model uses a numeric "monthly spend" feature. A few enterprise customers spend hundreds of times more than everyone else, and the team must keep those rows. When the feature is scaled, the extreme values should not squash all ordinary customers into a tiny part of the range. Which transform in SageMaker Canvas data preparation fits best?

  1. ARobust Scaler
  2. BMin Max Scaler
  3. COrdinal encode
  4. DDrop missing
Show the answer and why
  • ARobust Scaler

    Correct

    The Robust Scaler scales the column with statistics that are robust to outliers, so a handful of extreme values do not dominate how the ordinary values are spread.

  • BMin Max Scaler

    Incorrect

    Min-max scaling maps the column to a fixed range from its minimum and maximum. A few huge values set the maximum, so nearly every other customer ends up packed near the bottom of the range.

  • COrdinal encode

    Incorrect

    Ordinal encoding turns categories into integers. Monthly spend is already numeric, and the problem is its scale, not its type.

  • DDrop missing

    Incorrect

    Drop missing removes rows that have missing values in the input column. It neither scales the feature nor addresses the outliers.

Distance-based models such as k-NN need features on comparable scales. When a feature has extreme values that must stay in the data, scale it with statistics such as the median and quantiles that are not pulled by those values.

Question 2 · choose 1

A dataset for a linear model has a "country" column with 40 values such as Brazil, Japan and Kenya. The values have no natural order. How should the ML engineer encode this column?

  1. AOrdinal encode it, so each country becomes an integer from 0 to 39
  2. BApply the Standard Scaler to the column
  3. COne-hot encode it, one indicator column per country
  4. DApply the Quantile Numeric Outliers transform to the column
Show the answer and why
  • AOrdinal encode it, so each country becomes an integer from 0 to 39

    Incorrect

    Ordinal encoding suits categories with a natural order, such as education levels. For a linear model, integers would imply that country 39 is "larger" than country 1.

  • BApply the Standard Scaler to the column

    Incorrect

    The Standard Scaler centers and scales numeric columns. A string category has to be encoded before any scaling makes sense.

  • COne-hot encode it, one indicator column per country

    Correct

    Country is a nominal category with no inherent order. One-hot encoding gives every category its own indicator, so the model does not read any order or distance into the codes.

  • DApply the Quantile Numeric Outliers transform to the column

    Incorrect

    This transform detects and fixes outliers in numeric features. The country column holds text categories, so there are no numeric outliers to fix and it still needs encoding.

Nominal categories (no order) are usually one-hot encoded; ordinal categories (an inherent order) can use ordered integer codes. The distinction matters most for linear and distance-based models.

Question 3 · choose 1

A team is building an Amazon Bedrock knowledge base over long equipment manuals, using an Amazon OpenSearch Serverless vector store. Retrieval should match small, precise passages, but the model should then receive the broader section around each match so its answers have full context. Which chunking strategy should the team choose for the data source?

  1. AHierarchical chunking with a parent and a child chunk size
  2. BDefault chunking of about 300 tokens per chunk
  3. CNo chunking, so each manual is one chunk
  4. DSemantic chunking with a higher breakpoint percentile threshold
Show the answer and why
  • AHierarchical chunking with a parent and a child chunk size

    Correct

    Hierarchical chunking embeds small child chunks for precise matching and, at retrieval time, replaces them with their larger parent chunks so the model gets more complete context.

  • BDefault chunking of about 300 tokens per chunk

    Incorrect

    Default chunking produces chunks of roughly 300 tokens that honor sentence boundaries. What is retrieved is what is returned, so there is no wider section around the match.

  • CNo chunking, so each manual is one chunk

    Incorrect

    With no chunking each document becomes a single chunk. A whole manual as one embedding cannot match small passages precisely.

  • DSemantic chunking with a higher breakpoint percentile threshold

    Incorrect

    Semantic chunking splits text where meaning changes, and a higher threshold gives fewer, larger chunks. It does not retrieve on small chunks and then return a larger parent.

Small chunks embed precisely; large chunks give context. Hierarchical chunking gets both by searching on child chunks and returning their parents. The documentation does not recommend it with an S3 vector bucket as the store, which is not the case here.

Question 4 · choose 1

Before fine-tuning a model on 2 million support tickets stored as text files in Amazon S3, a company must remove personal data. Each name, address and phone number must be replaced with a label of its type, such as [NAME], so the model still learns where such details appear in a ticket. What should the ML engineer use?

  1. AAn Amazon Macie sensitive data discovery job on the bucket
  2. BAn Amazon Comprehend sentiment analysis job on the tickets
  3. CA Comprehend PII redaction job that replaces each entity with its type
  4. DA Comprehend PII redaction job that masks each character with an asterisk
Show the answer and why
  • AAn Amazon Macie sensitive data discovery job on the bucket

    Incorrect

    Macie discovers sensitive data in S3 objects and reports it as findings. It does not write a redacted copy of the text.

  • BAn Amazon Comprehend sentiment analysis job on the tickets

    Incorrect

    Sentiment analysis returns whether text is positive, negative, neutral or mixed. It does not find or replace personal data.

  • CA Comprehend PII redaction job that replaces each entity with its type

    Correct

    Comprehend redaction jobs read text from S3 and write a redacted copy. With MaskMode set to REPLACE_WITH_PII_ENTITY_TYPE, an entity such as "Jane Doe" becomes "[NAME]".

  • DA Comprehend PII redaction job that masks each character with an asterisk

    Incorrect

    MASK mode replaces every character of an entity with the mask character, such as "**** ***". The text is protected, but the type label the team wants to keep is lost.

Comprehend can redact PII in batch from S3 in two ways: mask the characters or replace each entity with its type. Keeping the type preserves useful structure for training while removing the personal value.

Question 5 · choose 2

A company wants to use Amazon Bedrock Model Distillation to create a smaller, cheaper model that answers questions about its insurance policies as well as a large teacher model does. Which inputs can the ML engineer use to prepare the training data for the distillation job? (Choose TWO.)

  1. AA JSONL file of representative prompts, from which the teacher model generates the responses
  2. BModel invocation logs from production traffic to the teacher model, stored in Amazon S3
  3. CA CSV file with one numeric feature per column and a label column
  4. DThousands of human-written answers, because distillation cannot use model-generated responses
  5. EThe teacher model weights, exported from Amazon Bedrock to Amazon S3
Show the answer and why
  • AA JSONL file of representative prompts, from which the teacher model generates the responses

    Correct

    Distillation input data is a collection of prompts. Bedrock sends them to the teacher model and uses the generated responses to fine-tune the student.

  • BModel invocation logs from production traffic to the teacher model, stored in Amazon S3

    Correct

    If invocation logging is enabled, the job can use the logged prompts, or the logged prompt-response pairs when the teacher in the job matches the logged model.

  • CA CSV file with one numeric feature per column and a label column

    Incorrect

    Distillation datasets are JSONL records in a conversation format for text-to-text models, not tabular feature files.

  • DThousands of human-written answers, because distillation cannot use model-generated responses

    Incorrect

    The opposite is true: distillation is built on responses generated by the teacher model, so human-written answers are not required.

  • EThe teacher model weights, exported from Amazon Bedrock to Amazon S3

    Incorrect

    Distillation needs prompts (and optionally logged responses), not the teacher's weights. The teacher is invoked by Bedrock during the job.

In distillation the teacher writes the training answers. You supply prompts, either directly as JSONL or from invocation logs, and Bedrock creates the synthetic data and fine-tunes the student on it.

Question 6 · choose 1

A data engineering team prepares customer data in AWS Glue DataBrew for several ML teams. Customer email addresses must not be readable in the output, but the ML teams still need to join three datasets on the email column, and an authorized auditor must be able to recover the original value. Which DataBrew technique should the team apply to the email column?

  1. AProbabilistic encryption
  2. BShuffling the values within the column
  3. CDeterministic encryption
  4. DNulling out the column
Show the answer and why
  • AProbabilistic encryption

    Incorrect

    Probabilistic encryption produces different ciphertext each time it is applied, so the same email would not match across the three datasets.

  • BShuffling the values within the column

    Incorrect

    Shuffling moves values between rows, which breaks the link between each email and its own customer record, so joins return wrong matches.

  • CDeterministic encryption

    Correct

    Deterministic encryption always produces the same ciphertext for the same value, so the encrypted column still joins across datasets, and holders of the key can decrypt it.

  • DNulling out the column

    Incorrect

    Replacing the values with nulls hides them but leaves nothing to join on and nothing for an auditor to recover.

Pick the masking method from the requirements: joinable and reversible points to deterministic encryption; irreversible but joinable points to hashing; not needed at all points to nulling out or deleting the column.

Question 7 · choose 1

A knowledge base uses default chunking, which suits the team's documents. The team now wants each chunk tagged with the section heading it came from, so queries can be filtered by section. What should the ML engineer configure?

  1. ANo chunking with the whole document as one chunk
  2. BDefault chunking plus a metadata Lambda function
  3. CSemantic chunking with a larger buffer size
  4. DA reranker model applied at query time
Show the answer and why
  • ANo chunking with the whole document as one chunk

    Incorrect

    No chunking makes each document one chunk, which loses the section-level granularity and cannot tag sections separately.

  • BDefault chunking plus a metadata Lambda function

    Correct

    You can keep a built-in chunking strategy and provide a custom transformation Lambda function that receives the pre-chunked files and adds chunk-level metadata before the knowledge base continues processing.

  • CSemantic chunking with a larger buffer size

    Incorrect

    Buffer size changes how semantic boundaries are found; it does not add section metadata to chunks.

  • DA reranker model applied at query time

    Incorrect

    Reranking reorders results; it does not attach metadata for filtering.

Custom transformation Lambda functions can replace chunking entirely or just enrich built-in chunks with metadata.

Question 8 · choose 1

An analyst built a set of cleaning and transformation steps for January's sales file in an AWS Glue DataBrew project. The same steps must now run on every new monthly file without being rebuilt by hand. What should the analyst do?

  1. ARecreate the steps manually each month
  2. BRun a profile job on each new monthly file
  3. CSave the steps as a recipe for a recipe job
  4. DUse a Glue crawler to transform the file
Show the answer and why
  • ARecreate the steps manually each month

    Incorrect

    Rebuilding the steps by hand is slow and error-prone when a saved recipe can be reused.

  • BRun a profile job on each new monthly file

    Incorrect

    Profile jobs evaluate and summarize data; they do not apply transformation steps.

  • CSave the steps as a recipe for a recipe job

    Correct

    DataBrew saves transformations as steps in a recipe, which you can update or reuse later with other datasets and deploy on a continuing basis.

  • DUse a Glue crawler to transform the file

    Incorrect

    Crawlers catalog schemas; they do not transform data.

Make data preparation repeatable: capture steps once as a recipe and apply it to each new batch.

Question 9 · choose 1

A team consumes events continuously from an Apache Kafka cluster on Amazon MSK and must clean and transform them with Spark code into an S3 data lake for later training. The job should run continuously, and the team prefers a serverless Spark environment. What should the ML engineer use?

  1. AAn AWS Glue crawler on the topic
  2. BAn Athena CTAS query on the topic
  3. CA DataBrew profile job on the stream
  4. DAn AWS Glue streaming ETL job
Show the answer and why
  • AAn AWS Glue crawler on the topic

    Incorrect

    Crawlers catalog stored data; they do not consume or transform streams.

  • BAn Athena CTAS query on the topic

    Incorrect

    Athena queries data at rest in S3; it does not read Kafka topics continuously.

  • CA DataBrew profile job on the stream

    Incorrect

    Profile jobs summarize datasets; they do not run continuous stream transformations.

  • DAn AWS Glue streaming ETL job

    Correct

    Glue streaming ETL jobs run continuously, consume sources such as Kinesis Data Streams, Apache Kafka and Amazon MSK, cleanse and transform the data, and load it into S3 data lakes.

For continuous Spark-based stream transformation without managing clusters, Glue streaming ETL is a managed option alongside Flink-based processing.

Question 10 · choose 1

A dataset has an "education" column with values such as High school, Bachelors, Masters and Doctorate. For a linear model, the team wants these mapped to the ordered numbers 1 to 4, with any unexpected value replaced by 0, in an AWS Glue DataBrew recipe. Which recipe step fits?

  1. ACATEGORICAL_MAPPING with an other value
  2. BONE_HOT_ENCODING of the education column
  3. CBUCKETIZATION of the education column
  4. DTOKENIZATION of the education column
Show the answer and why
  • ACATEGORICAL_MAPPING with an other value

    Correct

    CATEGORICAL_MAPPING maps categorical values to numeric or other values through a category map, and its "other" parameter replaces all non-mapped values.

  • BONE_HOT_ENCODING of the education column

    Incorrect

    One-hot encoding creates a column per category and discards the order the team wants to keep.

  • CBUCKETIZATION of the education column

    Incorrect

    Bucketization groups numeric values into ranges; the column holds text categories.

  • DTOKENIZATION of the education column

    Incorrect

    Tokenization splits text into words; it does not assign ordered numbers.

Ordinal categories can be mapped to ordered numbers explicitly; nominal categories are usually one-hot encoded instead.

Question 11 · choose 2

Chat transcripts flow into a pipeline that will create training data for an intent classifier. Personal details in each message must be detected and masked as messages arrive, one message at a time, and a nightly batch of older transcripts in S3 must also be redacted. Which Amazon Comprehend options fit these two needs? (Choose TWO.)

  1. AComprehend sentiment analysis on each message
  2. BReal-time PII detection on each message, then masking
  3. CComprehend topic modeling on the transcripts
  4. DComprehend language detection on each message
  5. EAn asynchronous PII redaction job for the S3 batch
Show the answer and why
  • AComprehend sentiment analysis on each message

    Incorrect

    Sentiment classifies opinion; it does not find personal data.

  • BReal-time PII detection on each message, then masking

    Correct

    Comprehend can detect PII entities in a single document in real time, returning their types and offsets so the application can mask them as messages arrive.

  • CComprehend topic modeling on the transcripts

    Incorrect

    Topic modeling groups documents by topics, does not redact PII, and is not available to new customers.

  • DComprehend language detection on each message

    Incorrect

    Language detection identifies the language; it does not find or mask personal data.

  • EAn asynchronous PII redaction job for the S3 batch

    Correct

    Comprehend asynchronous jobs read text from S3 and write a redacted copy, which fits the nightly batch.

Use synchronous PII detection for streaming, per-message masking, and asynchronous redaction jobs for large batches in S3.

Practise domain 1 →Practise all domains →