Skip to content
BytePatterns

MLA-C02 · Domain 1: Data Preparation for ML and AI · 28% of the exam

Task 1.1: Collect and store data.

Getting data out of S3, databases and streams, choosing a storage service and a file format for how the data will be read, merging sources, storing text, images and audio, setting up vector stores for AI applications, and loading features into SageMaker Feature Store.

Study it

  • Where ML data lives: S3, EFS, FSx, databases and choosing storage by cost, performance and compliance

    Partly covered by: S3: Consistency, Classes, Lifecycle, RDS vs DynamoDB

  • File formats and ingestion: Parquet, ORC, JSON, CSV, Kinesis, Firehose and Managed Service for Apache Flink

    Lesson coming

  • Merging sources with AWS Glue and Spark on Amazon EMR

    Lesson coming

  • Vector stores for AI applications: OpenSearch Service, RDS with pgvector and Amazon S3 Vectors

    Partly covered by: Vector Databases, Approximate Neighbours

  • SageMaker Feature Store: online and offline stores, ingestion and feature groups

    Lesson coming

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

A payments company is building a fraud model. At inference time the endpoint must look up each card's latest aggregated features within milliseconds. Data scientists also need every historical value of those features, with the time each value was valid, to build training sets. Which Amazon SageMaker Feature Store configuration meets both needs?

  1. AA feature group with only the offline store, queried with Amazon Athena at inference time
  2. BA feature group with only the online store, exported nightly for training
  3. CTwo feature groups that use the same record identifier, one per store
  4. DA feature group with both the online store and the offline store enabled
Show the answer and why
  • AA feature group with only the offline store, queried with Amazon Athena at inference time

    Incorrect

    The offline store is meant for cases where sub-second reads are not needed, such as exploration, training and batch inference. It is not a millisecond lookup path for a live endpoint.

  • BA feature group with only the online store, exported nightly for training

    Incorrect

    The online store keeps only the record with the latest event time for each record identifier, so the earlier values that training needs are not there to export.

  • CTwo feature groups that use the same record identifier, one per store

    Incorrect

    One feature group can enable both stores, and Feature Store keeps them in sync. Two separate groups would have to be written and kept consistent by the team, which adds work and risks training-serving skew.

  • DA feature group with both the online store and the offline store enabled

    Correct

    The online store serves the latest record for each identifier with low latency through GetRecord, and the offline store in Amazon S3 keeps every historical record for training. When both are enabled they stay in sync.

Online store for the newest value at low latency, offline store for full history: enabling both on one feature group is the standard way to serve and train from the same feature definitions.

Question 2 · choose 1

An Amazon Data Firehose stream delivers JSON clickstream events to Amazon S3. Analysts query the data with Amazon Athena, and a nightly job reads a few of the many columns to build training data. Query costs are growing because every query scans whole JSON objects. What should the ML engineer change to reduce the data scanned with the least custom code?

  1. ATurn on Firehose record format conversion to Apache Parquet
  2. BAdd a Lambda transformation to the stream that rewrites each JSON record as a CSV line
  3. CTurn on dynamic partitioning and keep writing JSON objects
  4. DDeliver the stream to an S3 Express One Zone directory bucket instead
Show the answer and why
  • ATurn on Firehose record format conversion to Apache Parquet

    Correct

    Firehose can convert JSON input to Apache Parquet or Apache ORC before writing to S3, using a deserializer, a schema from the Glue Data Catalog and a serializer. Columnar files let queries read only the columns they need.

  • BAdd a Lambda transformation to the stream that rewrites each JSON record as a CSV line

    Incorrect

    CSV is still a row-oriented format, so a query that needs three columns still reads every row in full. It also adds a function to write and maintain.

  • CTurn on dynamic partitioning and keep writing JSON objects

    Incorrect

    Dynamic partitioning groups objects under prefixes by keys in the data, which helps queries that filter on those keys. Each object is still JSON, so a query reads every field of every row it touches.

  • DDeliver the stream to an S3 Express One Zone directory bucket instead

    Incorrect

    S3 Express One Zone is the lowest-latency storage class, but the objects would still be JSON and Athena would still scan whole records.

Parquet and ORC are columnar formats that take less space and let engines read only the referenced columns. Firehose converts JSON to either one natively when given a schema from the Glue Data Catalog.

Question 3 · choose 2

Sensor gateways write about 6 MB per second to an Amazon Kinesis Data Streams stream in provisioned mode with 4 shards. Every record uses the same partition key, "telemetry". Producers receive ProvisionedThroughputExceededException errors, and the feature pipeline that reads the stream falls behind. Which actions will resolve the write throttling? (Choose TWO.)

  1. AIncrease the data retention period of the stream from 24 hours to 7 days
  2. BUse a high-cardinality partition key, such as the device ID, so records spread across shards
  3. CRegister the feature pipeline as an enhanced fan-out consumer
  4. DRaise the stream capacity to at least 6 shards, or switch the stream to on-demand mode
  5. ESwitch the producers from PutRecords batches to individual PutRecord calls
Show the answer and why
  • AIncrease the data retention period of the stream from 24 hours to 7 days

    Incorrect

    Retention controls how long records stay readable after they are written. It does not change how much data a shard accepts per second.

  • BUse a high-cardinality partition key, such as the device ID, so records spread across shards

    Correct

    Kinesis maps each partition key to one shard. With a single key value every record lands on one shard, which accepts at most 1 MB per second of writes. Many distinct keys spread the puts across all shards.

  • CRegister the feature pipeline as an enhanced fan-out consumer

    Incorrect

    Enhanced fan-out gives a consumer dedicated read throughput of up to 2 MB per second per shard. It helps readers, not producers that are throttled on writes.

  • DRaise the stream capacity to at least 6 shards, or switch the stream to on-demand mode

    Correct

    Each shard accepts up to 1 MB per second or 1,000 records per second of writes, so 4 shards cannot take 6 MB per second even with an even spread. More shards, or on-demand mode that manages shards automatically, adds the capacity.

  • ESwitch the producers from PutRecords batches to individual PutRecord calls

    Incorrect

    Batching with PutRecords, or the Kinesis Producer Library, is the recommended way to raise producer throughput. Single-record calls add overhead and do not lift the per-shard limit.

Two limits are hit at once: one hot shard caused by a constant partition key, and a stream that is too small for the total rate. Spreading keys fixes the first, and adding shards (or on-demand mode) fixes the second.

Question 4 · choose 1

A legal research team has 900 million document embeddings of 1,024 dimensions for a retrieval application. The archive is queried only a few times per hour, subsecond query latency is acceptable, and the main goal is the lowest cost to store and query the vectors without managing any infrastructure. Which vector store should the ML engineer use?

  1. AAn Amazon OpenSearch Service domain sized to keep the HNSW graphs in memory
  2. BAn Amazon Aurora PostgreSQL-Compatible cluster with the pgvector extension
  3. CAmazon S3 Vectors, with the embeddings in a vector index inside a vector bucket
  4. DEmbedding files in S3 Glacier Flexible Retrieval, restored when a query arrives
Show the answer and why
  • AAn Amazon OpenSearch Service domain sized to keep the HNSW graphs in memory

    Incorrect

    Memory-based k-NN search loads the vector graphs into the RAM of the data nodes, so the domain must be sized and paid for around that memory all the time. It fits high-query-rate, low-latency search better than a rarely queried archive.

  • BAn Amazon Aurora PostgreSQL-Compatible cluster with the pgvector extension

    Incorrect

    pgvector adds vector search to PostgreSQL, but the team would run and size a database cluster for the whole index, which does not match the goals of no infrastructure to manage and lowest cost.

  • CAmazon S3 Vectors, with the embeddings in a vector index inside a vector bucket

    Correct

    S3 Vectors is purpose-built, cost-optimized vector storage with its own query API and no infrastructure to provision. It delivers subsecond latency for infrequent queries, and an index can hold billions of vectors of up to 4,096 dimensions.

  • DEmbedding files in S3 Glacier Flexible Retrieval, restored when a query arrives

    Incorrect

    Objects in S3 Glacier Flexible Retrieval are archived and not available for real-time access; a copy must be restored before it can be read, so queries could not be answered in under a second.

S3 Vectors trades the very low latency of an in-memory engine for much lower storage and query cost at large scale. AWS positions OpenSearch for high query rates and low latency, and S3 Vectors for cost-effective storage and infrequent queries.

Question 5 · choose 1

A SageMaker AI training job reads 2 TB of image files from an Amazon S3 prefix. With the default input mode, each run spends about 40 minutes before the training script processes the first batch, and the team had to attach a very large storage volume to every training instance. Which change addresses both problems with the least effort?

  1. AKeep file mode and attach an even larger volume to each instance
  2. BSet the S3 data distribution type to ShardedByS3Key
  3. CPoint the channel at an augmented manifest file in file mode
  4. DSwitch the input channel to fast file mode
Show the answer and why
  • AKeep file mode and attach an even larger volume to each instance

    Incorrect

    In file mode SageMaker AI downloads the whole dataset before training starts, so a larger volume makes room for the data but does not remove the wait.

  • BSet the S3 data distribution type to ShardedByS3Key

    Incorrect

    ShardedByS3Key splits the dataset across the instances of a distributed job. Each instance still downloads its share before starting, and it does not help a single-instance job.

  • CPoint the channel at an augmented manifest file in file mode

    Incorrect

    A manifest only lists which objects to use. In file mode the listed data is still downloaded in full before training begins.

  • DSwitch the input channel to fast file mode

    Correct

    Fast file mode exposes the S3 objects through a file system interface and streams them on demand, so training starts without waiting for a full download and the dataset no longer has to fit on the instance.

File mode downloads first and trains second; fast file mode starts immediately and streams objects as the script reads them. For S3 prefixes that are read mostly sequentially, fast file mode is the simple fix.

Question 6 · choose 1

A contact center stores thousands of recorded customer calls as audio files in Amazon S3. The ML team wants text it can use to train an intent classifier, and the callers' personal information must not appear in that text. Which AWS service should the team use first?

  1. AAmazon Comprehend, with PII detection on the audio files
  2. BAmazon Textract, to extract the spoken words
  3. CAmazon Transcribe, with PII redaction turned on
  4. DAmazon Rekognition, to analyze the recordings
Show the answer and why
  • AAmazon Comprehend, with PII detection on the audio files

    Incorrect

    Comprehend is a natural language processing service that works on text. It cannot take audio files as input, so the calls must be transcribed first.

  • BAmazon Textract, to extract the spoken words

    Incorrect

    Textract extracts text and data from scanned documents and images, not from audio recordings.

  • CAmazon Transcribe, with PII redaction turned on

    Correct

    Transcribe converts speech in audio files to text, and its PII redaction option replaces personal information in the transcript, so the output is ready to use as training text.

  • DAmazon Rekognition, to analyze the recordings

    Incorrect

    Rekognition analyzes images and videos. It does not turn speech into text.

Match the modality first: speech goes through Transcribe, documents through Textract, images and video through Rekognition, and text through Comprehend. Transcribe can redact PII as part of the transcription job.

Question 7 · choose 1

A churn model needs customer attributes from an on-premises PostgreSQL database joined with clickstream files that are already in Amazon S3. The team wants a managed, serverless way to connect to both sources and produce one training table on a schedule. What should the ML engineer use?

  1. AAn Amazon Data Firehose stream per source
  2. BAn AWS Glue ETL job with a JDBC connection
  3. CAn S3 Lifecycle rule on the clickstream bucket
  4. DAn Amazon Kinesis Video Streams pipeline
Show the answer and why
  • AAn Amazon Data Firehose stream per source

    Incorrect

    Firehose delivers streaming records to destinations; it does not join a database table with files already in S3.

  • BAn AWS Glue ETL job with a JDBC connection

    Correct

    AWS Glue is a serverless data integration service that connects to many data sources, on premises and on AWS, through Glue connections, and runs ETL jobs that combine them into data lake tables.

  • CAn S3 Lifecycle rule on the clickstream bucket

    Incorrect

    Lifecycle rules move or expire objects; they do not read databases or join data.

  • DAn Amazon Kinesis Video Streams pipeline

    Incorrect

    Kinesis Video Streams ingests video from devices; it has nothing to do with joining tables.

Merging data from several sources into one training set is classic ETL; Glue connections reach databases on premises or in AWS, and Glue jobs do the joins.

Question 8 · choose 1

A multi-tenant RAG service stores all tenants' vectors in one Amazon S3 Vectors index, and every query must return only the calling tenant's vectors. Each vector also carries a long text snippet that is never used in filters. How should the ML engineer define the metadata when creating the index?

  1. AMark both of the fields as non-filterable
  2. BKeep both of the fields filterable (default)
  3. Ctenant_id filterable, snippet non-filterable
  4. DCreate the index first and mark fields later
Show the answer and why
  • AMark both of the fields as non-filterable

    Incorrect

    Non-filterable metadata cannot be used in query filters, so queries could not be restricted to one tenant.

  • BKeep both of the fields filterable (default)

    Incorrect

    Filterable metadata has stricter size limits, and long snippets that are never filtered on fit better as non-filterable metadata.

  • Ctenant_id filterable, snippet non-filterable

    Correct

    Metadata is filterable by default unless declared non-filterable at index creation. Filterable metadata can be used in query filters but has stricter size limits, while non-filterable metadata cannot be filtered on but can hold larger data.

  • DCreate the index first and mark fields later

    Incorrect

    Non-filterable metadata keys cannot be changed after the index is created; changing them requires a new index.

Plan vector metadata up front: fields used in filters stay filterable, bulky fields used only for display are non-filterable.

Question 9 · choose 1

Regulations require a company to keep the raw data behind every production model for seven years. The raw files are almost never read after the first 90 days, and retrieval within hours is acceptable when an audit asks for them. How should the ML engineer manage the storage cost?

  1. AS3 Express One Zone for all of the raw files
  2. BKeep everything in S3 Standard for seven years
  3. CDelete the raw files once training completes
  4. DA Lifecycle rule that archives files after 90 days
Show the answer and why
  • AS3 Express One Zone for all of the raw files

    Incorrect

    Express One Zone is the lowest-latency class for frequently accessed data, which is the opposite of rarely read archives.

  • BKeep everything in S3 Standard for seven years

    Incorrect

    Standard storage is priced for frequent access, so paying it for seven years of rarely read data wastes money.

  • CDelete the raw files once training completes

    Incorrect

    Deleting them breaks the seven-year retention requirement.

  • DA Lifecycle rule that archives files after 90 days

    Correct

    S3 Lifecycle transition actions move objects to lower-cost storage classes on a schedule, such as archiving them after a set number of days, and expiration actions can delete them after the retention period.

Storage decisions combine access pattern, retrieval time and compliance; lifecycle rules automate moving cold data to archive classes.

Question 10 · choose 1

An insurer wants to turn incoming claim documents, photos and recorded phone calls into structured outputs, such as document summaries and call transcriptions, that can be stored and used for model training, without building a separate pipeline for each data type. Which AWS capability fits?

  1. AAmazon Bedrock Data Automation
  2. BAmazon Kinesis Data Streams
  3. CAWS DataSync transfer tasks
  4. DAWS Glue crawlers and classifiers
Show the answer and why
  • AAmazon Bedrock Data Automation

    Correct

    Bedrock Data Automation processes documents, images, video and audio, returning standard outputs such as audio transcriptions and document summaries, or custom outputs that you define.

  • BAmazon Kinesis Data Streams

    Incorrect

    Kinesis Data Streams carries records; it does not extract summaries or transcriptions from files.

  • CAWS DataSync transfer tasks

    Incorrect

    DataSync moves files between storage systems; it does not analyze their content.

  • DAWS Glue crawlers and classifiers

    Incorrect

    Crawlers catalog schemas of structured data; they do not transcribe audio or summarize documents.

For mixed unstructured inputs, a multimodal extraction service turns each data type into structured, storable output in one place.

Question 11 · choose 2

Events from a multi-tenant application flow through Amazon Data Firehose into Amazon S3. Training jobs and Athena queries almost always read one customer's data for a date range, but today they scan everything. Which Firehose features should the ML engineer use so the data lands in a layout that suits these reads? (Choose TWO.)

  1. ADynamic partitioning on customer_id and date
  2. BA larger buffer interval only
  3. CDelivery to S3 Glacier Flexible Retrieval
  4. DRecord format conversion to Apache Parquet
  5. EA second Firehose stream with the same settings
Show the answer and why
  • ADynamic partitioning on customer_id and date

    Correct

    Dynamic partitioning uses keys within the data, such as customer_id, to deliver records grouped into matching S3 prefixes, which minimizes the data scanned by queries.

  • BA larger buffer interval only

    Incorrect

    Buffering changes how often files are written, not how they are organized for per-customer reads.

  • CDelivery to S3 Glacier Flexible Retrieval

    Incorrect

    Archived objects are not available for real-time access, which training jobs and queries need.

  • DRecord format conversion to Apache Parquet

    Correct

    Converting JSON to columnar Parquet lets queries and jobs read only the columns they need, which also reduces scanning.

  • EA second Firehose stream with the same settings

    Incorrect

    Duplicating the stream doubles the data without changing its layout.

Design the landing layout for the reads: partition by the common filters and store columns, not rows.

Practise domain 1 →Practise all domains →