DEA-C01 · Domain 1: Data Ingestion and Transformation · 34% of the exam
Task 1.2: Transform and process data
Turning raw data into useful data: picking Glue, EMR, Lambda or Redshift for a transformation, JDBC and ODBC connections, joining sources, converting CSV to Parquet, keeping processing cost down, fixing failed or slow jobs, data APIs for other systems, and LLMs that read or enrich data.
Study it
Batch ingestion: S3, AWS DMS, AppFlow and JDBC sources
Lesson coming
Choosing the engine: AWS Glue, Amazon EMR, Lambda or Redshift
Lesson coming
File formats: CSV to Parquet, compression and partitioned output
Lesson coming
Fixing slow and failing Spark jobs on Glue and EMR
Lesson coming
Data APIs with API Gateway, Lambda and the Redshift Data API
Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.
Question 1 · choose 2
Analysts use Amazon Athena to query 3 years of web logs that are stored in Amazon S3 as gzip-compressed CSV files. Almost every query filters on one day or one week and reads 6 of the 80 columns. Query cost and run time keep growing. Which actions will reduce the amount of data that Athena scans? (Choose TWO.)
AConvert the data to Apache Parquet with an AWS Glue ETL job
BMove the log objects to the S3 Intelligent-Tiering storage class
CSet a per-query data usage control on the analysts' workgroup
DPartition the data in Amazon S3 by event date and add the partitions to the table
EStore the files uncompressed so that Athena can read them directly
Show the answer and why
AConvert the data to Apache Parquet with an AWS Glue ETL job
Correct
Parquet stores each column separately, so Athena reads only the columns a query uses and can skip data through predicate pushdown.
BMove the log objects to the S3 Intelligent-Tiering storage class
Incorrect
Intelligent-Tiering moves objects between access tiers to lower storage cost. Athena still reads the same bytes for each query.
CSet a per-query data usage control on the analysts' workgroup
Incorrect
A per-query control cancels a query that scans more than the limit. It caps spending; it does not make queries read less data.
DPartition the data in Amazon S3 by event date and add the partitions to the table
Correct
With date partitions, a query that filters on a day or a week reads only those partitions instead of all 3 years.
EStore the files uncompressed so that Athena can read them directly
Incorrect
Athena charges for bytes scanned before decompression, so removing compression makes every query scan more data.
Athena cost follows the bytes it scans. A columnar format cuts the columns read and date partitions cut the rows read; storage classes and spending limits leave the scanned bytes unchanged.
Mobile apps send JSON events to an Amazon Data Firehose stream that delivers them to Amazon S3. The analytics team wants the delivered files to be in Apache Parquet so that Athena queries scan less data. The team does not want to write or run any conversion code. What should a data engineer do?
ATurn on dynamic partitioning in the Firehose stream and partition the output by event type
BTurn on record format conversion in Firehose, using a table in the AWS Glue Data Catalog
CSet the Firehose S3 compression to Snappy for every object delivered to the bucket
DAdd a nightly AWS Glue job that reads the delivered JSON and rewrites it as Parquet
Show the answer and why
ATurn on dynamic partitioning in the Firehose stream and partition the output by event type
Incorrect
Dynamic partitioning places records under S3 prefixes based on keys in the data. The records stay in their original JSON format.
BTurn on record format conversion in Firehose, using a table in the AWS Glue Data Catalog
Correct
Firehose can convert JSON input to Apache Parquet or Apache ORC before it writes to Amazon S3, using the table schema it reads from the AWS Glue Data Catalog.
CSet the Firehose S3 compression to Snappy for every object delivered to the bucket
Incorrect
Compression shrinks the delivered files, but they remain JSON text rather than a columnar format.
DAdd a nightly AWS Glue job that reads the delivered JSON and rewrites it as Parquet
Incorrect
A Glue job can write Parquet, but it is conversion code that the team would have to write, schedule and run.
Firehose has format conversion built in: JSON in, Parquet or ORC out, with the schema taken from a Glue table. Partitioning and compression are useful settings, but neither changes the file format.
An AWS Glue ETL job must read tables from an Amazon RDS for PostgreSQL DB instance in a private subnet. The data engineer created a JDBC connection that uses the DB instance's subnet and its security group. That security group allows PostgreSQL traffic only from the application servers. Job runs fail because the AWS Glue components cannot communicate with each other. What should the data engineer do?
AAdd an inbound rule for all TCP ports whose source is the same security group
BMake the DB instance publicly accessible and allow traffic from the internet
CAdd a route to an internet gateway in the route table of the private subnet
DCreate an interface VPC endpoint for the Amazon RDS API in the same VPC
Show the answer and why
AAdd an inbound rule for all TCP ports whose source is the same security group
Correct
AWS Glue needs a security group with a self-referencing inbound rule for all TCP ports so that its components can talk to each other. The rule is limited to that group and does not open the VPC to other networks.
BMake the DB instance publicly accessible and allow traffic from the internet
Incorrect
This exposes the database and still leaves the security group without the self-referencing rule that the Glue components need.
CAdd a route to an internet gateway in the route table of the private subnet
Incorrect
A route changes where traffic leaves the subnet. Security group rules still block the traffic between the Glue components.
DCreate an interface VPC endpoint for the Amazon RDS API in the same VPC
Incorrect
That endpoint carries RDS API calls such as creating or modifying DB instances privately. It does not carry database connections.
A Glue connection runs inside the VPC with the security group you give it. That group needs a self-referencing all-TCP inbound rule, and the database must accept traffic from it.
A nightly Apache Spark job runs on an Amazon EMR cluster that uses uniform instance groups. Intermediate data is kept in HDFS on the core nodes, and losing it would fail the whole run. Finishing an hour later than today is acceptable. Which change will reduce cost while protecting the run?
ARun the core instance group on Spot Instances and keep the task nodes On-Demand
BMove every instance group, including the primary and core nodes, to Spot Instances
CKeep primary and core nodes On-Demand and add Spot task nodes
DAdd more On-Demand core nodes so that the job finishes sooner each night
Show the answer and why
ARun the core instance group on Spot Instances and keep the task nodes On-Demand
Incorrect
Core nodes store HDFS data. Terminating a core node risks data loss, which is exactly what this run cannot tolerate.
BMove every instance group, including the primary and core nodes, to Spot Instances
Incorrect
All-Spot clusters suit cost-driven work where losing partial work is acceptable. Here a Spot interruption of a core node could fail the run.
CKeep primary and core nodes On-Demand and add Spot task nodes
Correct
Task nodes process data but hold no HDFS data, so a Spot interruption loses no data. AWS recommends On-Demand primary and core nodes with Spot task nodes when losing partial work is not acceptable.
DAdd more On-Demand core nodes so that the job finishes sooner each night
Incorrect
Extra core nodes add capacity at the On-Demand price. They do not use the cheaper Spot capacity that suits nodes without HDFS data.
Put Spot capacity where an interruption costs only time: task nodes. Keep the nodes that store HDFS data and run the cluster On-Demand when the data matters.
Once a month a company must extract the product name, issue type and sentiment from about 2 million customer emails stored as text files in Amazon S3, using a foundation model in Amazon Bedrock. Results are needed within a few days. The team wants to keep cost low and avoid handling request throttling in its own code. What should a data engineer do?
ASubmit a Bedrock batch inference job that reads its prompts from S3
BCall InvokeModel for every email from an AWS Lambda function that S3 events trigger
CBuy Provisioned Throughput for the model and send the requests from an EC2 instance
DIndex the emails in Amazon OpenSearch Service and run a search query for each field
Show the answer and why
ASubmit a Bedrock batch inference job that reads its prompts from S3
Correct
Batch inference takes the prompts from files in S3 and writes the responses back to S3 asynchronously. Select models are priced 50% below on-demand inference for batch.
BCall InvokeModel for every email from an AWS Lambda function that S3 events trigger
Incorrect
This uses on-demand pricing for 2 million calls, and the function code has to handle throttling and retries.
CBuy Provisioned Throughput for the model and send the requests from an EC2 instance
Incorrect
Provisioned Throughput is billed by the hour for committed model units. It suits steady traffic, not a once-a-month job that can wait.
DIndex the emails in Amazon OpenSearch Service and run a search query for each field
Incorrect
OpenSearch Service searches and analyzes indexed documents. A search returns matching emails; it does not read each email and produce the structured fields.
Large, non-urgent model workloads are what batch inference is for: one job, inputs and outputs in S3, no throttling logic, and a lower price for select models.
Several AWS Glue Spark jobs reprocess historical data every weekend. Nobody waits for them, and it does not matter if a run starts later or takes longer than usual. The jobs use G.1X workers on AWS Glue 4.0. How can a data engineer reduce the cost of these runs?
AChange the worker type to G.2X
BLower the job timeout so that each run is stopped after a shorter time
CSet the jobs' execution class to FLEX
DRaise the number of workers so that each weekend run completes sooner
Show the answer and why
AChange the worker type to G.2X
Incorrect
A G.2X worker maps to 2 DPU instead of 1, so each worker costs more. The jobs need more speed only if someone waits for them.
BLower the job timeout so that each run is stopped after a shorter time
Incorrect
The timeout ends a run that takes too long, which would leave the reprocessing unfinished. It does not change the price per DPU-hour.
CSet the jobs' execution class to FLEX
Correct
The flexible execution class is meant for non-urgent jobs and is priced lower per DPU-hour than the standard class. It supports Glue 3.0 or later with G.1X or G.2X workers.
DRaise the number of workers so that each weekend run completes sooner
Incorrect
More workers can shorten a run, but they do not lower the rate charged for each DPU-hour, which is what the flexible class changes.
When time does not matter, the flexible execution class trades start-up and run-time certainty for a lower DPU-hour price. Worker size, worker count and timeouts change how fast a job runs, not what each unit costs.
A nightly Amazon EMR cluster uses instance groups. Its task group runs on Spot Instances of a single instance type, and the cluster often waits for Spot capacity, so the job finishes late. The team wants to keep using Spot Instances. What should a data engineer do?
ATurn on EMR managed scaling with higher minimum and maximum capacity limits
BUse instance fleets with several instance types and an allocation strategy
CMove the task group to On-Demand Instances of the same instance type
DTurn on termination protection for the cluster before each run starts
Show the answer and why
ATurn on EMR managed scaling with higher minimum and maximum capacity limits
Incorrect
Managed scaling changes how many instances the cluster uses. The task group still asks for the same single Spot instance type.
BUse instance fleets with several instance types and an allocation strategy
Correct
Instance fleets accept up to five instance types, or up to 30 with an allocation strategy, and EMR fills the target capacity from any of them, which lowers the chance of insufficient capacity.
CMove the task group to On-Demand Instances of the same instance type
Incorrect
This could avoid Spot shortages, but it drops the Spot Instances the team wants to keep.
DTurn on termination protection for the cluster before each run starts
Incorrect
Termination protection guards against accidental termination. It does not help EMR obtain Spot capacity.
Diversifying instance types is the main lever for Spot availability. Instance fleets with an allocation strategy let EMR pick from many pools.
In an AWS Glue ETL job, the amount column of a DynamicFrame is a choice type: most values are numbers, but some are strings such as "N/A". Analysts want to keep both kinds of values, each in its own column. Which transform should a data engineer apply?
AResolveChoice with the project:double action
BRelationalize on the DynamicFrame
CResolveChoice with the make_cols action
DDropNullFields on the DynamicFrame
Show the answer and why
AResolveChoice with the project:double action
Incorrect
project keeps only values of the named type and drops the others, so the string values would be lost.
BRelationalize on the DynamicFrame
Incorrect
Relationalize flattens nested structures and pivots arrays into separate tables. It does not split a column by data type.
CResolveChoice with the make_cols action
Correct
make_cols flattens the ambiguity into one column per type, such as amount_int and amount_string, so no values are lost.
DDropNullFields on the DynamicFrame
Incorrect
DropNullFields removes fields whose type is NullType. The amount column holds values, so nothing changes.
ResolveChoice settles choice types: cast or project pick one type, make_cols splits by type into columns, and make_struct nests both in one column.
Analysts want to explore an Amazon Kinesis data stream with SQL and see results within seconds before they decide what a long-running streaming application should do. They do not want to provision servers. Which solution should a data engineer provide?
AAn AWS Glue crawler that crawls the stream into the Data Catalog
BAn AWS Glue DataBrew project with the stream as its dataset
CA Kinesis Client Library application that prints each record
DA Managed Service for Apache Flink Studio notebook
Show the answer and why
AAn AWS Glue crawler that crawls the stream into the Data Catalog
Incorrect
Crawlers read file-based and table-based data stores such as Amazon S3, DynamoDB, and JDBC databases. A Kinesis stream is not among them.
BAn AWS Glue DataBrew project with the stream as its dataset
Incorrect
A DataBrew dataset points to a file, Amazon S3, a JDBC source, or the Data Catalog, not to a live stream.
CA Kinesis Client Library application that prints each record
Incorrect
The KCL helps developers build consumer applications in code. It gives analysts no interactive SQL.
DA Managed Service for Apache Flink Studio notebook
Correct
Studio notebooks let users query data streams interactively in real time with SQL, Python, or Scala in a serverless notebook.
For ad hoc SQL on a live stream, Flink Studio notebooks give interactive results; a long-running Managed Service for Apache Flink application can come later.
An AWS Glue job reads a Data Catalog table partitioned by year, month, and day that holds 3 years of data. The script loads the whole table into a DynamicFrame and then keeps only yesterday's rows. Each run lists and reads every file. What should a data engineer change?
AReplace the row filter with a Filter transform that runs after the read
BRaise the job's number of workers so that the full read runs in parallel
CPass a push_down_predicate on the partition columns to the read
DRun an AWS Glue crawler on the table before each job run starts
Show the answer and why
AReplace the row filter with a Filter transform that runs after the read
Incorrect
Filtering after the read still loads the entire dataset first, which is exactly what the pushdown predicate avoids.
BRaise the job's number of workers so that the full read runs in parallel
Incorrect
More workers add capacity, but the job still lists and reads every file in the table.
CPass a push_down_predicate on the partition columns to the read
Correct
A pushdown predicate filters on partition metadata in the Data Catalog, so the job lists and reads only the matching partitions.
DRun an AWS Glue crawler on the table before each job run starts
Incorrect
A crawler updates table metadata in the Data Catalog. It does not limit what the job reads.
Partition pruning belongs at the read: push_down_predicate, and for tables with very many partitions also catalogPartitionPredicate, keep the job from touching old partitions.
A pandas script cleans one 300 MB CSV file each night. It runs for about 40 minutes and keeps failing as an AWS Lambda function. The team does not want to rewrite it for a distributed engine or to manage a cluster. Which option should a data engineer use?
AAn AWS Glue Python shell job
BThe same Lambda function with its timeout raised to 60 minutes
CAn Amazon EMR cluster that runs the script as a step
DAn AWS Glue streaming ETL job that reads the file
Show the answer and why
AAn AWS Glue Python shell job
Correct
Python shell jobs run Python scripts in AWS Glue with common analytics libraries such as pandas, on 0.0625 or 1 DPU, without Spark.
BThe same Lambda function with its timeout raised to 60 minutes
Incorrect
The Lambda timeout can be set to at most 900 seconds (15 minutes).
CAn Amazon EMR cluster that runs the script as a step
Incorrect
EMR runs big data frameworks on clusters that the team would have to size and manage, which it wants to avoid.
DAn AWS Glue streaming ETL job that reads the file
Incorrect
Streaming ETL jobs process data continuously from streaming sources such as Kinesis and Kafka, not a nightly file.
Lambda stops at 15 minutes. For a single-node Python script that needs longer, a Glue Python shell job is the managed fit.