Skip to content
BytePatterns

DEA-C01 · Domain 1: Data Ingestion and Transformation · 34% of the exam

Task 1.2: Transform and process data

Turning raw data into useful data: picking Glue, EMR, Lambda or Redshift for a transformation, JDBC and ODBC connections, joining sources, converting CSV to Parquet, keeping processing cost down, fixing failed or slow jobs, data APIs for other systems, and LLMs that read or enrich data.

Study it

  • Batch ingestion: S3, AWS DMS, AppFlow and JDBC sources

    Lesson coming

  • Choosing the engine: AWS Glue, Amazon EMR, Lambda or Redshift

    Lesson coming

  • File formats: CSV to Parquet, compression and partitioned output

    Lesson coming

  • Fixing slow and failing Spark jobs on Glue and EMR

    Lesson coming

  • Data APIs with API Gateway, Lambda and the Redshift Data API

    Partly covered by: Designing a REST API

  • LLMs in a pipeline: Amazon Bedrock for extraction and enrichment

    Partly covered by: What Is an LLM, The Tool-Use Loop

  • Containers for data processing: ECS, EKS and EMR on EKS

    Partly covered by: ECS vs EKS vs Fargate

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 2

Analysts use Amazon Athena to query 3 years of web logs that are stored in Amazon S3 as gzip-compressed CSV files. Almost every query filters on one day or one week and reads 6 of the 80 columns. Query cost and run time keep growing. Which actions will reduce the amount of data that Athena scans? (Choose TWO.)

  1. AConvert the data to Apache Parquet with an AWS Glue ETL job
  2. BMove the log objects to the S3 Intelligent-Tiering storage class
  3. CSet a per-query data usage control on the analysts' workgroup
  4. DPartition the data in Amazon S3 by event date and add the partitions to the table
  5. EStore the files uncompressed so that Athena can read them directly
Show the answer and why
  • AConvert the data to Apache Parquet with an AWS Glue ETL job

    Correct

    Parquet stores each column separately, so Athena reads only the columns a query uses and can skip data through predicate pushdown.

  • BMove the log objects to the S3 Intelligent-Tiering storage class

    Incorrect

    Intelligent-Tiering moves objects between access tiers to lower storage cost. Athena still reads the same bytes for each query.

  • CSet a per-query data usage control on the analysts' workgroup

    Incorrect

    A per-query control cancels a query that scans more than the limit. It caps spending; it does not make queries read less data.

  • DPartition the data in Amazon S3 by event date and add the partitions to the table

    Correct

    With date partitions, a query that filters on a day or a week reads only those partitions instead of all 3 years.

  • EStore the files uncompressed so that Athena can read them directly

    Incorrect

    Athena charges for bytes scanned before decompression, so removing compression makes every query scan more data.

Athena cost follows the bytes it scans. A columnar format cuts the columns read and date partitions cut the rows read; storage classes and spending limits leave the scanned bytes unchanged.

Question 2 · choose 1

Mobile apps send JSON events to an Amazon Data Firehose stream that delivers them to Amazon S3. The analytics team wants the delivered files to be in Apache Parquet so that Athena queries scan less data. The team does not want to write or run any conversion code. What should a data engineer do?

  1. ATurn on dynamic partitioning in the Firehose stream and partition the output by event type
  2. BTurn on record format conversion in Firehose, using a table in the AWS Glue Data Catalog
  3. CSet the Firehose S3 compression to Snappy for every object delivered to the bucket
  4. DAdd a nightly AWS Glue job that reads the delivered JSON and rewrites it as Parquet
Show the answer and why
  • ATurn on dynamic partitioning in the Firehose stream and partition the output by event type

    Incorrect

    Dynamic partitioning places records under S3 prefixes based on keys in the data. The records stay in their original JSON format.

  • BTurn on record format conversion in Firehose, using a table in the AWS Glue Data Catalog

    Correct

    Firehose can convert JSON input to Apache Parquet or Apache ORC before it writes to Amazon S3, using the table schema it reads from the AWS Glue Data Catalog.

  • CSet the Firehose S3 compression to Snappy for every object delivered to the bucket

    Incorrect

    Compression shrinks the delivered files, but they remain JSON text rather than a columnar format.

  • DAdd a nightly AWS Glue job that reads the delivered JSON and rewrites it as Parquet

    Incorrect

    A Glue job can write Parquet, but it is conversion code that the team would have to write, schedule and run.

Firehose has format conversion built in: JSON in, Parquet or ORC out, with the schema taken from a Glue table. Partitioning and compression are useful settings, but neither changes the file format.

Question 3 · choose 1

An AWS Glue ETL job must read tables from an Amazon RDS for PostgreSQL DB instance in a private subnet. The data engineer created a JDBC connection that uses the DB instance's subnet and its security group. That security group allows PostgreSQL traffic only from the application servers. Job runs fail because the AWS Glue components cannot communicate with each other. What should the data engineer do?

  1. AAdd an inbound rule for all TCP ports whose source is the same security group
  2. BMake the DB instance publicly accessible and allow traffic from the internet
  3. CAdd a route to an internet gateway in the route table of the private subnet
  4. DCreate an interface VPC endpoint for the Amazon RDS API in the same VPC
Show the answer and why
  • AAdd an inbound rule for all TCP ports whose source is the same security group

    Correct

    AWS Glue needs a security group with a self-referencing inbound rule for all TCP ports so that its components can talk to each other. The rule is limited to that group and does not open the VPC to other networks.

  • BMake the DB instance publicly accessible and allow traffic from the internet

    Incorrect

    This exposes the database and still leaves the security group without the self-referencing rule that the Glue components need.

  • CAdd a route to an internet gateway in the route table of the private subnet

    Incorrect

    A route changes where traffic leaves the subnet. Security group rules still block the traffic between the Glue components.

  • DCreate an interface VPC endpoint for the Amazon RDS API in the same VPC

    Incorrect

    That endpoint carries RDS API calls such as creating or modifying DB instances privately. It does not carry database connections.

A Glue connection runs inside the VPC with the security group you give it. That group needs a self-referencing all-TCP inbound rule, and the database must accept traffic from it.

Question 4 · choose 1

A nightly Apache Spark job runs on an Amazon EMR cluster that uses uniform instance groups. Intermediate data is kept in HDFS on the core nodes, and losing it would fail the whole run. Finishing an hour later than today is acceptable. Which change will reduce cost while protecting the run?

  1. ARun the core instance group on Spot Instances and keep the task nodes On-Demand
  2. BMove every instance group, including the primary and core nodes, to Spot Instances
  3. CKeep primary and core nodes On-Demand and add Spot task nodes
  4. DAdd more On-Demand core nodes so that the job finishes sooner each night
Show the answer and why
  • ARun the core instance group on Spot Instances and keep the task nodes On-Demand

    Incorrect

    Core nodes store HDFS data. Terminating a core node risks data loss, which is exactly what this run cannot tolerate.

  • BMove every instance group, including the primary and core nodes, to Spot Instances

    Incorrect

    All-Spot clusters suit cost-driven work where losing partial work is acceptable. Here a Spot interruption of a core node could fail the run.

  • CKeep primary and core nodes On-Demand and add Spot task nodes

    Correct

    Task nodes process data but hold no HDFS data, so a Spot interruption loses no data. AWS recommends On-Demand primary and core nodes with Spot task nodes when losing partial work is not acceptable.

  • DAdd more On-Demand core nodes so that the job finishes sooner each night

    Incorrect

    Extra core nodes add capacity at the On-Demand price. They do not use the cheaper Spot capacity that suits nodes without HDFS data.

Put Spot capacity where an interruption costs only time: task nodes. Keep the nodes that store HDFS data and run the cluster On-Demand when the data matters.

Question 5 · choose 1

Once a month a company must extract the product name, issue type and sentiment from about 2 million customer emails stored as text files in Amazon S3, using a foundation model in Amazon Bedrock. Results are needed within a few days. The team wants to keep cost low and avoid handling request throttling in its own code. What should a data engineer do?

  1. ASubmit a Bedrock batch inference job that reads its prompts from S3
  2. BCall InvokeModel for every email from an AWS Lambda function that S3 events trigger
  3. CBuy Provisioned Throughput for the model and send the requests from an EC2 instance
  4. DIndex the emails in Amazon OpenSearch Service and run a search query for each field
Show the answer and why
  • ASubmit a Bedrock batch inference job that reads its prompts from S3

    Correct

    Batch inference takes the prompts from files in S3 and writes the responses back to S3 asynchronously. Select models are priced 50% below on-demand inference for batch.

  • BCall InvokeModel for every email from an AWS Lambda function that S3 events trigger

    Incorrect

    This uses on-demand pricing for 2 million calls, and the function code has to handle throttling and retries.

  • CBuy Provisioned Throughput for the model and send the requests from an EC2 instance

    Incorrect

    Provisioned Throughput is billed by the hour for committed model units. It suits steady traffic, not a once-a-month job that can wait.

  • DIndex the emails in Amazon OpenSearch Service and run a search query for each field

    Incorrect

    OpenSearch Service searches and analyzes indexed documents. A search returns matching emails; it does not read each email and produce the structured fields.

Large, non-urgent model workloads are what batch inference is for: one job, inputs and outputs in S3, no throttling logic, and a lower price for select models.

Question 6 · choose 1

Several AWS Glue Spark jobs reprocess historical data every weekend. Nobody waits for them, and it does not matter if a run starts later or takes longer than usual. The jobs use G.1X workers on AWS Glue 4.0. How can a data engineer reduce the cost of these runs?

  1. AChange the worker type to G.2X
  2. BLower the job timeout so that each run is stopped after a shorter time
  3. CSet the jobs' execution class to FLEX
  4. DRaise the number of workers so that each weekend run completes sooner
Show the answer and why
  • AChange the worker type to G.2X

    Incorrect

    A G.2X worker maps to 2 DPU instead of 1, so each worker costs more. The jobs need more speed only if someone waits for them.

  • BLower the job timeout so that each run is stopped after a shorter time

    Incorrect

    The timeout ends a run that takes too long, which would leave the reprocessing unfinished. It does not change the price per DPU-hour.

  • CSet the jobs' execution class to FLEX

    Correct

    The flexible execution class is meant for non-urgent jobs and is priced lower per DPU-hour than the standard class. It supports Glue 3.0 or later with G.1X or G.2X workers.

  • DRaise the number of workers so that each weekend run completes sooner

    Incorrect

    More workers can shorten a run, but they do not lower the rate charged for each DPU-hour, which is what the flexible class changes.

When time does not matter, the flexible execution class trades start-up and run-time certainty for a lower DPU-hour price. Worker size, worker count and timeouts change how fast a job runs, not what each unit costs.

Question 7 · choose 1

A nightly Amazon EMR cluster uses instance groups. Its task group runs on Spot Instances of a single instance type, and the cluster often waits for Spot capacity, so the job finishes late. The team wants to keep using Spot Instances. What should a data engineer do?

  1. ATurn on EMR managed scaling with higher minimum and maximum capacity limits
  2. BUse instance fleets with several instance types and an allocation strategy
  3. CMove the task group to On-Demand Instances of the same instance type
  4. DTurn on termination protection for the cluster before each run starts
Show the answer and why
  • ATurn on EMR managed scaling with higher minimum and maximum capacity limits

    Incorrect

    Managed scaling changes how many instances the cluster uses. The task group still asks for the same single Spot instance type.

  • BUse instance fleets with several instance types and an allocation strategy

    Correct

    Instance fleets accept up to five instance types, or up to 30 with an allocation strategy, and EMR fills the target capacity from any of them, which lowers the chance of insufficient capacity.

  • CMove the task group to On-Demand Instances of the same instance type

    Incorrect

    This could avoid Spot shortages, but it drops the Spot Instances the team wants to keep.

  • DTurn on termination protection for the cluster before each run starts

    Incorrect

    Termination protection guards against accidental termination. It does not help EMR obtain Spot capacity.

Diversifying instance types is the main lever for Spot availability. Instance fleets with an allocation strategy let EMR pick from many pools.

Question 8 · choose 1

In an AWS Glue ETL job, the amount column of a DynamicFrame is a choice type: most values are numbers, but some are strings such as "N/A". Analysts want to keep both kinds of values, each in its own column. Which transform should a data engineer apply?

  1. AResolveChoice with the project:double action
  2. BRelationalize on the DynamicFrame
  3. CResolveChoice with the make_cols action
  4. DDropNullFields on the DynamicFrame
Show the answer and why
  • AResolveChoice with the project:double action

    Incorrect

    project keeps only values of the named type and drops the others, so the string values would be lost.

  • BRelationalize on the DynamicFrame

    Incorrect

    Relationalize flattens nested structures and pivots arrays into separate tables. It does not split a column by data type.

  • CResolveChoice with the make_cols action

    Correct

    make_cols flattens the ambiguity into one column per type, such as amount_int and amount_string, so no values are lost.

  • DDropNullFields on the DynamicFrame

    Incorrect

    DropNullFields removes fields whose type is NullType. The amount column holds values, so nothing changes.

ResolveChoice settles choice types: cast or project pick one type, make_cols splits by type into columns, and make_struct nests both in one column.

Question 9 · choose 1

Analysts want to explore an Amazon Kinesis data stream with SQL and see results within seconds before they decide what a long-running streaming application should do. They do not want to provision servers. Which solution should a data engineer provide?

  1. AAn AWS Glue crawler that crawls the stream into the Data Catalog
  2. BAn AWS Glue DataBrew project with the stream as its dataset
  3. CA Kinesis Client Library application that prints each record
  4. DA Managed Service for Apache Flink Studio notebook
Show the answer and why
  • AAn AWS Glue crawler that crawls the stream into the Data Catalog

    Incorrect

    Crawlers read file-based and table-based data stores such as Amazon S3, DynamoDB, and JDBC databases. A Kinesis stream is not among them.

  • BAn AWS Glue DataBrew project with the stream as its dataset

    Incorrect

    A DataBrew dataset points to a file, Amazon S3, a JDBC source, or the Data Catalog, not to a live stream.

  • CA Kinesis Client Library application that prints each record

    Incorrect

    The KCL helps developers build consumer applications in code. It gives analysts no interactive SQL.

  • DA Managed Service for Apache Flink Studio notebook

    Correct

    Studio notebooks let users query data streams interactively in real time with SQL, Python, or Scala in a serverless notebook.

For ad hoc SQL on a live stream, Flink Studio notebooks give interactive results; a long-running Managed Service for Apache Flink application can come later.

Question 10 · choose 1

An AWS Glue job reads a Data Catalog table partitioned by year, month, and day that holds 3 years of data. The script loads the whole table into a DynamicFrame and then keeps only yesterday's rows. Each run lists and reads every file. What should a data engineer change?

  1. AReplace the row filter with a Filter transform that runs after the read
  2. BRaise the job's number of workers so that the full read runs in parallel
  3. CPass a push_down_predicate on the partition columns to the read
  4. DRun an AWS Glue crawler on the table before each job run starts
Show the answer and why
  • AReplace the row filter with a Filter transform that runs after the read

    Incorrect

    Filtering after the read still loads the entire dataset first, which is exactly what the pushdown predicate avoids.

  • BRaise the job's number of workers so that the full read runs in parallel

    Incorrect

    More workers add capacity, but the job still lists and reads every file in the table.

  • CPass a push_down_predicate on the partition columns to the read

    Correct

    A pushdown predicate filters on partition metadata in the Data Catalog, so the job lists and reads only the matching partitions.

  • DRun an AWS Glue crawler on the table before each job run starts

    Incorrect

    A crawler updates table metadata in the Data Catalog. It does not limit what the job reads.

Partition pruning belongs at the read: push_down_predicate, and for tables with very many partitions also catalogPartitionPredicate, keep the job from touching old partitions.

Question 11 · choose 1

A pandas script cleans one 300 MB CSV file each night. It runs for about 40 minutes and keeps failing as an AWS Lambda function. The team does not want to rewrite it for a distributed engine or to manage a cluster. Which option should a data engineer use?

  1. AAn AWS Glue Python shell job
  2. BThe same Lambda function with its timeout raised to 60 minutes
  3. CAn Amazon EMR cluster that runs the script as a step
  4. DAn AWS Glue streaming ETL job that reads the file
Show the answer and why
  • AAn AWS Glue Python shell job

    Correct

    Python shell jobs run Python scripts in AWS Glue with common analytics libraries such as pandas, on 0.0625 or 1 DPU, without Spark.

  • BThe same Lambda function with its timeout raised to 60 minutes

    Incorrect

    The Lambda timeout can be set to at most 900 seconds (15 minutes).

  • CAn Amazon EMR cluster that runs the script as a step

    Incorrect

    EMR runs big data frameworks on clusters that the team would have to size and manage, which it wants to avoid.

  • DAn AWS Glue streaming ETL job that reads the file

    Incorrect

    Streaming ETL jobs process data continuously from streaming sources such as Kinesis and Kafka, not a nightly file.

Lambda stops at 15 minutes. For a single-node Python script that needs longer, a Glue Python shell job is the managed fit.

Practise domain 1 →Practise all domains →