Skip to content
BytePatterns

DEA-C01 · Domain 1: Data Ingestion and Transformation · 34% of the exam

Task 1.4: Apply programming concepts

The engineering around the pipeline: faster code, Lambda concurrency and storage, SQL and other languages, version control and tests, infrastructure as code with CloudFormation, the CDK and AWS SAM, CI/CD for pipelines, distributed computing, and data structures such as trees and graphs.

Study it

  • Lambda as a stream consumer: batching, retries and partial failures

    Partly covered by: Lambda & Event-Driven Design

  • Lambda concurrency, memory and storage for data jobs

    Partly covered by: Lambda & Event-Driven Design

  • Infrastructure as code and CI/CD for pipelines: CloudFormation, the CDK and AWS SAM

    Lesson coming

  • Distributed computing, trees and graphs for data engineers

    Partly covered by: Graph Basics, Tree Basics

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

An AWS Lambda function is invoked by Amazon S3 event notifications and writes each new file's rows to an Amazon RDS for MySQL database. When thousands of files arrive at once, the database runs out of connections. The function must never run more than 20 instances at the same time, and other functions in the account must not be affected. What should a data engineer do?

  1. ASet the function's provisioned concurrency to 20
  2. BRaise the function's memory
  3. CRaise the function's timeout to 15 minutes so that fewer retries happen
  4. DSet the function's reserved concurrency to 20
Show the answer and why
  • ASet the function's provisioned concurrency to 20

    Incorrect

    Provisioned concurrency keeps instances initialized to cut cold starts. It does not stop the function from scaling beyond that number.

  • BRaise the function's memory

    Incorrect

    More memory brings more CPU and shorter runs, but a burst of files still starts many instances at once.

  • CRaise the function's timeout to 15 minutes so that fewer retries happen

    Incorrect

    The timeout limits how long one invocation runs. It does not limit how many instances run at the same time.

  • DSet the function's reserved concurrency to 20

    Correct

    Reserved concurrency is the maximum number of concurrent instances for the function, and it is set aside for that function alone. AWS suggests it to protect downstream resources such as database connections.

To protect a database, cap the writer. Reserved concurrency is both the upper limit and a guaranteed share for one function, while provisioned concurrency is about start-up latency.

Question 2 · choose 1

A Lambda function downloads a 4 GB file from Amazon S3, unpacks it to about 5 GB of temporary files, and writes a summary back to S3. Nothing on disk is needed after the invocation ends or by any other function. The function fails with "No space left on device". What is the simplest fix?

  1. APackage the large reference files in a Lambda layer attached to the function
  2. BRaise the function's timeout to 15 minutes so the unpacking can finish
  3. CSet provisioned concurrency so that the temporary files are kept between invocations
  4. DIncrease the function's ephemeral storage for /tmp to 10,240 MB
Show the answer and why
  • APackage the large reference files in a Lambda layer attached to the function

    Incorrect

    A deployment package with its layers is limited to 250 MB unzipped. The problem is space for data the function creates at run time.

  • BRaise the function's timeout to 15 minutes so the unpacking can finish

    Incorrect

    The error is about disk space, not time. A longer timeout leaves /tmp at the same size.

  • CSet provisioned concurrency so that the temporary files are kept between invocations

    Incorrect

    Provisioned concurrency keeps environments initialized; it does not enlarge /tmp, and the files need not survive the invocation anyway.

  • DIncrease the function's ephemeral storage for /tmp to 10,240 MB

    Correct

    Ephemeral storage in /tmp can be set between 512 MB and 10,240 MB. It is temporary and private to each execution environment, which is all this function needs.

Scratch space that lives only for the invocation belongs in /tmp, which can grow to 10,240 MB. Shared or persistent storage would only be needed if other invocations had to read the files.

Question 3 · choose 1

A team builds a serverless pipeline from three Lambda functions, a Step Functions state machine and a DynamoDB table. It wants a short template with serverless resource types, a command-line tool that builds and packages the code, local test invocations of the functions, and deployment as an AWS CloudFormation stack. What should the team use?

  1. AAWS CodeDeploy for each function
  2. BAWS SAM, with sam build and sam deploy
  3. CCloudFormation StackSets, with one stack instance for the account
  4. DA shell script that calls aws lambda create-function for each function
Show the answer and why
  • AAWS CodeDeploy for each function

    Incorrect

    CodeDeploy automates deployments of application code to Lambda, EC2 or ECS. It does not define the state machine or the table as infrastructure.

  • BAWS SAM, with sam build and sam deploy

    Correct

    AWS SAM templates extend CloudFormation with shorthand serverless resources. The SAM CLI builds the code, invokes functions locally with sam local, and sam deploy deploys through CloudFormation.

  • CCloudFormation StackSets, with one stack instance for the account

    Incorrect

    StackSets manage stacks across many accounts and Regions. They do not add shorthand serverless resources, builds or local testing.

  • DA shell script that calls aws lambda create-function for each function

    Incorrect

    This creates functions one call at a time. It is not infrastructure as code and does not cover the other resources.

AWS SAM is the serverless layer on top of CloudFormation: a compact template plus a CLI that builds, tests locally and deploys the stack.

Question 4 · choose 1

A Python Lambda function decompresses and re-encodes files. Profiling shows that it spends nearly all of its time on CPU work, and with 512 MB of memory a typical file takes 9 minutes. The team wants each file processed faster without changing the code. What should a data engineer do?

  1. AConfigure provisioned concurrency for the function
  2. BIncrease the ephemeral storage that is mounted at /tmp
  3. CIncrease the memory setting of the function
  4. DSet reserved concurrency to 1,000
Show the answer and why
  • AConfigure provisioned concurrency for the function

    Incorrect

    Provisioned concurrency removes cold-start time. It does not add CPU to a running invocation.

  • BIncrease the ephemeral storage that is mounted at /tmp

    Incorrect

    More /tmp space helps when the disk is full. It gives the function no extra processing power.

  • CIncrease the memory setting of the function

    Correct

    Lambda allocates CPU power in proportion to memory; at 1,769 MB a function has the equivalent of one vCPU. CPU-bound code runs faster with more memory.

  • DSet reserved concurrency to 1,000

    Incorrect

    Reserved concurrency controls how many instances run at once. Each instance keeps the same CPU share.

Memory is the single compute dial in Lambda: raising it raises CPU in proportion, which shortens CPU-bound work.

Question 5 · choose 1

A fraud team wants to find groups of customer accounts that are connected through shared devices, cards and addresses, following links up to four steps away. The links change all day, and investigators expect answers in milliseconds. How should a data engineer model and store this data?

  1. AAs one DynamoDB item per account that holds lists of the related devices and cards
  2. BAs fact and dimension tables in Amazon Redshift that analysts join nightly
  3. CAs vertices and edges in Amazon Neptune, traversed with Gremlin or openCypher
  4. DAs JSON documents in Amazon OpenSearch Service with a keyword field per link
Show the answer and why
  • AAs one DynamoDB item per account that holds lists of the related devices and cards

    Incorrect

    Each extra step would need more lookups in application code. DynamoDB stores items by key; it does not traverse relationships.

  • BAs fact and dimension tables in Amazon Redshift that analysts join nightly

    Incorrect

    A warehouse fits analytical queries over large tables. Repeated self-joins for each hop and nightly loads do not give live, millisecond answers.

  • CAs vertices and edges in Amazon Neptune, traversed with Gremlin or openCypher

    Correct

    Neptune is a graph database built to store billions of relationships and to traverse them with millisecond latency, with fraud detection as a typical use case.

  • DAs JSON documents in Amazon OpenSearch Service with a keyword field per link

    Incorrect

    OpenSearch Service searches and analyzes documents. It does not follow multi-step connections between them.

When the questions are about connections (who links to whom, in how many steps), the data is a graph. A graph database stores the edges directly and walks them quickly.

Question 6 · choose 1

An AWS Glue PySpark job reads two large datasets, joins them into a DataFrame named df, and then writes three different aggregations of df. The Spark UI shows that the expensive join runs three times. What should a data engineer do?

  1. ACall df.coalesce(1) after the join and before each of the aggregations
  2. BAdd a broadcast hint for one of the two large datasets
  3. CCall df.persist() after the join and unpersist it after the writes
  4. DTurn on job bookmarks for the job and rerun it
Show the answer and why
  • ACall df.coalesce(1) after the join and before each of the aggregations

    Incorrect

    Coalescing to one partition reduces parallelism. Each action would still recompute the join from the source data.

  • BAdd a broadcast hint for one of the two large datasets

    Incorrect

    Broadcasting suits a small table that fits in memory. It does not stop later actions from repeating the join.

  • CCall df.persist() after the join and unpersist it after the writes

    Correct

    Caching a DataFrame that is used repeatedly keeps the computed result in executor memory and on disk, so later actions reuse it.

  • DTurn on job bookmarks for the job and rerun it

    Incorrect

    Job bookmarks track data already processed across runs. They do not reuse results within one run.

Spark evaluates lazily, so every action rebuilds its lineage. Persist a DataFrame that several actions reuse, and release it when done.

Question 7 · choose 1

A Python Lambda function needs machine learning libraries that take 1.5 GB when unzipped. Deployment fails because the package is too large. How should a data engineer package the function?

  1. AAs a .zip file split across several Lambda layers
  2. BAs a .zip file uploaded through Amazon S3
  3. CAs a container image pushed to Amazon ECR
  4. DAs a .zip file, with memory raised to 10,240 MB
Show the answer and why
  • AAs a .zip file split across several Lambda layers

    Incorrect

    The 250 MB unzipped limit covers the function and all of its layers together, so splitting does not help.

  • BAs a .zip file uploaded through Amazon S3

    Incorrect

    Uploading through S3 avoids the 50 MB direct upload limit, but the 250 MB unzipped limit still applies.

  • CAs a container image pushed to Amazon ECR

    Correct

    Lambda container images can be up to 10 GB uncompressed, including all layers, which fits 1.5 GB of libraries.

  • DAs a .zip file, with memory raised to 10,240 MB

    Incorrect

    Memory is a runtime setting. It does not change the limit on deployment package size.

Zip packages and layers stop at 250 MB unzipped. Large dependency sets go in a container image, up to 10 GB.

Question 8 · choose 2

An AWS CloudFormation stack defines the S3 bucket that holds a data lake's raw zone. The bucket and its data must remain if the stack is deleted and if a later template change makes CloudFormation replace the bucket. Which TWO settings should a data engineer add to the bucket resource? (Choose TWO.)

  1. ADeletionPolicy: Snapshot
  2. BDeletionPolicy: Retain
  3. CA DependsOn attribute that names the stack's other resources
  4. DA CreationPolicy attribute with a resource signal count of 1
  5. EUpdateReplacePolicy: Retain
Show the answer and why
  • ADeletionPolicy: Snapshot

    Incorrect

    Snapshot applies only to resources that support snapshots, such as volumes and database clusters; S3 buckets are not among them.

  • BDeletionPolicy: Retain

    Correct

    With Retain, CloudFormation keeps the resource and its contents when the stack is deleted.

  • CA DependsOn attribute that names the stack's other resources

    Incorrect

    DependsOn controls the order in which resources are created. It does not keep a resource after deletion.

  • DA CreationPolicy attribute with a resource signal count of 1

    Incorrect

    CreationPolicy holds a resource's creation until success signals arrive. It does not affect deletion or replacement.

  • EUpdateReplacePolicy: Retain

    Correct

    UpdateReplacePolicy keeps the existing physical resource when a stack update replaces it; by default the old resource is deleted.

Deletion and replacement are separate paths in CloudFormation; protecting stateful resources such as data buckets takes both Retain policies.

Question 9 · choose 1

An Amazon Athena table is partitioned by dt, a string such as 2026-10-01. An analyst's query filters with WHERE date(event_time) = DATE '2026-10-01' and scans the entire table. How should a data engineer change the query to scan less data?

  1. AAdd a filter on the partition column, such as dt = '2026-10-01'
  2. BRun MSCK REPAIR TABLE on the table before the query runs
  3. CTurn on partition projection for the table and keep the query as it is
  4. DRewrite the query as CREATE TABLE AS SELECT (CTAS)
Show the answer and why
  • AAdd a filter on the partition column, such as dt = '2026-10-01'

    Correct

    When a query on a partitioned table specifies the partition in the WHERE clause, Athena scans only that partition.

  • BRun MSCK REPAIR TABLE on the table before the query runs

    Incorrect

    MSCK REPAIR TABLE adds partitions found in S3 to the catalog metadata. It does not restrict what a query reads.

  • CTurn on partition projection for the table and keep the query as it is

    Incorrect

    Projection calculates partition values and locations from table properties. The query still needs a filter on dt to skip partitions.

  • DRewrite the query as CREATE TABLE AS SELECT (CTAS)

    Incorrect

    CTAS writes a query's results to a new table in S3. The SELECT inside it would scan the same data.

Partitions help only when the query filters on the partition column itself. A filter on another column leaves Athena reading every partition.

Question 10 · choose 1

A data platform is deployed with AWS CloudFormation. The team suspects that someone used the console to change the settings of an AWS Glue job and the lifecycle rules of an Amazon S3 bucket, both defined in the stack's template. A data engineer must find which resources no longer match the template without changing anything. What should the data engineer do?

  1. AValidate the template with the ValidateTemplate operation
  2. BTurn on termination protection for the stack
  3. CDeploy the template again with CloudFormation StackSets
  4. DRun drift detection on the stack
Show the answer and why
  • AValidate the template with the ValidateTemplate operation

    Incorrect

    ValidateTemplate checks that the template is valid JSON or YAML. It reads only the template, not the live resources.

  • BTurn on termination protection for the stack

    Incorrect

    Termination protection stops the stack from being deleted. It reports nothing about resource settings.

  • CDeploy the template again with CloudFormation StackSets

    Incorrect

    StackSets create, update, or delete stacks across accounts and Regions. Deploying again is a change, not a check.

  • DRun drift detection on the stack

    Correct

    Drift detection compares the actual property values of the stack's resources with the values the template expects and returns the drift status of each supported resource. Glue jobs and S3 buckets both support it.

Changes made outside CloudFormation are drift. Drift detection only reads and reports, listing each supported resource as in sync, modified or deleted; correcting the drift is a separate, deliberate step.

Practise domain 1 →Practise all domains →