Skip to content
BytePatterns

DEA-C01 · Domain 1: Data Ingestion and Transformation · 34% of the exam

Task 1.3: Orchestrate data pipelines

Running the steps in order: workflows in Step Functions, Amazon MWAA, Glue workflows and EventBridge, serverless pipelines that retry and survive failures, and alerts through SNS and SQS.

Study it

  • Orchestration: Step Functions, Amazon MWAA, Glue workflows and EventBridge

    Partly covered by: SQS vs SNS vs EventBridge

  • Resilient pipelines: retries, dead-letter queues and alerts with SNS and SQS

    Partly covered by: Message Queues

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

An AWS Step Functions state machine starts an AWS Glue job and then runs a Lambda function that validates the job's output. The Lambda function runs seconds after the job starts and finds no output. Job durations vary from 10 minutes to 2 hours. How should a data engineer fix the workflow?

  1. AUse the glue:startJobRun.sync resource for the Glue task
  2. BAdd a Wait state of 2 hours before the Lambda task
  3. CConvert the state machine to an Express workflow so the steps run in sequence
  4. DRaise the timeout of the validation Lambda function to its 15-minute maximum
Show the answer and why
  • AUse the glue:startJobRun.sync resource for the Glue task

    Correct

    With the Run a Job (.sync) pattern, Step Functions waits for the Glue job run to finish before it moves to the next state.

  • BAdd a Wait state of 2 hours before the Lambda task

    Incorrect

    A fixed wait wastes up to 2 hours on short runs and still breaks the day a run takes longer.

  • CConvert the state machine to an Express workflow so the steps run in sequence

    Incorrect

    Express workflows support only the Request Response pattern and run for at most five minutes, so they cannot wait for a 2-hour job.

  • DRaise the timeout of the validation Lambda function to its 15-minute maximum

    Incorrect

    A longer timeout does not delay the function's start, and 15 minutes is shorter than many of the job runs.

The default Request Response integration returns as soon as the job has started. The .sync suffix makes the Task state wait for the job to complete, however long it takes.

Question 2 · choose 1

A company runs about 200 Apache Airflow DAGs on a self-managed Airflow installation on Amazon EC2. Operating the scheduler, workers and upgrades takes too much time. The company wants a managed service and wants to keep its DAGs with as few changes as possible. Which service should a data engineer choose?

  1. AAWS Step Functions, with one state machine for each existing DAG
  2. BAWS Glue workflows, with triggers that replace the DAG dependencies
  3. CAmazon Managed Workflows for Apache Airflow (Amazon MWAA)
  4. DAmazon EventBridge Scheduler, with one schedule for each DAG task
Show the answer and why
  • AAWS Step Functions, with one state machine for each existing DAG

    Incorrect

    Step Functions defines workflows in its own language, so every DAG would have to be rewritten as a state machine.

  • BAWS Glue workflows, with triggers that replace the DAG dependencies

    Incorrect

    Glue workflows chain Glue jobs and crawlers. Moving 200 DAGs into them means rebuilding the pipelines rather than keeping them.

  • CAmazon Managed Workflows for Apache Airflow (Amazon MWAA)

    Correct

    Amazon MWAA runs Apache Airflow as a managed service with the same open-source code and user interface, and scales workers automatically.

  • DAmazon EventBridge Scheduler, with one schedule for each DAG task

    Incorrect

    Schedules start targets at set times. They do not model the dependencies between tasks that a DAG describes.

When the workflows already exist as Airflow DAGs, the managed Airflow service removes the operations work without a rewrite. The other orchestrators are good choices for new pipelines.

Question 3 · choose 2

A Task state in an AWS Step Functions workflow calls a Lambda function that loads data into a partner API. The API sometimes throttles requests for a few seconds. The workflow must retry automatically with growing waits, and if the calls still fail, it must notify the on-call team and end without an unhandled error. Which configurations meet these requirements? (Choose TWO.)

  1. AA TimeoutSeconds value on the Task state that is twice the usual run time
  2. BA Parallel state that runs two copies of the Task state at the same time
  3. CAn Express workflow type, which retries failed states without any settings
  4. DA Retry field with IntervalSeconds, MaxAttempts and BackoffRate set to 2
  5. EA Catch field that moves to a state publishing to an SNS topic
Show the answer and why
  • AA TimeoutSeconds value on the Task state that is twice the usual run time

    Incorrect

    A timeout fails a state that runs too long. It neither retries the throttled calls nor notifies anyone.

  • BA Parallel state that runs two copies of the Task state at the same time

    Incorrect

    Running the call twice in parallel sends more requests to an API that is already throttling, and it adds no retry or fallback.

  • CAn Express workflow type, which retries failed states without any settings

    Incorrect

    Workflow type does not add retries. Retry and Catch must be defined in the state machine for either type.

  • DA Retry field with IntervalSeconds, MaxAttempts and BackoffRate set to 2

    Correct

    A retrier repeats the state after an error; BackoffRate multiplies the interval after each attempt, so the waits grow.

  • EA Catch field that moves to a state publishing to an SNS topic

    Correct

    A catcher runs after the retries are exhausted and sends the execution to a fallback state, here one that notifies the team through Amazon SNS.

Step Functions fails the whole execution on an unhandled error. Retry handles transient errors with exponential backoff; Catch handles what remains by routing to a fallback state.

Question 4 · choose 1

A data platform team runs dozens of AWS Glue jobs on schedules. The team wants an email whenever any job run fails or times out, without polling the Glue API and without changing the job scripts. What should a data engineer do?

  1. AAn EventBridge rule on Glue Job State Change events in FAILED or TIMEOUT state, targeting SNS
  2. BJob bookmarks on every job, with the team's email address subscribed to bookmark updates
  3. CA CloudWatch metric filter on StartJobRun calls logged by CloudTrail, with an alarm on the count
  4. DA Lambda function that runs every 5 minutes, calls GetJobRuns for each job and emails failures
Show the answer and why
  • AAn EventBridge rule on Glue Job State Change events in FAILED or TIMEOUT state, targeting SNS

    Correct

    Glue sends Glue Job State Change events for SUCCEEDED, FAILED, TIMEOUT and STOPPED. A rule can match the failed states and notify an SNS topic with email subscribers.

  • BJob bookmarks on every job, with the team's email address subscribed to bookmark updates

    Incorrect

    Job bookmarks track which data a job has already processed. They do not report run failures to anyone.

  • CA CloudWatch metric filter on StartJobRun calls logged by CloudTrail, with an alarm on the count

    Incorrect

    StartJobRun records that a run was started. It says nothing about how the run ended.

  • DA Lambda function that runs every 5 minutes, calls GetJobRuns for each job and emails failures

    Incorrect

    This works, but it is the polling the team wants to avoid, and it is code to maintain.

Glue publishes job state changes to EventBridge, so a rule with a pattern on the state field turns failures into notifications without any polling.

Question 5 · choose 1

A Step Functions workflow receives a list of S3 object keys in its input and uses a Map state to run a Lambda function on each object. A new dataset has about 3 million objects under one prefix. The workflow now fails because the input is too large, and the team needs far more than 40 objects processed at a time. What should a data engineer change?

  1. AKeep the Map state in Inline mode and set MaxConcurrency to 1000
  2. BSwitch the Map state to Distributed mode, reading items from the S3 prefix
  3. CReplace the Map state with a Parallel state that has one branch per object
  4. DRun the workflow as an Express workflow so that it can process more items
Show the answer and why
  • AKeep the Map state in Inline mode and set MaxConcurrency to 1000

    Incorrect

    Inline mode supports up to 40 concurrent iterations and accepts only a JSON array from the state input, so the size problem remains.

  • BSwitch the Map state to Distributed mode, reading items from the S3 prefix

    Correct

    Distributed mode is meant for datasets over 256 KiB, histories over 25,000 events and more than 40 concurrent iterations. It can read items from S3 and runs each batch as a child workflow execution.

  • CReplace the Map state with a Parallel state that has one branch per object

    Incorrect

    A Parallel state runs a fixed set of branches defined in the state machine. It cannot create millions of branches from data.

  • DRun the workflow as an Express workflow so that it can process more items

    Incorrect

    Express workflows run for at most five minutes. The workflow type does not lift the Inline Map limits.

Inline Map works on a JSON array in the state input, up to 40 at a time. For large S3 datasets, Distributed Map reads the items itself and fans out to child workflows with much higher concurrency.

Question 6 · choose 1

A Step Functions workflow uses an Inline Map state to call a partner API once for each of 400 items. The partner allows at most 5 concurrent requests, and many calls are being throttled. What should a data engineer do so that the workflow never exceeds the partner's limit?

  1. ASwitch the Map state to Distributed mode
  2. BSet MaxConcurrency to 5 on the Map state
  3. CAdd a Wait state of 5 seconds before the Map state
  4. DAdd a Retry with a BackoffRate to the Task inside the Map
Show the answer and why
  • ASwitch the Map state to Distributed mode

    Incorrect

    Distributed mode runs each iteration as a child workflow for high concurrency. It does not, by itself, lower the number of parallel calls.

  • BSet MaxConcurrency to 5 on the Map state

    Correct

    MaxConcurrency caps how many iterations run at the same time; a value of 5 keeps at most 5 calls in flight.

  • CAdd a Wait state of 5 seconds before the Map state

    Incorrect

    A Wait state delays the workflow once. The Map iterations still start in parallel afterward.

  • DAdd a Retry with a BackoffRate to the Task inside the Map

    Incorrect

    Retry reruns a call after it fails. The first attempts still start in parallel, so the limit is still exceeded.

To respect a downstream concurrency limit, cap the fan-out itself with MaxConcurrency. Retries only soften the errors after the limit is hit.

Question 7 · choose 1

An Amazon EventBridge rule matches S3 Object Created events and starts a Step Functions state machine. The state machine expects an input of only two fields, bucket and key, but it receives the whole event. What should a data engineer configure?

  1. AA more specific event pattern on the rule
  2. BAn archive for the rule's event bus
  3. CA dead-letter queue on the rule's target
  4. DAn input transformer on the rule's target
Show the answer and why
  • AA more specific event pattern on the rule

    Incorrect

    An event pattern decides which events match the rule. Matched events still go to the target unchanged.

  • BAn archive for the rule's event bus

    Incorrect

    An archive stores events so that they can be replayed later. It does not reshape what the target receives.

  • CA dead-letter queue on the rule's target

    Incorrect

    A dead-letter queue keeps events that EventBridge could not deliver to a target. It does not change successful deliveries.

  • DAn input transformer on the rule's target

    Correct

    An input transformer extracts values from the event with an input path and places them in an input template before EventBridge passes the information to the target.

Patterns choose events; input transformers shape them for the target. Shaping the payload in EventBridge keeps the state machine's input contract simple.

Question 8 · choose 1

A Step Functions task uses .waitForTaskToken to wait for an on-premises worker that can run for up to 6 hours. If the worker crashes, the workflow must fail within 10 minutes instead of waiting for the full 6 hours. The worker can call AWS APIs while it runs. What should a data engineer do?

  1. ASet TimeoutSeconds to 600 on the Task state that waits for the worker
  2. BAdd a Retry on States.TaskFailed to the task with three attempts
  3. CAdd a Catch on States.ALL that moves the execution to a cleanup state
  4. DSet HeartbeatSeconds to 600; the worker calls SendTaskHeartbeat
Show the answer and why
  • ASet TimeoutSeconds to 600 on the Task state that waits for the worker

    Incorrect

    TimeoutSeconds limits the whole task, so healthy runs that take hours would also fail after 10 minutes.

  • BAdd a Retry on States.TaskFailed to the task with three attempts

    Incorrect

    Retry acts after an error is raised. A crashed worker raises no error, so the task keeps waiting.

  • CAdd a Catch on States.ALL that moves the execution to a cleanup state

    Incorrect

    Catch routes errors that occur. Without a heartbeat, no error occurs until the long task timeout expires.

  • DSet HeartbeatSeconds to 600; the worker calls SendTaskHeartbeat

    Correct

    If more than HeartbeatSeconds pass between heartbeats, the Task state fails with States.Timeout, so a crashed worker is detected within 10 minutes.

Long callback tasks need two clocks: a long TimeoutSeconds for the whole job and a short HeartbeatSeconds that proves the worker is still alive.

Question 9 · choose 1

An AWS Glue workflow starts on every S3 Object Created event and now runs hundreds of times a day. The team wants each run to start after 50 new files have arrived, or 15 minutes after the first new file, whichever comes first. What should a data engineer configure?

  1. AA scheduled trigger that starts the workflow every 15 minutes of the day
  2. BA conditional trigger that watches the crawler that runs first
  3. CAn on-demand trigger that the upload application calls itself
  4. DAn event trigger with batch size 50 and batch window 900 seconds
Show the answer and why
  • AA scheduled trigger that starts the workflow every 15 minutes of the day

    Incorrect

    A schedule fires on the clock, whether or not files arrived, and never starts early when 50 files are waiting.

  • BA conditional trigger that watches the crawler that runs first

    Incorrect

    Conditional triggers fire on the state of jobs or crawlers in the workflow, not on a count of arriving files.

  • CAn on-demand trigger that the upload application calls itself

    Incorrect

    An on-demand trigger starts only when someone or something calls it, so the batching logic would move into the uploader.

  • DAn event trigger with batch size 50 and batch window 900 seconds

    Correct

    Event triggers can wait for a batch of events and start the workflow when the batch size is reached or the window after the first event ends, whichever occurs first.

Glue event triggers batch EventBridge events by count and time window, turning a flood of object events into fewer, larger workflow runs.

Question 10 · choose 1

In a Step Functions workflow that uses JSONPath, a Task state calls a Lambda function to look up a customer's tier. The next state needs both the original order fields and the tier, but the task output contains only the tier. What should a data engineer set on the Task state?

  1. A"ResultPath": "$.tier"
  2. B"ResultPath": null
  3. C"OutputPath": "$.tier"
  4. D"InputPath": "$.order"
Show the answer and why
  • A"ResultPath": "$.tier"

    Correct

    ResultPath can include the task result with the input, so the output keeps the order fields and adds the tier under $.tier.

  • B"ResultPath": null

    Incorrect

    With ResultPath set to null, the state passes the original input to the output and discards the task result, so the tier is lost.

  • C"OutputPath": "$.tier"

    Incorrect

    OutputPath filters the state's output before passing it on, so it would keep only the tier, not the order.

  • D"InputPath": "$.order"

    Incorrect

    InputPath limits the input passed to the task. It does not decide how the result combines with the input.

InputPath picks what the task sees, ResultPath places the result, and OutputPath trims what moves on. ResultPath to a field merges the two.

Practise domain 1 →Practise all domains →