Skip to content
BytePatterns

DEA-C01 · Domain 3: Data Operations and Support · 22% of the exam

Task 3.1: Automate data processing by using AWS services

Letting the platform do the work: orchestration in MWAA and Step Functions and fixing it when it fails, SDK calls from code, features of EMR, Redshift and Glue, preparing data in DataBrew and SageMaker Unified Studio, querying with Athena, Lambda, and EventBridge schedules.

Study it

  • Automating with MWAA, Step Functions, Lambda and EventBridge, and fixing failed runs

    Partly covered by: Lambda & Event-Driven Design

  • Preparing data: AWS Glue DataBrew and SageMaker Unified Studio

    Lesson coming

  • Querying S3 with Athena: partitions, formats and workgroups

    Lesson coming

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

A data engineer uploads a new DAG file to the dags folder of an Amazon MWAA environment's S3 bucket. Other DAGs keep running, but the new DAG never appears in the Airflow UI. Where should the engineer look first to find out why?

  1. AThe task logs in CloudWatch Logs for the new DAG
  2. BThe webserver logs in CloudWatch Logs
  3. CThe CloudTrail event history for the S3 bucket
  4. DThe DAG processing logs in CloudWatch Logs
Show the answer and why
  • AThe task logs in CloudWatch Logs for the new DAG

    Incorrect

    Task logs are written when a DAG's tasks run. A DAG that was never parsed has no task runs to log.

  • BThe webserver logs in CloudWatch Logs

    Incorrect

    The webserver log group holds logs from the Airflow web interface. The UI only shows DAGs after the scheduler has processed them.

  • CThe CloudTrail event history for the S3 bucket

    Incorrect

    CloudTrail records API calls such as the upload itself. It does not show how Airflow processed the file afterwards.

  • DThe DAG processing logs in CloudWatch Logs

    Correct

    MWAA can send the logs of the DAG processor manager, the part of the scheduler that processes DAG files, to a CloudWatch log group. Errors while parsing a new file show up there.

A missing DAG is a parsing problem until proven otherwise. Turn on the DAG processing log group (it is off unless enabled) and look for the import or syntax error of the new file.

Question 2 · choose 1

A Lambda function calls the Athena StartQueryExecution API and then immediately calls GetQueryResults with the returned QueryExecutionId. The second call often fails because the query has not finished. How should the code be changed?

  1. APass a ClientRequestToken so that StartQueryExecution waits for the query to finish
  2. BCall GetQueryResults with a larger MaxResults value so that it waits for the rows
  3. CCall GetQueryExecution until the state is SUCCEEDED, then call GetQueryResults
  4. DRead the result file from the S3 output location as soon as StartQueryExecution returns
Show the answer and why
  • APass a ClientRequestToken so that StartQueryExecution waits for the query to finish

    Incorrect

    The token makes the start request idempotent so a retry does not run the query twice. It does not make the call wait.

  • BCall GetQueryResults with a larger MaxResults value so that it waits for the rows

    Incorrect

    MaxResults sets how many rows come back per page. GetQueryResults does not run or wait for the query.

  • CCall GetQueryExecution until the state is SUCCEEDED, then call GetQueryResults

    Correct

    StartQueryExecution only starts the query. GetQueryExecution returns its current state (QUEUED, RUNNING, SUCCEEDED and so on), and results can be read once it has succeeded.

  • DRead the result file from the S3 output location as soon as StartQueryExecution returns

    Incorrect

    Athena writes results to the output location as the query completes, so right after the start call there is nothing complete to read.

Athena queries are asynchronous: start, check the state, then fetch. In a Step Functions workflow the athena:startQueryExecution.sync integration does the waiting for you.

Question 3 · choose 1

Business analysts receive a CSV extract in Amazon S3 every day. They must standardize date formats, trim spaces and split a full-name column before the data is written to S3 as Apache Parquet. The analysts do not write code, and the same steps must run on every new extract. What should a data engineer set up for them?

  1. AAn AWS Glue DataBrew profile job that runs on every new extract
  2. BAn AWS Glue crawler that runs on the extract folder every day
  3. CAn Amazon Athena CTAS query that the analysts edit for each extract
  4. DAn AWS Glue DataBrew recipe, run daily by a scheduled recipe job
Show the answer and why
  • AAn AWS Glue DataBrew profile job that runs on every new extract

    Incorrect

    A profile job evaluates a dataset and reports statistics. It does not change the data or write cleaned output.

  • BAn AWS Glue crawler that runs on the extract folder every day

    Incorrect

    A crawler records the schema in the Data Catalog. It does not clean or transform any values.

  • CAn Amazon Athena CTAS query that the analysts edit for each extract

    Incorrect

    CTAS can write Parquet, but the cleaning would be SQL that someone has to write and run, which the analysts cannot do.

  • DAn AWS Glue DataBrew recipe, run daily by a scheduled recipe job

    Correct

    DataBrew offers more than 250 point-and-click transformations saved as steps in a recipe. A recipe job applies the recipe to the data and writes the output, including Parquet, and jobs can be scheduled.

DataBrew is the no-code preparation tool: build the steps once as a recipe on a sample, then let a scheduled recipe job apply them to each new file.

Question 4 · choose 1

An AWS Glue job must start at 06:00 Europe/Berlin local time every weekday, and the start time must stay at 06:00 local time when daylight saving time begins and ends. The team wants no custom code. Which approach meets these requirements?

  1. AA scheduled rule in the default EventBridge event bus with a cron expression
  2. BAn EventBridge Scheduler cron schedule with the Europe/Berlin time zone
  3. CA time-based trigger on the AWS Glue job with a cron expression
  4. DA Lambda function that runs every minute and starts the job at 06:00
Show the answer and why
  • AA scheduled rule in the default EventBridge event bus with a cron expression

    Incorrect

    Scheduled rules use the UTC+0 time zone, so a fixed cron expression drifts by an hour in local time when daylight saving time changes.

  • BAn EventBridge Scheduler cron schedule with the Europe/Berlin time zone

    Correct

    Scheduler evaluates cron and rate expressions in the time zone you specify and handles daylight saving time, and it can start the Glue job directly as a target.

  • CA time-based trigger on the AWS Glue job with a cron expression

    Incorrect

    Glue time-based schedules are defined in UTC, so the local start time moves with daylight saving time.

  • DA Lambda function that runs every minute and starts the job at 06:00

    Incorrect

    This is custom code and wastes thousands of invocations a day for one start.

For schedules tied to a local clock, use EventBridge Scheduler and give it the time zone. Legacy scheduled rules and Glue triggers work in UTC.

Question 5 · choose 1

A Step Functions state machine must start an AWS Glue crawler as one of its steps. The team does not want to write or maintain any function code for this step. How should a data engineer define the step?

  1. AA Task state that invokes a Lambda function that calls StartCrawler
  2. BA Task state that uses the aws-sdk:glue:startCrawler resource
  3. CAn activity task with a worker that polls and starts the crawler
  4. DAn HTTP Task that calls the AWS Glue API endpoint for the crawler
Show the answer and why
  • AA Task state that invokes a Lambda function that calls StartCrawler

    Incorrect

    This works, but it adds a function whose code the team would have to write and maintain.

  • BA Task state that uses the aws-sdk:glue:startCrawler resource

    Correct

    AWS SDK service integrations let a Task state call an AWS API action directly, with a resource such as arn:aws:states:::aws-sdk:glue:startCrawler.

  • CAn activity task with a worker that polls and starts the crawler

    Incorrect

    With activities, a worker program running outside Step Functions does the work, which is code the team would have to run.

  • DAn HTTP Task that calls the AWS Glue API endpoint for the crawler

    Incorrect

    HTTP Tasks are for calling HTTPS APIs such as third-party SaaS applications. AWS APIs have SDK integrations.

Step Functions reaches most AWS APIs through SDK integrations, so glue code in Lambda is needed only when a step does real work of its own.

Question 6 · choose 1

An EMR Serverless application uses pre-initialized capacity so that hourly jobs start quickly during business hours. The team sees charges all night, when no jobs run, and finds that someone turned off a default setting. What should a data engineer change?

  1. ATurn auto-stop back on, with a suitable idle timeout
  2. BLower the application's maximum capacity for CPU and memory
  3. CTurn off auto-start so jobs cannot start the application at night
  4. DAdd more pre-initialized workers so each job finishes sooner at night
Show the answer and why
  • ATurn auto-stop back on, with a suitable idle timeout

    Correct

    Auto-stop stops an idle application, 15 minutes by default, and a stopped application releases its pre-initialized capacity.

  • BLower the application's maximum capacity for CPU and memory

    Incorrect

    Maximum capacity caps how far the application can scale. The pre-initialized workers are still paid for while they sit idle.

  • CTurn off auto-start so jobs cannot start the application at night

    Incorrect

    Auto-start starts the application when a job is submitted. It does not stop an application that is already running and idle.

  • DAdd more pre-initialized workers so each job finishes sooner at night

    Incorrect

    Pre-initialized workers are paid for even when the application is idle, so adding more raises the overnight cost.

Pre-initialized capacity trades idle cost for fast starts; keeping auto-stop on bounds that idle cost.

Question 7 · choose 1

Every morning, a scheduled job runs an Amazon Athena query that summarizes yesterday's orders. The summary rows must be added to an existing summary table that dashboards already read. Which statement should the query use?

  1. ACREATE TABLE AS SELECT, writing a new table each day
  2. BUNLOAD the SELECT query results to the table's S3 path
  3. CINSERT INTO the summary table with the SELECT query
  4. DCREATE VIEW over the orders table with the SELECT query
Show the answer and why
  • ACREATE TABLE AS SELECT, writing a new table each day

    Incorrect

    CTAS creates a new table from the query results, so each run would produce a separate table instead of adding to the existing one.

  • BUNLOAD the SELECT query results to the table's S3 path

    Incorrect

    UNLOAD writes query results to S3 in a chosen format. It is not a statement for adding rows to a table.

  • CINSERT INTO the summary table with the SELECT query

    Correct

    INSERT INTO adds new rows to an existing destination table based on a SELECT statement.

  • DCREATE VIEW over the orders table with the SELECT query

    Incorrect

    A view is a logical table whose query runs each time it is used. It stores no daily summary rows.

CTAS creates, INSERT INTO appends: scheduled jobs that grow an existing table use INSERT INTO with a SELECT.

Question 8 · choose 1

A Lambda function uses the Amazon Redshift Data API to run three SQL statements: delete a day's rows, insert the corrected rows, and update a load log. Either all three changes must be committed or none of them. Which approach should a data engineer use?

  1. AOne BatchExecuteStatement call with the AUTO_COMMIT run mode
  2. BThree ExecuteStatement calls, issued one after another in the same order
  3. COne BatchExecuteStatement call with the default TRANSACTION run mode
  4. DThree ExecuteStatement calls, issued in parallel from three threads
Show the answer and why
  • AOne BatchExecuteStatement call with the AUTO_COMMIT run mode

    Incorrect

    With AUTO_COMMIT, each statement is committed on its own, and one failure does not affect the others.

  • BThree ExecuteStatement calls, issued one after another in the same order

    Incorrect

    Each call runs its own statement, so a failure in the third leaves the first two committed.

  • COne BatchExecuteStatement call with the default TRANSACTION run mode

    Correct

    By default, the statements in a batch run as a single transaction, and if any statement fails, all of the work is rolled back.

  • DThree ExecuteStatement calls, issued in parallel from three threads

    Incorrect

    Parallel, separate statements are not one transaction, and the delete and insert could also run out of order.

Use BatchExecuteStatement in its default transaction mode when several statements must succeed or fail together.

Question 9 · choose 1

Every night, thousands of containerized simulation jobs must run, each for 2 to 3 hours. The team wants a managed queue for the jobs and wants compute provisioned for them automatically. Which AWS service should a data engineer use?

  1. AAWS Batch
  2. BAWS Lambda with container images
  3. CAn Amazon ECS service with a set desired count
  4. DStep Functions Express workflows that run each job
Show the answer and why
  • AAWS Batch

    Correct

    AWS Batch runs batch computing workloads and provisions the compute resources for submitted jobs, removing the work of managing that infrastructure.

  • BAWS Lambda with container images

    Incorrect

    Lambda functions can run for at most 900 seconds, far short of 2 to 3 hours.

  • CAn Amazon ECS service with a set desired count

    Incorrect

    An ECS service keeps a specified number of tasks running and replaces failed ones, which suits long-running services, not a queue of batch jobs.

  • DStep Functions Express workflows that run each job

    Incorrect

    Express workflows run for at most 5 minutes, so they cannot wait for a job that takes hours.

Queued, long-running container work is AWS Batch's job: submit jobs to a queue, and Batch provisions the compute to run them.

Practise domain 3 →Practise all domains →