Skip to content
BytePatterns

DEA-C01 · Domain 3: Data Operations and Support · 22% of the exam

Task 3.4: Ensure data quality

Trusting the numbers: quality rules and checks while data is processed, DataBrew and Glue Data Quality rules, finding inconsistent data, sampling, and dealing with skewed data.

Study it

  • Data quality rules: AWS Glue Data Quality and DataBrew

    Lesson coming

  • Sampling and data skew

    Partly covered by: Database Sharding

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 2

An AWS Glue ETL job loads orders into a data lake. While the job runs, every record must be checked that order_id is not null and that amount is between 0 and 100000. Records that fail must go to a quarantine location, and good records must still be loaded. Which actions meet these requirements? (Choose TWO.)

  1. AAdd an Evaluate Data Quality transform with IsComplete and ColumnValues rules
  2. BAdd the row-level outcome columns and route rows whose DataQualityEvaluationResult is Failed
  3. CSet the data quality action to fail the job without loading the target data at all
  4. DRun an AWS Glue DataBrew profile job on the data lake tables after every load finishes
  5. ECreate a CloudWatch alarm on the run time of the ETL job
Show the answer and why
  • AAdd an Evaluate Data Quality transform with IsComplete and ColumnValues rules

    Correct

    Glue Data Quality evaluates DQDL rules inside the ETL job; IsComplete checks for nulls and ColumnValues checks values against an expression.

  • BAdd the row-level outcome columns and route rows whose DataQualityEvaluationResult is Failed

    Correct

    The rowLevelOutcomes option adds columns that mark each record Passed or Failed, so the job can split failing rows from good ones.

  • CSet the data quality action to fail the job without loading the target data at all

    Incorrect

    That stops the whole load when rules fail, so the good records would not be loaded either.

  • DRun an AWS Glue DataBrew profile job on the data lake tables after every load finishes

    Incorrect

    A profile job reports statistics after the fact. Bad records would already be in the lake.

  • ECreate a CloudWatch alarm on the run time of the ETL job

    Incorrect

    Run time says nothing about whether individual records are valid.

Data quality checks belong inside the pipeline: rules evaluate each record, and row-level outcomes let the job load the good rows and quarantine the rest.

Question 2 · choose 1

Before designing cleaning steps for a new 50 GB dataset in Amazon S3, analysts want a report of missing values, duplicate rows, value distributions and correlations between columns. They do not write code. What should a data engineer run?

  1. AAn AWS Glue DataBrew recipe job on the dataset
  2. BAn AWS Glue DataBrew profile job on the dataset
  3. CAn AWS Glue crawler on the dataset's S3 prefix
  4. DAn Amazon Macie sensitive data discovery job on the bucket
Show the answer and why
  • AAn AWS Glue DataBrew recipe job on the dataset

    Incorrect

    A recipe job applies cleaning steps and writes transformed data. It does not produce a profile of the data.

  • BAn AWS Glue DataBrew profile job on the dataset

    Correct

    A profile job runs a series of evaluations on a dataset and writes the results to S3; it can count duplicate rows and build a correlations matrix.

  • CAn AWS Glue crawler on the dataset's S3 prefix

    Incorrect

    A crawler infers the schema for the Data Catalog. It does not measure missing values or distributions.

  • DAn Amazon Macie sensitive data discovery job on the bucket

    Incorrect

    Macie looks for sensitive data such as PII. It does not report data quality statistics.

Profile first, then clean: a DataBrew profile shows where the data is incomplete or inconsistent, and a recipe then fixes it.

Question 3 · choose 1

An AWS Glue Spark job joins orders with customers on customer_id. A few customers own most of the orders. In the Spark UI, the longest task of the join stage processes about 100 times more data than the 75th percentile. The job parameters set spark.sql.adaptive.enabled to false. Which change addresses this with the least code change?

  1. ATurn on Adaptive Query Execution and its skew join settings
  2. BAdd more workers so that the stage gets more executors
  3. CCoalesce the orders DataFrame into fewer partitions before the join
  4. DTurn on job bookmarks so that each run reads less data
Show the answer and why
  • ATurn on Adaptive Query Execution and its skew join settings

    Correct

    AQE re-plans using runtime statistics; its skew join feature splits oversized partitions in sort-merge joins. AWS suggests it when the Spark UI shows data skew.

  • BAdd more workers so that the stage gets more executors

    Incorrect

    The oversized partition is still processed by one task, so more executors mostly sit idle while it runs.

  • CCoalesce the orders DataFrame into fewer partitions before the join

    Incorrect

    Fewer partitions make each one larger. The skewed key still ends up in one task.

  • DTurn on job bookmarks so that each run reads less data

    Incorrect

    Bookmarks skip data processed in earlier runs. The skew in the data that is read stays the same.

Skew means one partition holds far more than its share. Spark's AQE skew join splits those partitions at run time; a higher-cardinality or composite join key is the design-level fix.

Question 4 · choose 1

Analysts prepare a dataset in AWS Glue DataBrew. Before it is shared, the data must pass checks such as APY values between 0 and 100 and fewer than 1% missing values in key columns, with a result for each rule. What should a data engineer set up?

  1. AA DataBrew recipe with transformation steps
  2. BA custom classifier on an AWS Glue crawler
  3. CA Filter transform in an AWS Glue job
  4. DA DataBrew ruleset added to a profile job
Show the answer and why
  • AA DataBrew recipe with transformation steps

    Incorrect

    A recipe is a set of data transformation steps. It changes the data rather than validating it rule by rule.

  • BA custom classifier on an AWS Glue crawler

    Incorrect

    A classifier recognizes a data format and generates a schema. It does not check value ranges.

  • CA Filter transform in an AWS Glue job

    Incorrect

    Filter removes rows that do not meet a condition. It reports no per-rule results.

  • DA DataBrew ruleset added to a profile job

    Correct

    A ruleset compares data metrics against expected values, fails if any rule's criteria are not met, and is validated by adding it to a profile job.

DataBrew keeps cleaning (recipes) and checking (rulesets in profile jobs) separate, and the ruleset gives a pass or fail for each rule.

Question 5 · choose 2

AWS Glue Data Quality rulesets run every night against Data Catalog tables. The on-call engineer must get an Amazon SNS notification whenever a ruleset evaluation fails. Which TWO steps should a data engineer take? (Choose TWO.)

  1. ACreate an EventBridge rule for Data Quality Evaluation Results Available events in the FAILED state
  2. BSet the SNS topic as the target of that EventBridge rule
  3. CTurn on AWS Glue job run insights for the evaluation jobs
  4. DGenerate Data Catalog column statistics for the evaluated tables
  5. EOpen a CloudWatch Logs Live Tail session on the job's log group
Show the answer and why
  • ACreate an EventBridge rule for Data Quality Evaluation Results Available events in the FAILED state

    Correct

    Glue Data Quality publishes an EventBridge event when a ruleset evaluation run completes, and the event pattern can select the FAILED state.

  • BSet the SNS topic as the target of that EventBridge rule

    Correct

    With the matching event routed to the topic, alerts go out whenever a rule fails.

  • CTurn on AWS Glue job run insights for the evaluation jobs

    Incorrect

    Job run insights help debug and optimize jobs. They do not send notifications about data quality results.

  • DGenerate Data Catalog column statistics for the evaluated tables

    Incorrect

    Column statistics describe values for query planning. They do not alert on ruleset failures.

  • EOpen a CloudWatch Logs Live Tail session on the job's log group

    Incorrect

    Live Tail shows log events to someone watching in near real time. It does not notify anyone.

Data quality results are events: route the failed ones through EventBridge to SNS so that people hear about bad data before consumers do.

Question 6 · choose 1

A Glue Data Quality ruleset must fail when any order has a ship_date earlier than its order_date. Which DQDL rule should a data engineer add?

  1. AIsComplete "ship_date"
  2. BColumnCorrelation "ship_date" "order_date" > 0.9
  3. CCustomSql "select count(*) from primary where ship_date < order_date" = 0
  4. DDataFreshness "order_date" <= 24 hours
Show the answer and why
  • AIsComplete "ship_date"

    Incorrect

    IsComplete checks that all values in a column are non-null. It does not compare two columns.

  • BColumnCorrelation "ship_date" "order_date" > 0.9

    Incorrect

    ColumnCorrelation measures the linear correlation between two columns. Highly correlated dates can still be out of order.

  • CCustomSql "select count(*) from primary where ship_date < order_date" = 0

    Correct

    CustomSql runs a SQL statement that returns a single numeric value and checks it against an expression, here that no such rows exist.

  • DDataFreshness "order_date" <= 24 hours

    Incorrect

    DataFreshness checks how recent the values in a date column are. It does not compare two columns.

Cross-column business rules that no built-in rule type covers fit CustomSql, which checks a SQL result against an expression.

Question 7 · choose 1

Before a full data quality review, an analyst wants a quick check on about 1% of the rows in a large Athena table, with each row chosen independently at random. Which clause should the query use?

  1. ATABLESAMPLE SYSTEM (1)
  2. BTABLESAMPLE BERNOULLI (100)
  3. CLIMIT 1000 after the WHERE clause
  4. DTABLESAMPLE BERNOULLI (1)
Show the answer and why
  • ATABLESAMPLE SYSTEM (1)

    Incorrect

    SYSTEM samples whole logical segments and does not guarantee independent sampling probabilities.

  • BTABLESAMPLE BERNOULLI (100)

    Incorrect

    The value is the sampling percentage, so 100 would select every row.

  • CLIMIT 1000 after the WHERE clause

    Incorrect

    LIMIT caps how many rows are returned. It does not select rows with a sampling probability.

  • DTABLESAMPLE BERNOULLI (1)

    Correct

    BERNOULLI selects each row with a probability of the given percentage, so each row is chosen independently.

For a statistically fair sample, BERNOULLI picks rows independently; SYSTEM is cheaper but samples in blocks.

Question 8 · choose 1

An AWS Glue Studio visual job reads orders where the same order_id appears several times with different ingestion timestamps. Only one row per order_id must remain. Which transform configuration should a data engineer use?

  1. ADrop Duplicates, matching rows that are completely the same
  2. BFilter, keeping rows whose order_id is not empty
  3. CDropNullFields on the order data before writing
  4. DDrop Duplicates, matching on the order_id field only
Show the answer and why
  • ADrop Duplicates, matching rows that are completely the same

    Incorrect

    Full-row matching removes only identical rows. Rows that differ in their timestamps would all remain.

  • BFilter, keeping rows whose order_id is not empty

    Incorrect

    Filter removes rows that fail a condition. Repeated order IDs all pass this condition.

  • CDropNullFields on the order data before writing

    Incorrect

    DropNullFields removes fields of NullType. It does not remove repeated rows.

  • DDrop Duplicates, matching on the order_id field only

    Correct

    Drop Duplicates can match on chosen fields and remove the rows that repeat those fields, keeping one row per order_id.

Deduplication depends on the key: full-row matching only catches exact copies, while a field match enforces one row per business key.

Practise domain 3 →Practise all domains →