Question 1 · choose 2
An AWS Glue ETL job loads orders into a data lake. While the job runs, every record must be checked that order_id is not null and that amount is between 0 and 100000. Records that fail must go to a quarantine location, and good records must still be loaded. Which actions meet these requirements? (Choose TWO.)
- AAdd an Evaluate Data Quality transform with IsComplete and ColumnValues rules
- BAdd the row-level outcome columns and route rows whose DataQualityEvaluationResult is Failed
- CSet the data quality action to fail the job without loading the target data at all
- DRun an AWS Glue DataBrew profile job on the data lake tables after every load finishes
- ECreate a CloudWatch alarm on the run time of the ETL job
Show the answer and why
AAdd an Evaluate Data Quality transform with IsComplete and ColumnValues rules
Correct
Glue Data Quality evaluates DQDL rules inside the ETL job; IsComplete checks for nulls and ColumnValues checks values against an expression.
BAdd the row-level outcome columns and route rows whose DataQualityEvaluationResult is Failed
Correct
The rowLevelOutcomes option adds columns that mark each record Passed or Failed, so the job can split failing rows from good ones.
CSet the data quality action to fail the job without loading the target data at all
Incorrect
That stops the whole load when rules fail, so the good records would not be loaded either.
DRun an AWS Glue DataBrew profile job on the data lake tables after every load finishes
Incorrect
A profile job reports statistics after the fact. Bad records would already be in the lake.
ECreate a CloudWatch alarm on the run time of the ETL job
Incorrect
Run time says nothing about whether individual records are valid.
Data quality checks belong inside the pipeline: rules evaluate each record, and row-level outcomes let the job load the good rows and quarantine the rest.
AWS documentation