Question 1 · choose 1
An AWS Glue ETL job builds the daily training table for a churn model. The team wants the job itself to check the data in transit, for example that customer_id is always present and unique and that tenure is within a valid range, and to stop the pipeline before a bad table reaches training. What should the ML engineer add?
- AAn AWS Glue crawler that runs on the output prefix after each job
- BAWS Glue Data Quality rules in DQDL, evaluated in the ETL job
- CAmazon CloudWatch alarms on the Glue job metrics
- DAn Amazon Macie job on the output bucket
Show the answer and why
AAn AWS Glue crawler that runs on the output prefix after each job
Incorrect
A crawler infers schemas and adds or updates tables in the Data Catalog. It does not test the values in the data against rules.
BAWS Glue Data Quality rules in DQDL, evaluated in the ETL job
Correct
Glue Data Quality evaluates rules such as IsComplete, IsUnique and range checks written in the Data Quality Definition Language, both on Data Catalog tables and inside Glue ETL jobs, and the job can act on the results.
CAmazon CloudWatch alarms on the Glue job metrics
Incorrect
Glue job metrics describe the job's execution, such as memory and data moved, not whether each value in the table is valid.
DAn Amazon Macie job on the output bucket
Incorrect
Macie discovers sensitive data such as PII in S3 objects. It does not check completeness, uniqueness or value ranges.
Glue Data Quality, built on the open-source Deequ framework, brings declarative data checks into the pipeline so bad data can be stopped or quarantined before it is used for training.
AWS documentation