Question 1 · choose 1
A company loads product records from several suppliers into Amazon S3 with an AWS Glue ETL job, and the records are then ingested into a knowledge base. Some supplier files arrive with empty descriptions or duplicate SKUs, which leads to poor answers. The data engineers want declarative rules instead of custom code, a quality score for each run, and no data written to the target when the rules fail, with an alert when that happens. Which solution meets these requirements?
- AWrite a Lambda function that is triggered by each new S3 object and checks every record before the knowledge base ingests it
- BAdd a DQDL ruleset with AWS Glue Data Quality to the ETL job, fail it without loading the target, and publish the results
- CCall Amazon Comprehend on each description during the job and drop records whose entity list is empty
- DRun Amazon Macie discovery jobs on the landing bucket and stop the ingestion pipeline whenever Macie reports new findings
Show the answer and why
AWrite a Lambda function that is triggered by each new S3 object and checks every record before the knowledge base ingests it
Incorrect
A Lambda function can validate files, but it is custom code without declarative rules or a quality score, which the engineers want to avoid.
BAdd a DQDL ruleset with AWS Glue Data Quality to the ETL job, fail it without loading the target, and publish the results
Correct
Glue Data Quality evaluates DQDL rules inside the ETL job, reports a data quality score, can fail the job without loading data, and publishes metrics to CloudWatch and EventBridge for alerting.
CCall Amazon Comprehend on each description during the job and drop records whose entity list is empty
Incorrect
Comprehend extracts entities from text, which is useful for enrichment, but it is not a rule engine for empty fields or duplicate keys and adds per-record cost.
DRun Amazon Macie discovery jobs on the landing bucket and stop the ingestion pipeline whenever Macie reports new findings
Incorrect
Macie discovers sensitive data such as PII in S3. It does not check completeness or uniqueness of business fields.
Validation belongs before the data reaches the vector store. Glue Data Quality gives declarative DQDL rules, rule recommendations, a score per run, configurable job failure behavior and CloudWatch and EventBridge integration, so bad supplier files never reach the knowledge base.
AWS documentation