Question 1 · choose 1
An AWS Glue crawler runs on s3://example-bucket/orders/, where Parquet files are stored under one prefix per day (dt=2026-10-01/ and so on). Some days have extra columns. Instead of one table partitioned by dt, the crawler creates a separate table for several days. What should a data engineer change?
- AConfigure the crawler to crawl only new folders on every later scheduled run
- BAdd an exclude pattern that skips the daily folders that have extra columns
- CRun MSCK REPAIR TABLE in Athena after every crawler run
- DTurn on the crawler option to create a single schema for each S3 path
Show the answer and why
AConfigure the crawler to crawl only new folders on every later scheduled run
Incorrect
Incremental crawls add new partitions after the first full crawl. They do not change how the crawler groups schemas into tables.
BAdd an exclude pattern that skips the daily folders that have extra columns
Incorrect
Excluding the files hides data from the catalog. It would leave those days out of every query.
CRun MSCK REPAIR TABLE in Athena after every crawler run
Incorrect
MSCK REPAIR TABLE adds missing partitions to an existing table. It does not merge the separate tables the crawler already created.
DTurn on the crawler option to create a single schema for each S3 path
Correct
With this grouping option (CombineCompatibleSchemas), a crawler that finds compatible data under the include path creates one table with the combined columns and the folder as a partition key.
By default a crawler weighs schema similarity, so folders with different columns can become different tables. The single-schema option tells it to combine compatible data under the path into one partitioned table.
AWS documentation