Skip to content
BytePatterns

DEA-C01 · Domain 2: Data Store Management · 26% of the exam

Task 2.2: Understand data cataloging systems

Knowing what data exists and where: the Glue Data Catalog and Hive metastores, crawlers that discover schemas, keeping partitions in step with the catalog, connections for sources and targets, and business catalogs in Amazon SageMaker Catalog.

Study it

  • The Glue Data Catalog: crawlers, partitions and connections

    Lesson coming

  • Business catalogs: Amazon SageMaker Catalog

    Lesson coming

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

An AWS Glue crawler runs on s3://example-bucket/orders/, where Parquet files are stored under one prefix per day (dt=2026-10-01/ and so on). Some days have extra columns. Instead of one table partitioned by dt, the crawler creates a separate table for several days. What should a data engineer change?

  1. AConfigure the crawler to crawl only new folders on every later scheduled run
  2. BAdd an exclude pattern that skips the daily folders that have extra columns
  3. CRun MSCK REPAIR TABLE in Athena after every crawler run
  4. DTurn on the crawler option to create a single schema for each S3 path
Show the answer and why
  • AConfigure the crawler to crawl only new folders on every later scheduled run

    Incorrect

    Incremental crawls add new partitions after the first full crawl. They do not change how the crawler groups schemas into tables.

  • BAdd an exclude pattern that skips the daily folders that have extra columns

    Incorrect

    Excluding the files hides data from the catalog. It would leave those days out of every query.

  • CRun MSCK REPAIR TABLE in Athena after every crawler run

    Incorrect

    MSCK REPAIR TABLE adds missing partitions to an existing table. It does not merge the separate tables the crawler already created.

  • DTurn on the crawler option to create a single schema for each S3 path

    Correct

    With this grouping option (CombineCompatibleSchemas), a crawler that finds compatible data under the include path creates one table with the combined columns and the folder as a partition key.

By default a crawler weighs schema similarity, so folders with different columns can become different tables. The single-schema option tells it to combine compatible data under the path into one partitioned table.

Question 2 · choose 2

A job writes a new Hive-style partition (dt=YYYY-MM-DD) to an Amazon S3 table location every day. The table is defined in the AWS Glue Data Catalog, and Athena queries do not return data from the new days. The team does not want to run a crawler. Which actions will make the new partitions queryable? (Choose TWO.)

  1. ATurn on query result reuse in the analysts' Athena workgroup
  2. BCreate an Athena view that selects all columns from the table
  3. CConfigure partition projection on the table for the dt column
  4. DRun MSCK REPAIR TABLE after each daily write
  5. ETurn on S3 Versioning for the bucket that holds the table data
Show the answer and why
  • ATurn on query result reuse in the analysts' Athena workgroup

    Incorrect

    Result reuse returns stored results of an earlier identical query. It cannot show partitions that the table metadata does not contain.

  • BCreate an Athena view that selects all columns from the table

    Incorrect

    A view is a stored query that runs against the same table. It sees the same partitions the table has.

  • CConfigure partition projection on the table for the dt column

    Correct

    With partition projection, Athena calculates partition values and locations from table properties, so new partitions need no metadata updates at all.

  • DRun MSCK REPAIR TABLE after each daily write

    Correct

    MSCK REPAIR TABLE scans the table location for Hive-compatible partitions added after the table was created and adds them to the metadata.

  • ETurn on S3 Versioning for the bucket that holds the table data

    Incorrect

    Versioning keeps earlier versions of objects. It does not register partitions in the Data Catalog.

A Hive-style table only reads the partitions its metadata knows about. Either add them (MSCK REPAIR TABLE or ALTER TABLE ADD PARTITION) or let Athena project them from a rule.

Question 3 · choose 1

A team runs Apache Hive and Spark on transient Amazon EMR clusters that terminate after each job. Table definitions disappear with every cluster, and the same tables must also be queryable from Amazon Athena. What should a data engineer do?

  1. AKeep the default metastore and turn on termination protection for each cluster
  2. BGive the primary node a larger EBS volume for the MySQL metastore database
  3. CRecreate the tables with a bootstrap action script when each cluster starts
  4. DConfigure the clusters to use the AWS Glue Data Catalog as the Hive metastore
Show the answer and why
  • AKeep the default metastore and turn on termination protection for each cluster

    Incorrect

    Termination protection guards against accidental termination. These clusters are meant to terminate, and the default metastore is still local to one cluster.

  • BGive the primary node a larger EBS volume for the MySQL metastore database

    Incorrect

    By default the metastore is a MySQL database on the primary node, and local data is lost when the cluster terminates, whatever its size.

  • CRecreate the tables with a bootstrap action script when each cluster starts

    Incorrect

    A script rebuilds definitions on each cluster but keeps no shared catalog, so Athena still has nothing to query.

  • DConfigure the clusters to use the AWS Glue Data Catalog as the Hive metastore

    Correct

    From EMR release 5.8.0, Hive and Spark can use the Glue Data Catalog as their metastore. It lives outside the cluster and is the catalog that Athena uses.

A metastore that must outlive clusters and be shared with other engines has to be external. The Glue Data Catalog is that shared catalog for EMR, Athena and Redshift Spectrum.

Question 4 · choose 1

Analysts in several business units cannot tell which tables hold "active customer" data, and each unit defines the term differently. The company wants data owners to publish curated assets with agreed business terms, analysts to search by those terms and request access, and owners to approve each request. Which solution meets these requirements?

  1. ATable and column descriptions in the AWS Glue Data Catalog, filled in by the data owners
  2. BAmazon SageMaker Catalog, with a business glossary and subscription requests
  3. CLF-Tags in AWS Lake Formation, with one tag value per business unit on each table
  4. DAmazon Macie classification jobs that label the tables containing customer data
Show the answer and why
  • ATable and column descriptions in the AWS Glue Data Catalog, filled in by the data owners

    Incorrect

    The Glue Data Catalog is a technical catalog of databases and tables. It has no glossary of business terms or request-and-approve workflow.

  • BAmazon SageMaker Catalog, with a business glossary and subscription requests

    Correct

    A business glossary keeps shared terms and attaches them to assets and columns for search. Owners publish assets, and a project member requests a subscription that the owner approves or rejects.

  • CLF-Tags in AWS Lake Formation, with one tag value per business unit on each table

    Incorrect

    LF-Tags grant permissions on catalog resources. They do not provide business definitions, search by term, or access requests.

  • DAmazon Macie classification jobs that label the tables containing customer data

    Incorrect

    Macie finds sensitive data such as PII in S3. It does not publish assets or handle access requests.

A business catalog sits above the technical one: it adds agreed meaning, ownership and a request-and-approve path to data that the technical catalog already describes.

Question 5 · choose 1

An AWS Glue crawler uses a JDBC connection to a PostgreSQL database named erp. The crawler must catalog every table in the sales schema and nothing else. Which include path should a data engineer set?

  1. Aerp/sales/%
  2. B%/sales/%
  3. Cerp/public/%
  4. Dsales/%
Show the answer and why
  • Aerp/sales/%

    Correct

    For engines with schemas, the path is database/schema/table, and the percent sign stands for all tables in that schema.

  • B%/sales/%

    Incorrect

    The percent sign can stand for a schema or a table, but not for the database.

  • Cerp/public/%

    Incorrect

    This path names the public schema, so the crawler would catalog the tables there instead of the sales schema.

  • Dsales/%

    Incorrect

    The path starts with the database name, so this would look for a database named sales.

JDBC include paths follow database/schema/table (or database/table for engines without schemas); % can replace a schema or table, never the database.

Question 6 · choose 2

An Amazon Athena table uses the OpenX JSON SerDe. Queries fail because a few lines in the files are not valid JSON, and some keys contain dots, such as device.model, which Athena does not allow in column names. Which TWO SerDe properties should a data engineer set? (Choose TWO.)

  1. A"ignore.malformed.json" = "TRUE"
  2. B"case.insensitive" = "FALSE"
  3. C"mapping.device_model" = "device_model"
  4. D"dots.in.keys" = "TRUE"
  5. ERecreate the table with the Lazy Simple SerDe
Show the answer and why
  • A"ignore.malformed.json" = "TRUE"

    Correct

    When set to TRUE, the SerDe skips malformed JSON syntax instead of failing the query.

  • B"case.insensitive" = "FALSE"

    Incorrect

    This makes key matching case sensitive, which then needs a mapping for each mixed-case key. It does not handle dots or bad lines.

  • C"mapping.device_model" = "device_model"

    Incorrect

    A mapping property ties a column to a key with a different name, but here it names the same string twice and skips no bad lines.

  • D"dots.in.keys" = "TRUE"

    Correct

    When set to TRUE, the SerDe replaces dots in key names with underscores, so device.model can be a column named device_model.

  • ERecreate the table with the Lazy Simple SerDe

    Incorrect

    The Lazy Simple SerDe is for CSV, TSV, and custom-delimited files, not JSON.

The OpenX JSON SerDe has properties for messy JSON: ignore.malformed.json for bad lines, dots.in.keys for dotted keys, and case.insensitive with mapping for mixed-case keys.

Question 7 · choose 1

A team stores Apache Iceberg tables in several Amazon S3 table buckets and creates new table buckets often. Analysts must see all current and future tables in Amazon Athena and Amazon Redshift without per-bucket setup. What should a data engineer do?

  1. ACreate an AWS Glue crawler for each new table bucket and run it hourly
  2. BCreate a Glue database by hand for each namespace in every table bucket
  3. CIntegrate the table buckets with AWS analytics services in the Region
  4. DTurn on table maintenance, such as compaction, for every table bucket
Show the answer and why
  • ACreate an AWS Glue crawler for each new table bucket and run it hourly

    Incorrect

    After the integration, current and future table buckets are added to the Data Catalog automatically, so per-bucket crawlers are the setup the team wants to avoid.

  • BCreate a Glue database by hand for each namespace in every table bucket

    Incorrect

    The integration populates each namespace as a database automatically; creating them by hand is per-bucket work.

  • CIntegrate the table buckets with AWS analytics services in the Region

    Correct

    The integration adds the s3tablescatalog to the AWS Glue Data Catalog, and all current and future table buckets in that Region appear there.

  • DTurn on table maintenance, such as compaction, for every table bucket

    Incorrect

    Maintenance operations improve the management and performance of individual tables. They do not make tables visible to query engines.

S3 Tables reach Athena, Redshift, EMR, and other services through one Regional integration with the Glue Data Catalog, not through crawlers.

Question 8 · choose 1

In Amazon SageMaker Catalog, a producer team publishes six related tables for sales analytics. Consumers should find the tables as one unit and get access to all of them with a single request. What should the producer team do?

  1. APublish each table as an asset that consumers subscribe to one by one
  2. BAttach one business glossary term to all six tables
  3. CGroup the six assets into one data product and publish it
  4. DAdd a metadata enforcement rule for the producer's domain unit
Show the answer and why
  • APublish each table as an asset that consumers subscribe to one by one

    Incorrect

    Separate assets mean separate subscription requests, which data products are designed to avoid.

  • BAttach one business glossary term to all six tables

    Incorrect

    A glossary term classifies assets and helps search, but each table still needs its own access request.

  • CGroup the six assets into one data product and publish it

    Correct

    Data products group related assets into one package that consumers find as a single unit and access with a single request.

  • DAdd a metadata enforcement rule for the producer's domain unit

    Incorrect

    Enforcement rules set metadata requirements for publishing. They do not bundle assets or their access.

Data products are the catalog's unit for business-aligned bundles: one listing, one subscription, one access model.

Question 9 · choose 1

Hundreds of assets in Amazon SageMaker Catalog have cryptic column names and no business descriptions. Data stewards want suggested names and descriptions that they can review and approve instead of writing every one by hand. What should a data engineer use?

  1. AAI recommendations for names and descriptions
  2. BA custom classifier attached to the AWS Glue crawler
  3. CAWS Glue Data Catalog column statistics
  4. DA metadata enforcement rule for publishing
Show the answer and why
  • AAI recommendations for names and descriptions

    Correct

    SageMaker Unified Studio can generate business names and descriptions for assets, which users can edit, accept, or reject.

  • BA custom classifier attached to the AWS Glue crawler

    Incorrect

    A classifier recognizes the format of data and generates a schema. It does not write business descriptions.

  • CAWS Glue Data Catalog column statistics

    Incorrect

    Column statistics hold values such as null and distinct counts for query planning, not business descriptions.

  • DA metadata enforcement rule for publishing

    Incorrect

    An enforcement rule requires metadata before publishing, but someone still has to write it.

Generated metadata speeds up cataloging, and the review step keeps a steward accountable for what consumers read.

Practise domain 2 →Practise all domains →