Skip to content
BytePatterns

DEA-C01 · Domain 2: Data Store Management · 26% of the exam

Task 2.3: Manage the lifecycle of data

Moving, aging and deleting data: COPY and UNLOAD between S3 and Redshift, S3 Lifecycle transitions and expirations, versioning and DynamoDB TTL, deleting data for legal reasons, and keeping data durable and available.

Study it

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

Every night a single 40 GB gzip-compressed CSV file is loaded into an Amazon Redshift table with one COPY command. The load takes hours, and most slices sit idle during it. How can a data engineer make the load faster?

  1. ASplit the export into many files and start one concurrent COPY command for each file
  2. BSplit the export into similar-sized gzip files, a multiple of the slice count, and run one COPY
  3. CReplace the COPY command with batched INSERT statements sent from a client script
  4. DConvert the export to one gzip-compressed JSON file and keep a single COPY command for the load
Show the answer and why
  • ASplit the export into many files and start one concurrent COPY command for each file

    Incorrect

    Several concurrent COPY commands into one table force a serialized load, which is much slower and may need a VACUUM afterwards.

  • BSplit the export into similar-sized gzip files, a multiple of the slice count, and run one COPY

    Correct

    Gzip-compressed CSV cannot be split automatically, so one file loads serially. AWS recommends files of similar size, 1 MB to 1 GB after compression, in a multiple of the number of slices, loaded by one COPY.

  • CReplace the COPY command with batched INSERT statements sent from a client script

    Incorrect

    COPY loads large amounts of data much more efficiently than INSERT statements.

  • DConvert the export to one gzip-compressed JSON file and keep a single COPY command for the load

    Incorrect

    Compressed JSON is not split automatically either, so one file is still loaded serially.

Redshift loads in parallel across slices only when there is more than one piece to load. A splittable format or several similar-sized files, loaded with one COPY, keeps every slice busy.

Question 2 · choose 2

Raw files land in a versioning-enabled S3 bucket. After 90 days they are rarely read, but they must be kept for 7 years from creation and then be removed permanently, including any older versions created by overwrites. Storage cost should stay low. Which S3 Lifecycle configuration elements meet these requirements? (Choose TWO.)

  1. AA rule that transitions current versions to S3 Glacier Flexible Retrieval after 90 days and expires them after 7 years
  2. BA single expiration action at 7 years, which permanently deletes every version in a versioning-enabled bucket
  3. CSuspend versioning on the bucket so that the 7-year expiration action deletes the data
  4. DAn S3 Object Lock retention period of 7 years in compliance mode on every object
  5. EA NoncurrentVersionExpiration action that permanently deletes noncurrent versions after a set number of days
Show the answer and why
  • AA rule that transitions current versions to S3 Glacier Flexible Retrieval after 90 days and expires them after 7 years

    Correct

    Lifecycle transitions move objects to a cheaper class, and an expiration action ends the current version after the retention period.

  • BA single expiration action at 7 years, which permanently deletes every version in a versioning-enabled bucket

    Incorrect

    In a versioning-enabled bucket the expiration action adds a delete marker and makes the current version noncurrent. The data is kept.

  • CSuspend versioning on the bucket so that the 7-year expiration action deletes the data

    Incorrect

    Suspending versioning does not remove the versions already stored, and expiration in a suspended bucket still creates a delete marker.

  • DAn S3 Object Lock retention period of 7 years in compliance mode on every object

    Incorrect

    Object Lock prevents deletion for the retention period. It neither moves data to a cheaper class nor deletes it afterwards.

  • EA NoncurrentVersionExpiration action that permanently deletes noncurrent versions after a set number of days

    Correct

    In a versioned bucket, expiring a current version only adds a delete marker. NoncurrentVersionExpiration permanently deletes the noncurrent versions.

On a versioned bucket, plan for both kinds of versions: transition and expiration rules for the current version, and NoncurrentVersionExpiration so that overwritten and expired data really goes away.

Question 3 · choose 1

A DynamoDB table stores user sessions with an expires_at attribute in Unix epoch seconds, and Time to Live (TTL) is turned on for that attribute. Users report that the application still shows sessions that expired hours ago. Expired sessions must never be shown, and deletion should not use write capacity. What should a data engineer do?

  1. ALower the TTL deletion delay on the table to a few seconds
  2. BChange the expires_at attribute to an ISO 8601 date string
  3. CKeep TTL and add a filter expression that hides items whose expires_at has passed
  4. DRun a scheduled Lambda function that deletes expired items with DeleteItem calls
Show the answer and why
  • ALower the TTL deletion delay on the table to a few seconds

    Incorrect

    There is no such setting. DynamoDB deletes expired items at any time, typically within a few days of expiry.

  • BChange the expires_at attribute to an ISO 8601 date string

    Incorrect

    TTL needs a Number in Unix epoch time. A string attribute would stop TTL from deleting the items at all.

  • CKeep TTL and add a filter expression that hides items whose expires_at has passed

    Correct

    TTL deletes expired items within a few days and without consuming write throughput. Until then they can still be read, so AWS recommends filter expressions to drop them from results.

  • DRun a scheduled Lambda function that deletes expired items with DeleteItem calls

    Incorrect

    These deletes consume write capacity, which the requirement rules out, and there is still a gap between runs.

TTL is a free, background clean-up, not an exact expiry. Applications that must hide expired data filter on the timestamp themselves.

Question 4 · choose 1

A company wants to move 5 years of sales history from an Amazon Redshift table into its S3 data lake so that Athena can query it efficiently by year and region. What is the MOST efficient way to export the data?

  1. AUNLOAD the query result to S3 as CSV with PARALLEL OFF so that it lands in one file
  2. BTake a manual snapshot of the cluster and point Athena at the snapshot
  3. CUNLOAD the result to S3 with FORMAT AS PARQUET and PARTITION BY (year, region)
  4. DRun the query in the query editor, download the result, and upload it to S3
Show the answer and why
  • AUNLOAD the query result to S3 as CSV with PARALLEL OFF so that it lands in one file

    Incorrect

    A serial unload writes one row-based file, which Athena must scan completely for every query.

  • BTake a manual snapshot of the cluster and point Athena at the snapshot

    Incorrect

    Snapshots are point-in-time backups for restoring a cluster. They are not files in the data lake that Athena can query.

  • CUNLOAD the result to S3 with FORMAT AS PARQUET and PARTITION BY (year, region)

    Correct

    UNLOAD writes a query result to S3 as text, JSON or Apache Parquet, and PARTITION BY places the files in folders by the given columns.

  • DRun the query in the query editor, download the result, and upload it to S3

    Incorrect

    This is manual, limited by one client, and produces an unpartitioned file.

UNLOAD is the bulk path out of Redshift. Writing Parquet partitioned by the columns that queries filter on gives Athena a lake layout it can prune.

Question 5 · choose 1

A bucket holds files from many pipelines under mixed prefixes. Every temporary file carries the object tag classification=temp. Temporary files must be deleted 7 days after creation, and all other files must be kept. What should a data engineer configure?

  1. AA Lifecycle rule with a prefix filter and a 7-day expiration
  2. BS3 Object Lock with a 7-day retention period on the bucket
  3. CS3 Intelligent-Tiering for every object in the bucket
  4. DA Lifecycle rule with a tag filter and a 7-day expiration
Show the answer and why
  • AA Lifecycle rule with a prefix filter and a 7-day expiration

    Incorrect

    The temporary files are spread across mixed prefixes, so a prefix filter cannot select only them.

  • BS3 Object Lock with a 7-day retention period on the bucket

    Incorrect

    Object Lock prevents objects from being deleted or overwritten during retention. It does not delete them afterward.

  • CS3 Intelligent-Tiering for every object in the bucket

    Incorrect

    Intelligent-Tiering moves objects between access tiers to save cost. It does not delete objects.

  • DA Lifecycle rule with a tag filter and a 7-day expiration

    Correct

    Lifecycle rules can filter on object tags, so the expiration applies only to objects tagged classification=temp.

Tags let lifecycle rules follow what an object is rather than where it lives, which matters when prefixes are shared.

Question 6 · choose 1

A data engineer added a replication rule to a bucket that already holds 50 million objects. New objects now replicate, but the existing objects do not. All objects must reach the destination bucket. What should the data engineer do?

  1. ARun an S3 Batch Replication job for the existing objects
  2. BTurn on S3 Replication Time Control for the replication rule
  3. CTurn on an S3 Inventory report for the source bucket
  4. DTurn on S3 Transfer Acceleration for the source bucket
Show the answer and why
  • ARun an S3 Batch Replication job for the existing objects

    Correct

    Live replication handles new objects. Batch Replication replicates existing objects on demand, including objects that existed before the rule.

  • BTurn on S3 Replication Time Control for the replication rule

    Incorrect

    Replication Time Control sets a time target for replicating new objects. It does not reach objects that existed before the rule.

  • CTurn on an S3 Inventory report for the source bucket

    Incorrect

    S3 Inventory lists objects and their metadata. It does not copy them.

  • DTurn on S3 Transfer Acceleration for the source bucket

    Incorrect

    Transfer Acceleration speeds up transfers between clients and a bucket over long distances. It does not replicate objects.

Replication rules look forward. Objects that predate the rule, or that failed to replicate, need a Batch Replication job.

Question 7 · choose 1

Files land in an Amazon S3 prefix throughout the day. They must be loaded into an Amazon Redshift table automatically as they arrive, without building an ingestion pipeline and without loading the same file twice. What should a data engineer set up?

  1. AAn external table that Redshift Spectrum queries in place
  2. BA query editor v2 scheduled query that runs COPY every 5 minutes
  3. CAn UNLOAD command that runs each time a file arrives in the prefix
  4. DAn S3 event integration with an auto-copy job (COPY JOB)
Show the answer and why
  • AAn external table that Redshift Spectrum queries in place

    Incorrect

    Spectrum queries files in S3 without loading them into Redshift tables, and the requirement is to load the table.

  • BA query editor v2 scheduled query that runs COPY every 5 minutes

    Incorrect

    A scheduled query runs on a timer and leaves it to the team to work out which files are new.

  • CAn UNLOAD command that runs each time a file arrives in the prefix

    Incorrect

    UNLOAD writes query results from Redshift to S3, the opposite direction.

  • DAn S3 event integration with an auto-copy job (COPY JOB)

    Correct

    An auto-copy job runs COPY automatically for new files without an external pipeline, and Redshift keeps track of which files it has loaded.

Auto-copy turns S3 arrivals into tracked COPY runs; the team writes no scheduling or bookkeeping code.

Question 8 · choose 2

A Lifecycle rule moved last year's partitions of an Amazon Athena Hive table to S3 Glacier Flexible Retrieval. An auditor now needs to query those partitions with Athena engine version 3, and the query returns no rows from them. Which TWO actions should a data engineer take? (Choose TWO.)

  1. ARun MSCK REPAIR TABLE on the table to reload its partitions
  2. BRestore the archived objects in Amazon S3
  3. CSet the table property read_restored_glacier_objects to true
  4. DTurn on partition projection for the table's year column
  5. ERaise the workgroup's per-query data usage limit
Show the answer and why
  • ARun MSCK REPAIR TABLE on the table to reload its partitions

    Incorrect

    MSCK REPAIR TABLE adds partitions found in S3 to the catalog. The archived objects would still be skipped.

  • BRestore the archived objects in Amazon S3

    Correct

    Athena does not restore archived objects; they must be restored before they can be queried.

  • CSet the table property read_restored_glacier_objects to true

    Correct

    Without this property, Athena skips the table's Glacier Flexible Retrieval and Deep Archive objects, even after they are restored.

  • DTurn on partition projection for the table's year column

    Incorrect

    Partition projection calculates partition locations from table properties. It does not make archived objects readable.

  • ERaise the workgroup's per-query data usage limit

    Incorrect

    Data usage limits cap how much a query can scan. They do not change which objects Athena reads.

Querying archives takes two steps: restore the objects in S3, then tell Athena to read restored objects through the table property.

Question 9 · choose 1

Analysts need a full copy of a 2 TB Amazon DynamoDB table in Amazon S3 every day to query with Athena. The copy must not consume the table's read capacity or affect its performance. What should a data engineer use?

  1. AAn AWS Glue job that reads the table with dynamodb.throughput.read.percent
  2. BDynamoDB export to Amazon S3, with point-in-time recovery turned on
  3. CDynamoDB Streams with a Lambda function that writes each change to S3
  4. DA DynamoDB import from S3 that runs on a daily schedule
Show the answer and why
  • AAn AWS Glue job that reads the table with dynamodb.throughput.read.percent

    Incorrect

    That setting is the share of the table's read capacity the job uses, so the job consumes read capacity.

  • BDynamoDB export to Amazon S3, with point-in-time recovery turned on

    Correct

    Exports use point-in-time recovery data, are asynchronous, do not consume read capacity units, and do not affect table performance.

  • CDynamoDB Streams with a Lambda function that writes each change to S3

    Incorrect

    A stream captures item-level modifications. It does not produce a full copy of the existing table.

  • DA DynamoDB import from S3 that runs on a daily schedule

    Incorrect

    Import from S3 creates a new DynamoDB table from S3 data, the opposite direction.

Exports are the hands-off way to snapshot DynamoDB into the data lake: they read from PITR data instead of the live table.

Practise domain 2 →Practise all domains →