Skip to content
BytePatterns

DEA-C01 · Domain 4: Data Security and Governance · 18% of the exam

Task 4.5: Understand data privacy and governance

Rules about where data lives and who sees it: Redshift data sharing, finding PII with Macie, keeping backups and copies out of disallowed Regions, configuration history in AWS Config, data sovereignty, and sharing through SageMaker Catalog projects.

Study it

  • SageMaker Unified Studio domains, domain units and projects

    Lesson coming

  • Governance: Redshift data sharing, Macie, Region controls, AWS Config and data sovereignty

    Lesson coming

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

A producer Amazon Redshift cluster in account A holds a sales schema. An analytics team in account B uses Redshift Serverless and needs read access to the live sales data without copying it or building a pipeline. What should a data engineer set up?

  1. AA datashare in account A that is authorized for account B
  2. BAn UNLOAD of the schema to S3 followed by a COPY in account B
  3. CA manual snapshot shared with account B and restored there
  4. DA federated query from account B that points to the producer cluster
Show the answer and why
  • AA datashare in account A that is authorized for account B

    Correct

    Data sharing gives live access to data across clusters, workgroups and accounts without copying or moving it; the consumer account associates the datashare.

  • BAn UNLOAD of the schema to S3 followed by a COPY in account B

    Incorrect

    This copies the data, and the copy is only as fresh as the last run.

  • CA manual snapshot shared with account B and restored there

    Incorrect

    A restored snapshot is a separate copy of the data as of the snapshot time, not live data.

  • DA federated query from account B that points to the producer cluster

    Incorrect

    Federated queries reach PostgreSQL and MySQL databases on RDS and Aurora, not another Redshift cluster.

Redshift data sharing separates storage from consumers: the producer shares objects, and consumers in other accounts query them live.

Question 2 · choose 1

A company has about 300 S3 buckets in its data lake. The privacy team must find out which buckets and objects contain personal data such as names and passport numbers, and keep that view current as data changes, without building custom code. Which service should a data engineer use?

  1. AAWS Glue crawlers with custom classifiers on each bucket
  2. BCloudTrail data events for all S3 buckets
  3. CAWS Glue DataBrew profile jobs on each bucket
  4. DAmazon Macie with automated sensitive data discovery
Show the answer and why
  • AAWS Glue crawlers with custom classifiers on each bucket

    Incorrect

    Crawlers infer schemas for the Data Catalog. They do not identify personal data inside the objects.

  • BCloudTrail data events for all S3 buckets

    Incorrect

    Data events record who accessed objects. They do not look at what the objects contain.

  • CAWS Glue DataBrew profile jobs on each bucket

    Incorrect

    Profile jobs report statistics such as missing values and distributions for one dataset; they are not a managed scan of 300 buckets for PII.

  • DAmazon Macie with automated sensitive data discovery

    Correct

    Macie discovers sensitive data such as PII with machine learning and pattern matching; automated discovery gives broad, ongoing visibility across buckets.

Finding PII across a large S3 estate is Macie's purpose: automated discovery for broad coverage, targeted discovery jobs for depth.

Question 3 · choose 1

Amazon Macie scans the company's data lake buckets. Internal employee IDs, such as EMP-482913, count as sensitive data, but Macie's managed data identifiers do not detect them. What should a data engineer configure?

  1. AA Macie custom data identifier with a regex and keywords
  2. BA Macie allow list that contains the EMP- employee ID text pattern
  3. CA CloudWatch Logs data protection policy for the IDs
  4. DAn AWS Config managed rule on the data lake buckets
Show the answer and why
  • AA Macie custom data identifier with a regex and keywords

    Correct

    A custom data identifier uses a regular expression, optionally refined with keywords and a proximity rule, to detect sensitive data that you define.

  • BA Macie allow list that contains the EMP- employee ID text pattern

    Incorrect

    Allow lists define text that Macie should ignore, the opposite of what is needed.

  • CA CloudWatch Logs data protection policy for the IDs

    Incorrect

    These policies mask sensitive data in log groups, not in S3 objects that Macie inspects.

  • DAn AWS Config managed rule on the data lake buckets

    Incorrect

    Config rules evaluate resource configurations, not the contents of objects.

Managed identifiers cover common data types; company-specific formats need a custom data identifier, and allow lists handle the false positives.

Question 4 · choose 1

When Amazon Macie produces a new sensitive data finding for a data lake bucket, a Lambda function must tag the bucket and the governance team must be notified, with no manual steps. What should a data engineer set up?

  1. AS3 Event Notifications on the bucket for every new object upload
  2. BCloudTrail data events for the objects in the data lake bucket
  3. CAn EventBridge rule for Macie findings, targeting Lambda and SNS
  4. DA Macie custom data identifier for each type of sensitive data
Show the answer and why
  • AS3 Event Notifications on the bucket for every new object upload

    Incorrect

    These fire when objects are created, not when Macie reports sensitive data.

  • BCloudTrail data events for the objects in the data lake bucket

    Incorrect

    Data events record object-level API activity. They do not carry Macie's findings.

  • CAn EventBridge rule for Macie findings, targeting Lambda and SNS

    Correct

    Macie publishes finding events to EventBridge, which can send specific types of findings to services such as Lambda for automated processing.

  • DA Macie custom data identifier for each type of sensitive data

    Incorrect

    Custom data identifiers define what Macie detects. They do not act on findings.

Detection and response are separate: Macie finds, EventBridge routes, and targets such as Lambda and SNS act.

Question 5 · choose 1

A governance team must continuously check whether any S3 bucket in an account allows public read access and see the compliance status of each bucket. Which solution meets this requirement?

  1. ACloudTrail Insights events for API calls to the Amazon S3 service
  2. BS3 server access logging turned on for every bucket
  3. CVPC Flow Logs for the subnets that access Amazon S3
  4. DThe AWS Config managed rule s3-bucket-public-read-prohibited
Show the answer and why
  • ACloudTrail Insights events for API calls to the Amazon S3 service

    Incorrect

    Insights flags unusual API call or error rates. It does not evaluate bucket configurations.

  • BS3 server access logging turned on for every bucket

    Incorrect

    Server access logs record requests to a bucket. They do not report whether its settings allow public access.

  • CVPC Flow Logs for the subnets that access Amazon S3

    Incorrect

    Flow logs capture IP traffic for network interfaces, not bucket access settings.

  • DThe AWS Config managed rule s3-bucket-public-read-prohibited

    Correct

    This managed rule checks Block Public Access settings, bucket policies, and ACLs and reports whether each bucket allows public read access.

Config rules turn policies into continuous compliance checks; managed rules cover common cases like public buckets.

Question 6 · choose 1

A compliance team needs one view of AWS Config configuration and compliance data for all 40 accounts in the company's organization across three Regions. What should a data engineer set up?

  1. AAn organization trail in AWS CloudTrail for all accounts
  2. BA conformance pack deployed to each account
  3. CAn AWS Config aggregator for the organization
  4. DThe Config resource timeline in each account
Show the answer and why
  • AAn organization trail in AWS CloudTrail for all accounts

    Incorrect

    An organization trail records API activity across accounts. It does not collect Config compliance data.

  • BA conformance pack deployed to each account

    Incorrect

    Conformance packs deploy rules and remediation actions. They do not combine results from many accounts into one view.

  • CAn AWS Config aggregator for the organization

    Correct

    An aggregator collects Config configuration and compliance data from an organization's accounts and Regions into one view.

  • DThe Config resource timeline in each account

    Incorrect

    The timeline shows one resource's history in its own account, not a combined view of all accounts.

Rules evaluate, conformance packs deploy rules at scale, and aggregators bring the results together across accounts and Regions.

Question 7 · choose 1

A company must make sure that no S3 bucket in a data lake account can be made public, including buckets created later and even if someone adds a public bucket policy. What should a data engineer do?

  1. ATurn on all S3 Block Public Access settings at the account level
  2. BAdd a bucket policy that denies public access to each existing bucket
  3. CAdd the Config rule s3-bucket-public-read-prohibited to the account
  4. DSet Object Ownership to bucket owner enforced on every bucket
Show the answer and why
  • ATurn on all S3 Block Public Access settings at the account level

    Correct

    Block Public Access settings override bucket policies and permissions that would allow public access, and account-level settings cover the account's buckets.

  • BAdd a bucket policy that denies public access to each existing bucket

    Incorrect

    Per-bucket policies do not cover buckets created later, and they can be edited by the same people the control must stop.

  • CAdd the Config rule s3-bucket-public-read-prohibited to the account

    Incorrect

    The rule reports whether buckets allow public read access. It detects the problem rather than preventing it.

  • DSet Object Ownership to bucket owner enforced on every bucket

    Incorrect

    This disables ACLs, but a bucket policy could still grant public access.

Prevention beats detection for public data: account-level Block Public Access overrides any policy or ACL that would make a bucket public.

Practise domain 4 →Practise all domains →