Question 1 · choose 1
A data platform account holds a data lake in Amazon S3 that is cataloged in the AWS Glue Data Catalog. Analysts in 30 other accounts of the same organization query it with Amazon Athena from their own accounts. Today each team receives a replicated copy of the S3 data it needs. A new governance standard requires that the data is not duplicated anywhere, that access follows classification labels that data stewards assign to databases, tables and individual columns, that a table classified later reaches the right teams without any new grant, and that each consumer account's own data lake administrator decides which of its users get the shared data. Which solution meets these requirements?
- AKeep replicating the S3 data into a bucket in each consumer account and have every account catalog its copy with its own Glue crawler
- BGrant the consumer accounts access to the tables in a Data Catalog resource policy and to the data prefixes in the S3 bucket policy
- CLoad the tables into an Amazon Redshift cluster in the platform account and share them with each consumer account through Redshift datashares
- DRegister the S3 locations with Lake Formation, assign LF-Tags to the catalog resources, and grant tag-based permissions to the accounts
Show the answer and why
AKeep replicating the S3 data into a bucket in each consumer account and have every account catalog its copy with its own Glue crawler
Incorrect
S3 Replication copies objects asynchronously into other buckets, including buckets owned by other accounts, so every consumer account would keep its own duplicate of the data, which the standard forbids.
BGrant the consumer accounts access to the tables in a Data Catalog resource policy and to the data prefixes in the S3 bucket policy
Incorrect
Catalog resource policies and bucket policies grant access to named Data Catalog resources and S3 objects. The grants do not follow classification labels or reach individual columns, and every newly classified table would mean editing both policies.
CLoad the tables into an Amazon Redshift cluster in the platform account and share them with each consumer account through Redshift datashares
Incorrect
Datashares give other accounts live access without copying data between warehouses, but loading the lake with COPY puts a second copy of the data into Redshift tables, and the analysts would no longer query the lake with Athena.
DRegister the S3 locations with Lake Formation, assign LF-Tags to the catalog resources, and grant tag-based permissions to the accounts
Correct
LF-Tags can be assigned to databases, tables and columns, and a grant on an LF-Tag expression covers every resource with a matching tag, including tables tagged later. Cross-account grants keep the data in place, the consumer's data lake administrator grants the shared data to its own principals, and Athena queries it through resource links.
The constraints are no duplicate data, access by classification down to the column, automatic coverage of newly classified tables, and consumer-side control of who sees the shared data. Replication duplicates the lake, Redshift loading duplicates it too, and catalog and bucket policies grant access by resource name rather than by label. Lake Formation tag-based access control shares the catalog resources across accounts by LF-Tag, and each consumer's data lake administrator grants them on to its own analysts.
AWS documentation
- Adding an Amazon S3 location to your data lake (opens in a new tab)
- Lake Formation tag-based access control (opens in a new tab)
- Cross-account data sharing in Lake Formation (opens in a new tab)
- Granting permissions on a database or table shared with your account (opens in a new tab)
- How resource links work in Lake Formation (opens in a new tab)
- Replicating objects within and across Regions (opens in a new tab)
- Granting cross-account access (opens in a new tab)
- Data sharing in Amazon Redshift (opens in a new tab)
- COPY (opens in a new tab)