Skip to content
BytePatterns

SAA-C03 · Domain 3: Design High-Performing Architectures · 24% of the exam

Task 3.5: Determine high-performing data ingestion and transformation solutions.

Getting data in and making it useful: streaming versus batch ingestion, transfer services, data lakes on S3, format conversion, and query and visualization services.

Study it

  • Hybrid storage: Storage Gateway and DataSync

    Lesson coming

  • Streaming ingestion: Kinesis Data Streams, Amazon Data Firehose and Amazon MSK

    Lesson coming

  • Data lakes: S3, Glue, Athena, Lake Formation, EMR and columnar formats

    Lesson coming

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

A company collects JSON clickstream events from its website and wants them in Amazon S3 as Apache Parquet files within a few minutes of arrival, for querying with Athena. The team wants a fully managed solution with no consumer code to write or run. Which solution meets these requirements?

  1. AAmazon Kinesis Data Streams with consumer applications on EC2 that write Parquet to S3
  2. BAmazon SQS with a Lambda function that writes every event to S3 as its own JSON object
  3. CAmazon Data Firehose with record format conversion to Parquet, using a Glue schema
  4. DAWS DataSync with a task that copies the web servers' log folders to S3 every hour
Show the answer and why
  • AAmazon Kinesis Data Streams with consumer applications on EC2 that write Parquet to S3

    Incorrect

    Kinesis Data Streams needs consumer applications that the team builds and runs, which the requirement rules out.

  • BAmazon SQS with a Lambda function that writes every event to S3 as its own JSON object

    Incorrect

    This writes one small JSON object per event, not Parquet, and it is custom consumer code.

  • CAmazon Data Firehose with record format conversion to Parquet, using a Glue schema

    Correct

    Firehose delivers streaming data to S3 and can convert JSON to Parquet, using a schema from the AWS Glue Data Catalog.

  • DAWS DataSync with a task that copies the web servers' log folders to S3 every hour

    Incorrect

    DataSync copies files as they are on a schedule. It neither converts them to Parquet nor delivers within minutes.

Managed delivery to S3 plus JSON-to-Parquet conversion is a built-in Firehose feature.

Question 2 · choose 2

Analysts use Amazon Athena to query three years of application logs stored in S3 as gzip-compressed CSV files under one prefix. Most queries filter on a single day and read only a few columns, yet every query scans the whole dataset. Which changes will reduce query cost and run time? (Choose TWO.)

  1. AMove objects older than 90 days to S3 Glacier Deep Archive
  2. BTurn on S3 Transfer Acceleration for the logs bucket
  3. CConvert the data to Apache Parquet with an AWS Glue job
  4. DPartition the data by date and filter on the partition key
  5. ESave the logs as plain CSV files without gzip compression
Show the answer and why
  • AMove objects older than 90 days to S3 Glacier Deep Archive

    Incorrect

    Athena cannot read archived objects until they are restored, so this breaks queries on older data instead of speeding them up.

  • BTurn on S3 Transfer Acceleration for the logs bucket

    Incorrect

    Transfer Acceleration speeds up long-distance transfers to and from a bucket; it does not reduce the data a query scans.

  • CConvert the data to Apache Parquet with an AWS Glue job

    Correct

    In a columnar format such as Parquet, Athena reads only the columns a query needs.

  • DPartition the data by date and filter on the partition key

    Correct

    When a query filters on partition key columns, Athena reads only the matching partitions.

  • ESave the logs as plain CSV files without gzip compression

    Incorrect

    Removing compression makes Athena read more bytes, which raises both cost and run time.

Read less data: partitions cut the rows, a columnar format cuts the columns.

Question 3 · choose 1

A company is building a data lake on Amazon S3, cataloged in the AWS Glue Data Catalog. Analysts in several AWS accounts query it with Athena. The security team must grant access at the database, table and column level, for example hiding salary columns from most analysts, and manage those grants in one place. Which solution meets these requirements?

  1. AWrite S3 bucket policies that let each analyst role read only certain object prefixes
  2. BUse AWS Lake Formation permissions to grant analysts table- and column-level access
  3. CRun Amazon Macie on the data lake buckets and block analysts from objects it flags
  4. DKeep a copy of each table without the salary columns in a separate bucket per account
Show the answer and why
  • AWrite S3 bucket policies that let each analyst role read only certain object prefixes

    Incorrect

    Bucket policies grant access to buckets and objects. They cannot hide individual columns inside the data files.

  • BUse AWS Lake Formation permissions to grant analysts table- and column-level access

    Correct

    Lake Formation manages fine-grained grants down to column, row and cell level, enforces them in Athena, and shares data across accounts.

  • CRun Amazon Macie on the data lake buckets and block analysts from objects it flags

    Incorrect

    Macie discovers sensitive data and reports findings. It does not enforce query permissions on tables and columns.

  • DKeep a copy of each table without the salary columns in a separate bucket per account

    Incorrect

    Copies multiply storage and pipelines and spread the grants across many buckets instead of managing them in one place.

Column-level, cross-account grants managed centrally for a Glue-cataloged lake is what Lake Formation does.

Question 4 · choose 1

A ride-sharing app sends driver location updates at about 20,000 records per second. A pricing service, a fraud service and an archive service must each read every record independently within a second of arrival, and the pricing service must be able to reprocess the last 24 hours after a bug fix. Which ingestion service fits best?

  1. AAmazon Kinesis Data Streams, with each service reading the stream as its own consumer
  2. BAn Amazon SQS standard queue that the three services poll for new records
  3. CAmazon Data Firehose delivering the records to an S3 bucket that the services read
  4. DAn Amazon SNS standard topic with each service subscribed through an HTTPS endpoint and keeping its own copy
Show the answer and why
  • AAmazon Kinesis Data Streams, with each service reading the stream as its own consumer

    Correct

    A data stream keeps records for at least 24 hours, lets several consumers read the same records independently in real time, and lets a consumer read again from an earlier point.

  • BAn Amazon SQS standard queue that the three services poll for new records

    Incorrect

    A queue hands each message to one consumer, which deletes it, so the services would compete for records, and nothing can be replayed.

  • CAmazon Data Firehose delivering the records to an S3 bucket that the services read

    Incorrect

    Firehose buffers records and delivers them to destinations such as S3. The services would read files after delivery, not each record within a second.

  • DAn Amazon SNS standard topic with each service subscribed through an HTTPS endpoint and keeping its own copy

    Incorrect

    A standard topic pushes each message once to its subscribers and keeps no history, so the pricing service could not reprocess the past day.

Multiple independent real-time readers plus replay is the stream model, and Kinesis Data Streams retains records so that consumers can read them again.

Question 5 · choose 1

A company must migrate 60 TB from an on-premises NFS file server to Amazon EFS over its existing 1 Gbps AWS Direct Connect link within two weeks. File permissions and timestamps must be preserved, the copied data must be verified, and a final incremental copy will run during the cutover weekend. Which service should a solutions architect use?

  1. AAn Amazon S3 File Gateway on premises that the NFS clients write through
  2. BAWS DataSync, with an agent on premises and the EFS file system as the destination
  3. CAWS Transfer Family with an SFTP endpoint that writes to the EFS file system
  4. DA script that runs rsync on an EC2 instance that mounts both the NFS export and the EFS file system
Show the answer and why
  • AAn Amazon S3 File Gateway on premises that the NFS clients write through

    Incorrect

    S3 File Gateway stores files as objects in Amazon S3, not in EFS, and it gives ongoing hybrid access rather than a one-time migration.

  • BAWS DataSync, with an agent on premises and the EFS file system as the destination

    Correct

    DataSync copies data online from NFS to EFS, preserves file metadata, verifies the transferred data, and on later runs copies only what changed.

  • CAWS Transfer Family with an SFTP endpoint that writes to the EFS file system

    Incorrect

    Transfer Family accepts file uploads over SFTP into EFS, but someone would still have to script the uploads, check them and work out what changed for the final copy.

  • DA script that runs rsync on an EC2 instance that mounts both the NFS export and the EFS file system

    Incorrect

    This can work, but the company would build, run, verify and tune the transfer itself, which a managed transfer service does for it.

An online, verified, incremental migration from NFS to an AWS file system is the job DataSync is built for.

Question 6 · choose 1

Business users want interactive dashboards with filters and drill-downs on sales data that analysts already query with Amazon Athena from an S3 data lake. The users must view and share the dashboards in a web browser, with no software to install and no servers to manage. Which solution meets these requirements?

  1. AAmazon OpenSearch Service with OpenSearch Dashboards, after the sales data is copied into a domain
  2. BAn open-source BI tool on an EC2 instance that connects to Athena through JDBC
  3. CAmazon Quick Sight dashboards that use Athena as their data source
  4. DAthena saved queries shared with the users through the Athena console
Show the answer and why
  • AAmazon OpenSearch Service with OpenSearch Dashboards, after the sales data is copied into a domain

    Incorrect

    OpenSearch Dashboards could chart the data, but only after it is copied and kept in sync in an OpenSearch domain, which is extra work and cost.

  • BAn open-source BI tool on an EC2 instance that connects to Athena through JDBC

    Incorrect

    This gives dashboards, but the company would run, patch and scale the server itself.

  • CAmazon Quick Sight dashboards that use Athena as their data source

    Correct

    Quick Sight is a fully managed BI service that connects to Athena and publishes interactive dashboards that users open in a browser.

  • DAthena saved queries shared with the users through the Athena console

    Incorrect

    Saved queries return tables of results, not interactive visual dashboards, and users would have to work in SQL.

Managed dashboards on top of Athena is a visualization job for Quick Sight, the business intelligence feature of Amazon Quick.

Question 7 · choose 1

Producer applications on EC2 instances in private subnets write records to an Amazon Kinesis data stream. A security review requires that this traffic stays on the AWS network, never uses the internet or a NAT device, and is allowed only for this one stream. What should a solutions architect configure?

  1. AA gateway VPC endpoint for Kinesis Data Streams in the private subnets' route tables
  2. BA NAT gateway in a public subnet, with a security group on the NAT gateway that allows only the Kinesis IP ranges
  3. CAn AWS Direct Connect public virtual interface to the Kinesis endpoints
  4. DAn interface VPC endpoint for Kinesis Data Streams, with an endpoint policy for that stream only
Show the answer and why
  • AA gateway VPC endpoint for Kinesis Data Streams in the private subnets' route tables

    Incorrect

    Gateway endpoints exist only for Amazon S3 and DynamoDB. Kinesis Data Streams is reached privately through an interface endpoint.

  • BA NAT gateway in a public subnet, with a security group on the NAT gateway that allows only the Kinesis IP ranges

    Incorrect

    The traffic would still leave through a NAT device to public endpoints, and a security group cannot be associated with a NAT gateway at all.

  • CAn AWS Direct Connect public virtual interface to the Kinesis endpoints

    Incorrect

    A public virtual interface connects an on-premises network to public AWS endpoints. It does not give instances in a VPC a private path.

  • DAn interface VPC endpoint for Kinesis Data Streams, with an endpoint policy for that stream only

    Correct

    An interface endpoint keeps the traffic between the VPC and Kinesis on the AWS network, and its endpoint policy can allow access to just this stream.

Private access from a VPC to Kinesis is an interface endpoint (AWS PrivateLink), and the endpoint policy narrows it to one stream.

Question 8 · choose 1

Producers write about 8 MB of data per second, in roughly 3,000 records per second, to an Amazon Kinesis data stream in provisioned mode with 4 shards. Partition keys are random, so the load is spread evenly, yet the producers keep receiving ProvisionedThroughputExceededException errors. What should a solutions architect do?

  1. AAggregate several user records into each Kinesis record with the Kinesis Producer Library
  2. BIncrease the stream's data retention period from 24 hours to 7 days
  3. CIncrease the number of shards in the stream to at least 8
  4. DRegister the downstream applications as enhanced fan-out consumers
Show the answer and why
  • AAggregate several user records into each Kinesis record with the Kinesis Producer Library

    Incorrect

    Aggregation lowers the number of records per second, which is already under the limit. The bytes per second stay the same and still exceed what 4 shards accept.

  • BIncrease the stream's data retention period from 24 hours to 7 days

    Incorrect

    Retention controls how long records are kept, not how fast they can be written.

  • CIncrease the number of shards in the stream to at least 8

    Correct

    Each shard accepts up to 1 MB per second of writes, so 4 shards take about 4 MB per second. At least 8 shards are needed for 8 MB per second.

  • DRegister the downstream applications as enhanced fan-out consumers

    Incorrect

    Enhanced fan-out adds read throughput for consumers. The errors come from the write side.

Kinesis capacity is sized per shard: 1 MB or 1,000 records per second for writes. Check both limits against the producers' traffic.

Practise domain 3 →Practise all domains →