Skip to content
BytePatterns

SAA-C03 · Domain 3: Design High-Performing Architectures · 24% of the exam

Task 3.1: Determine high-performing and/or scalable storage solutions.

Matching storage to the workload's speed and growth: object, file and block storage, EBS volume types and IOPS, shared file systems, and hybrid storage that keeps on-premises applications working.

Study it

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

A company is moving a mission-critical Oracle database to a single EC2 instance built on the Nitro System. The database needs 150,000 IOPS of sustained, consistent performance on one EBS volume. Which volume type meets this requirement?

  1. AGeneral Purpose SSD (gp3) with its IOPS provisioned to the maximum allowed
  2. BProvisioned IOPS SSD (io2) Block Express with 150,000 IOPS provisioned
  3. CThroughput Optimized HDD (st1) at its largest possible volume size
  4. DProvisioned IOPS SSD (io1) with its IOPS provisioned to the maximum allowed
Show the answer and why
  • AGeneral Purpose SSD (gp3) with its IOPS provisioned to the maximum allowed

    Incorrect

    A gp3 volume can be provisioned with at most 80,000 IOPS, short of the requirement.

  • BProvisioned IOPS SSD (io2) Block Express with 150,000 IOPS provisioned

    Correct

    io2 Block Express volumes support up to 256,000 IOPS and are aimed at demanding databases such as Oracle on Nitro-based instances.

  • CThroughput Optimized HDD (st1) at its largest possible volume size

    Incorrect

    st1 is built for large sequential throughput and tops out at 500 IOPS per volume; it is not for random database I/O.

  • DProvisioned IOPS SSD (io1) with its IOPS provisioned to the maximum allowed

    Incorrect

    An io1 volume supports at most 64,000 IOPS, well below 150,000.

Past 80,000 IOPS on a single volume, only io2 Block Express is left. Read the maximum IOPS row of the volume types table.

Question 2 · choose 1

A containerized application runs as Amazon ECS tasks in three Availability Zones. All tasks must read and write the same files through a shared file system with POSIX permissions that grows and shrinks automatically as files are added and removed. Which storage solution meets these requirements?

  1. AAn io2 EBS volume with Multi-Attach enabled and attached to the hosts of every task
  2. BAn S3 bucket that every task reads and writes through the Amazon S3 object API
  3. CAn Amazon FSx for Windows File Server share that the tasks mount over SMB
  4. DAn Amazon EFS file system mounted by the tasks in all three Availability Zones
Show the answer and why
  • AAn io2 EBS volume with Multi-Attach enabled and attached to the hosts of every task

    Incorrect

    Multi-Attach works only for instances in the same Availability Zone, and the volume has a fixed provisioned size.

  • BAn S3 bucket that every task reads and writes through the Amazon S3 object API

    Incorrect

    Amazon S3 is object storage accessed through an API, not a shared file system with POSIX permissions.

  • CAn Amazon FSx for Windows File Server share that the tasks mount over SMB

    Incorrect

    FSx for Windows File Server provides Windows-native SMB file shares with Windows access control lists, not POSIX permissions.

  • DAn Amazon EFS file system mounted by the tasks in all three Availability Zones

    Correct

    EFS is a shared file system that supports POSIX permissions and grows and shrinks automatically as files are added and removed.

Shared, POSIX, across zones and elastic: that is Amazon EFS. Multi-Attach stays inside one zone.

Question 3 · choose 1

An on-premises application writes files to an NFS share. The company wants each file stored durably in Amazon S3 within minutes so that cloud analytics can use it, while the application keeps low-latency access to recently used files and needs no code changes. Which solution meets these requirements?

  1. ADeploy Amazon S3 File Gateway on premises and mount its NFS share from the application
  2. BDeploy a Volume Gateway in cached mode and present an iSCSI block volume to the application
  3. CSchedule an AWS DataSync task every night that copies new files from the NFS share to S3
  4. DTurn on S3 Transfer Acceleration so that the application uploads its files to S3 faster
Show the answer and why
  • ADeploy Amazon S3 File Gateway on premises and mount its NFS share from the application

    Correct

    S3 File Gateway stores files as objects in S3 and serves them over NFS or SMB, with a local cache for low-latency access to recently used data.

  • BDeploy a Volume Gateway in cached mode and present an iSCSI block volume to the application

    Incorrect

    Volume Gateway presents block storage over iSCSI. The data is not stored as individual S3 objects that analytics services can read as files.

  • CSchedule an AWS DataSync task every night that copies new files from the NFS share to S3

    Incorrect

    DataSync copies data between storage systems, but a nightly schedule leaves new files out of S3 for up to a day.

  • DTurn on S3 Transfer Acceleration so that the application uploads its files to S3 faster

    Incorrect

    Transfer Acceleration speeds up transfers to a bucket through the S3 API. The application writes to NFS, so it would need code changes.

File protocol on premises, objects in S3 and a local cache: that is the S3 File Gateway.

Question 4 · choose 1

A genomics company processes 200 TB of input files stored in Amazon S3 with thousands of Linux compute instances. Each run needs a shared POSIX file system with sub-millisecond latency and very high aggregate throughput, and the results must end up back in S3. Which storage solution fits best?

  1. AAmazon FSx for Lustre, linked to the S3 bucket through a data repository association
  2. BAmazon EFS with Elastic throughput, filled from S3 by a copy job before each run and copied back afterward
  3. CProvisioned IOPS SSD (io2) volumes with Multi-Attach, attached to all of the compute instances
  4. DAn Amazon FSx for Windows File Server file system that the instances mount over SMB
Show the answer and why
  • AAmazon FSx for Lustre, linked to the S3 bucket through a data repository association

    Correct

    FSx for Lustre is a high-performance parallel file system for compute workloads such as HPC, with sub-millisecond latency. Linked to S3, it presents the objects as files and can export results back to the bucket.

  • BAmazon EFS with Elastic throughput, filled from S3 by a copy job before each run and copied back afterward

    Incorrect

    EFS is a shared file system, but it has no link to S3, so every run needs custom copy jobs in both directions, and it is not the service built for parallel HPC throughput.

  • CProvisioned IOPS SSD (io2) volumes with Multi-Attach, attached to all of the compute instances

    Incorrect

    Multi-Attach shares one volume with at most 16 instances in a single Availability Zone, and it needs a clustered file system. It cannot serve thousands of instances.

  • DAn Amazon FSx for Windows File Server file system that the instances mount over SMB

    Incorrect

    FSx for Windows File Server is built for Windows workloads over SMB, not for Linux HPC at this scale, and it has no link to S3.

HPC on Linux with data that lives in S3 is the use case FSx for Lustre is designed for: a fast parallel file system that loads from and exports to an S3 bucket.

Question 5 · choose 1

A company is moving a Windows-based document management application to EC2 instances in two Availability Zones. The instances need a shared file system that uses the SMB protocol, applies NTFS permissions from the company's Active Directory, is fully managed, and stays available if one zone fails. Which storage service meets these requirements?

  1. AAmazon EFS with a mount target in each Availability Zone, mounted by every instance
  2. BAmazon FSx for Windows File Server with a Multi-AZ deployment joined to the directory
  3. CAmazon FSx for Lustre with a persistent deployment type in one of the two zones
  4. DA Provisioned IOPS SSD (io2) volume with Multi-Attach, shared by the instances in both zones
Show the answer and why
  • AAmazon EFS with a mount target in each Availability Zone, mounted by every instance

    Incorrect

    EFS is an NFS file system for Linux. It is not supported on Windows instances and does not use SMB or NTFS permissions.

  • BAmazon FSx for Windows File Server with a Multi-AZ deployment joined to the directory

    Correct

    FSx for Windows File Server provides fully managed SMB shares with NTFS access control from Active Directory, and a Multi-AZ file system fails over to a standby in another zone.

  • CAmazon FSx for Lustre with a persistent deployment type in one of the two zones

    Incorrect

    FSx for Lustre is a high-performance file system for Linux compute clients, not an SMB file share for Windows.

  • DA Provisioned IOPS SSD (io2) volume with Multi-Attach, shared by the instances in both zones

    Incorrect

    Multi-Attach works only for instances in the same Availability Zone as the volume, and the volume is block storage, not a managed SMB file share.

SMB, Active Directory and NTFS permissions point to FSx for Windows File Server, and its Multi-AZ deployment covers the loss of a zone.

Question 6 · choose 1

Users on several continents upload video files of 2 to 10 GB to an S3 bucket in us-east-1 through presigned URLs from a web app. Uploads from distant countries are slow and sometimes fail near the end, so users must start again. Which change improves upload speed and reliability the MOST?

  1. AStore the uploads in an S3 Express One Zone directory bucket in us-east-1
  2. BSpread the uploaded objects across many prefixes so that the bucket can accept more requests per second
  3. CTurn on S3 Transfer Acceleration and upload each file in parts with multipart upload
  4. DTurn on S3 Versioning so that an interrupted upload resumes from the last saved version
Show the answer and why
  • AStore the uploads in an S3 Express One Zone directory bucket in us-east-1

    Incorrect

    S3 Express One Zone gives very low latency to compute in the same Availability Zone. It does not shorten the long internet path from distant users.

  • BSpread the uploaded objects across many prefixes so that the bucket can accept more requests per second

    Incorrect

    Prefixes raise the request rate a bucket can take. The problem here is distance and large single uploads, not the number of requests.

  • CTurn on S3 Transfer Acceleration and upload each file in parts with multipart upload

    Correct

    Transfer Acceleration routes uploads through the nearest edge location over the AWS network, which helps over long distances. Multipart upload sends parts in parallel and retries only a failed part instead of the whole file.

  • DTurn on S3 Versioning so that an interrupted upload resumes from the last saved version

    Incorrect

    Versioning keeps versions of complete objects. It does not resume an interrupted upload.

Long distance calls for the edge network of Transfer Acceleration, and large files call for multipart upload, which makes a failure cost one part instead of the whole file.

Question 7 · choose 1

A distributed cache runs on a cluster of EC2 instances and keeps three copies of every item on different nodes. Each node needs the lowest possible latency and a very high rate of random I/O for its local working data. Data on a failed or stopped node can be lost, because the cluster rebuilds it from the other copies. Which storage option fits best?

  1. AProvisioned IOPS SSD (io2) Block Express volumes with the maximum IOPS provisioned
  2. BAmazon EFS file systems mounted by the nodes, with Elastic throughput turned on
  3. CGeneral Purpose SSD (gp3) volumes striped in a RAID 0 array on each node to add up their IOPS
  4. DInstance store volumes on a storage optimized instance type with local NVMe SSDs
Show the answer and why
  • AProvisioned IOPS SSD (io2) Block Express volumes with the maximum IOPS provisioned

    Incorrect

    io2 Block Express is the fastest EBS volume type, but EBS is network storage with a per-volume IOPS ceiling, and its durability is paid for although this data does not need it.

  • BAmazon EFS file systems mounted by the nodes, with Elastic throughput turned on

    Incorrect

    EFS is a network file system shared across instances. Its latency is well above that of local disks.

  • CGeneral Purpose SSD (gp3) volumes striped in a RAID 0 array on each node to add up their IOPS

    Incorrect

    Striping adds up IOPS, but the volumes stay network-attached and are limited by the instance's EBS bandwidth, so latency stays higher than on local disks.

  • DInstance store volumes on a storage optimized instance type with local NVMe SSDs

    Correct

    Instance store is physically attached to the host and gives the lowest latency and very high random IOPS. Its data is temporary, which this replicated cache accepts.

When the application already replicates its data and wants raw local speed, temporary instance store on storage optimized instances is the intended fit.

Question 8 · choose 2

An analytics application in the same Region as its S3 bucket reads objects of about 2 GB each. Single-object downloads are slower than the job needs, and at peak the application sends about 15,000 GET requests per second to objects under one prefix and receives HTTP 503 Slow Down errors. Which changes improve performance? (Choose TWO.)

  1. ADownload each object with several parallel byte-range GET requests
  2. BTurn on S3 Transfer Acceleration for the bucket
  3. CSpread the objects across several prefixes and read them in parallel
  4. DMove the objects to the S3 Glacier Instant Retrieval storage class
  5. ETurn on S3 Versioning for the bucket
Show the answer and why
  • ADownload each object with several parallel byte-range GET requests

    Correct

    Fetching byte ranges of one object in parallel raises the throughput of a large download, as the S3 performance guidelines recommend.

  • BTurn on S3 Transfer Acceleration for the bucket

    Incorrect

    Transfer Acceleration speeds up transfers over long distances. The application is in the same Region, so it gains little.

  • CSpread the objects across several prefixes and read them in parallel

    Correct

    S3 supports at least 5,500 GET requests per second per prefix, and there is no limit to the number of prefixes, so more prefixes raise the request rate the bucket can serve.

  • DMove the objects to the S3 Glacier Instant Retrieval storage class

    Incorrect

    The class lowers storage cost for rarely read data and adds retrieval fees. It does not raise throughput or request rate.

  • ETurn on S3 Versioning for the bucket

    Incorrect

    Versioning keeps earlier versions of objects. It changes neither the throughput of a download nor the request rate a prefix can serve.

Per-object speed comes from parallel byte-range requests, and request rate scales with the number of prefixes. Distance-based acceleration and storage classes do not address either problem.

Practise domain 3 →Practise all domains →