Skip to content
BytePatterns

SAA-C03 · Domain 4: Design Cost-Optimized Architectures · 20% of the exam

Task 4.1: Design cost-optimized storage solutions.

Paying only for the storage a workload needs: S3 storage classes and lifecycle rules, the right EBS volume type and size, archival and backup choices, and the cheapest way to move data into AWS.

Study it

  • Storage cost: S3 classes and lifecycle, Requester Pays, EBS volume choice and archival

    Partly covered by: S3: Consistency, Classes, Lifecycle

  • Moving data in for less: DataSync, Transfer Family and Storage Gateway

    Lesson coming

  • Cost visibility: Cost Explorer, Budgets, Cost and Usage Report and cost allocation tags

    Lesson coming

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

A media company keeps millions of objects, most larger than 1 MB, in S3 Standard. Access patterns are unknown and change over time: some objects stay popular for months, others are never read again. The company wants automatic savings with no retrieval fees and no change in performance when an old object is read again. Which option should a solutions architect choose?

  1. AS3 Standard-IA, applied to every object by a lifecycle rule after 30 days
  2. BS3 One Zone-IA, applied to every object by a lifecycle rule after 30 days
  3. CS3 Glacier Flexible Retrieval, applied by a lifecycle rule after 90 days
  4. DS3 Intelligent-Tiering, with its optional archive access tiers left off
Show the answer and why
  • AS3 Standard-IA, applied to every object by a lifecycle rule after 30 days

    Incorrect

    Standard-IA charges a per-GB retrieval fee, so objects that become popular again cost more to read.

  • BS3 One Zone-IA, applied to every object by a lifecycle rule after 30 days

    Incorrect

    One Zone-IA also charges retrieval fees, and it keeps data in a single Availability Zone.

  • CS3 Glacier Flexible Retrieval, applied by a lifecycle rule after 90 days

    Incorrect

    Objects in this class must be restored before they can be read, which takes minutes to hours.

  • DS3 Intelligent-Tiering, with its optional archive access tiers left off

    Correct

    Intelligent-Tiering moves objects between access tiers by usage, without performance impact and with no retrieval fees; only a small monitoring fee per object applies.

Unknown or changing access patterns are the stated use case for S3 Intelligent-Tiering. Leaving the archive tiers off keeps every read at millisecond speed.

Question 2 · choose 2

A company stores scanned contracts in S3 Standard. A contract is read often in its first 30 days and about once a quarter for the rest of the year, and it must be deleted 365 days after it was created. Every read must return in milliseconds, and the data must survive the loss of an Availability Zone. Which lifecycle configuration steps are the MOST cost-effective? (Choose TWO.)

  1. ATransition objects to S3 Glacier Flexible Retrieval 30 days after creation
  2. BTransition objects to S3 One Zone-IA 30 days after creation
  3. CRun a daily Lambda function that deletes objects older than 365 days
  4. DTransition objects to S3 Glacier Instant Retrieval 30 days after creation
  5. EAdd an expiration action that deletes objects 365 days after creation
Show the answer and why
  • ATransition objects to S3 Glacier Flexible Retrieval 30 days after creation

    Incorrect

    Objects in Glacier Flexible Retrieval must be restored before reading, which takes minutes to hours, not milliseconds.

  • BTransition objects to S3 One Zone-IA 30 days after creation

    Incorrect

    One Zone-IA keeps data in one Availability Zone, so it is not resilient to the loss of that zone.

  • CRun a daily Lambda function that deletes objects older than 365 days

    Incorrect

    It works, but it is custom code and extra requests for something a lifecycle expiration action does for free.

  • DTransition objects to S3 Glacier Instant Retrieval 30 days after creation

    Correct

    Glacier Instant Retrieval is meant for long-lived data read about once a quarter with millisecond access, stored across at least three Availability Zones.

  • EAdd an expiration action that deletes objects 365 days after creation

    Correct

    Lifecycle expiration actions delete objects on your behalf at the set age, with no code to run.

Quarterly reads with millisecond access and multi-zone resilience point to Glacier Instant Retrieval; deletion at a fixed age is a lifecycle expiration action.

Question 3 · choose 1

A research institute shares a 400 TB genomics dataset in an S3 bucket with other universities, which download large parts of it. The institute will pay to store the data but wants each university to pay for its own requests and downloads. What should a solutions architect do?

  1. ATurn on Requester Pays for the bucket; universities then send authenticated requests
  2. BServe the dataset through a CloudFront distribution and give each university signed URLs
  3. CCreate presigned URLs for each university so that their downloads are billed to them
  4. DTurn on S3 Transfer Acceleration so the universities download faster and more cheaply
Show the answer and why
  • ATurn on Requester Pays for the bucket; universities then send authenticated requests

    Correct

    With Requester Pays, the requester pays for requests and downloads while the bucket owner pays for storage. Anonymous access is not allowed, so requests must be authenticated.

  • BServe the dataset through a CloudFront distribution and give each university signed URLs

    Incorrect

    Signed URLs control who can access content and for how long. The distribution's owner still pays for the delivery.

  • CCreate presigned URLs for each university so that their downloads are billed to them

    Incorrect

    A presigned URL grants temporary access with its creator's permissions. It does not move request and download costs to whoever uses it.

  • DTurn on S3 Transfer Acceleration so the universities download faster and more cheaply

    Incorrect

    Transfer Acceleration speeds up long-distance transfers and can add data transfer charges; it does not shift costs to the downloader.

Shifting request and download costs to the reader is exactly what a Requester Pays bucket does.

Question 4 · choose 1

A log analytics cluster on EC2 reads and writes multi-terabyte files in large sequential streams all day and needs about 400 MiB/s per volume. Small random I/O is rare. The company wants the lowest-cost EBS volume type that meets this throughput. Which volume type should it use?

  1. ACold HDD (sc1) volumes sized to hold the full dataset
  2. BGeneral Purpose SSD (gp3) with extra throughput provisioned
  3. CThroughput Optimized HDD (st1) volumes sized for the dataset
  4. DProvisioned IOPS SSD (io2) volumes with high IOPS provisioned
Show the answer and why
  • ACold HDD (sc1) volumes sized to hold the full dataset

    Incorrect

    sc1 is for less frequently accessed data and tops out at 250 MiB/s per volume, below the requirement.

  • BGeneral Purpose SSD (gp3) with extra throughput provisioned

    Incorrect

    gp3 can reach the throughput, but SSD storage costs more per GB than the low-cost HDD types built for this pattern.

  • CThroughput Optimized HDD (st1) volumes sized for the dataset

    Correct

    st1 is a low-cost HDD designed for frequently accessed, throughput-intensive sequential workloads, with up to 500 MiB/s per volume.

  • DProvisioned IOPS SSD (io2) volumes with high IOPS provisioned

    Incorrect

    io2 is priced for I/O-intensive workloads such as databases. Paying for IOPS does not suit large sequential reads.

Large, sequential and frequently accessed at a few hundred MiB/s is the st1 profile. sc1 is cheaper but too slow and meant for colder data.

Question 5 · choose 1

An application uses twenty 1 TiB General Purpose SSD (gp2) volumes. Monitoring shows that each volume needs at most 3,000 IOPS and 200 MiB/s of throughput. The company wants to lower its EBS cost without losing performance and without downtime. What should a solutions architect do?

  1. AChange the volumes to Throughput Optimized HDD (st1) to lower the price per GiB
  2. BModify the volumes in place to gp3 and provision 200 MiB/s of throughput on each
  3. CChange the volumes to Provisioned IOPS SSD (io2) with 3,000 IOPS provisioned on each one
  4. DSnapshot the volumes and restore them as smaller gp2 volumes to pay for less storage
Show the answer and why
  • AChange the volumes to Throughput Optimized HDD (st1) to lower the price per GiB

    Incorrect

    st1 is built for large sequential I/O. Its performance on small random I/O is far below 3,000 IOPS, so the application would slow down.

  • BModify the volumes in place to gp3 and provision 200 MiB/s of throughput on each

    Correct

    gp3 includes 3,000 IOPS and 125 MiB/s at any size, extra throughput can be provisioned separately, and its price per GB is lower than gp2. Elastic Volumes changes the type while the volume stays in use.

  • CChange the volumes to Provisioned IOPS SSD (io2) with 3,000 IOPS provisioned on each one

    Incorrect

    io2 delivers the performance, but it charges for storage and for every provisioned IOPS at higher rates, so the bill goes up.

  • DSnapshot the volumes and restore them as smaller gp2 volumes to pay for less storage

    Incorrect

    gp2 IOPS grow with volume size, so smaller gp2 volumes would get fewer IOPS, and restoring and switching volumes means downtime.

gp3 decouples performance from size and costs less per GB than gp2, and the change is an online volume modification.

Question 6 · choose 1

A company must keep scanned tax records for 10 years. After the first 90 days the records are almost never read, and when an auditor asks for one, a wait of up to 48 hours is acceptable. The records must survive the loss of an Availability Zone. Which storage approach is the MOST cost-effective?

  1. AAn S3 Lifecycle rule that moves the objects to S3 Glacier Deep Archive after 90 days
  2. BAn S3 Lifecycle rule that moves the objects to S3 Glacier Flexible Retrieval after 90 days
  3. CAn S3 Lifecycle rule that moves the objects to S3 One Zone-IA after 90 days
  4. DA nightly job that copies objects older than 90 days to a Cold HDD (sc1) EBS volume
Show the answer and why
  • AAn S3 Lifecycle rule that moves the objects to S3 Glacier Deep Archive after 90 days

    Correct

    Glacier Deep Archive is the lowest-cost S3 storage class, for data read rarely and retrieved within hours, and it stores data across multiple Availability Zones.

  • BAn S3 Lifecycle rule that moves the objects to S3 Glacier Flexible Retrieval after 90 days

    Incorrect

    Flexible Retrieval meets the retrieval time and survives a zone loss, but it costs more per GB than Deep Archive, which also fits a 48-hour wait.

  • CAn S3 Lifecycle rule that moves the objects to S3 One Zone-IA after 90 days

    Incorrect

    One Zone-IA keeps data in a single Availability Zone, so it fails the resilience requirement, and it costs more than the archive classes.

  • DA nightly job that copies objects older than 90 days to a Cold HDD (sc1) EBS volume

    Incorrect

    An EBS volume lives in one Availability Zone and costs far more per GB than archive storage in S3.

Rare access, a retrieval window of days and a 10-year retention point to the cheapest archive class, reached automatically with a lifecycle rule.

Question 7 · choose 1

A content team keeps 40 TB of project files on Amazon EFS in the EFS Standard storage class. Files are used heavily for about a month, after which most are rarely opened, but they must stay at the same file system paths for the applications. What is the MOST cost-effective change?

  1. ACopy files older than 30 days to an S3 bucket with a nightly script and delete them from EFS
  2. BMove the data to an Amazon FSx for Lustre file system with a scratch deployment type
  3. CSwitch the file system to Provisioned throughput mode with a lower throughput setting than today
  4. DTurn on EFS lifecycle management to move files not accessed for 30 days to Infrequent Access
Show the answer and why
  • ACopy files older than 30 days to an S3 bucket with a nightly script and delete them from EFS

    Incorrect

    The files would leave the file system paths the applications use, and the script is custom code to build and run.

  • BMove the data to an Amazon FSx for Lustre file system with a scratch deployment type

    Incorrect

    Scratch file systems are for temporary data and do not replicate it, and moving to Lustre does not lower the cost of rarely used files.

  • CSwitch the file system to Provisioned throughput mode with a lower throughput setting than today

    Incorrect

    Throughput mode sets how throughput is charged. It does not change the cost of storing 40 TB of mostly idle files.

  • DTurn on EFS lifecycle management to move files not accessed for 30 days to Infrequent Access

    Correct

    Lifecycle management moves cold files to the cheaper Infrequent Access class automatically, and the files keep their paths in the same file system.

Tiering inside the file system cuts the storage cost of cold files without moving them out from under the applications.

Question 8 · choose 2

Finance wants to see each project's monthly S3 storage cost in AWS Cost Explorer. Each of the company's buckets is used by exactly one project. Which steps should a solutions architect take? (Choose TWO.)

  1. ATurn on Requester Pays for every bucket so that each project is billed on its own
  2. BAdd a project tag to each bucket
  3. CActivate the project tag key as a user-defined cost allocation tag
  4. DTurn on S3 Storage Lens advanced metrics to see each bucket's cost
  5. ECreate one AWS Budgets budget for each bucket name
Show the answer and why
  • ATurn on Requester Pays for every bucket so that each project is billed on its own

    Incorrect

    Requester Pays makes the requester pay for requests and downloads. The bucket owner still pays for storage, and nothing is grouped by project.

  • BAdd a project tag to each bucket

    Correct

    Tags on the buckets carry the project name into the billing data once the tag key is activated for cost allocation.

  • CActivate the project tag key as a user-defined cost allocation tag

    Correct

    Only activated cost allocation tags appear in Cost Explorer and the cost reports, where costs can then be grouped by project.

  • DTurn on S3 Storage Lens advanced metrics to see each bucket's cost

    Incorrect

    Storage Lens reports storage usage and activity metrics, not costs broken down in Cost Explorer.

  • ECreate one AWS Budgets budget for each bucket name

    Incorrect

    Budgets alert on spend against a threshold. They do not break storage cost down by bucket without tags, and they do not change what Cost Explorer shows.

Cost by business dimension means tagging the resources and then activating the tag key as a cost allocation tag.

Question 9 · choose 1

An application writes sensor readings of about 2 KB each to Amazon S3, one object per reading, about 50 million objects per day. Analysts query the data in daily batches with Amazon Athena. S3 PUT request charges have grown larger than the storage charges. Which change reduces cost the MOST?

  1. AWrite each reading to the S3 Standard-IA storage class instead of S3 Standard to lower its price
  2. BTurn on S3 Transfer Acceleration for the bucket that receives the readings
  3. CBuffer the readings and write them as larger batched objects, for example with Amazon Data Firehose
  4. DStore the readings of each sensor in its own S3 bucket
Show the answer and why
  • AWrite each reading to the S3 Standard-IA storage class instead of S3 Standard to lower its price

    Incorrect

    Standard-IA bills small objects as at least 128 KB and charges more per request, so this raises the cost of millions of tiny objects.

  • BTurn on S3 Transfer Acceleration for the bucket that receives the readings

    Incorrect

    Transfer Acceleration adds a per-GB charge for faster long-distance transfers and does nothing about the number of PUT requests.

  • CBuffer the readings and write them as larger batched objects, for example with Amazon Data Firehose

    Correct

    Batching turns many tiny PUTs into a few large ones, which cuts request charges, and Athena also reads fewer, larger files faster.

  • DStore the readings of each sensor in its own S3 bucket

    Incorrect

    The number of objects and PUT requests stays exactly the same, now spread over more buckets.

Request charges scale with the number of objects. Batching small records into larger objects before upload is the lever, and it helps queries too.

Practise domain 4 →Practise all domains →