Skip to content
BytePatterns

SAP-C02 · Domain 3: Continuous Improvement for Existing Solutions · 25% of the exam

Task 3.5: Identify opportunities for cost optimizations.

Cutting the bill of an existing system: idle and oversized resources found in usage reports, Spot and Savings Plans, data transfer charges, billing alarms, and cost allocation tags read at a granular level in the Cost and Usage Report.

Study it

  • Cost cleanup: idle resources, rightsizing, Spot, Savings Plans and data transfer

    Partly covered by: AWS Cost Levers

  • Billing alarms, cost allocation tags and the Cost and Usage Report

    Lesson coming

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

A FinOps team for an organization with 70 accounts gets cost saving advice from several places: Compute Optimizer, Savings Plans recommendations and checks for idle resources. It wants one view across all accounts and Regions that combines these recommendations, removes double counting of savings and ranks them. Which AWS feature should the team use?

  1. AAWS Budgets reports sent to the team every week
  2. BAWS Cost Anomaly Detection monitors for each account
  3. CCost Optimization Hub, enabled for the organization
  4. DThe Cost Explorer cost and usage graphs grouped by linked account
Show the answer and why
  • AAWS Budgets reports sent to the team every week

    Incorrect

    Budgets reports show performance against budgets. They do not collect or rank savings recommendations.

  • BAWS Cost Anomaly Detection monitors for each account

    Incorrect

    Cost Anomaly Detection finds unusual spend. It does not recommend rightsizing, idle deletion or commitments.

  • CCost Optimization Hub, enabled for the organization

    Correct

    Cost Optimization Hub consolidates and prioritizes recommendations for rightsizing, idle resources, Savings Plans and Reserved Instances across accounts and Regions, and it deduplicates savings.

  • DThe Cost Explorer cost and usage graphs grouped by linked account

    Incorrect

    Cost Explorer shows where money went. Grouping by account does not combine or rank savings opportunities.

"One ranked list of savings, deduplicated, across accounts" is what Cost Optimization Hub is for.

Question 2 · choose 1

Finance asks for the cost of every individual EC2 instance, EBS volume and S3 bucket over the past year, grouped by cost allocation tag, so it can find waste. The data must be queryable with SQL and refreshed automatically. Which approach meets these requirements?

  1. ATurn on resource-level data in Cost Explorer and export the grouped cost and usage view every month
  2. BExport CUR 2.0 with resource IDs through AWS Data Exports to Amazon S3 and query it with Amazon Athena
  3. CDownload the monthly invoice PDFs from the Billing console and load them into a shared spreadsheet
  4. DCreate an AWS Budgets cost budget for each tag value and compare the budgets with each other every month
Show the answer and why
  • ATurn on resource-level data in Cost Explorer and export the grouped cost and usage view every month

    Incorrect

    Resource-level data in Cost Explorer covers only the past 14 days at daily granularity, so it cannot show a full year per resource.

  • BExport CUR 2.0 with resource IDs through AWS Data Exports to Amazon S3 and query it with Amazon Athena

    Correct

    Data Exports creates recurring exports of CUR 2.0 in Amazon S3, a standard export can include the ID of each resource, and the exports can be analyzed with tools such as Amazon Athena.

  • CDownload the monthly invoice PDFs from the Billing console and load them into a shared spreadsheet

    Incorrect

    Invoices summarize charges. They do not list each resource or its tags.

  • DCreate an AWS Budgets cost budget for each tag value and compare the budgets with each other every month

    Incorrect

    Budgets track spend against thresholds. They do not give per-resource costs that can be queried with SQL.

Per-resource, per-tag cost data over long periods comes from the Cost and Usage Report, delivered by Data Exports and queried with Athena.

Question 3 · choose 1

Continuous integration build agents run around the clock on 40 On-Demand EC2 instances of one type. Builds are stateless, take 5-20 minutes, can be retried if interrupted, and come from a queue that is nearly empty at night and on weekends. Which change gives the largest cost reduction?

  1. ARun the agents on Spot Instances of several types in an Auto Scaling group that scales on queue depth
  2. BBuy a three-year Compute Savings Plan that covers all 40 build instances around the clock, every day of the week
  3. CBuy Standard Reserved Instances for the 40 instances of the current type for a three-year term
  4. DMove the agents to Dedicated Hosts so that the company controls instance placement and host utilization
Show the answer and why
  • ARun the agents on Spot Instances of several types in an Auto Scaling group that scales on queue depth

    Correct

    Spot Instances suit stateless, interruptible work, diversifying instance types with the price-capacity-optimized strategy reduces interruptions, and scaling on the queue removes idle agents at night.

  • BBuy a three-year Compute Savings Plan that covers all 40 build instances around the clock, every day of the week

    Incorrect

    A commitment discounts the hours, but the company would keep paying for 40 instances that are mostly idle at night and on weekends.

  • CBuy Standard Reserved Instances for the 40 instances of the current type for a three-year term

    Incorrect

    Reserved Instances lower the hourly price but lock in the full fleet around the clock, including the idle hours.

  • DMove the agents to Dedicated Hosts so that the company controls instance placement and host utilization

    Incorrect

    Dedicated Hosts address needs such as bringing existing per-socket or per-core licenses. They do nothing about the idle hours of the fleet.

Interruptible, queue-driven work that is idle half the time is Spot plus scaling to demand, not a commitment.

Question 4 · choose 1

Every Monday a QA team restores the latest snapshot of a 6 TB production Aurora PostgreSQL cluster into 10 test environments. Each environment has its own AWS account in the same organization as the production account, and the key policy of the cluster's KMS key already lets these accounts use the key. The restores take hours, and storage for the 10 full copies is now one of the largest items on the QA bill, although each week's tests change less than 2% of the data. The testers need writable copies of recent production data in their own accounts, ready within minutes, and the tests must not run on the production instances. Finance wants the copies' storage cost cut sharply. Which solution meets these requirements?

  1. AShare the production cluster with the 10 test accounts through AWS RAM and have each account create its own Aurora clone of the cluster
  2. BKeep restoring the latest snapshot every week, but on smaller instance classes and with the copies kept for only five days
  3. CAdd Aurora Replicas to the production cluster and point each test environment at the instance endpoint of its own replica
  4. DExport the latest snapshot to Amazon S3 in Parquet each week and let the test environments query the export with Amazon Athena
Show the answer and why
  • AShare the production cluster with the 10 test accounts through AWS RAM and have each account create its own Aurora clone of the cluster

    Correct

    A clone starts by sharing data pages with its source through copy-on-write, so new storage is allocated only for changed data, and cloning is faster and more space-efficient than restoring a snapshot. A cluster shared through AWS RAM can be cloned by each authorized account, one cross-account clone per account and up to 15 per cluster, and each clone gets its own DB instances.

  • BKeep restoring the latest snapshot every week, but on smaller instance classes and with the copies kept for only five days

    Incorrect

    A restore creates a new DB cluster by physically copying the data, which is slower and uses more space than cloning, so each environment still stores a full copy while it exists. Smaller instances cut compute, not storage.

  • CAdd Aurora Replicas to the production cluster and point each test environment at the instance endpoint of its own replica

    Incorrect

    Replicas share the cluster volume, so they add no storage copies, but they are read-only instances inside the production cluster and account, while the testers need writable copies in their own accounts.

  • DExport the latest snapshot to Amazon S3 in Parquet each week and let the test environments query the export with Amazon Athena

    Incorrect

    Snapshot export writes compressed Parquet files that tools such as Athena can analyze, but tests that write to a PostgreSQL database cannot run against exported files.

The constraints are writable copies of recent data in each test account, ready in minutes, off the production instances, and far less storage. Snapshot restores are slow full copies, replicas are read-only and stay in the production cluster, and exported Parquet files are not a database. Sharing the cluster through AWS RAM lets each of the 10 accounts create one cross-account clone, well within the 15 allowed per cluster, and each clone stores only the pages its environment changes.

Question 5 · choose 1

A cost review shows that CloudWatch Logs storage charges in a platform account keep growing. About 1,200 log groups were created years ago with default settings and still hold every event since then. Teams query only the last 30 days during incidents, and audit evidence is already kept elsewhere. The fix must stop paying for old events in all existing log groups, keep 30 days queryable in CloudWatch, leave today's application logging unchanged, and need no recurring jobs. Which change meets these requirements?

  1. ASet a 30-day retention period on every existing log group so that CloudWatch Logs deletes older events automatically
  2. BExport each log group to Amazon S3 every day and let S3 lifecycle rules archive the exported files after 30 days
  3. CAdd data protection policies to the log groups so that sensitive values are masked before the events are stored
  4. DLower the applications' log level to WARN so that much less log data is written to the log groups from now on
Show the answer and why
  • ASet a 30-day retention period on every existing log group so that CloudWatch Logs deletes older events automatically

    Correct

    By default, log data is stored indefinitely. With a retention setting, data older than the setting is deleted automatically, typically within 72 hours, and the last 30 days stay in the log group for queries.

  • BExport each log group to Amazon S3 every day and let S3 lifecycle rules archive the exported files after 30 days

    Incorrect

    Exported files can be archived or deleted by lifecycle rules, but exporting does not remove anything from the log groups, and AWS recommends against regular exports as a way to archive logs.

  • CAdd data protection policies to the log groups so that sensitive values are masked before the events are stored

    Incorrect

    Data protection policies help protect sensitive log data by masking it. They do not shorten how long events are kept.

  • DLower the applications' log level to WARN so that much less log data is written to the log groups from now on

    Incorrect

    Writing less would slow the growth, but it changes today's logging, which must stay as it is, and every old event would still be stored indefinitely.

The constraints are deleting old events, keeping 30 days in CloudWatch, no change to application logging and no recurring jobs. Exports add copies without removing data, masking only protects content, and a lower log level changes the applications. A 30-day retention setting on each log group lets CloudWatch Logs delete older events on its own.

Question 6 · choose 1

A data lake bucket encrypts objects with SSE-KMS under a customer managed key. Ingestion jobs write about a billion small objects a day, and analytics jobs read most of them within a week, so AWS KMS request charges are now larger than the S3 storage bill. Compliance requires that the data stay encrypted with that customer managed key under the company's key policy. The team wants the KMS request cost cut sharply without changing the ingestion or analytics code. Which change meets these requirements?

  1. AChange the bucket's default encryption to SSE-S3 so that new objects no longer need any requests to AWS KMS
  2. BSwitch the bucket's default encryption to DSSE-KMS with the same customer managed key for stronger protection
  3. CRequest a higher AWS KMS request quota for the account so that the ingestion jobs are not throttled at peak
  4. DTurn on S3 Bucket Keys for SSE-KMS in the bucket's default encryption with the same customer managed key
Show the answer and why
  • AChange the bucket's default encryption to SSE-S3 so that new objects no longer need any requests to AWS KMS

    Incorrect

    SSE-S3 adds no encryption fees, so it would remove the KMS charges, but the objects would no longer be protected by the company's customer managed key, which compliance requires.

  • BSwitch the bucket's default encryption to DSSE-KMS with the same customer managed key for stronger protection

    Incorrect

    DSSE-KMS applies two layers of encryption for standards that require multilayer encryption. It brings more KMS API calls and a higher cost than SSE-KMS, and S3 Bucket Keys are not supported with it.

  • CRequest a higher AWS KMS request quota for the account so that the ingestion jobs are not throttled at peak

    Incorrect

    KMS request quotas are adjustable through Service Quotas, which helps when calls are throttled. A higher quota does not lower the number of requests or their cost.

  • DTurn on S3 Bucket Keys for SSE-KMS in the bucket's default encryption with the same customer managed key

    Correct

    A bucket-level key reduces SSE-KMS request costs by up to 99 percent by cutting the request traffic from S3 to AWS KMS, while objects stay encrypted under the same KMS key. It needs no change to client applications and applies to new objects.

The constraints are the same customer managed key, a large cut in KMS request cost, and no code changes. SSE-S3 drops the required key, DSSE-KMS costs more and cannot use Bucket Keys, and a quota increase only allows more requests. S3 Bucket Keys keep SSE-KMS with the same key while sharply reducing calls to AWS KMS for the new objects that dominate this workload.

Question 7 · choose 1

An inference platform runs 12 Amazon ECS services on EC2 instances with several GPUs each, and every task reserves one GPU. The cluster's Auto Scaling group has a fixed desired capacity of 40 instances, sized for last year's launch peak, and the services spread their tasks across instances. Average GPU reservation across the cluster is 30%, and demand changes by the hour. The company wants to stop paying for idle instances, keep each service spread across Availability Zones, and avoid new long-term commitments while a platform redesign is planned. Which change meets these requirements?

  1. AKeep the fixed group and buy an EC2 Instance Savings Plan that covers the hourly cost of all 40 instances
  2. BSpread tasks across zones and then binpack on memory, with a capacity provider that uses managed scaling for the group
  3. CAdd a distinctInstance placement constraint so that every task gets an instance of its own and GPUs are not shared
  4. DMove the services to AWS Fargate so that capacity follows the tasks and no instances are left idle
Show the answer and why
  • AKeep the fixed group and buy an EC2 Instance Savings Plan that covers the hourly cost of all 40 instances

    Incorrect

    A Savings Plan lowers the rate in exchange for a one- or three-year commitment, which the company wants to avoid, and the idle instances keep running and are still paid for.

  • BSpread tasks across zones and then binpack on memory, with a capacity provider that uses managed scaling for the group

    Correct

    Binpack places tasks to leave the least unused CPU or memory, which minimizes the number of instances in use, while the zone spread keeps each service across Availability Zones. Managed scaling then adds and removes instances to match what the tasks reserve.

  • CAdd a distinctInstance placement constraint so that every task gets an instance of its own and GPUs are not shared

    Incorrect

    distinctInstance places each task on a different container instance, which needs at least as many instances as tasks and leaves the other GPUs on each instance idle.

  • DMove the services to AWS Fargate so that capacity follows the tasks and no instances are left idle

    Incorrect

    Fargate removes instance management, but the gpu task definition parameter is not valid for Fargate tasks, so GPU tasks cannot run there.

The waste comes from two choices: spreading tasks over every instance and a fixed group size. Packing tasks with binpack after a zone spread uses as few instances as possible, and ECS cluster auto scaling with managed scaling at the default 100% target capacity removes the instances that no longer hold tasks and adds them back when demand rises. A Savings Plan only discounts the waste, and Fargate cannot run GPU tasks.

Question 8 · choose 1

A provisioned DynamoDB table holds web session items that are useless 24 hours after they are created. Every night a job scans the whole table and deletes old items, and the read and write capacity it consumes has become a large part of the table's cost. The team wants to stop paying for these deletions and the scan, run no clean-up job at all, and make sure the application never treats an expired session as valid, even if an expired item is still in the table. Which solution meets these requirements?

  1. ASwitch the table to on-demand capacity mode so that the nightly job no longer needs provisioned capacity set aside for it
  2. BAdd a global secondary index on the creation date so that the job queries only old items instead of scanning the table
  3. CTurn on TTL with an expiration timestamp on each item, and filter out expired items in the application's queries
  4. DProcess the table's stream with a Lambda function that deletes each session item once it is older than 24 hours
Show the answer and why
  • ASwitch the table to on-demand capacity mode so that the nightly job no longer needs provisioned capacity set aside for it

    Incorrect

    On-demand mode removes capacity planning for the nightly burst, but the job would still run, and every read and delete it makes would still be billed as a request.

  • BAdd a global secondary index on the creation date so that the job queries only old items instead of scanning the table

    Incorrect

    A query on an index avoids reading the whole table, but the job still runs and issues every delete, and the index adds its own storage and write costs.

  • CTurn on TTL with an expiration timestamp on each item, and filter out expired items in the application's queries

    Correct

    DynamoDB deletes expired items automatically, without consuming write throughput, typically within a few days of expiration. Because expired items can be read until then, AWS advises filter expressions to remove them from query and scan results.

  • DProcess the table's stream with a Lambda function that deletes each session item once it is older than 24 hours

    Incorrect

    DynamoDB Streams captures item-level changes for processing, so the function could find old items without a scan, but its deletes still consume write capacity like any other delete.

The constraints are no paid deletes, no scan, no clean-up job, and no expired session accepted. On-demand mode, an index and a stream-driven function all keep a component that issues paid deletes. TTL deletes expired items without using write throughput, and filtering on the expiration attribute covers the days before an expired item is removed.

Practise domain 3 →Practise all domains →