SOA-C03 · Domain 1: Monitoring, Logging, Analysis, Remediation, and Performance Optimization · 22% of the exam
Task 1.3: Implement performance optimization strategies for compute, storage, and database resources.
Making resources faster and cheaper with evidence: right-sizing with Compute Optimizer, EBS volume types and their metrics, faster S3 transfers, choosing a shared file system, RDS performance data (Performance Insights, now part of CloudWatch Database Insights) and RDS Proxy, and placement groups for EC2.
Study it
Compute performance: right-sizing, Compute Optimizer and placement groups
Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.
Question 1 · choose 1
A database on an Amazon EC2 instance stores its data on a 300 GiB gp2 Amazon EBS volume. A nightly batch job needs about 4,000 IOPS of small, random I/O for two hours. During the job the volume's BurstBalance metric falls to zero and latency rises sharply. The team wants to fix this at the lowest cost and without downtime. What should the team do?
AGrow the gp2 volume to about 1,400 GiB so that its baseline performance reaches about 4,200 IOPS
BChange the volume to gp3 with Elastic Volumes and provision 4,000 IOPS
CChange the volume to Throughput Optimized HDD (st1)
DTurn on fast snapshot restore for the snapshots of the volume
Show the answer and why
AGrow the gp2 volume to about 1,400 GiB so that its baseline performance reaches about 4,200 IOPS
Incorrect
A larger gp2 volume does get a higher baseline, but the team would pay for more than a terabyte of storage it does not need. gp3 sets IOPS independently of size at a lower price per GiB.
BChange the volume to gp3 with Elastic Volumes and provision 4,000 IOPS
Correct
gp3 does not use burst credits and sustains its provisioned IOPS. Elastic Volumes changes the type while the volume stays attached and in use.
CChange the volume to Throughput Optimized HDD (st1)
Incorrect
st1 defines performance by throughput and suits large, sequential I/O. AWS recommends SSD volumes for small, random I/O.
DTurn on fast snapshot restore for the snapshots of the volume
Incorrect
Fast snapshot restore removes first-access latency on volumes created from a snapshot. It does not add IOPS to a running volume.
BurstBalance at zero means a gp2 volume has spent its I/O credits. Moving to gp3 buys the needed IOPS directly, without buying capacity to get them.
An API built on AWS Lambda writes orders to an Amazon RDS for PostgreSQL DB instance. During traffic surges, thousands of concurrent function instances each open a new database connection, and the database starts refusing connections while its CPU stays moderate. The team must fix this without changing the DB instance class and with minimal code changes. What should the team do?
AConvert the DB instance to a Multi-AZ DB instance deployment
BCreate a read replica and point the function at the replica's endpoint instead
CPut an RDS Proxy in front of the database and use the proxy endpoint
DTurn on storage autoscaling for the DB instance
Show the answer and why
AConvert the DB instance to a Multi-AZ DB instance deployment
Incorrect
A Multi-AZ DB instance adds a synchronous standby for failover. The standby does not serve traffic, so the connection limit is unchanged.
BCreate a read replica and point the function at the replica's endpoint instead
Incorrect
A read replica allows only read-only connections, so the function's writes would fail there.
CPut an RDS Proxy in front of the database and use the proxy endpoint
Correct
RDS Proxy pools and reuses database connections and limits how many it opens, so traffic surges no longer oversubscribe the database.
DTurn on storage autoscaling for the DB instance
Incorrect
Storage autoscaling grows allocated storage when free space runs low. It has no effect on connections.
Many short-lived clients opening their own connections is the problem RDS Proxy solves: a shared pool in front of the database, usually with no code change beyond the endpoint.
A team runs a tightly coupled simulation on eight Amazon EC2 instances that exchange data constantly. Profiling shows that network latency between the instances limits the job. The job may run in a single Availability Zone. How should the team launch the instances?
AIn a spread placement group, so that each instance runs on distinct hardware
BIn a partition placement group, with the instances divided across partitions
CIn an Auto Scaling group whose subnets span three Availability Zones
DIn a cluster placement group, with every instance in one Availability Zone
Show the answer and why
AIn a spread placement group, so that each instance runs on distinct hardware
Incorrect
Spread placement groups keep a small number of critical instances apart from each other. They do not bring instances closer together.
BIn a partition placement group, with the instances divided across partitions
Incorrect
Partition placement groups reduce the chance of correlated hardware failures for large distributed workloads, not network latency.
CIn an Auto Scaling group whose subnets span three Availability Zones
Incorrect
Spreading the nodes across zones works against the goal. AWS places a low-latency group in a single Availability Zone.
DIn a cluster placement group, with every instance in one Availability Zone
Correct
Cluster placement groups pack instances close together and are recommended for applications that need low network latency and high network throughput between instances.
Cluster is for speed between nodes, partition for limiting correlated failures in big clusters, spread for keeping a few critical instances apart.
Video creators in Asia, Europe and South America upload files of 2 to 8 GB through a web application straight into one Amazon S3 bucket in us-east-1. Uploads are slow, and when a connection drops the upload starts again from the beginning. Which change improves both the speed and the resilience of the uploads?
ATurn on Transfer Acceleration and send multipart uploads to the accelerate endpoint
BUpload the objects in the S3 Intelligent-Tiering storage class
CReplicate the bucket with S3 Cross-Region Replication to new buckets in Asia and Europe
DAdd an S3 Lifecycle rule that stops incomplete multipart uploads after 7 days
Show the answer and why
ATurn on Transfer Acceleration and send multipart uploads to the accelerate endpoint
Correct
Transfer Acceleration routes data through CloudFront edge locations over an optimized path, and multipart upload sends parts in parallel and retransmits only a failed part.
BUpload the objects in the S3 Intelligent-Tiering storage class
Incorrect
Intelligent-Tiering moves objects between access tiers to save storage cost. It does not change how fast objects arrive.
CReplicate the bucket with S3 Cross-Region Replication to new buckets in Asia and Europe
Incorrect
Replication copies objects asynchronously after they are written to the source bucket. The long upload to us-east-1 still happens first.
DAdd an S3 Lifecycle rule that stops incomplete multipart uploads after 7 days
Incorrect
This rule cleans up the parts of uploads that never finished, to save storage cost. It does not speed up or resume an upload.
Long-distance uploads of large files call for Transfer Acceleration plus multipart upload; AWS recommends multipart for objects of 100 MB or more.