Skip to content
BytePatterns

SAA-C03 · Domain 3: Design High-Performing Architectures · 24% of the exam

Task 3.2: Design high-performing and elastic compute solutions.

Picking and sizing compute: instance families, Lambda memory, containers, batch and big-data clusters, and the metrics that should make a fleet grow or shrink.

Study it

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

A Lambda function resizes and compresses images. It is configured with 512 MB of memory, uses less than 200 MB of it, and takes 40 seconds per image because it is CPU-bound. The team wants each invocation to finish faster with the smallest possible change. What should a solutions architect recommend?

  1. ATurn on provisioned concurrency so that execution environments stay initialized
  2. BSet reserved concurrency on the function so that it always has capacity available
  3. CRaise the function's memory setting, which also raises its allocated CPU power
  4. DIncrease the function's timeout so that each image has more time to finish processing
Show the answer and why
  • ATurn on provisioned concurrency so that execution environments stay initialized

    Incorrect

    Provisioned concurrency removes cold-start initialization. It does not make the code inside an invocation run faster.

  • BSet reserved concurrency on the function so that it always has capacity available

    Incorrect

    Reserved concurrency guarantees and caps the number of concurrent instances. It does not add CPU to an invocation.

  • CRaise the function's memory setting, which also raises its allocated CPU power

    Correct

    Lambda allocates CPU power in proportion to memory, so a CPU-bound function gets faster when memory is raised even if it does not need the memory.

  • DIncrease the function's timeout so that each image has more time to finish processing

    Incorrect

    A longer timeout only lets a slow invocation run longer; it does not speed it up.

In Lambda, memory is the CPU dial. Concurrency settings change how many invocations run, not how fast each one runs.

Question 2 · choose 1

A research team runs thousands of containerized simulation jobs every night. Each job takes 1 to 3 hours, some jobs may start only after others succeed, and the team wants Spot capacity used where possible without building its own scheduler. Which service should a solutions architect recommend?

  1. AAWS Lambda functions started by Amazon EventBridge Scheduler, one invocation per job
  2. BAWS Batch with job queues and a managed compute environment that uses Spot capacity
  3. CAmazon EMR clusters that run each of the simulations as a separate cluster step
  4. DAn Amazon SQS queue polled by a fleet of EC2 Spot Instances running each job
Show the answer and why
  • AAWS Lambda functions started by Amazon EventBridge Scheduler, one invocation per job

    Incorrect

    A Lambda function can run for at most 15 minutes, far less than a 1 to 3 hour job.

  • BAWS Batch with job queues and a managed compute environment that uses Spot capacity

    Correct

    AWS Batch provisions compute, queues jobs, runs a job only after its dependencies succeed, and can use Spot or On-Demand Instances.

  • CAmazon EMR clusters that run each of the simulations as a separate cluster step

    Incorrect

    EMR runs big data frameworks such as Apache Spark and Hadoop. It is not a general scheduler for containerized batch jobs.

  • DAn Amazon SQS queue polled by a fleet of EC2 Spot Instances running each job

    Incorrect

    It can work, but the team would build dependency handling, retries and capacity management itself.

Many long container jobs with dependencies and Spot is AWS Batch's home ground. Lambda's 15-minute limit rules it out.

Question 3 · choose 2

A web application runs in an Auto Scaling group with a target tracking policy on average CPU. Every weekday at 08:00 traffic quadruples within minutes. New instances need about 10 minutes to start and warm up, so users see slow responses each morning until capacity catches up. Which actions should a solutions architect take? (Choose TWO.)

  1. ALower the CPU target of the target tracking policy from 50 to 45 percent
  2. BAdd a scheduled action that raises desired capacity before 08:00 on weekdays
  3. CRaise the health check grace period of the group to 15 minutes for new instances
  4. DAdd a lifecycle hook that keeps new instances waiting until they finish warming up
  5. EAdd a predictive scaling policy so capacity launches ahead of the forecast load
Show the answer and why
  • ALower the CPU target of the target tracking policy from 50 to 45 percent

    Incorrect

    Target tracking still reacts after the load arrives, so the 10-minute warm-up gap remains each morning.

  • BAdd a scheduled action that raises desired capacity before 08:00 on weekdays

    Correct

    Scheduled scaling changes capacity at set times for predictable load, so instances are ready before traffic rises.

  • CRaise the health check grace period of the group to 15 minutes for new instances

    Incorrect

    The grace period only delays health checks on new instances. It does not launch capacity any earlier.

  • DAdd a lifecycle hook that keeps new instances waiting until they finish warming up

    Incorrect

    A lifecycle hook pauses instances during launch. It does not make the group launch them before the load arrives.

  • EAdd a predictive scaling policy so capacity launches ahead of the forecast load

    Correct

    Predictive scaling learns daily and weekly patterns and launches capacity in advance, which suits instances that take long to initialize.

Long warm-up plus a predictable spike calls for scaling ahead of time: a schedule, or a forecast. Reactive policies always lag behind.

Question 4 · choose 1

A company runs an in-memory analytics engine on EC2. Each node must hold a dataset of about 600 GiB in RAM, and the workload uses relatively few CPU cores. Which instance family is designed for this profile?

  1. AMemory optimized instances, such as the R family
  2. BCompute optimized instances, such as the C family
  3. CBurstable general purpose instances, such as the T family
  4. DStorage optimized instances, such as the I family
Show the answer and why
  • AMemory optimized instances, such as the R family

    Correct

    Memory optimized instances are designed for workloads that process large datasets in memory, with a high ratio of memory to vCPUs.

  • BCompute optimized instances, such as the C family

    Incorrect

    Compute optimized instances suit compute-bound work and have a low ratio of memory to vCPUs.

  • CBurstable general purpose instances, such as the T family

    Incorrect

    Burstable instances give a baseline CPU level for spiky, light work, and their sizes are far too small for a 600 GiB dataset.

  • DStorage optimized instances, such as the I family

    Incorrect

    Storage optimized instances are built for high, sequential and random I/O to large local datasets, not for holding data in RAM.

Large in-memory datasets map to the memory optimized families, whose instances offer the most memory per vCPU.

Question 5 · choose 1

A web application runs in an Auto Scaling group behind an Application Load Balancer. It is I/O-bound: response times rise sharply when an instance handles more than about 1,000 requests per minute, while CPU stays under 30 percent. A target tracking policy on average CPU therefore never adds capacity. Which metric should the policy track instead?

  1. AAverage network bytes received by each instance in the group
  2. BApplication Load Balancer request count per target
  3. CAverage CPU utilization of the group, with the target value lowered to 10 percent
  4. DThe number of healthy hosts in the load balancer's target group
Show the answer and why
  • AAverage network bytes received by each instance in the group

    Incorrect

    Bytes received depend on the size of each request as much as on the number of requests, so the metric does not follow the load that slows the application.

  • BApplication Load Balancer request count per target

    Correct

    Request count per target is a predefined target tracking metric that measures exactly the per-instance load the application is sensitive to, so the group adds instances to hold it near the target.

  • CAverage CPU utilization of the group, with the target value lowered to 10 percent

    Incorrect

    CPU does not track this application's load, so a lower CPU target would scale at the wrong times, or still not at all.

  • DThe number of healthy hosts in the load balancer's target group

    Incorrect

    Healthy host count measures capacity that is up, not the load on it, so it cannot tell the group when to add instances.

Scale on the metric that limits the application. For a request-bound web tier behind an ALB, that is request count per target.

Question 6 · choose 1

Data engineers run nightly jobs that use Apache Spark, Apache Hive and Apache HBase with specific open-source versions and custom configuration, and they sometimes connect to the cluster nodes to tune them. The input is 50 TB in Amazon S3. The company wants a managed big data platform and clusters that exist only while the jobs run. Which solution meets these requirements?

  1. AAWS Glue ETL jobs that run the Spark code on serverless capacity
  2. BAWS Lambda functions invoked in parallel, one for each input object, that write their results back to S3
  3. CTransient Amazon EMR clusters that read from S3 and terminate when their steps finish
  4. DAWS Batch jobs on AWS Fargate that run the Spark code in containers
Show the answer and why
  • AAWS Glue ETL jobs that run the Spark code on serverless capacity

    Incorrect

    Glue runs Spark jobs without servers, but it does not host Hive or HBase clusters or give engineers access to the nodes.

  • BAWS Lambda functions invoked in parallel, one for each input object, that write their results back to S3

    Incorrect

    Lambda functions run for at most 15 minutes each and do not run Spark, Hive or HBase, so the jobs would have to be rewritten.

  • CTransient Amazon EMR clusters that read from S3 and terminate when their steps finish

    Correct

    EMR runs Spark, Hive and HBase as a managed cluster with node access and custom configuration, and a transient cluster shuts down after its steps, so it costs nothing between runs.

  • DAWS Batch jobs on AWS Fargate that run the Spark code in containers

    Incorrect

    Batch schedules container jobs, but it does not provide the Hadoop ecosystem, and Fargate gives no access to the underlying nodes.

Several open-source big data frameworks, custom versions and node access add up to Amazon EMR, and transient clusters match "only while the jobs run".

Question 7 · choose 1

A weather-modeling application runs a tightly coupled MPI job across 64 EC2 instances. The nodes exchange data constantly, and the job's speed is limited by node-to-node latency and network throughput. Which configuration gives the BEST network performance?

  1. AA spread placement group across three Availability Zones
  2. BA partition placement group with one partition per rack
  3. CInstances spread across several Availability Zones that talk to each other through a Network Load Balancer
  4. DA cluster placement group in one Availability Zone, with Elastic Fabric Adapter on the instances
Show the answer and why
  • AA spread placement group across three Availability Zones

    Incorrect

    A spread placement group puts each instance on distinct hardware to limit correlated failures, and it holds at most seven running instances per Availability Zone. It does not lower latency.

  • BA partition placement group with one partition per rack

    Incorrect

    Partition placement groups isolate groups of instances for large distributed systems such as HDFS. They do not pack instances for the lowest latency.

  • CInstances spread across several Availability Zones that talk to each other through a Network Load Balancer

    Incorrect

    Traffic between zones and through a load balancer adds latency, which is the opposite of what a tightly coupled job needs.

  • DA cluster placement group in one Availability Zone, with Elastic Fabric Adapter on the instances

    Correct

    A cluster placement group packs instances close together for low latency and high throughput, and EFA gives MPI applications low-latency, high-bandwidth communication between nodes.

Tightly coupled HPC wants instances physically close together and a network interface built for MPI: a cluster placement group with EFA.

Question 8 · choose 2

A Lambda function processes records from an Amazon Kinesis data stream in provisioned mode with 4 shards. Each batch takes a stable 200 ms, but the stream's iterator age keeps growing during the day because records arrive faster than they are processed. Records with the same partition key must stay in order. Which changes increase processing throughput? (Choose TWO.)

  1. ARaise the parallelization factor of the event source mapping
  2. BIncrease the number of shards in the stream
  3. CRaise the function's timeout from 1 minute to 15 minutes
  4. DTurn on provisioned concurrency for the function
  5. ELower the batch size of the event source mapping to 1 record
Show the answer and why
  • ARaise the parallelization factor of the event source mapping

    Correct

    The parallelization factor lets Lambda process up to 10 batches from each shard at the same time, while keeping records with the same partition key in order.

  • BIncrease the number of shards in the stream

    Correct

    Lambda processes each shard separately, so more shards mean more batches processed in parallel.

  • CRaise the function's timeout from 1 minute to 15 minutes

    Incorrect

    Batches already finish in 200 ms, so a longer timeout changes nothing about how many batches run at once.

  • DTurn on provisioned concurrency for the function

    Incorrect

    For a Kinesis source, the number of concurrent batches comes from the shards and the parallelization factor. Preinitialized environments do not let Lambda read more batches at once.

  • ELower the batch size of the event source mapping to 1 record

    Incorrect

    One record per invocation adds per-invocation overhead and lowers throughput.

Stream consumers scale with the number of shards and, in Lambda, with the parallelization factor per shard. Both keep per-key ordering.

Practise domain 3 →Practise all domains →