Skip to content
BytePatterns

SAP-C02 · Domain 3: Continuous Improvement for Existing Solutions · 25% of the exam

Task 3.1: Determine a strategy to improve overall operational excellence.

Making a running system easier to operate: better logging and alarms, automatic remediation, safer deployment strategies, configuration management at scale, and game days that rehearse failure.

Study it

  • Observability that finds problems: CloudWatch, X-Ray and centralized logs

    Partly covered by: CloudWatch, Alarms & X-Ray, Observability Basics, Tracing a Request

  • Automation and remediation: EventBridge, Systems Manager Automation and Config remediation

    Lesson coming

  • Improving deployments and rehearsing failure with AWS Fault Injection Service

    Lesson coming

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

A company runs microservices in 50 accounts. During incidents, engineers sign in to several accounts one after another to read metrics, logs and traces, which slows down diagnosis. The company wants one account in which engineers can search and correlate this telemetry from all 50 accounts without copying the data. What should a solutions architect do?

  1. ACreate CloudWatch dashboards in each of the 50 accounts and share them publicly with the operations team
  2. BUse CloudWatch cross-account observability with a monitoring account and the 50 accounts linked as sources
  3. CCreate an organization trail in CloudTrail and search its log files during each incident
  4. DExport all log groups from the 50 accounts to one S3 bucket every hour and query them with Amazon Athena
Show the answer and why
  • ACreate CloudWatch dashboards in each of the 50 accounts and share them publicly with the operations team

    Incorrect

    Sharing dashboards lets people view them, but engineers would still move between 50 separate dashboards and could not query across accounts.

  • BUse CloudWatch cross-account observability with a monitoring account and the 50 accounts linked as sources

    Correct

    With cross-account observability, a central monitoring account can view and interact with the metrics, logs and traces generated in the source accounts linked to it, and search across them.

  • CCreate an organization trail in CloudTrail and search its log files during each incident

    Incorrect

    CloudTrail records API activity in the accounts. It does not hold application metrics, logs or traces.

  • DExport all log groups from the 50 accounts to one S3 bucket every hour and query them with Amazon Athena

    Incorrect

    Exports copy the data and arrive late, they do not include metrics or traces, and they are a poor fit for live incident diagnosis.

"One place, all accounts, no copies" is CloudWatch cross-account observability with one monitoring account.

Question 2 · choose 1

A web application on AWS Elastic Beanstalk is deployed with the all at once policy, and each release causes a few minutes of downtime. The team wants releases with no loss of capacity, wants to test each new version on 10% of production traffic for 30 minutes before the full rollout, and wants a quick way back if the test fails. Which deployment policy meets these requirements?

  1. ARolling, with a batch size of 10% of the instances
  2. BRolling with an additional batch, with a batch size of 10%
  3. CImmutable, so that new instances are launched in a separate Auto Scaling group
  4. DTraffic splitting, with 10% of traffic and a 30-minute evaluation time
Show the answer and why
  • ARolling, with a batch size of 10% of the instances

    Incorrect

    Rolling deployments take batches of instances out of service, which reduces capacity, and they do not hold a fixed share of traffic on the new version for an evaluation period.

  • BRolling with an additional batch, with a batch size of 10%

    Incorrect

    An additional batch keeps full capacity, but the rollout still moves batch by batch without an evaluation period on a share of traffic.

  • CImmutable, so that new instances are launched in a separate Auto Scaling group

    Incorrect

    Immutable deployments keep capacity and are safe to roll back, but they do not send a chosen percentage of traffic to the new version for a set evaluation time.

  • DTraffic splitting, with 10% of traffic and a 30-minute evaluation time

    Correct

    Traffic-splitting deployments send a set percentage of client traffic to new instances for an evaluation period and shift back to the old version if the new one is unhealthy, without reducing capacity.

Canary testing on Elastic Beanstalk is the traffic-splitting policy; the other policies differ in capacity and rollback but not in a timed traffic share.

Question 3 · choose 1

An operations team is paged several times a week because the data volume of a log-heavy application fills up on one of 300 Amazon EC2 instances. The fix is always the same: compress old logs and move them to Amazon S3. The team wants this handled automatically, with each run recorded, and without giving anyone SSH access. Which solution meets these requirements?

  1. AWrite a Lambda function that opens SSH sessions to each instance every hour and runs the clean-up commands
  2. BAdd the AWS Trusted Advisor checks to the weekly operations review and clean up the instances that it lists
  3. CAlarm on agent disk metrics and let EventBridge start a Systems Manager Automation runbook for the instance
  4. DReplace the instances with larger instance types that have twice the disk space for the application logs
Show the answer and why
  • AWrite a Lambda function that opens SSH sessions to each instance every hour and runs the clean-up commands

    Incorrect

    This relies on SSH access and keys, which the team wants to avoid, and it runs on every instance instead of the ones that need it.

  • BAdd the AWS Trusted Advisor checks to the weekly operations review and clean up the instances that it lists

    Incorrect

    Trusted Advisor checks do not watch disk usage inside instances, and a weekly review is not automatic remediation.

  • CAlarm on agent disk metrics and let EventBridge start a Systems Manager Automation runbook for the instance

    Correct

    The CloudWatch agent publishes disk use from inside the instance, alarm state changes reach EventBridge, and a rule can start an Automation runbook whose executions are recorded and need no SSH access.

  • DReplace the instances with larger instance types that have twice the disk space for the application logs

    Incorrect

    More disk only delays the problem and raises cost; the volume would fill up again later.

A repeated manual fix is a candidate for automation: a metric from the agent, an alarm, and an Automation runbook started by EventBridge.

Question 4 · choose 2

A company claims that its order service survives the loss of an Availability Zone, but it has never tested this. Leadership wants a repeatable experiment in a production-like environment that simulates a zone outage and stops on its own if customer-facing error rates climb too high. Which actions meet these requirements? (Choose TWO.)

  1. ABuild an AWS Fault Injection Service experiment from the AZ Availability Power Interruption scenario
  2. BRun an AWS Resilience Hub assessment and treat its resiliency score as the result of the test
  3. CAsk an engineer to stop instances in one zone from the console during a quiet period
  4. DRun the experiment in a separate account that has no traffic and no copy of the service
  5. EAdd a stop condition to the experiment template that uses the CloudWatch alarm for the error rate
Show the answer and why
  • ABuild an AWS Fault Injection Service experiment from the AZ Availability Power Interruption scenario

    Correct

    This scenario induces the expected symptoms of a complete loss of power in one Availability Zone across the affected resource types, which is what the claim needs to be tested against.

  • BRun an AWS Resilience Hub assessment and treat its resiliency score as the result of the test

    Incorrect

    A Resilience Hub assessment evaluates the application against its resiliency policy. On its own, it does not inject a zone failure into the running system.

  • CAsk an engineer to stop instances in one zone from the console during a quiet period

    Incorrect

    A manual action is not repeatable, covers only some resource types and has no automatic stop if errors climb.

  • DRun the experiment in a separate account that has no traffic and no copy of the service

    Incorrect

    An experiment against an environment without the service and its traffic shows nothing about how the order service behaves.

  • EAdd a stop condition to the experiment template that uses the CloudWatch alarm for the error rate

    Correct

    Stop conditions use CloudWatch alarms to stop a running experiment when the alarm goes into the ALARM state.

Fault Injection Service gives a repeatable zone-outage scenario, and its stop conditions tie the blast radius to a CloudWatch alarm.

Question 5 · choose 1

During an incident, engineers need to search the last two hours of JSON application logs in CloudWatch Logs for the slowest requests and group them by endpoint, without exporting the logs anywhere. Which tool fits?

  1. AA metric filter created after the incident
  2. BAn export of the log group to S3 followed by Athena queries
  3. CCloudWatch Logs Insights queries on the log group
  4. DAWS Config advanced queries
Show the answer and why
  • AA metric filter created after the incident

    Incorrect

    Metric filters count matches in new log data; they do not search past logs by endpoint.

  • BAn export of the log group to S3 followed by Athena queries

    Incorrect

    Exporting adds delay and is what the engineers want to avoid.

  • CCloudWatch Logs Insights queries on the log group

    Correct

    Logs Insights lets you interactively search and analyze log data in CloudWatch Logs with queries.

  • DAWS Config advanced queries

    Incorrect

    Config queries resource configuration, not application logs.

Ad hoc log analysis in place is CloudWatch Logs Insights.

Question 6 · choose 1

Driver updates and software installs on 600 managed instances must only run during an approved weekly window on Sunday nights, and they currently run whenever an engineer starts them. Which Systems Manager capability enforces the schedule?

  1. ASession Manager sessions opened on Sundays
  2. BMaintenance Windows with the update tasks registered
  3. CParameter Store parameters that hold the schedule for each update
  4. DInventory collection every Sunday
Show the answer and why
  • ASession Manager sessions opened on Sundays

    Incorrect

    Interactive sessions still depend on engineers acting by hand.

  • BMaintenance Windows with the update tasks registered

    Correct

    Maintenance Windows define a schedule for potentially disruptive actions such as updating drivers or installing software.

  • CParameter Store parameters that hold the schedule for each update

    Incorrect

    Parameters store values; they do not run tasks on a schedule.

  • DInventory collection every Sunday

    Incorrect

    Inventory gathers data about instances; it does not install software.

Disruptive changes on a schedule belong in a maintenance window.

Question 7 · choose 1

When the checkout service degrades, its CPU, latency, 5xx and health check alarms each page the on-call engineer, often a dozen times per incident. The team wants a single page only when the 5xx alarm and at least one of the latency or health check alarms are in ALARM at the same time. No page may go out while a deployment alarm signals a planned release, the individual alarms must stay in place for diagnosis, and the team will not write code for this. Which approach meets these requirements?

  1. ACreate an EventBridge rule that matches state changes of the four alarms and sends each matching event to the pager integration
  2. BRaise the evaluation periods and datapoints to alarm on each of the four alarms so that they fire less often
  3. CPage from a composite alarm whose rule requires the 5xx alarm and either of the other two, with the deployment alarm as suppressor
  4. DSend all four alarms to one SNS topic and give the pager subscription a filter policy that passes only messages from the 5xx alarm
Show the answer and why
  • ACreate an EventBridge rule that matches state changes of the four alarms and sends each matching event to the pager integration

    Incorrect

    CloudWatch sends an event to EventBridge whenever an alarm changes state, so a rule can route those events to a pager. Each matching event is still a separate page, and combining alarm states would need code.

  • BRaise the evaluation periods and datapoints to alarm on each of the four alarms so that they fire less often

    Incorrect

    Datapoints to alarm sets how many breaching data points within the evaluation periods put an alarm into ALARM, which filters out short spikes. During a real incident, every alarm still pages on its own.

  • CPage from a composite alarm whose rule requires the 5xx alarm and either of the other two, with the deployment alarm as suppressor

    Correct

    Composite alarms combine the states of other alarms with Boolean rule expressions and act only at the aggregated level, which reduces alarm noise. A suppressor alarm stops the composite alarm's actions during planned deployments, and the underlying alarms stay in place for diagnosis.

  • DSend all four alarms to one SNS topic and give the pager subscription a filter policy that passes only messages from the 5xx alarm

    Incorrect

    A subscription filter policy delivers only the subset of messages that match it. Filtering on one alarm pages on 5xx errors alone, ignores the required combination and still pages during releases.

The constraints are one page for a specific combination of alarms, silence during planned releases, the original alarms kept, and no code. EventBridge routing and SNS filtering each pass single alarm notifications, and tuning datapoints only makes each alarm less sensitive. A composite alarm with a rule expression and a suppressor alarm meets all of them.

Question 8 · choose 1

A multi-tenant API serves about 4,000 tenants and writes JSON access logs with tenantId, route and status fields to CloudWatch Logs. Twice last month one tenant's faulty integration flooded the API with failing requests and slowed it for everyone, and finding that tenant took an hour. Operations wants a continuously updated ranking of the tenants that cause the most 5xx responses, graphed over time, and an alarm when any single tenant exceeds 500 errors per minute. The team does not want a custom metric for each tenant and cannot change the application. Which solution meets these requirements?

  1. ACreate a Contributor Insights rule on the log group for 5xx responses keyed on tenantId, and alarm on its top contributor value
  2. BCreate an anomaly detection alarm on the API's total 5xx metric so that unusual error levels page the on-call team
  3. CAdd a metric filter for 5xx responses with tenantId as a dimension, and alarm on the metric that the filter publishes
  4. DSave a CloudWatch Logs Insights query that groups 5xx responses by tenantId, and run it when the API slows down
Show the answer and why
  • ACreate a Contributor Insights rule on the log group for 5xx responses keyed on tenantId, and alarm on its top contributor value

    Correct

    Contributor Insights analyzes high-cardinality log data into time series of the top-N contributors. The INSIGHT_RULE_METRIC function with MaxContributorValue returns the top contributor's count for each period, and an alarm can be set on it.

  • BCreate an anomaly detection alarm on the API's total 5xx metric so that unusual error levels page the on-call team

    Incorrect

    Anomaly detection compares a metric with a band of expected values. On the total error metric it can page during a flood, but it cannot rank tenants or say which tenant is responsible.

  • CAdd a metric filter for 5xx responses with tenantId as a dimension, and alarm on the metric that the filter publishes

    Incorrect

    Each tenantId value becomes its own custom metric, which the team wants to avoid. AWS warns against high-cardinality dimensions and may disable a metric filter that generates 1,000 different dimension values.

  • DSave a CloudWatch Logs Insights query that groups 5xx responses by tenantId, and run it when the API slows down

    Incorrect

    Logs Insights is for interactive searches during an investigation. A query that someone runs by hand is not a continuously updated ranking, and it pages nobody when a tenant crosses the limit.

The requirement is a top-talkers view over thousands of keys, which is what Contributor Insights is for: one rule on the existing JSON logs ranks tenants continuously, and an alarm on MaxContributorValue fires when the worst tenant crosses 500 errors per minute, all without per-tenant metrics or code changes. Metric filters with a tenant dimension create the per-tenant metrics the team wants to avoid.

Question 9 · choose 1

A shared services account keeps about 300 configuration values for each of dev, test and prod, many of them long connection strings, as a flat list of standard Parameter Store parameters with names such as prod-app-db-url. Engineers and services often read another environment's value by mistake. The platform team wants each environment's application role to be able to read only that environment's values, each application to load all values for its environment with one query instead of naming every parameter, and no move to a paid parameter tier. Which change meets these requirements?

  1. AAttach labels such as dev, test and prod to the parameter versions and have applications request values by label
  2. BRename the parameters into hierarchies such as /prod/app/db-url, grant each role its path, and read with GetParametersByPath
  3. CMove the parameters to the advanced tier and add parameter policies that notify owners when a value is not changed for a long time
  4. DStore each environment's values as one JSON document in a single parameter per environment and parse it in the applications
Show the answer and why
  • AAttach labels such as dev, test and prod to the parameter versions and have applications request values by label

    Incorrect

    A parameter label is a user-defined alias for a version of one parameter, which helps track which version is in use. It does not group different parameters by environment for IAM or for bulk retrieval.

  • BRename the parameters into hierarchies such as /prod/app/db-url, grant each role its path, and read with GetParametersByPath

    Correct

    Hierarchies organize parameters by a path and let GetParametersByPath return everything under one level. IAM policies grant access by parameter ARN, which carries the path, so each role can be limited to its own environment, all on the standard tier.

  • CMove the parameters to the advanced tier and add parameter policies that notify owners when a value is not changed for a long time

    Incorrect

    Parameter policies such as NoChangeNotification need the advanced tier, which is charged. They help keep values current; they do not separate environments or allow bulk retrieval.

  • DStore each environment's values as one JSON document in a single parameter per environment and parse it in the applications

    Incorrect

    One parameter per environment would allow one read and one IAM grant, but a standard parameter holds at most 4 KB and an advanced one 8 KB, which is charged, so 300 long values do not fit in the free tier.

The constraints are per-environment access control, a single query per environment, and no paid tier. Labels track versions of one parameter, parameter policies need the paid tier, and one JSON parameter per environment exceeds the size limits. Path-based hierarchies give IAM scoping by path and GetParametersByPath on standard parameters.

Practise domain 3 →Practise all domains →