Skip to content
BytePatterns

DOP-C02 · Domain 2: Configuration Management and IaC · 17% of the exam

Task 2.3: Design and build automated solutions for complex tasks and large-scale environments.

Running large fleets by automation: Systems Manager inventory, patching and State Manager associations, Lambda and Step Functions for multi-step jobs, and keeping software compliant without logging in to servers.

Study it

  • Fleet automation: Systems Manager inventory, Patch Manager and State Manager

    Lesson coming

  • Multi-step automation: Lambda, the SDKs and Step Functions

    Partly covered by: Lambda & Event-Driven Design

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

A company has 2,000 managed nodes, Amazon EC2 instances and on-premises servers, spread across 60 accounts and four AWS Regions in one organization. All nodes must be scanned daily and patched weekly against a defined baseline, and the operations team wants one configuration for the whole organization instead of setting up each account and Region. What should the DevOps engineer do?

  1. ACreate a maintenance window with an AWS-RunPatchBaseline task in every account and Region, deployed by a StackSet
  2. BCreate an EC2 Image Builder pipeline that bakes patched AMIs weekly and replace the instances from the new AMIs
  3. CCreate one Quick Setup patch policy for the organization with scan and install schedules and the patch baseline
  4. DRun AWS-RunPatchBaseline with Run Command from the management account every week, targeting every node by tag
Show the answer and why
  • ACreate a maintenance window with an AWS-RunPatchBaseline task in every account and Region, deployed by a StackSet

    Incorrect

    This works but means a maintenance window and task in each of 240 account and Region pairs, which is the per-account setup the team wants to avoid.

  • BCreate an EC2 Image Builder pipeline that bakes patched AMIs weekly and replace the instances from the new AMIs

    Incorrect

    New AMIs do not patch running on-premises servers, and replacing every instance weekly is not a scan-and-patch process.

  • CCreate one Quick Setup patch policy for the organization with scan and install schedules and the patch baseline

    Correct

    A patch policy defines the schedule and baseline, and a single patch policy configuration can cover all accounts and Regions in an organization.

  • DRun AWS-RunPatchBaseline with Run Command from the management account every week, targeting every node by tag

    Incorrect

    Run Command performs one-time changes when someone starts it. Running it weekly in every account and Region, plus daily scans, would need custom scheduling around it.

Quick Setup patch policies turn Patch Manager into one organization-wide configuration: the baseline and the scan and install schedules apply to the chosen accounts and Regions. Nodes must be managed nodes for the policy to reach them.

Question 2 · choose 1

Every Amazon EC2 instance tagged Tier=web must run a vendor monitoring agent with a specific configuration file. Engineers sometimes stop the agent or edit the file while troubleshooting and forget to undo it. The company wants the agent installed and running with the approved configuration on all current and future web instances, with drift corrected within an hour and no custom scheduler. What should the DevOps engineer do?

  1. ACreate a State Manager association for Tier=web with an SSM document that installs, configures and starts the agent every 30 minutes
  2. BAdd commands to the launch template's user data that install, configure and start the agent when each web instance boots
  3. CRun an SSM document once with Run Command against all instances tagged Tier=web, and repeat it after each incident
  4. DPackage the agent with Distributor and ask engineers to install the package from Fleet Manager whenever they notice that it is missing
Show the answer and why
  • ACreate a State Manager association for Tier=web with an SSM document that installs, configures and starts the agent every 30 minutes

    Correct

    An association applies a defined state to its targets on a schedule; if the software is missing or its service stopped, the next run installs or starts it again.

  • BAdd commands to the launch template's user data that install, configure and start the agent when each web instance boots

    Incorrect

    User data covers new instances at launch only. Nothing corrects an agent that is stopped or a file that is edited later.

  • CRun an SSM document once with Run Command against all instances tagged Tier=web, and repeat it after each incident

    Incorrect

    Run Command makes one-time changes. Repeating it by hand after incidents does not correct drift within an hour.

  • DPackage the agent with Distributor and ask engineers to install the package from Fleet Manager whenever they notice that it is missing

    Incorrect

    Distributor packages and publishes software; installing it on a schedule is done through State Manager. Installing on request leaves drift to people.

State Manager exists to keep managed nodes in a state you define and to reduce configuration drift. An association pairs an SSM document with targets and a schedule, so new instances with the tag are covered and changes made by hand are reverted on the next run.

Question 3 · choose 1

A quarterly maintenance process creates an AMI from each of 40 database hosts, waits for each AMI to become available, opens a change ticket, and then waits up to two days for a manager to approve before it replaces the hosts. Each step must retry on throttling errors. Operators want to see the progress of every run as a graph in the console, and the team wants to call the Amazon EC2 APIs without writing Lambda code for those calls. Which solution meets these requirements?

  1. AAn Express workflow in Step Functions that calls the EC2 APIs through AWS SDK integrations, retries throttling errors, and waits for the approval with a task token
  2. BA Standard workflow in Step Functions that calls the EC2 APIs through AWS SDK integrations with Retry rules and waits for the approval with a task token
  3. CA Lambda durable function that calls the EC2 APIs through the AWS SDK, retries in code, and waits for the approval as a callback
  4. DAn Amazon EventBridge Scheduler schedule for each step, with each target calling one EC2 API action directly
Show the answer and why
  • AAn Express workflow in Step Functions that calls the EC2 APIs through AWS SDK integrations, retries throttling errors, and waits for the approval with a task token

    Incorrect

    Express workflows run for at most five minutes and do not support the callback (.waitForTaskToken) integration pattern.

  • BA Standard workflow in Step Functions that calls the EC2 APIs through AWS SDK integrations with Retry rules and waits for the approval with a task token

    Correct

    Standard workflows run for up to a year, AWS SDK integrations call almost any AWS API without code, and the callback pattern waits until the approver returns the task token.

  • CA Lambda durable function that calls the EC2 APIs through the AWS SDK, retries in code, and waits for the approval as a callback

    Incorrect

    Durable functions can wait for days, but every EC2 call is written in function code and runs are not shown as a visual workflow graph.

  • DAn Amazon EventBridge Scheduler schedule for each step, with each target calling one EC2 API action directly

    Incorrect

    Separate schedules do not pass state from one step to the next, cannot wait for an approval, and give operators no view of a run.

Long waits for people and many calls to AWS APIs point to a Standard Step Functions workflow. Durable functions suit workflows tightly coupled with code in Lambda; Step Functions suits orchestration across AWS services with visual design and native integrations.

Question 4 · choose 1

For a software license audit, a company must report which versions of certain commercial applications are installed on about 3,000 managed nodes in 25 accounts and three AWS Regions, including on-premises servers. The report must stay current as nodes change, and auditors want to run SQL queries against the data in one place. What should the DevOps engineer do?

  1. ARun a shell script on all nodes with Run Command each month that lists installed packages and writes the output to a central S3 bucket
  2. BCall the EC2 DescribeInstances API in every account and Region each day and load the instance descriptions into S3 for Amazon Athena
  3. CExport the AWS CloudTrail logs of all accounts to S3 and query the RunInstances events with Amazon Athena
  4. DCollect application inventory with Systems Manager Inventory and use a resource data sync to one S3 bucket for Amazon Athena
Show the answer and why
  • ARun a shell script on all nodes with Run Command each month that lists installed packages and writes the output to a central S3 bucket

    Incorrect

    Run Command makes one-time changes, so the data is stale between runs, and the team must write and maintain the collection script.

  • BCall the EC2 DescribeInstances API in every account and Region each day and load the instance descriptions into S3 for Amazon Athena

    Incorrect

    Instance descriptions cover attributes such as the AMI and instance type, not software installed in the operating system, and they leave out the on-premises servers.

  • CExport the AWS CloudTrail logs of all accounts to S3 and query the RunInstances events with Amazon Athena

    Incorrect

    CloudTrail records API activity in the account. It says nothing about which applications are installed inside an operating system.

  • DCollect application inventory with Systems Manager Inventory and use a resource data sync to one S3 bucket for Amazon Athena

    Correct

    Inventory collects metadata such as installed applications from managed nodes, and resource data sync keeps the data from all nodes current in one S3 bucket for Athena queries.

Systems Manager Inventory gathers OS and application metadata from managed nodes, on-premises ones included. Resource data sync sends it to a single S3 bucket and updates it as new data is collected, so Athena can query the whole fleet across accounts and Regions.

Question 5 · choose 1

An Automation runbook creates an AMI from an instance and must then launch a test instance from it. The launch step sometimes fails because the AMI is still pending. The team wants the runbook to wait until the image state is available, without a fixed sleep and without custom code. Which action should the DevOps engineer add before the launch step?

  1. Aaws:sleep with a duration of 30 minutes before every launch
  2. Baws:executeScript with a Python loop that polls DescribeImages until the state changes
  3. Caws:waitForAwsResourceProperty on DescribeImages until the image state is available
  4. Daws:approve so that an operator confirms the AMI is ready before the launch
Show the answer and why
  • Aaws:sleep with a duration of 30 minutes before every launch

    Incorrect

    A fixed sleep waits too long or too short instead of watching the state.

  • Baws:executeScript with a Python loop that polls DescribeImages until the state changes

    Incorrect

    This is custom polling code that a built-in action replaces.

  • Caws:waitForAwsResourceProperty on DescribeImages until the image state is available

    Correct

    This action makes the automation wait for a specific resource state before it continues.

  • Daws:approve so that an operator confirms the AMI is ready before the launch

    Incorrect

    A manual approval adds a person instead of waiting on the state.

aws:waitForAwsResourceProperty polls an API operation until a property reaches the expected value, which avoids fixed delays and custom scripts.

Question 6 · choose 1

A company wants to move 4,000 attached gp2 data volumes across many accounts to gp3 to lower costs. The volumes serve running applications, so the change must not detach volumes or restart instances. What should the DevOps engineer automate?

  1. ASnapshot each volume, create a gp3 volume from the snapshot, and swap the volumes during a maintenance window
  2. BStop each instance, change the volume type, and start the instance again
  3. CCreate a new launch template that uses gp3 and replace all instances through instance refreshes
  4. DCall ModifyVolume with the gp3 type for each volume through an automation that runs in every account
Show the answer and why
  • ASnapshot each volume, create a gp3 volume from the snapshot, and swap the volumes during a maintenance window

    Incorrect

    Swapping volumes requires detaching them, which is not allowed.

  • BStop each instance, change the volume type, and start the instance again

    Incorrect

    Restarting instances is not allowed and is not required for the change.

  • CCreate a new launch template that uses gp3 and replace all instances through instance refreshes

    Incorrect

    Replacing instances is far more disruptive than modifying the volumes.

  • DCall ModifyVolume with the gp3 type for each volume through an automation that runs in every account

    Correct

    Elastic Volumes changes the volume type in place without detaching the volume or restarting the instance.

Elastic Volumes operations modify size, type and performance of attached volumes in place. When gp2 becomes gp3 without new settings, EBS keeps equivalent or baseline performance.

Question 7 · choose 1

Containerized services on Amazon ECS keep their settings in Parameter Store, one parameter per setting, under names such as /prod/payments/db/host and /prod/payments/feature/flags. Some are SecureString parameters. At startup a service must load all of its settings, including settings added later, without code or task definition changes. The database team and the product team must be able to change different settings of the same service under separate IAM permissions. What should the DevOps engineer recommend?

  1. AKeep the hierarchy and call GetParametersByPath recursively, with decryption, for the service's path
  2. BStore all settings of each service as one SecureString JSON value in a single parameter
  3. CKeep a list of parameter names in the service's configuration and call GetParameters with the list
  4. DReference each parameter in the secrets section of the task definition so that ECS injects it as an environment variable
Show the answer and why
  • AKeep the hierarchy and call GetParametersByPath recursively, with decryption, for the service's path

    Correct

    A path query returns every parameter below the path, including new ones, and can decrypt SecureString values, while each team's IAM policy covers only its own branch of the hierarchy.

  • BStore all settings of each service as one SecureString JSON value in a single parameter

    Incorrect

    One parameter loads easily, but IAM can no longer separate the database settings from the feature flags inside the same value.

  • CKeep a list of parameter names in the service's configuration and call GetParameters with the list

    Incorrect

    Names work with separate permissions, but a setting added later is missed until someone updates the list.

  • DReference each parameter in the secrets section of the task definition so that ECS injects it as an environment variable

    Incorrect

    ECS injects named parameters at task start, but every new setting needs a task definition change and a new deployment.

Parameter hierarchies use path-based names. GetParametersByPath returns all parameters under a path, recursively and decrypted if requested, so new settings are picked up, and IAM policies can grant write access by path.

Question 8 · choose 1

An operations team must run a data cleanup Step Functions workflow once, on the night of a planned migration, sometime between 01:00 and 02:00. The exact minute does not matter, and they want nothing left running or scheduled afterwards. What should the DevOps engineer create?

  1. AA one-time Scheduler schedule at 01:00 with a 60-minute flexible time window for the workflow
  2. BAn EventBridge Scheduler rate schedule of 1 day that starts the workflow every night at 01:00
  3. CA cron job on an EC2 instance that starts the workflow at 01:30 on the migration night
  4. DA scheduled rule (legacy) in EventBridge with a cron expression for the migration night
Show the answer and why
  • AA one-time Scheduler schedule at 01:00 with a 60-minute flexible time window for the workflow

    Correct

    One-time schedules invoke a target once, and a flexible time window lets Scheduler invoke it within the window.

  • BAn EventBridge Scheduler rate schedule of 1 day that starts the workflow every night at 01:00

    Incorrect

    A rate schedule repeats every day and stays in place afterwards.

  • CA cron job on an EC2 instance that starts the workflow at 01:30 on the migration night

    Incorrect

    This leaves an instance and a cron entry to manage.

  • DA scheduled rule (legacy) in EventBridge with a cron expression for the migration night

    Incorrect

    Scheduled rules are legacy; AWS recommends Scheduler, which supports one-time schedules.

EventBridge Scheduler supports rate-based, cron-based and one-time schedules. Flexible time windows spread invocations when an exact time is not required.

Question 9 · choose 1

AWS Config records all resources continuously in 60 development accounts, where resources change hundreds of times a day and costs are high. The team needs only a daily view of the latest state in these accounts, and Firewall Manager is not used there. Production accounts must keep real-time detection. What should the DevOps engineer do?

  1. ATurn off AWS Config in the development accounts and run a weekly script that lists resources
  2. BSet daily recording in the development accounts and keep continuous recording in production
  3. CKeep continuous recording but delete configuration items older than one day from the S3 bucket
  4. DExclude all resource types from recording in the development accounts
Show the answer and why
  • ATurn off AWS Config in the development accounts and run a weekly script that lists resources

    Incorrect

    This loses Config's history and compliance data instead of reducing recording.

  • BSet daily recording in the development accounts and keep continuous recording in production

    Correct

    Daily recording records the latest state once a day if it changed, which can reduce cost, while continuous recording detects changes at once.

  • CKeep continuous recording but delete configuration items older than one day from the S3 bucket

    Incorrect

    Costs come from recording changes, which this does not reduce.

  • DExclude all resource types from recording in the development accounts

    Incorrect

    Nothing would be recorded, so the daily view would be lost.

AWS Config offers continuous and daily recording frequencies. Daily recording lowers the number of recorded changes, while Firewall Manager relies on continuous recording.

Practise domain 2 →Practise all domains →