50 AWS Interview Questions, Answered Visually (Part 2: Architecture & Scenarios)
10 min readBytePatterns
25 AWS interview questions on messaging, CloudFront, containers, CloudWatch, cost, disaster recovery and a full design walkthrough, checked against AWS docs.
Part 1 covered what the core services promise. This half is where AWS interviews usually go next: how the services talk to each other, where caches sit, how you would know it broke, what it costs, what happens when a whole Region goes away — and finally one design, drawn end to end.
The rules are the same as before. Answers are short enough to say aloud; every behaviour described was checked against the AWS documentation listed at the end; and numbers are documented defaults or limits as of September 2026, so say "by default" when you quote them.
Messaging and events
The messaging lesson animates questions 1 to 6.
1. SQS, SNS or EventBridge — how do you choose?
SQS is a queue: consumers pull messages and delete them when done, and the queue buffers work when consumers are slow or down. SNS is pub/sub: a topic pushes a copy of each message to every subscriber — SQS queues, Lambda functions, HTTP endpoints, email, SMS and more. EventBridge is an event bus: rules match events by their content and route them to targets, from AWS services, your own applications and SaaS partners.
2. Standard queue or FIFO queue?
Standard queues deliver at least once — a message can occasionally arrive twice — with best-effort ordering. FIFO queues preserve order within a message group and provide exactly-once processing, deduplicating within a five-minute interval. Choose FIFO when order is part of correctness, and design everything else to be idempotent.
3. What is the visibility timeout?
When a consumer receives a message, the message stays in the queue but is hidden from other consumers. If it is not deleted before the timeout expires, it becomes visible again and another consumer can take it. The default is 30 seconds and the maximum is 12 hours. Set it longer than your processing time, or the same message will be worked on twice.
4. What is a dead-letter queue, and when does a message end up there?
A queue for messages that keep failing. The source queue's redrive policy names the DLQ and a maxReceiveCount; once a message's receive count exceeds it, SQS moves the message aside so it stops blocking the rest. A FIFO queue's DLQ must also be FIFO, and a redrive can later move messages back to the source queue. Alarm on DLQ depth, or nobody notices.
5. How do you deliver one event to several independent services?
Fan-out: publish to an SNS topic and subscribe one SQS queue per service. Each service gets its own copy and its own pace, and an outage in one only grows its own queue. The queue's access policy must allow the topic to send to it, and subscription filter policies let a subscriber receive only the messages it cares about.
6. Why must consumers be idempotent on AWS?
Because at-least-once delivery is the norm: standard SQS queues and S3 event notifications can both deliver the same event more than once. Give each message a key — an order ID, an event ID — and record that you processed it, so a duplicate is detected and skipped instead of charging a customer twice.
Edge and caching
The CloudFront lesson animates questions 7 to 9.
7. How does CloudFront decide what to cache, and for how long?
A request is served from an edge location; on a miss it goes to a regional edge cache, then to the origin. The cache policy decides the cache key — which headers, cookies and query strings count — and the minimum, maximum and default TTLs. Within those bounds the origin's Cache-Control max-age or s-maxage sets the lifetime. Keep the cache key small: every extra value splits the cache and lowers the hit rate.
8. Invalidation or versioned file names?
For frequently updated assets, AWS recommends versioned names: app.3f9a.js is a new cache key, so nothing stale has to be removed. Invalidations work for the odd urgent fix, but paths beyond a monthly free allowance are billed per path (a wildcard path counts as one).
9. How do you keep an S3 origin private behind CloudFront?
Origin access control. The bucket policy allows reads only from the CloudFront service principal for that distribution, so viewers must come through CloudFront. The older origin access identity is documented as legacy.
Containers
The containers lesson animates questions 10 and 11.
10. ECS or EKS?
ECS is AWS's own fully managed orchestrator, built around task definitions, tasks, services and clusters. EKS is managed Kubernetes: AWS runs the control plane and you get the Kubernetes API and its ecosystem. Pick EKS when the team already runs Kubernetes or needs its portability and tooling; pick ECS when you want fewer moving parts and are all-in on AWS.
11. What is Fargate, and when would you not use it?
Fargate is serverless compute for containers that works with both ECS and EKS: no instances to provision or patch, and each task or pod runs in its own isolation boundary. It is the wrong choice for GPU workloads, because Fargate does not offer GPUs. Tasks use the awsvpc network mode, so tasks in private subnets need a NAT gateway or VPC endpoints to pull images.
Observability
The observability lesson animates questions 12 to 14.
12. Metrics, logs and traces — what does each answer on AWS?
CloudWatch metrics answer that something is wrong: numbers per period, named by namespace and dimensions. X-Ray traces answer where: one request's segments and subsegments across services, linked by a trace ID header. CloudWatch Logs answers why: log groups you search with Logs Insights. X-Ray's SDKs and daemon entered maintenance mode on 25 February 2026, and AWS recommends OpenTelemetry for new instrumentation.
13. What should you alarm on, and how do you avoid noisy pages?
Alarm on symptoms users feel — error rates, latency — rather than on every metric. An alarm evaluates a metric over several periods, with "datapoints to alarm" deciding how many must breach, and moves between OK, ALARM and INSUFFICIENT_DATA. Decide how missing data is treated: an Application Load Balancer reports no request metrics when there is no traffic.
14. CloudWatch or CloudTrail?
CloudWatch monitors the performance and health of resources and applications. CloudTrail records API activity — which user, role or service did what, to which resource, and when. "Who deleted the bucket?" is a CloudTrail question.
Cost
The cost lesson animates questions 15 to 18.
15. What are the main cost levers?
In order: switch off what idles, right-size what remains (Compute Optimizer flags over-provisioned instances), scale with demand, commit the steady baseline with Savings Plans or Reserved Instances, move interruptible work to Spot, and then look at data transfer.
16. When is Spot the right answer?
For flexible, fault-tolerant work — batch jobs, CI, rendering — that can resume after an interruption. Spot uses spare EC2 capacity at steep discounts, and EC2 can reclaim it with a two-minute interruption notice. Savings Plans do not cover Spot usage.
17. What does a Savings Plan commit you to?
A consistent amount of usage, measured in dollars per hour, for a one- or three-year term. Compute Savings Plans are the most flexible: they apply regardless of instance family, size, Region or operating system, and also cover Fargate and Lambda. EC2 Instance Savings Plans commit to one instance family in one Region. Commit to the floor you have measured, not the peak you expect.
18. Where do hidden data transfer costs come from?
NAT gateways are billed per hour and per GB processed, so private instances pulling from S3 through NAT pay for every byte; a gateway endpoint for S3 or DynamoDB removes that path at no extra charge. Keeping resources in the same Availability Zone as their NAT gateway reduces cross-zone transfer, and CloudFront caching reduces the load your origin has to serve.
Reliability and architecture
19. What are the six pillars of the Well-Architected Framework?
Operational excellence, security, reliability, performance efficiency, cost optimization and sustainability. In an interview, use them as a checklist of follow-up questions after you have drawn the request path.
20. What are RTO and RPO?
Recovery time objective: the maximum acceptable delay between an interruption and restored service. Recovery point objective: the maximum acceptable time since the last recovery point — how much data you can afford to lose. Both are business decisions first; the architecture follows them.
21. What are the four disaster recovery strategies?
Backup and restore: back up data and redeploy infrastructure, configuration and code in the recovery Region when needed. Pilot light: data is replicated and core infrastructure provisioned, but application servers are switched off until failover. Warm standby: a scaled-down but fully functional copy runs and can take traffic immediately, then scales up. Multi-site active/active: the workload serves traffic from several Regions at once. Cost and complexity rise in that order as recovery time falls toward zero.
22. Multi-AZ or multi-Region?
A Region contains multiple isolated Availability Zones placed a meaningful distance apart, so spreading across zones already protects against the loss of a data centre. AWS's disaster recovery guidance says that for that kind of event a highly available workload may only need backup and restore; go multi-Region when your definition of disaster includes losing a Region, or when regulation requires it.
23. Which Route 53 routing policies matter for resilience?
Failover (active/passive with health checks), latency (send users to the Region with the lowest latency), weighted (split traffic by percentage), geolocation and geoproximity (route by where users are), plus simple, IP-based and multivalue answer. Failover and latency are the two most interviews ask about.
24. Design an image upload pipeline on AWS.
The browser asks your API for a presigned URL and uploads directly to a raw S3 bucket. An S3 event notification feeds an SQS queue, which absorbs bursts. A Lambda function polls the queue, writes thumbnails to a second bucket — writing back into the triggering bucket risks a recursive loop — and records status in DynamoDB. CloudFront serves thumbnails from a private bucket through origin access control, and failures land in a dead-letter queue with an alarm. Then walk the pillars: signed uploads and private buckets (security), queue and DLQ (reliability), edge caching (performance), nothing idling between uploads (cost and sustainability), alarms (operational excellence).
25. Where do you store secrets?
AWS Secrets Manager, which supports automatic rotation. Systems Manager Parameter Store can hold SecureString parameters encrypted with KMS but does not rotate them, and AWS recommends Secrets Manager for secrets. Underneath both is envelope encryption: data is encrypted with a data key, and the data key is encrypted under a KMS key.
Watch it run
The animation below is question 24, drawn one box at a time: the presigned upload, the event into the queue, the worker writing to a separate bucket, the status row, CloudFront in front, the dead-letter queue for the file that never resizes — and then the six pillars read back against the drawing.
Design a System on AWS
Step 1 of 12
The prompt: users upload photos and see a thumbnail within seconds. Start from the request path, one box at a time.
The same interactive animation as the lesson — step through it with the controls.
How to say it in an interview
Draw the request path before naming a single optimisation, then let the pillars generate your follow-ups before the interviewer does. For every box, say what it guarantees and what it costs: "the queue gives me retries and burst absorption; the price is that a thumbnail is ready in seconds, not instantly". An answer that names the trade-off first is the one that sounds like it has been run in production.
Sources
- Amazon SQS standard queues
- SQS exactly-once processing (FIFO)
- SQS visibility timeout
- SQS dead-letter queues
- What is Amazon SNS?
- Subscribing an SQS queue to an SNS topic
- What is Amazon EventBridge?
- How CloudFront delivers content
- Control the cache key with a policy
- Invalidate files
- Restrict access to an S3 origin
- Choosing an AWS container service
- ECS task definition differences for Fargate
- AWS Fargate on Amazon EKS
- CloudWatch alarm evaluation
- AWS X-Ray concepts
- X-Ray SDK and daemon support timeline
- ALB CloudWatch metrics
- What is AWS CloudTrail?
- EC2 billing and purchasing options
- Spot Instance interruption notices
- Savings Plans types
- Pricing for NAT gateways
- Well-Architected: the pillars of the framework
- Disaster recovery: business continuity plan
- Disaster recovery options in the cloud
- Regions and Zones
- Route 53 routing policies
- S3 Event Notifications
- Lambda recursive loop detection
- Rotate Secrets Manager secrets
- Systems Manager Parameter Store
- AWS KMS cryptography essentials