50 AWS Interview Questions, Answered Visually (Part 1: Core Services)
10 min readBytePatterns
25 AWS interview questions on IAM, S3, EC2, Auto Scaling, Lambda, VPC, RDS and DynamoDB, with short model answers checked against the AWS documentation.
An AWS interview round rarely asks you to recite a service list. It asks what a service promises, what it costs you, and what breaks first. This is the first half of fifty questions: the core services every other answer is built on. Part 2 covers messaging, caching, containers, observability, cost and full architecture scenarios.
Every answer below is short on purpose — the length you can say out loud in under a minute — and every behaviour it describes was checked against the AWS documentation linked at the end. Where a number appears, it is a documented default or limit as of September 2026; limits change, so say "by default" in the room and check before you build.
How to use this list
Read the question, answer it aloud, then compare. If your answer names a service but not a trade-off, it is half an answer: the follow-up is always "and what does that cost you?". Each group links to the lesson that animates it.
Security and IAM
1. What is the shared responsibility model?
AWS is responsible for security of the cloud: facilities, hardware, network and the virtualization layer. You are responsible for security in the cloud: your data, IAM configuration, guest operating system patches on EC2, security group rules and encryption choices. The line moves with the service. On RDS, AWS maintains the underlying operating system and engine versions, while you still own database logins, network access and when the maintenance window runs.
2. What is the difference between an IAM user and an IAM role?
A user is an identity with long-term credentials such as a password or access keys. A role has no long-term credentials: whoever assumes it receives temporary security credentials. Workloads — EC2 instances, Lambda functions, containers — should use roles, which is what AWS's IAM best practices recommend.
3. How does IAM decide whether a request is allowed?
Every request starts as an implicit deny. It is allowed only if an applicable policy explicitly allows it, and an explicit Deny in any applicable policy overrides every Allow. The interview version: "deny by default, allow explicitly, deny wins".
4. Identity-based versus resource-based policies?
Identity-based policies attach to users, groups and roles and say what that identity may do. Resource-based policies attach to a resource, such as an S3 bucket policy, and say who may act on it — which is why they must name a Principal, an element identity-based policies cannot use.
5. How should an application on EC2 get AWS credentials?
Through an IAM role attached with an instance profile. The application reads temporary credentials from the instance metadata, usually via the SDK's credential provider, and those credentials are updated automatically before they expire. No access key is ever written to disk or to an environment file.
6. How do you actually reach least privilege?
Start with the specific actions the code calls on the specific resources it touches, and widen only with evidence. IAM Access Analyzer can generate a policy from the access activity recorded in CloudTrail. Add MFA for people, and keep the root user for the few tasks that require it.
Storage: S3
The S3 lesson animates questions 7 to 10.
7. What consistency does S3 give you?
Strong read-after-write consistency for PUT and DELETE of objects in all Regions: after a successful write, a following GET or LIST returns the new data. Two caveats worth naming: bucket configuration changes are eventually consistent, and S3 does not lock an object for concurrent writers — the request with the latest timestamp wins.
8. Which storage class would you choose, and when?
S3 Standard for frequently accessed data. Standard-IA (or One Zone-IA for data you can re-create) for data read rarely but needed in milliseconds. Glacier Instant Retrieval for archives read about once a quarter that still need millisecond access. Glacier Flexible Retrieval and Deep Archive for archives that are restored before reading. Express One Zone for single-digit-millisecond access in one Availability Zone.
9. What do lifecycle rules do, and what are the traps?
They transition objects to other storage classes and expire them, and can also expire noncurrent versions and abort incomplete multipart uploads. The traps are billing ones: the IA and Glacier classes have minimum storage durations — 30 days for the IA classes — and by default objects smaller than 128 KB are not transitioned, so moving small or short-lived objects can cost more than it saves.
10. How do users upload files to S3 without the bytes passing through your servers?
Your API returns a presigned URL: a time-limited PUT for one key, signed with the API's own credentials and permissions. The browser uploads straight to S3. With the SDK or CLI the expiry can be up to seven days, but a URL signed with temporary credentials stops working when those credentials expire. If you sign a Content-Type, the upload must send the same one.
11. How do you stop a bucket from becoming public?
New buckets already have S3 Block Public Access turned on and Object Ownership set to "bucket owner enforced", which disables ACLs. Keep both, grant access with IAM and bucket policies, and put CloudFront with origin access control in front when content must be public.
12. When would you pick S3 Intelligent-Tiering?
When you cannot predict access. It moves each object between access tiers based on its own access pattern, charges a small per-object monitoring fee and no retrieval fees. Objects under 128 KB are not monitored and stay in the frequent tier, and the optional archive tiers require a restore before reading.
Compute: EC2, Auto Scaling and load balancers
13. What does an Auto Scaling group do?
It keeps a fleet of EC2 instances between a minimum and a maximum capacity, launched from a launch template and balanced across the Availability Zones you give it. It adjusts desired capacity through scaling policies and replaces instances that fail health checks.
14. Which scaling policies exist, and which is the default choice?
Dynamic scaling (target tracking, step and simple policies), scheduled scaling, and predictive scaling. Target tracking is the usual answer: you name a metric and a target, such as average CPU at 50%, and the group adds or removes capacity to stay near it. A default instance warmup keeps a booting instance's metrics out of the average until it is ready.
15. ALB, NLB or Gateway Load Balancer?
Application Load Balancer works at layer 7 — HTTP and HTTPS, routing by path and host. Network Load Balancer works at layer 4 — TCP, UDP and TLS. Gateway Load Balancer works at layer 3 and is for fleets of virtual appliances such as firewalls.
16. How do health checks interact between the load balancer and the group?
The load balancer routes only to healthy targets in a target group (if every target is unhealthy, an ALB fails open and routes to all of them). The Auto Scaling group always uses EC2 status checks and can also use the load balancer's health checks after a grace period; an unhealthy instance is terminated and replaced.
Serverless: Lambda
The Lambda lesson animates questions 17 to 20.
17. What is a cold start, and how do you reduce it?
When no initialised execution environment is free, Lambda creates one: it downloads your code, starts the runtime and runs your initialisation code outside the handler. Later requests can reuse the environment warm. To reduce cold starts, keep init work small, use provisioned concurrency to keep environments initialised, or use SnapStart on supported runtimes (Java 11+, Python 3.12+, .NET 8+). SnapStart and provisioned concurrency cannot be combined on one function.
18. Reserved concurrency versus provisioned concurrency?
Concurrency is the number of requests a function is handling at the same moment. Reserved concurrency is both a maximum and a guaranteed minimum for one function, at no extra charge. Provisioned concurrency pre-initialises environments so requests skip the cold start, and it is billed additionally.
19. How does Lambda consume an SQS queue, and what happens on failure?
An event source mapping polls the queue and invokes the function synchronously with batches. By default, if the function errors, the whole batch becomes visible again after the queue's visibility timeout. With ReportBatchItemFailures, the handler returns only the failed message IDs and the rest count as processed. Set a dead-letter queue in the source queue's redrive policy; AWS recommends a maxReceiveCount of at least 5 and a visibility timeout of at least six times the function timeout.
20. When is Lambda the wrong tool?
When a single job runs longer than the maximum function timeout of 15 minutes, or when a burst would exceed the Region's shared concurrency quota — 1,000 concurrent executions by default, raisable on request. It is also a poor fit when you need state to survive between invocations, because an execution environment may be reused but is never guaranteed.
Networking: VPC
The VPC lesson animates questions 21 to 23.
21. What makes a subnet public or private?
Its route table. A public subnet has a route to an internet gateway; a private subnet has no direct route to one. A VPC spans all Availability Zones in its Region, and each subnet lives entirely in one zone.
22. What does a NAT gateway do, and what does it cost you?
A public NAT gateway sits in a public subnet with an Elastic IP. Private subnets route 0.0.0.0/0 to it, so their instances can open connections to the internet while nothing outside can open a connection to them. It is billed per hour and per GB processed, so send S3 and DynamoDB traffic through gateway endpoints, which need no NAT and have no extra charge.
23. Security groups versus network ACLs?
Security groups apply at the instance level, support allow rules only, evaluate all rules, and are stateful: a reply to an allowed request is allowed automatically. Network ACLs apply at the subnet level, support allow and deny rules, are evaluated from the lowest rule number with the first match applying, and are stateless, so replies need their own rules.
Databases
24. RDS Multi-AZ versus read replicas?
A Multi-AZ DB instance deployment keeps a synchronous standby in another Availability Zone for failover; that standby does not serve reads. Read replicas are updated asynchronously and do serve read-only traffic. A Multi-AZ DB cluster deployment is a third shape: one writer and two readable standbys in three zones.
25. How does DynamoDB partition data, and what is a hot partition?
The partition key is hashed, and the hash picks the partition. Each partition delivers at most 3,000 read units and 1,000 write units per second, so a key that takes most of the traffic throttles even when the table has spare capacity. Burst and adaptive capacity help only within those limits; write sharding — a random or calculated suffix on the key — spreads the load. One read capacity unit is one strongly consistent (or two eventually consistent) reads per second for items up to 4 KB; one write unit is one write per second up to 1 KB. On-demand mode, the default, removes capacity planning. Reads from a global secondary index are always eventually consistent.
Watch it run
The animation below is the S3 lesson's own: an overwrite that is immediately readable, then one object walking down the storage classes under a lifecycle rule until the expiration action deletes it. Step through it and say each transition out loud — that is questions 7 to 9 in one picture.
S3: Consistency, Classes, Lifecycle
Step 1 of 10
A bucket is a flat namespace of keys. Every object also has a storage class; Standard is the default.
The same interactive animation as the lesson — step through it with the controls.
How to say it in an interview
Name the service, then the promise, then the price. "S3 is strongly consistent for object writes, so I don't need a read-after-write workaround; the cost I watch is retrieval and minimum-duration charges if I tier too aggressively." That shape — mechanism, guarantee, trade-off — is what separates an answer from a definition, and it is the shape Part 2 uses for whole architectures.
Sources
- Shared Responsibility Model
- IAM policy evaluation: allow or deny
- IAM roles
- Identity-based and resource-based policies
- Use an IAM role for applications on EC2
- Security best practices in IAM
- Maintaining an RDS DB instance
- What is Amazon S3?
- S3 storage classes
- Managing the lifecycle of objects
- Presigned URLs
- Blocking public access to S3
- Auto Scaling groups
- Choose your scaling method
- Auto Scaling health checks
- What is an Application Load Balancer?
- ALB target group health checks
- Lambda execution environment lifecycle
- Lambda function scaling
- Lambda SnapStart
- SQS event source mappings
- Lambda quotas
- Subnets for your VPC
- NAT gateways
- Gateway endpoints
- Security groups versus network ACLs
- RDS Multi-AZ DB instance deployments
- RDS read replicas
- DynamoDB partition key design
- DynamoDB provisioned capacity mode
- DynamoDB read consistency