What's changed: In-scope coverage: added machine learning (SageMaker AI/Comprehend/Kendra/Polly/Rekognition/Textract/Transcribe/Translate), storage (FSx/Storage Gateway), media (Elastic Transcoder/Kinesis Video Streams), frontend/API (API Gateway/Amplify/Device Farm), analytics (Kinesis Data Streams/Data Firehose/MSK/Glue/Lake Formation/Athena/EMR/OpenSearch/QuickSight/Data Exchange), and additional databases (Aurora Serverless/DocumentDB/Keyspaces) across sections
3.3Choosing and Scaling Databases
Understand RDS/Aurora (relational), DynamoDB (serverless NoSQL), scaling reads with read replicas, and choosing purpose-built databases.
High performance means choosing a database that fits the data: RDS/Aurora for relational, DynamoDB for low-latency NoSQL at scale.
3.3.1Choosing a database
- RDS / Aurora: relational (SQL); scale reads with read replicas, HA via Multi-AZ. Aurora is MySQL/PostgreSQL-compatible, high-performance and durable.
- DynamoDB: serverless NoSQL; single-digit-ms latency at scale, for key-value/document; auto-scaling, minimal ops.
- Purpose-built: Redshift (warehouse/analytics), ElastiCache (cache), Neptune (graph)—pick the best fit.
High performance means picking a purpose-built database that fits the data and use case. Structured with important relationships (joins) → RDS/Aurora (SQL); key-value/document with ultra-low latency at scale → DynamoDB (NoSQL); large-scale aggregation/analytics → Redshift (data warehouse); graph data → Neptune. Scaling differs by what you scale: read load = read replicas (add copies), availability = Multi-AZ (auto failover—different), and DynamoDB tunes throughput auto/provisioned. Scaling writes needs different thinking (sharding, DynamoDB partition design) than reads.
| Data/use | Service |
|---|---|
| Relational (joins, transactions) | RDS / Aurora |
| Key-value/doc, ultra-low latency, scale | DynamoDB |
| Large-scale analytics (warehouse) | Redshift |
| Graph (traverse relationships) | Neptune |
| Scale reads | Read replicas |
Scenario: choosing databases. Core order data needing consistency (joins, transactions) → Aurora (SQL), with read replicas if reads are heavy. Ultra-low-latency, large-scale key-value like sessions or carts → DynamoDB. Large-scale executive dashboards → Redshift. Ensure availability separately with Multi-AZ. Combine purpose-built DBs per use.
Watch the mix-ups: (1) read replica = scale reads (performance) / Multi-AZ = availability (auto failover)—different aims. (2) RDS/Aurora (SQL, joins) vs DynamoDB (NoSQL, key-value)—choose by use. (3) Large-scale analytics use Redshift, not OLTP RDS. (4) DynamoDB is designed around access patterns, not joins.
Q. Read replica vs Multi-AZ? Read replicas scale reads (performance); Multi-AZ provides availability (auto failover). Q. RDS or DynamoDB? RDS/Aurora for joins/transactions; DynamoDB for ultra-low-latency, large-scale key-value/document. Q. Analytics? Large-scale aggregation uses Redshift (data warehouse).
Common: scale reads = read replicas, availability = Multi-AZ, serverless low-latency NoSQL = DynamoDB, large-scale analytics = Redshift, graph = Neptune. Note read-scaling and availability use different means.
3.3.2Other key in-scope services (analytics and additional databases)
For collecting, transforming, and analyzing large-scale data there is a family of purpose-built managed services. For ingest (streaming): Amazon Kinesis Data Streams ingests real-time stream data for multiple consumers; Amazon Data Firehose auto-delivers streams to S3/Redshift/OpenSearch with buffering and transforms; Amazon MSK provides a managed Apache Kafka-compatible streaming platform. For transform/catalog: AWS Glue (serverless ETL and data catalog); for centralized data-lake permissions: AWS Lake Formation. For query/search/visualize: Amazon Athena runs SQL directly over data in S3 (serverless); Amazon EMR runs large-scale distributed processing with Spark/Hadoop; Amazon OpenSearch Service handles log/full-text search and analytics; Amazon QuickSight provides BI dashboards. To obtain external data, AWS Data Exchange lets you subscribe to third-party datasets.
Beyond relational/key-value, there are more purpose-built databases. Amazon Aurora Serverless auto-scales Aurora capacity with demand—choose it for intermittent or unpredictable workloads where you want to avoid capacity management. Amazon DocumentDB is a MongoDB-compatible document database, and Amazon Keyspaces is an Apache Cassandra-compatible wide-column database—pick them to run an existing MongoDB/Cassandra data model and API as a managed service on AWS.
| What you want | Service |
|---|---|
| Ingest real-time streams | Amazon Kinesis Data Streams / Amazon MSK |
| Auto-deliver streams to S3, etc. | Amazon Data Firehose |
| Serverless ETL and catalog | AWS Glue |
| Centralize data-lake permissions | AWS Lake Formation |
| Run SQL over data in S3 | Amazon Athena |
| Large-scale Spark/Hadoop | Amazon EMR |
| Log/full-text search and analytics | Amazon OpenSearch Service |
| BI dashboards | Amazon QuickSight |
| Obtain third-party datasets | AWS Data Exchange |
| Auto-scaling Aurora | Amazon Aurora Serverless |
| MongoDB / Cassandra compatible | Amazon DocumentDB / Amazon Keyspaces |
3.3.3Section summary
- Choose RDS/Aurora (SQL/joins) / DynamoDB (NoSQL/low latency) / Redshift (analytics) / Neptune (graph) by use
- Scale reads = read replicas; availability = Multi-AZ (different aims)
Sign in to track progress — Log in.
Quick check
(just a quick review)Q1. Which scales the "read load" of a relational database?
Q2. Which is a serverless NoSQL database with low latency at scale?
Q3. Which AWS service suits large-scale analytics (data warehouse)?
Keep track of your progress
The full study guide is free to read. Sign up free to practice with the question bank, track what you have read, review your mistakes, and highlight passages.

