Instiq
Chapter 2 · Design for New Solutions·v2.1.0·Updated 8/7/2026·~11 min

What's changed: Service coverage: added catalog sections (s4 compute/DB/storage/network, s5 analytics/ML/IoT/media/integration)

2.1High Availability and Disaster Recovery

Key points

Understand designs for availability and resilience—Multi-AZ/Multi-Region, RTO/RPO, the four DR strategies (backup & restore/pilot light/warm standby/multi-site), and Route 53 failover.

Availability starts with AZ redundancy; for disasters you pick among four strategies based on RTO/RPO. Cost and recovery speed trade off.

2.1.1Choosing a DR strategy

Diagram of the four disaster-recovery strategies along a cost vs recovery-speed (RTO/RPO) trade-off: from left, backup & restore (cheapest; RTO/RPO in hours), pilot light (core minimal running, data replicated; RTO in tens of minutes), warm standby (a scaled-down version always on; RTO in minutes), and multi-site/active-active (full capacity in multiple regions; near-zero RTO/RPO; most expensive), with Route 53 health checks and failover routing switching over.
The four DR strategies and RTO/RPO
  • RTO/RPO: RTO = tolerable time to recover, RPO = tolerable data loss (time); requirements drive the strategy.
  • Backup & restore: cheapest; RTO/RPO in hours.
  • Pilot light: only the core (e.g., replicated DB) runs minimally; scale out on disaster; RTO in tens of minutes.
  • Warm standby: a scaled-down copy always running; scale up fast; RTO in minutes.
  • Multi-site (active-active): full capacity in multiple regions; near-zero RTO/RPO; most expensive.
Exam point

Most common on SAP-C02: cheapest/slow = backup & restore, core-only = pilot light, scaled-down always-on = warm standby, full multi-region = multi-site, plus RTO = time / RPO = data. Shorter RTO/RPO requirements push you to higher (costlier) strategies.

Note

Route 53 health checks + failover routing auto-switch to standby on primary failure. Replicate data with RDS cross-region read replicas or DynamoDB global tables.

Establish availability first within one region via Multi-AZ (ELB + Auto Scaling across AZs, RDS Multi-AZ, 3-AZ-redundant S3/DynamoDB), then choose multi-region DR by requirement for regional failures. RTO = tolerable time to recover; RPO = tolerable data loss (time). The four strategies trade cost vs speed: backup & restore (AWS Backup, cross-region snapshot copy; RTO/RPO in hours) < pilot light (DB continuously replicated via cross-region read replicas, app stopped/minimal and started on disaster; tens of minutes) < warm standby (a scaled-down full stack always on, scale up fast; minutes) < multi-site/active-active (production in multiple regions, DynamoDB global tables; near-zero). Switch with Route 53 health checks + failover/latency/weighted routing, deliver via CloudFront, and reproduce the standby with IaC (CloudFormation/StackSets). For RPO≈0 use active-active or synchronous replication; for cost, a lower strategy—derive it from the RTO/RPO numbers.

StrategyStandbyRTO/RPOCost
Backup & restoreNone (backups only)HoursLowest
Pilot lightCore (DB) onlyTens of minLow
Warm standbyScaled-down, always onMinutesMedium
Multi-siteFull, multi-regionNear-zeroHighest
Note

Scenario: meet RPO of 1 min and RTO of 5 min for payments while limiting cost. → Choose warm standby: keep a scaled-down full stack always on in another region, continuously replicate the DB (cross-region replica or global tables), scale to production capacity with Auto Scaling on disaster, and switch via Route 53 failover. For stricter RPO≈0, move up to active-active.

Note

FAQ: Q. RTO vs RPO? → A. RTO is how fast you recover (time); RPO is how much data loss you tolerate (time span). Q. Pilot light vs warm standby? → A. Pilot light keeps only the core (mainly DB) on and starts the app on disaster; warm standby keeps a scaled-down full stack always on for fast scale-up.

Warning

Trap: “choose multi-site (active-active) when cost is the top priority” is wrong—cheapest is backup & restore (RTO/RPO in hours). Conversely “meet RPO≈0 with backup & restore” is wrong—near-zero needs synchronous replication/active-active. Derive the strategy from the RTO/RPO numbers.

2.1.2Section summary

  • DR = backup&restore < pilot light < warm standby < multi-site (cost/speed)
  • RTO = recovery time / RPO = data loss; failover = Route 53

Sign in to track progress — Log in.

Quick check

(just a quick review)

Q1. For a mission-critical system with near-zero RPO and RTO of minutes, which DR strategy fits?

Q2. Which DR strategy prioritizes cost and tolerates RTO/RPO of several hours?

Q3. On primary-region failure, you want to auto-direct user traffic to a standby region. What?

Check your understandingPractice questions for Chapter 2: Design for New Solutions

Keep track of your progress

The full study guide is free to read. Sign up free to practice with the question bank, track what you have read, review your mistakes, and highlight passages.