Instiq
Chapter 2 · AWS Global Infrastructure·v2.0.0·Updated 6/3/2026·~12 min

What's changed: Revamped CLF-C02 Chapter 2 to the new depth (Regions/AZs/edge/containment/selection, SPOF/multi-AZ+ELB+Auto Scaling/HA/fault tolerance/elasticity/DR)

2.2Designing for High Availability and Fault Tolerance

Key points

Understand the basics of avoiding single points of failure and achieving high availability and fault tolerance via multi-AZ deployment, load balancing, and Auto Scaling.

Good design avoids a single point of failure (SPOF) and keeps running through failures. A SPOF is a weak point whose single failure stops the whole system. On AWS, the classic high-availability pattern is three pieces: make components redundant across multiple AZs, distribute traffic with a load balancer at the entry, and let Auto Scaling adjust instance count automatically.

2.2.1Redundancy across AZs

Diagram of a high-availability setup where an Elastic Load Balancer distributes traffic to app instances in two Availability Zones, shifting to the other if one AZ fails.
High availability with multiple AZs and a load balancer
  • Place instances across multiple AZs (multi-AZ).
  • Use a load balancer (Elastic Load Balancing, ELB) to distribute traffic to healthy instances across AZs.
  • Use Auto Scaling to automatically adjust instance count with demand or failures (replacing lost capacity).

When these mesh, here is what happens: if an AZ fails, the ELB stops routing to it and keeps sending traffic to healthy instances in other AZs, while Auto Scaling launches replacements in another AZ to restore capacity—near-zero downtime for users. Conversely, a single-AZ, single-instance setup is itself the SPOF and goes down entirely on one failure.

2.2.2High availability, fault tolerance, elasticity

TermMeaningExample
High availabilityMinimize downtime, keep runningMulti-AZ + ELB
Fault toleranceWhole keeps working despite component failureRedundancy for no downtime
ElasticityAuto scale with demandAuto Scaling

2.2.3Peripheral knowledge: disaster recovery (DR)

Multi-AZ is high availability "within one Region." For wide-area disasters affecting an entire Region, plan disaster recovery (DR) with backup/replication to another Region. DR has targets—RPO (how much data loss is acceptable) and RTO (how fast to recover)—and you pick strategies (backup & restore / pilot light / warm standby / multi-site) balancing cost (detailed in higher certs). For CLF-C02, it is enough to grasp "another Region for wide-area disasters" and "RPO/RTO are the targets."

Example

Scenario: a Web API that must not go down. Run EC2 across two or more AZs, front it with an ELB, and use Auto Scaling to follow demand and auto-replace failures. This survives a single AZ outage (HA). For severe incidents, add periodic backups to another Region (DR). You eliminate SPOFs step by step.

Warning

Watch the mix-ups: (1) high availability (design to stay up) vs elasticity (auto scale with demand) are different (elasticity = Auto Scaling, HA = multi-AZ + ELB). (2) Don’t confuse multi-AZ (HA within a Region) with multi-Region (DR for wide-area disasters). (3) The load balancer itself is redundant (AWS-managed, automatically).

Exam point

Common points: the pattern multi-AZ + load balancer (ELB) + Auto Scaling = high availability and the principle of avoiding SPOF. Also note HA (multi-AZ within a Region) vs DR (another Region) and HA vs elasticity.

2.2.4Section summary

  • Core principle: avoid single points of failure (SPOF)
  • HA pattern: multiple AZs + load balancer (ELB) + Auto Scaling
  • Distinguish HA / fault tolerance / elasticity; wide-area disasters need another Region (DR, RPO/RTO)

Sign in to track progress — Log in.

Quick check

(just a quick review)

Q1. Which placement best improves high availability?

Q2. What is the main role of Auto Scaling?

Q3. Why should you avoid a single point of failure (SPOF)?

Q4. What is the main role of a load balancer (Elastic Load Balancing)?

Q5. What best prepares for a wide-area disaster affecting an entire Region?

Check your understandingPractice questions for Chapter 2: AWS Global Infrastructure