Dự phòng & High AvailabilityRedundancy & High Availability
Thuộc bộ kiến thức Network Engineer Roadmap.
Tổng quan
Mạng nào rồi cũng có lúc hỏng. Một bộ nguồn chết, một module quang tối đèn, một thay đổi cấu hình gõ nhầm làm hỏng bảng định tuyến, một switch tự reboot vì lỗi firmware lúc 3 giờ sáng. Khác biệt giữa một mạng chỉ giật nhẹ và một mạng sập hẳn nằm ở redundancy (dự phòng) — có sẵn đường đi thứ hai, thiết bị thứ hai, hoặc data center thứ hai sẵn sàng gánh tải — kết hợp với failover nhanh và tự động để thiết bị dự phòng tiếp quản trước khi người dùng (hoặc SLA) kịp nhận ra.
High Availability (HA) là kỷ luật thiết kế hệ thống để tiếp tục phục vụ dù có component hỏng. Nó không phải một tính năng bật lên là xong; nó là một thuộc tính hình thành từ các quyết định ở mọi tầng:
- Physical: hai đường nguồn (dual power feed), dual supervisor, dual NIC, chạy sợi quang theo tuyến vật lý khác nhau.
- Link layer: link aggregation, redundant uplink, topology không loop (STP/RSTP — xem ./08-switching.md).
- Network layer: nhiều router, First Hop Redundancy Protocol (FHRP), equal-cost multipath (ECMP), hội tụ (convergence) của dynamic routing.
- Application/service layer: load balancer phân phối traffic trên một pool server, health check loại bỏ node chết, geographic redundancy trên nhiều region.
Mục tiêu chung là loại bỏ single point of failure (SPOF). SPOF là bất kỳ thành phần nào mà khi nó hỏng thì cả dịch vụ sập theo. Nếu hai firewall dự phòng của bạn cùng cắm vào một switch, thì switch đó là SPOF. Nếu hai switch dùng chung một ổ điện, ổ điện đó là SPOF. Thiết kế HA về bản chất là việc tìm kiếm và loại bỏ SPOF một cách có hệ thống.
Note này bao gồm: toán học của availability, các First Hop Redundancy Protocol (HSRP, VRRP, GLBP), các khái niệm failover (active-passive vs active-active, heartbeat, split-brain, preemption), load balancing chuyên sâu (L4 vs L7, thuật toán, health check, session persistence, SSL termination), link/path redundancy, và cách lắp ghép chúng thành topology có sức chống chịu. Các note liên quan:
- Network devices (bao gồm load balancer như một appliance): ./04-network-devices.md
- Switching, VLAN, STP, EtherChannel: ./08-switching.md
- Traffic management và QoS: ./13-traffic-management-and-qos.md
- Cloud networking (managed LB, multi-AZ/region): ./16-cloud-networking.md
Kiến thức nền tảng
Vì sao HA quan trọng: SPOF, chi phí downtime và blast radius
Downtime rất đắt và bất đối xứng. Vài giây mất gói khi failover sạch sẽ thì gần như vô hình; nhưng một sự cố 30 phút trong giờ làm việc có thể mất doanh thu, vi phạm SLA và kích hoạt phạt hợp đồng. Các động lực nghiệp vụ cho HA:
- Bảo vệ doanh thu — e-commerce, thanh toán, SaaS mất tiền theo từng phút downtime.
- Tuân thủ SLA — hợp đồng cam kết “three nines” hay “four nines”; không đạt thì phải bồi thường/phạt.
- An toàn và tuân thủ pháp lý — bệnh viện, điều khiển công nghiệp, hệ thống tài chính có yêu cầu uptime bắt buộc.
- Uy tín — sự cố lặp đi lặp lại làm xói mòn niềm tin khách hàng nhanh hơn hầu hết mọi thứ khác.
Công việc kỹ thuật là thu nhỏ blast radius (bán kính ảnh hưởng) của mỗi sự cố để không một component nào hỏng có thể kéo sập tất cả.
”Các số 9”: toán học của availability
Availability là tỷ lệ thời gian hệ thống hoạt động, thường biểu diễn bằng phần trăm và nói vui là số lượng “số 9”:
Availability = Uptime / (Uptime + Downtime)
= MTBF / (MTBF + MTTR)
| Availability | ”Số 9” | Downtime / năm | Downtime / tháng | Downtime / ngày |
|---|---|---|---|---|
| 90% | one nine | 36.5 ngày | 72 giờ | 2.4 giờ |
| 99% | two nines | 3.65 ngày | 7.2 giờ | 14.4 phút |
| 99.9% | three nines | 8.77 giờ | 43.8 phút | 1.44 phút |
| 99.95% | three rưỡi | 4.38 giờ | 21.9 phút | 43.2 giây |
| 99.99% | four nines | 52.6 phút | 4.38 phút | 8.6 giây |
| 99.999% | five nines | 5.26 phút | 26.3 giây | 0.86 giây |
| 99.9999% | six nines | 31.5 giây | 2.63 giây | 0.086 giây |
Hai nhận định quan trọng:
- Các số 9 nhân với nhau khi mắc nối tiếp (series). Nếu một request phải đi qua firewall (99.99%), load balancer (99.99%) và app server (99.9%), thì availability đầu-cuối là
0.9999 × 0.9999 × 0.999 ≈ 99.88%— tệ hơn bất kỳ component đơn lẻ nào. Chuỗi phụ thuộc kéo availability xuống. - Redundancy mắc song song (parallel) nhân xác suất hỏng. Hai component độc lập, mỗi cái 99% availability (1% hỏng), đặt song song chỉ hỏng cùng lúc với xác suất
0.01 × 0.01 = 0.0001→ availability tổng hợp99.99%. Đây là lý do redundancy mạnh đến vậy — nhưng chỉ khi các component hỏng độc lập (không chung nguồn, không chung sợi quang, không chung một bug bị kích hoạt bởi cùng input).
MTBF và MTTR
- MTBF (Mean Time Between Failures) — thời gian hoạt động trung bình giữa hai lần hỏng. Cải thiện bằng phần cứng chất lượng cao hơn, làm mát tốt hơn, component nội bộ dự phòng. Bạn ít kiểm soát được MTBF sau khi đã mua thiết bị.
- MTTR (Mean Time To Repair/Recover) — thời gian trung bình để khôi phục dịch vụ sau khi hỏng. Đây là nơi người vận hành có đòn bẩy khổng lồ: automatic failover, monitoring tốt, có sẵn linh kiện tại chỗ, runbook và diễn tập.
Vì Availability = MTBF / (MTBF + MTTR), giảm một nửa MTTR cải thiện availability y hệt như tăng gấp đôi MTBF — và thường rẻ hơn nhiều. Failover nhanh, tự động là khoản đầu tư HA có ROI cao nhất. Một cú failover FHRP dưới một giây biến việc reboot thiết bị từ một sự cố thành một chuyện không đáng kể.
RTO và RPO (bối cảnh DR)
HA (giữ cho chạy) đi kèm với Disaster Recovery (DR) (khôi phục sau tổn thất lớn). Hai metric của DR đóng khung yêu cầu:
- RTO (Recovery Time Objective) — thời gian tối đa chấp nhận được để khôi phục dịch vụ sau thảm họa. “Phải trở lại trong 15 phút.”
- RPO (Recovery Point Objective) — lượng dữ liệu mất tối đa chấp nhận được, đo bằng thời gian. “Chỉ được phép mất tối đa 5 phút giao dịch.”
Cơ chế HA (FHRP, load balancer, clustering) nhắm tới RTO gần bằng 0 cho lỗi component. Cơ chế DR (backup, replication, standby site) nhắm tới RTO/RPO có giới hạn cho lỗi thảm khốc. Một thiết kế hoàn chỉnh dùng cả hai: HA cho trường hợp thường gặp (một thiết bị chết), DR cho trường hợp hiếm (mất cả một site).
Khái niệm chính
First Hop Redundancy Protocol (FHRP)
Host được cấu hình một default gateway IP duy nhất. Nếu gateway đó (router/L3 switch) chết, mọi host trong subnet mất kết nối ra ngoài subnet — một SPOF kinh điển. First Hop Redundancy Protocol giải quyết bằng cách cho hai hoặc nhiều router chia sẻ một virtual IP (VIP) và virtual MAC đóng vai gateway. Host trỏ vào VIP; các router bầu ra một active forwarder; nếu nó hỏng, một standby tiếp quản VIP/MAC mà không cần cấu hình lại host.
Điều kỳ diệu là virtual MAC di chuyển theo vai trò. Vì host có ánh xạ VIP→virtual-MAC trong ARP cache, và router active mới bắt đầu trả lời cho chính virtual MAC đó, nên các entry ARP của host vẫn hợp lệ — traffic chỉ đơn giản chảy sang một thiết bị vật lý khác. Failover trong suốt đối với end station.
HSRP (Hot Standby Router Protocol) — Cisco độc quyền
- Cisco độc quyền (mô tả trong RFC 2281 nhưng do Cisco sở hữu).
- Các router tạo thành một group; một cái là Active, một cái là Standby, còn lại là Listen.
- Bầu chọn theo priority (mặc định 100, cao thắng); hòa thì IP cao hơn thắng.
- Virtual MAC:
0000.0C07.ACxx(v1) vớixxlà số group; v2 dùng0000.0C9F.Fxxx. - Hello timer mặc định 3 s, hold timer mặc định 10 s → failover ~10 s theo mặc định (chỉnh được xuống dưới một giây).
- Giao tiếp qua multicast
224.0.0.2(v1) /224.0.0.102(v2), UDP 1985. - Hỗ trợ interface tracking (giảm priority nếu uplink chết) và preemption (router priority cao giành lại vai Active khi trở về).
- Load sharing chỉ đạt được bằng cách chạy nhiều group và trỏ các host/VLAN khác nhau vào các VIP khác nhau (Multiple HSRP / MHSRP).
! Cisco HSRP — hai router, VIP 10.1.1.1, group 10
! Router A (primary)
interface Vlan10
ip address 10.1.1.2 255.255.255.0
standby version 2
standby 10 ip 10.1.1.1
standby 10 priority 110
standby 10 preempt
standby 10 timers msec 250 msec 750 ! hello 250ms / hold 750ms
standby 10 track GigabitEthernet0/1 20 ! -20 priority nếu uplink down
!
! Router B (secondary) — cùng VIP, priority thấp hơn
interface Vlan10
ip address 10.1.1.3 255.255.255.0
standby version 2
standby 10 ip 10.1.1.1
standby 10 priority 100
standby 10 preempt
VRRP (Virtual Router Redundancy Protocol) — chuẩn mở, RFC 5798
- Chuẩn mở (RFC 5798 cho VRRPv3, bao gồm IPv4 và IPv6; RFC 3768 là VRRPv2). Đa hãng (multi-vendor).
- Vai trò là Master và Backup (tương ứng Active/Standby của HSRP).
- Bầu theo priority (1–254, mặc định 100, cao thắng); address owner (router có IP interface thật bằng VIP) dùng priority 255 và luôn thắng.
- Virtual MAC:
0000.5E00.01xxvớixxlà VRID (Virtual Router ID). - Advertisement interval mặc định 1 s; master-down interval ≈ 3 × interval + skew → failover mặc định ~3 s (chỉnh được xuống dưới giây với đơn vị centisecond ở v3).
- Giao tiếp qua multicast
224.0.0.18, IP protocol 112. - Preemption BẬT theo mặc định (khác HSRP phải tự bật).
- Load sharing cũng cần nhiều VRRP group (VRID).
! VRRP (chạy đa hãng; cú pháp Cisco), VIP 10.1.1.1, VRID 10
interface Vlan10
ip address 10.1.1.2 255.255.255.0
vrrp 10 ip 10.1.1.1
vrrp 10 priority 110
vrrp 10 timers advertise 1
vrrp 10 preempt ! bật mặc định; ghi ra cho rõ
GLBP (Gateway Load Balancing Protocol) — Cisco, active-active load sharing
- Cisco độc quyền. Điểm khác biệt: load balancing gốc trên nhiều router dùng chung một VIP.
- Một router được bầu làm AVG (Active Virtual Gateway); nó phát ra tối đa bốn virtual MAC khác nhau cho các host khi chúng ARP hỏi VIP.
- Mỗi router tham gia là một AVF (Active Virtual Forwarder) sở hữu một virtual MAC và thực sự forward traffic gửi tới nó.
- Phương pháp load balancing: round-robin (mặc định), weighted (tỷ lệ theo weight cấu hình), host-dependent (một host luôn nhận cùng MAC → sticky).
- Nếu một AVF hỏng, một AVF khác tiếp quản virtual MAC của nó → vẫn trong suốt với host.
- Multicast
224.0.0.102, UDP 3222.
! GLBP — active-active, một VIP 10.1.1.1, group 10
! Router A
interface Vlan10
ip address 10.1.1.2 255.255.255.0
glbp 10 ip 10.1.1.1
glbp 10 priority 110
glbp 10 preempt
glbp 10 load-balancing weighted
glbp 10 weighting 100
! Router B
interface Vlan10
ip address 10.1.1.3 255.255.255.0
glbp 10 ip 10.1.1.1
glbp 10 priority 100
glbp 10 preempt
glbp 10 load-balancing weighted
glbp 10 weighting 80
Bảng so sánh FHRP
| Đặc điểm | HSRP | VRRP | GLBP |
|---|---|---|---|
| Chuẩn | Cisco độc quyền (RFC 2281) | Mở (RFC 5798) | Cisco độc quyền |
| Vai trò | Active / Standby / Listen | Master / Backup | AVG / AVF |
| Bầu chọn | Priority (100), cao thắng | Priority (100), cao thắng; owner = 255 | AVG theo priority; forwarder chia tải |
| Virtual MAC | 0000.0C07.ACxx (v1) | 0000.5E00.01xx | tối đa 4 vMAC (0007.B4xx…) |
| Hello / hold mặc định | 3 s / 10 s | 1 s adv / ~3 s down | 3 s / 10 s |
| Preemption mặc định | Tắt (phải bật) | Bật | Tắt (phải bật) |
| Multicast / proto | 224.0.0.2 / 102, UDP 1985 | 224.0.0.18, IP proto 112 | 224.0.0.102, UDP 3222 |
| Load sharing | Nhiều group (MHSRP) | Nhiều group (VRID) | Gốc (một VIP, active-active) |
| Interface/object tracking | Có | Có | Có |
| Dùng khi | Toàn Cisco | Multi-vendor / chuẩn mở | Cisco, muốn chia tải gateway theo flow |
Mẹo: Cả ba đều hỗ trợ object/interface tracking để router mất link hướng lên (upstream) tự hạ priority và nhường vai, tránh tình huống “black hole” nơi gateway active còn sống ở phía LAN nhưng không có đường ra ngoài.
Các khái niệm failover
Active-passive vs active-active
| Active-Passive | Active-Active | |
|---|---|---|
| Xử lý traffic | Một node phục vụ; node kia đứng chờ standby | Tất cả node phục vụ đồng thời |
| Dùng tài nguyên | ~50% (standby ngồi không) | Gần 100% (tất cả làm việc) |
| Hành vi failover | Standby lên khi có lỗi (giật nhẹ) | Node còn sống hấp thụ tải ngay lập tức |
| Hoạch định năng lực | Standby phải bằng primary | Cần headroom N+1 (mỗi node ≤ (N-1)/N tải) |
| Độ phức tạp | Đơn giản hơn; không xung đột state | Khó hơn; cần shared/replicated state, session sync |
| Ví dụ | Cặp HSRP/VRRP, cặp firewall HA | GLBP, pool server sau LB, ECMP, anycast |
Một lưu ý quan trọng của active-active: nếu hai node mỗi node chạy 70% và một cái chết, node còn lại cần 140% — sẽ quá tải. Hãy hoạch định năng lực active-active sao cho một node bất kỳ có thể hấp thụ được tải failover.
Heartbeat
Các peer dự phòng trao đổi thông điệp heartbeat định kỳ (hello của FHRP, keepalive của cụm LB, link HA của firewall) để biết đối phương còn sống. Mất N heartbeat liên tiếp thì kết luận peer đã chết và kích hoạt failover. Việc tinh chỉnh là đánh đổi:
- Timer quyết liệt (dưới giây) → failover nhanh nhưng dễ false positive (báo peer khỏe-nhưng-bận là chết khi CPU tăng vọt hoặc link flap thoáng qua), gây flapping.
- Timer thận trọng → ổn định nhưng failover chậm.
Dùng BFD (Bidirectional Forwarding Detection) để phát hiện liveness nhanh, chi phí thấp mà nhiều FHRP và routing protocol có thể móc vào để hội tụ dưới giây mà không cần timer protocol quá quyết liệt.
Split-brain
Split-brain là kiểu lỗi nguy hiểm khi link heartbeat/giao tiếp giữa các peer đứt nhưng cả hai peer thật ra vẫn sống. Mỗi bên kết luận bên kia đã chết và cả hai đều lên active — hai master cùng trả lời cho một VIP, hoặc hai firewall cùng giữ active state, gây MAC flap, trùng IP, hỏng shared state và asymmetric routing. Cách giảm thiểu:
- Link heartbeat riêng, dự phòng (tách khỏi data path) để một lỗi data-link đơn lẻ không cô lập control plane.
- Quorum / witness / tie-breaker — một node bỏ phiếu thứ ba (số node lẻ, hoặc witness bên ngoài) để phân vùng thiểu số tự lùi xuống.
- Fencing / STONITH (“Shoot The Other Node In The Head”) — cưỡng bức tắt nguồn hoặc cô lập peer trước khi tiếp quản tài nguyên chia sẻ.
Preemption
Preemption quy định điều gì xảy ra khi một thiết bị priority cao hơn phục hồi:
- Bật preempt — primary vừa phục hồi giành lại vai active. Khôi phục phân phối traffic như thiết kế nhưng gây thêm một sự kiện failover nữa (lại giật nhẹ).
- Tắt preempt — ai đang active thì ở lại active cho tới khi chính nó hỏng. Ít gián đoạn hơn, nhưng traffic có thể chạy “ngược” trên node dự kiến làm secondary vô thời hạn.
Best practice: bật preemption kèm delay (preempt delay minimum) để router vừa reboot hội tụ xong bảng định tuyến và forwarding plane trước khi giành vai active — nếu không nó lên active trong khi vẫn còn black-hole traffic.
Load balancing (chuyên sâu)
Một load balancer (LB) trình ra cho client một virtual IP (VIP) duy nhất và phân phối các connection/request tới một pool server backend (“real server”). Nó mang lại quy mô (thêm server sau một địa chỉ), tính sẵn sàng (tự loại server hỏng qua health check), và sự linh hoạt (rolling deploy, SSL offload, routing).
L4 vs L7 load balancing
| Layer 4 (Transport) | Layer 7 (Application) | |
|---|---|---|
| Hoạt động trên | IP + TCP/UDP port | HTTP(S), header, URL, cookie, gRPC |
| Cơ sở quyết định | 5-tuple (src/dst IP, port, proto) | Nội dung: path, host, method, cookie, header |
| Tầm nhìn | Không thấy payload | Kiểm tra toàn bộ request |
| Tính năng | Nhanh, đơn giản, không phụ thuộc protocol | Content routing, rewrite header, WAF, caching, nén |
| Hiệu năng | Throughput rất cao, latency thấp | Tốn CPU hơn (parse/terminate) |
| Persistence | Dựa trên source IP | Dựa trên cookie/session (chính xác) |
| TLS | Pass-through (hoặc cơ bản) | Termination, re-encryption, SNI routing |
| Ví dụ | LVS/IPVS, AWS NLB, F5 (L4 VS) | NGINX, HAProxy (http mode), Envoy, AWS ALB |
Chọn L4 khi cần tốc độ thô và protocol không phải HTTP; chọn L7 khi cần route theo URL/host, làm TLS termination, rewrite header, hay áp policy hiểu ứng dụng.
Các thuật toán load balancing
| Thuật toán | Cách hoạt động | Khi nào dùng | Lưu ý |
|---|---|---|---|
| Round-robin | Xoay vòng qua các server theo thứ tự | Server đồng nhất, request ngắn đều nhau | Bỏ qua tải thực; request dài làm lệch |
| Weighted round-robin | Round-robin có thiên vị theo weight từng server | Server không đồng nhất (CPU/RAM khác) | Weight tĩnh không thích ứng tải thời gian thực |
| Least connections | Gửi tới server có ít connection active nhất | Connection dài / độ dài biến thiên (DB, WebSocket) | Cần theo dõi số connection |
| Weighted least connections | Least-connections có nhân theo capacity weight | Server sức chứa khác nhau, connection dài | Cần giữ thêm chút state |
| Least response time | Ít connection active nhất + latency thấp nhất | Dịch vụ nhạy latency | Cần probe response-time |
| IP hash / source hash | Hash IP client → server xác định | Stickiness rẻ, không cần cookie | Lệch nếu client nằm sau ít NAT/proxy |
| Consistent hashing | Hash key lên một vòng; remap tối thiểu khi thay đổi | Cache, backend shard, sticky ở quy mô lớn | Phức tạp hơn; cần virtual node để cân bằng |
| Random (+ two choices) | Chọn ngẫu nhiên (hoặc tốt hơn trong hai lần ngẫu nhiên) | Đơn giản, hiệu quả bất ngờ ở quy mô lớn | ”Power of two choices” cần thấy được tải |
Quy tắc chung: bắt đầu với round-robin (hoặc least-connections nếu thời lượng request rất biến thiên), thêm weight cho phần cứng không đều, và chỉ dùng tới hashing/consistent hashing khi cần stickiness hoặc cache affinity.
Health check
Health check là thứ biến LB thành công cụ availability chứ không chỉ là bộ phân phối. LB liên tục probe từng backend và loại các cái không khỏe khỏi vòng xoay:
- Passive — quan sát traffic thật; đánh dấu server down sau N request thật thất bại (phản ứng nhanh, nhưng vài user gặp lỗi trước).
- Active — LB gửi probe tổng hợp theo chu kỳ:
- L3: ICMP ping (server có tới được không?).
- L4: TCP connect tới port (service có đang listen không?).
- L7: HTTP GET tới endpoint
/healthzkỳ vọng200 OK(app có thực sự hoạt động, DB có tới được không?).
- Tham số chỉnh: interval, timeout, rise (số lần thành công liên tiếp để đánh dấu khỏe), fall (số lần thất bại liên tiếp để đánh dấu hỏng). Health check L7 sâu bắt được server “up nhưng hỏng” mà check L3/L4 vẫn giữ trong vòng xoay.
Session persistence (sticky session)
Một số ứng dụng giữ state theo từng user trong bộ nhớ server, nên các request của user phải quay lại đúng backend đó. Các phương pháp persistence:
- Source IP affinity — hash IP client → cùng server. Đơn giản, không cần app hợp tác, nhưng phá vỡ tính công bằng khi đứng sau NAT lớn và hỏng khi IP thay đổi (mobile).
- Cookie-based — LB chèn cookie riêng (ví dụ HAProxy
SRV=s1) hoặc học cookie có sẵn của app (ví dụJSESSIONID) để ghim session. Chính xác, theo từng session. - SSL session ID — ghim theo TLS session (nay ít dùng).
Best practice: ưu tiên backend stateless (lưu session vào store dùng chung như Redis hoặc token đã ký) để load balance thoải mái và không mất session khi failover. Sticky session là cái nạng làm hỏng phân phối đều và làm phức tạp rolling deploy.
SSL/TLS termination
- TLS termination — LB giữ certificate, giải mã TLS và nói HTTP plaintext với backend. Gỡ tải crypto CPU khỏi app server, tập trung quản lý cert, và cho phép kiểm tra/routing ở L7. Traffic giữa LB và backend nên nằm trong phân đoạn tin cậy.
- TLS re-encryption (bridging) — LB terminate, kiểm tra, rồi mã hóa lại tới backend. Cần cho mã hóa đầu-cuối / tuân thủ.
- TLS pass-through — LB chuyển tiếp byte đã mã hóa nguyên trạng (L4). Backend tự làm TLS; LB không kiểm tra được L7. Dùng khi LB không được thấy plaintext hoặc khi làm SNI-based routing ở L4.
Hardware vs software vs cloud load balancer
| Loại | Ví dụ | Ưu | Nhược |
|---|---|---|---|
| Hardware appliance | F5 BIG-IP, Citrix ADC (NetScaler), A10 | Throughput rất cao, ASIC SSL offload, tính năng chín muồi | Đắt, năng lực là một hộp phải scale/thay, khóa nhà cung cấp |
| Software | HAProxy, NGINX, Envoy, LVS/IPVS, Traefik | Rẻ/miễn phí, linh hoạt, chạy mọi nơi, hợp IaC | Bạn tự lo scale/HA/patch; hiệu năng phụ thuộc host |
| Cloud managed | AWS ALB/NLB, GCP Cloud Load Balancing, Azure LB | Co giãn, HA sẵn có, multi-AZ, trả theo dùng | Ít kiểm soát, chi phí theo request, tính năng riêng nhà cung cấp |
# HAProxy — L7, round-robin với health check, cookie stickiness, TLS termination
global
maxconn 20000
defaults
mode http
timeout connect 5s
timeout client 30s
timeout server 30s
option httpchk GET /healthz
frontend web_in
bind :80
bind :443 ssl crt /etc/haproxy/certs/site.pem # TLS termination
redirect scheme https if !{ ssl_fc }
default_backend web_pool
backend web_pool
balance roundrobin
cookie SRV insert indirect nocache # sticky session
server web1 10.0.1.11:8080 check inter 2s fall 3 rise 2 cookie s1 weight 100
server web2 10.0.1.12:8080 check inter 2s fall 3 rise 2 cookie s2 weight 100
server web3 10.0.1.13:8080 check inter 2s fall 3 rise 2 cookie s3 weight 50
# NGINX — upstream least-connections với các tùy chọn liên quan health
upstream app {
least_conn;
server 10.0.1.11:8080 max_fails=3 fail_timeout=10s;
server 10.0.1.12:8080 max_fails=3 fail_timeout=10s;
server 10.0.1.13:8080 backup; # chỉ dùng khi các server khác down
}
server {
listen 443 ssl;
ssl_certificate /etc/nginx/site.crt;
ssl_certificate_key /etc/nginx/site.key;
location / {
proxy_pass http://app;
proxy_next_upstream error timeout http_502 http_503;
}
}
Link và path redundancy
FHRP bảo vệ gateway; còn các đường đi giữa các thiết bị cần redundancy riêng của chúng.
Link aggregation (LAG / EtherChannel / bond)
Gộp nhiều link vật lý giữa hai thiết bị thành một link logic (IEEE 802.3ad / LACP; Cisco EtherChannel; Linux bonding). Lợi ích: băng thông tổng cao hơn và redundancy ở mức link — nếu một member hỏng, traffic phân phối lại qua các member còn sống mà không cần STP hội tụ lại. Traffic được hash trên các member (theo MAC/IP/port), nên một flow đơn không vượt quá tốc độ một member. Xem ./08-switching.md.
! Cisco LACP EtherChannel
interface range GigabitEthernet0/1 - 2
channel-group 1 mode active ! LACP
interface Port-channel1
switchport mode trunk
Redundant uplink và STP
Hai uplink từ một access switch tới hai distribution switch tạo thành loop vật lý. STP/RSTP chặn đường dự phòng để chống broadcast storm, rồi mở lại trong vài giây nếu đường chính hỏng (RSTP hội tụ nhanh hơn STP cổ điển nhiều). Thiết kế hiện đại ưa dùng MLAG / vPC / stacking để cả hai uplink cùng active (không có port bị STP chặn) mà vẫn không loop — uplink active-active thay vì active-passive. Chi tiết trong ./08-switching.md.
ECMP (Equal-Cost Multi-Path)
Ở Layer 3, khi bảng định tuyến có nhiều đường cùng cost tới một đích, router cài tất cả và hash các flow trên chúng — vừa load sharing active-active vừa redundancy ở tầng routing. Nếu một next-hop hỏng, phần của nó phân phối lại qua các đường còn lại ở lần cập nhật routing kế tiếp (hoặc ngay lập tức với BFD). ECMP là nền của fabric spine-leaf trong data center và của anycast. Flow được hash theo 5-tuple nên một flow ở lại một đường (tránh reorder gói).
So sánh các tầng redundancy
| Cơ chế | Tầng | Redundancy cho | Active-active? |
|---|---|---|---|
| STP/RSTP | L2 | Link switch dự phòng (không loop) | Không (chặn backup) |
| MLAG / vPC / stack | L2 | Uplink dual-homed | Có |
| LACP / EtherChannel | L2 | Các member link trong một bundle | Có |
| HSRP / VRRP | L3 | Default gateway | Không (standby ngồi không) |
| GLBP | L3 | Default gateway | Có |
| ECMP | L3 | Đường routed | Có |
| Load balancer | L4/L7 | Pool server | Có |
| Anycast | L3 | Endpoint địa lý/dịch vụ | Có |
Thiết kế topology dự phòng
HA được xếp lớp — mỗi tầng loại bỏ SPOF ở mức của nó:
- Component redundancy — hai bộ nguồn (trên PDU/feed riêng biệt), dual supervisor/route processor, dual NIC bond, quạt hot-swap.
- Device redundancy — cặp switch/router/firewall/LB, gắn với nhau bằng FHRP (gateway), HA clustering (firewall/LB) và MLAG (switch).
- Link/path redundancy — dual uplink (MLAG/ECMP), bundle LACP, và các tuyến sợi quang khác nhau về mặt vật lý để một nhát máy xúc không cắt đứt cả hai.
- Site/geographic redundancy — nhân bản dịch vụ qua nhiều data center / cloud region/AZ. Điều hướng traffic bằng GSLB / DNS-based load balancing, anycast, hoặc global LB của nhà cung cấp. Đây là nơi RTO/RPO và DR gặp HA.
Mẫu doanh nghiệp kinh điển — collapsed-core / hai tầng: cặp distribution/core switch (MLAG), mỗi access switch dual-home vào cả hai, FHRP (hoặc SVI trên cặp MLAG) cho gateway VIP, cặp firewall trong HA pair, cặp router tới hai ISP chạy BGP để dự phòng internet.
Quy tắc chung:
- Loại SPOF theo từng cặp; sau mỗi cặp, hãy hỏi “còn thứ đơn lẻ nào hỏng vẫn kéo ta sập không?”
- Đảm bảo lỗi độc lập: nguồn riêng, tuyến vật lý riêng, tránh shared fate. Hai link “dự phòng” trong cùng một ống là một link.
- Test failover thường xuyên (game day / chaos drill). Failover chưa test là một giả thuyết, không phải tính năng — cấu hình standby mục nát, cert hết hạn, và preempt lỗi chỉ lộ ra khi rút cáp thật.
- Sizing tầng active-active theo N+1: cái chết của một node bất kỳ phải được phần còn lại hấp thụ được.
- Tự động hóa và monitor: cảnh báo khi mất redundancy (một link hỏng nhưng đã được che), không chỉ khi có outage — nếu không bạn âm thầm chạy không lưới an toàn cho tới lần hỏng thứ hai.
Best Practices
- Loại bỏ SPOF một cách có hệ thống. Vẽ đường đi của traffic và, với mỗi thành phần, hỏi “nếu cái này chết, ta có sập không?” Nhân đôi (redundant pair) mọi thứ trả lời “có” — kể cả nguồn, cáp và làm mát, không chỉ các hộp thiết bị.
- Tối ưu MTTR, không chỉ MTBF. Failover nhanh tự động, monitoring tốt và runbook đã diễn tập mua được nhiều availability trên mỗi đồng hơn phần cứng cao cấp.
- Chỉnh timer FHRP có chủ đích. Mặc định (~3–10 s) là thận trọng. Dùng timer dưới giây hoặc BFD cho failover nhanh, nhưng đừng quyết liệt tới mức tải tăng thoáng qua gây false failover/flapping.
- Luôn bật interface/object tracking trên FHRP để gateway mất uplink tự lùi xuống thay vì black-hole traffic.
- Bật preemption kèm delay để thiết bị đang phục hồi hội tụ xong trước khi giành lại vai active.
- Phòng chống split-brain bằng link heartbeat riêng, dự phòng cộng với quorum/witness hoặc fencing.
- Làm cho backend stateless. Đưa session ra ngoài (Redis/token) để load balance thoải mái và không mất gì khi failover; coi sticky session là lựa chọn cuối cùng.
- Dùng health check sâu (L7) đánh vào endpoint hiểu được các dependency thật, để server “up nhưng hỏng” bị loại khỏi vòng xoay.
- Terminate TLS tại LB để gỡ tải CPU và quản lý cert tập trung; re-encrypt tới backend nơi tuân thủ yêu cầu mã hóa đầu-cuối.
- Hoạch định active-active với headroom N+1. Nếu mất một node làm quá tải phần còn lại, bạn có một sự cố chậm hơn, chứ không phải redundancy.
- Đảm bảo các failure domain độc lập. Nguồn riêng, sợi quang khác tuyến vật lý, và cảnh giác với component/bug dùng chung làm mất tính độc lập thống kê.
- Test failover thường xuyên. Lên lịch diễn tập failover; redundancy chưa test sẽ hỏng đúng lúc bạn cần nhất.
- Cảnh báo khi mất redundancy, không chỉ khi có outage, để sửa lưới an toàn trước khi lần hỏng thứ hai ập tới.
- Xếp lớp HA và DR cùng nhau. HA (FHRP, LB, clustering) cho lỗi component thường gặp với RTO gần bằng 0; DR (replication, standby site) cho thảm họa với RTO/RPO có giới hạn.
Tài liệu tham khảo
- RFC 5798 — Virtual Router Redundancy Protocol (VRRP) Version 3
- RFC 2281 — Cisco Hot Standby Router Protocol (HSRP)
- Cisco — First Hop Redundancy Protocols (HSRP/VRRP/GLBP) Configuration Guide
- Cisco — GLBP Overview and Configuration
- HAProxy Documentation
- NGINX — HTTP Load Balancing Guide
- AWS — Elastic Load Balancing (ALB/NLB) Documentation
- Google SRE Book — Availability, MTBF/MTTR, và “các số 9”
Part of the Network Engineer Roadmap knowledge base.
Overview
Every network eventually breaks. A power supply dies, an optic goes dark, a config change fat-fingers a routing table, a switch reboots for a firmware bug at 3 a.m. The difference between a network that hiccups and a network that goes down is redundancy — having a second path, a second device, or a second data center ready to carry the load — combined with fast, automatic failover so that the spare takes over before users (or SLAs) notice.
High Availability (HA) is the discipline of engineering systems to keep serving despite component failures. It is not a single feature you switch on; it is a property that emerges from decisions at every layer:
- Physical: dual power feeds, dual supervisors, dual NICs, diverse fiber runs.
- Link layer: link aggregation, redundant uplinks, loop-free topologies (STP/RSTP — see ./08-switching.md).
- Network layer: multiple routers, First Hop Redundancy Protocols (FHRP), equal-cost multipath (ECMP), dynamic routing convergence.
- Application/service layer: load balancers distributing across a pool of servers, health checks pulling dead nodes, geographic redundancy across regions.
The unifying goal is the elimination of single points of failure (SPOFs). A SPOF is any element whose failure takes down the whole service. If your two redundant firewalls both plug into one switch, that switch is a SPOF. If your two switches share one power strip, that strip is a SPOF. HA design is largely the systematic hunt for and removal of SPOFs.
This note covers the math of availability, First Hop Redundancy Protocols (HSRP, VRRP, GLBP), failover concepts (active-passive vs active-active, heartbeat, split-brain, preemption), load balancing in depth (L4 vs L7, algorithms, health checks, session persistence, SSL termination), link and path redundancy, and how to assemble these into resilient topologies. Related notes:
- Network devices (including load balancers as appliances): ./04-network-devices.md
- Switching, VLANs, STP, EtherChannel: ./08-switching.md
- Traffic management and QoS: ./13-traffic-management-and-qos.md
- Cloud networking (managed LBs, multi-AZ/region): ./16-cloud-networking.md
Fundamentals
Why HA matters: SPOFs, cost of downtime, and blast radius
Downtime is expensive and asymmetric. A few seconds of packet loss during a clean failover is invisible; a 30-minute outage during business hours can cost revenue, breach an SLA, and trigger contractual penalties. The business drivers for HA are:
- Revenue protection — e-commerce, payments, and SaaS lose money per minute of downtime.
- SLA compliance — contracts promise “three nines” or “four nines”; missing them incurs credits/penalties.
- Safety and compliance — hospitals, industrial control, and financial systems have regulatory uptime requirements.
- Reputation — repeated outages erode customer trust faster than almost anything else.
The engineering job is to shrink the blast radius of any single failure so that no one component failing can take everything down.
The “nines”: availability math
Availability is the fraction of time a system is operational, usually expressed as a percentage and colloquially as a number of “nines”:
Availability = Uptime / (Uptime + Downtime)
= MTBF / (MTBF + MTTR)
| Availability | ”Nines” | Downtime / year | Downtime / month | Downtime / day |
|---|---|---|---|---|
| 90% | one nine | 36.5 days | 72 hours | 2.4 hours |
| 99% | two nines | 3.65 days | 7.2 hours | 14.4 min |
| 99.9% | three nines | 8.77 hours | 43.8 min | 1.44 min |
| 99.95% | three and a half | 4.38 hours | 21.9 min | 43.2 s |
| 99.99% | four nines | 52.6 min | 4.38 min | 8.6 s |
| 99.999% | five nines | 5.26 min | 26.3 s | 0.86 s |
| 99.9999% | six nines | 31.5 s | 2.63 s | 0.086 s |
Two crucial insights:
- Nines are multiplicative in series. If a request must traverse a firewall (99.99%), a load balancer (99.99%), and an app server (99.9%), the end-to-end availability is
0.9999 × 0.9999 × 0.999 ≈ 99.88%— worse than any single component. Chains of dependencies drag availability down. - Redundancy in parallel multiplies the failure probability. Two independent components each with 99% availability (1% failure) placed in parallel fail together only
0.01 × 0.01 = 0.0001of the time → combined availability99.99%. This is why redundancy is so powerful — but only if the components fail independently (no shared power, no shared fiber, no shared bug triggered by the same input).
MTBF and MTTR
- MTBF (Mean Time Between Failures) — average operational time between failures. Improved by higher-quality hardware, better cooling, redundant internal components. You have limited control over MTBF once you buy the gear.
- MTTR (Mean Time To Repair/Recover) — average time to restore service after a failure. This is where operators have enormous leverage: automated failover, good monitoring, spare parts on site, runbooks, and practice all crush MTTR.
Because Availability = MTBF / (MTBF + MTTR), halving MTTR improves availability just as much as doubling MTBF — and is usually far cheaper. Fast, automatic failover is the highest-ROI HA investment. A sub-second FHRP failover turns a device reboot from an outage into a non-event.
RTO and RPO (the DR context)
HA (keeping running) is adjacent to Disaster Recovery (DR) (recovering after a major loss). Two DR metrics frame the requirements:
- RTO (Recovery Time Objective) — the maximum acceptable time to restore service after a disaster. “We must be back within 15 minutes.”
- RPO (Recovery Point Objective) — the maximum acceptable data loss measured in time. “We can lose at most 5 minutes of transactions.”
HA mechanisms (FHRP, load balancers, clustering) target near-zero RTO for component failures. DR mechanisms (backups, replication, standby sites) target bounded RTO/RPO for catastrophic failures. A complete design uses both: HA for the common case (a device dies), DR for the rare case (a site is lost).
Key Concepts
First Hop Redundancy Protocols (FHRP)
Hosts are configured with a single default gateway IP. If that gateway (router/L3 switch) dies, every host on the subnet loses off-subnet connectivity — a classic SPOF. First Hop Redundancy Protocols solve this by letting two or more routers share a virtual IP (VIP) and virtual MAC that acts as the gateway. Hosts point at the VIP; the routers elect an active forwarder; if it fails, a standby takes over the VIP/MAC with no reconfiguration on the hosts.
The magic is that the virtual MAC moves with the role. Because hosts have the VIP→virtual-MAC mapping in their ARP cache, and the new active router starts answering for that same virtual MAC, the hosts’ ARP entries stay valid — traffic just starts flowing to a different physical box. Failover is transparent to end stations.
HSRP (Hot Standby Router Protocol) — Cisco proprietary
- Cisco proprietary (defined in RFC 2281, but Cisco-owned).
- Routers form a group; one is Active, one is Standby, others are Listen.
- Election by priority (default 100, higher wins); tie broken by highest IP.
- Virtual MAC:
0000.0C07.ACxx(v1) wherexxis the group number; v2 uses0000.0C9F.Fxxx. - Hello timer default 3 s, hold timer default 10 s → failover ~10 s by default (tunable to sub-second).
- Communicates via multicast
224.0.0.2(v1) /224.0.0.102(v2), UDP 1985. - Supports interface tracking (decrement priority if an uplink dies) and preemption (higher-priority router reclaims Active when it returns).
- Load sharing is achieved only by running multiple groups and pointing different hosts/VLANs at different VIPs (Multiple HSRP / MHSRP).
! Cisco HSRP — two routers, VIP 10.1.1.1, group 10
! Router A (primary)
interface Vlan10
ip address 10.1.1.2 255.255.255.0
standby version 2
standby 10 ip 10.1.1.1
standby 10 priority 110
standby 10 preempt
standby 10 timers msec 250 msec 750 ! hello 250ms / hold 750ms
standby 10 track GigabitEthernet0/1 20 ! -20 priority if uplink down
!
! Router B (secondary) — same VIP, lower priority
interface Vlan10
ip address 10.1.1.3 255.255.255.0
standby version 2
standby 10 ip 10.1.1.1
standby 10 priority 100
standby 10 preempt
VRRP (Virtual Router Redundancy Protocol) — open standard, RFC 5798
- Open standard (RFC 5798 for VRRPv3, covering IPv4 and IPv6; RFC 3768 was VRRPv2). Multi-vendor.
- Roles are Master and Backup (analogous to HSRP Active/Standby).
- Election by priority (1–254, default 100, higher wins); the address owner (a router whose real interface IP equals the VIP) uses priority 255 and always wins.
- Virtual MAC:
0000.5E00.01xxwherexxis the VRID (Virtual Router ID). - Advertisement interval default 1 s; master-down interval ≈ 3 × interval + skew → default failover ~3 s (tunable to sub-second with centiseconds in v3).
- Communicates via multicast
224.0.0.18, IP protocol 112. - Preemption is ON by default (unlike HSRP where you must enable it).
- Load sharing again requires multiple VRRP groups (VRIDs).
! VRRP (works across vendors; Cisco syntax shown), VIP 10.1.1.1, VRID 10
interface Vlan10
ip address 10.1.1.2 255.255.255.0
vrrp 10 ip 10.1.1.1
vrrp 10 priority 110
vrrp 10 timers advertise 1
vrrp 10 preempt ! on by default; shown for clarity
GLBP (Gateway Load Balancing Protocol) — Cisco, active-active load sharing
- Cisco proprietary. Its differentiator: native load balancing across multiple routers using a single VIP.
- One router is elected AVG (Active Virtual Gateway); it hands out up to four different virtual MACs to hosts as they ARP for the VIP.
- Each participating router is an AVF (Active Virtual Forwarder) owning one virtual MAC and actually forwarding the traffic sent to it.
- Load-balancing methods: round-robin (default), weighted (proportional to configured weight), host-dependent (a given host always gets the same MAC → sticky).
- If an AVF fails, another AVF takes over its virtual MAC → still transparent to hosts.
- Multicast
224.0.0.102, UDP 3222.
! GLBP — active-active, single VIP 10.1.1.1, group 10
! Router A
interface Vlan10
ip address 10.1.1.2 255.255.255.0
glbp 10 ip 10.1.1.1
glbp 10 priority 110
glbp 10 preempt
glbp 10 load-balancing weighted
glbp 10 weighting 100
! Router B
interface Vlan10
ip address 10.1.1.3 255.255.255.0
glbp 10 ip 10.1.1.1
glbp 10 priority 100
glbp 10 preempt
glbp 10 load-balancing weighted
glbp 10 weighting 80
FHRP comparison
| Feature | HSRP | VRRP | GLBP |
|---|---|---|---|
| Standard | Cisco proprietary (RFC 2281) | Open (RFC 5798) | Cisco proprietary |
| Roles | Active / Standby / Listen | Master / Backup | AVG / AVF |
| Election | Priority (100), higher wins | Priority (100), higher wins; owner = 255 | AVG by priority; forwarders share |
| Virtual MAC | 0000.0C07.ACxx (v1) | 0000.5E00.01xx | up to 4 vMACs (0007.B4xx…) |
| Default hello / hold | 3 s / 10 s | 1 s adv / ~3 s down | 3 s / 10 s |
| Preemption default | Off (must enable) | On | Off (must enable) |
| Multicast / proto | 224.0.0.2 / 102, UDP 1985 | 224.0.0.18, IP proto 112 | 224.0.0.102, UDP 3222 |
| Load sharing | Multiple groups (MHSRP) | Multiple groups (VRIDs) | Native (single VIP, active-active) |
| Interface/object tracking | Yes | Yes | Yes |
| Use when | Cisco-only shop | Multi-vendor / standards-based | Cisco, want per-flow gateway load sharing |
Tip: All three support object/interface tracking so a router that loses its upstream link lowers its priority and steps aside, avoiding the “black hole” where the active gateway is up on the LAN side but has no path out.
Failover concepts
Active-passive vs active-active
| Active-Passive | Active-Active | |
|---|---|---|
| Traffic handling | One node serves; the other idles as standby | All nodes serve simultaneously |
| Resource use | ~50% (standby idle) | Near 100% (all working) |
| Failover behavior | Standby promotes on failure (brief blip) | Survivors absorb the load instantly |
| Capacity planning | Standby must equal primary | Must have N+1 headroom (each node ≤ (N-1)/N load) |
| Complexity | Simpler; no state conflicts | Harder; needs shared/replicated state, session sync |
| Examples | HSRP/VRRP pair, firewall HA pair | GLBP, LB server pool, ECMP, anycast |
A key active-active caveat: if two nodes each run at 70% and one dies, the survivor needs 140% — it will be overloaded. Plan active-active capacity so any single node can absorb the failover load.
Heartbeat
Redundant peers exchange periodic heartbeat messages (FHRP hellos, LB cluster keepalives, firewall HA links) to know each other is alive. Losing N consecutive heartbeats declares the peer dead and triggers failover. Tuning is a trade-off:
- Aggressive timers (sub-second) → fast failover but risk false positives (declaring a healthy-but-busy peer dead during a CPU spike or transient link flap), causing flapping.
- Conservative timers → stable but slow failover.
Use BFD (Bidirectional Forwarding Detection) for fast, low-overhead liveness detection that many FHRPs and routing protocols can hook into for sub-second convergence without over-aggressive protocol timers.
Split-brain
Split-brain is the dangerous failure where the heartbeat/communication link between peers breaks but both peers are actually alive. Each concludes the other is dead and both become active — two masters answering for the same VIP, or two firewalls both holding the active state, causing MAC flaps, duplicate IPs, corrupted shared state, and asymmetric routing. Mitigations:
- Dedicated, redundant heartbeat links (separate from the data path) so a single data-link failure doesn’t isolate the control plane.
- Quorum / witness / tie-breaker — a third voter (an odd number of nodes, or an external witness) so a minority partition steps down.
- Fencing / STONITH (“Shoot The Other Node In The Head”) — forcibly power off or isolate the peer before taking over shared resources.
Preemption
Preemption governs what happens when a higher-priority device recovers:
- Preempt enabled — the recovered primary reclaims the active role. Restores intended traffic distribution but causes a second failover event (another brief blip).
- Preempt disabled — whoever is currently active stays active until it fails. Fewer disruptions, but traffic may run “backwards” on the intended-secondary indefinitely.
Best practice: enable preemption with a delay (preempt delay minimum) so a rebooting router fully converges its routing table and forwarding plane before grabbing the active role — otherwise it becomes active while still black-holing traffic.
Load balancing (in depth)
A load balancer (LB) presents a single virtual IP (VIP) to clients and distributes incoming connections/requests across a pool of backend servers (“real servers”). It provides scale (add servers behind one address), availability (remove failed servers automatically via health checks), and flexibility (rolling deploys, SSL offload, routing).
L4 vs L7 load balancing
| Layer 4 (Transport) | Layer 7 (Application) | |
|---|---|---|
| Operates on | IP + TCP/UDP port | HTTP(S), headers, URL, cookies, gRPC |
| Decision basis | 5-tuple (src/dst IP, ports, proto) | Content: path, host, method, cookie, header |
| Visibility | Opaque payload | Full request inspection |
| Features | Fast, simple, protocol-agnostic | Content routing, header rewrite, WAF, caching, compression |
| Performance | Very high throughput, low latency | Higher CPU (parses/terminates) |
| Persistence | Source-IP based | Cookie/session based (precise) |
| TLS | Pass-through (or basic) | Termination, re-encryption, SNI routing |
| Examples | LVS/IPVS, AWS NLB, F5 (L4 VS) | NGINX, HAProxy (http mode), Envoy, AWS ALB |
L4 is chosen for raw speed and non-HTTP protocols; L7 is chosen when you need to route by URL/host, do TLS termination, rewrite headers, or apply application-aware policy.
Load balancing algorithms
| Algorithm | How it works | When to use | Caveats |
|---|---|---|---|
| Round-robin | Rotate through servers in order | Homogeneous servers, uniform short requests | Ignores actual load; long requests skew it |
| Weighted round-robin | Round-robin biased by per-server weight | Heterogeneous servers (different CPU/RAM) | Static weights don’t adapt to real-time load |
| Least connections | Send to server with fewest active connections | Long-lived / variable-length connections (DB, WebSocket) | Needs connection-count tracking |
| Weighted least connections | Least-connections scaled by capacity weight | Mixed-capacity servers with long connections | Slightly more state to maintain |
| Least response time | Fewest active connections + lowest latency | Latency-sensitive services | Needs response-time probing |
| IP hash / source hash | Hash client IP → deterministic server | Cheap session stickiness without cookies | Uneven if clients sit behind few NATs/proxies |
| Consistent hashing | Hash key onto a ring; minimal remap on change | Caches, sharded backends, sticky at scale | More complex; needs virtual nodes for balance |
| Random (+ two choices) | Pick random (or best of two random) | Simple, surprisingly good at scale | ”Power of two choices” needs load visibility |
Rule of thumb: start with round-robin (or least-connections if request durations vary a lot), add weights for uneven hardware, and reach for hashing/consistent hashing only when you need stickiness or cache affinity.
Health checks
Health checks are what make an LB an availability tool rather than just a distributor. The LB continuously probes each backend and removes unhealthy ones from rotation:
- Passive — observe live traffic; mark a server down after N failed real requests (fast to react, but a few users hit the failure first).
- Active — the LB sends synthetic probes on an interval:
- L3: ICMP ping (server reachable?).
- L4: TCP connect to the port (is the service listening?).
- L7: HTTP GET to a
/healthzendpoint expecting200 OK(is the app actually working, DB reachable, etc.?).
- Tunables: interval, timeout, rise (consecutive successes to mark healthy), fall (consecutive failures to mark unhealthy). Deep L7 checks catch “up but broken” servers that L3/L4 checks would keep in rotation.
Session persistence (sticky sessions)
Some applications keep per-user state in server memory, so a user’s requests must return to the same backend. Persistence methods:
- Source IP affinity — hash client IP → same server. Simple, no app cooperation, but breaks fairness behind large NATs and breaks on IP change (mobile).
- Cookie-based — LB inserts its own cookie (e.g., HAProxy
SRV=s1) or learns an existing app cookie (e.g.,JSESSIONID) to pin the session. Precise and per-session. - SSL session ID — pin by TLS session (less common now).
Best practice: prefer stateless backends (store session in a shared store like Redis or a signed token) so you can load balance freely and lose no sessions on failover. Sticky sessions are a crutch that hurts even distribution and complicates rolling deploys.
SSL/TLS termination
- TLS termination — the LB holds the certificate, decrypts TLS, and talks plaintext HTTP to backends. Offloads crypto CPU from app servers, centralizes cert management, and enables L7 inspection/routing. Traffic between LB and backends should stay on a trusted segment.
- TLS re-encryption (bridging) — LB terminates, inspects, then re-encrypts to the backend. Needed for end-to-end encryption / compliance.
- TLS pass-through — LB forwards encrypted bytes untouched (L4). Backends do their own TLS; LB cannot inspect L7. Used when the LB must not see plaintext or when using SNI-based routing at L4.
Hardware vs software vs cloud load balancers
| Type | Examples | Pros | Cons |
|---|---|---|---|
| Hardware appliance | F5 BIG-IP, Citrix ADC (NetScaler), A10 | Very high throughput, SSL offload ASICs, mature features | Expensive, capacity is a box you must scale/replace, vendor lock-in |
| Software | HAProxy, NGINX, Envoy, LVS/IPVS, Traefik | Cheap/free, flexible, run anywhere, IaC-friendly | You own scaling/HA/patching; performance depends on host |
| Cloud managed | AWS ALB/NLB, GCP Cloud Load Balancing, Azure LB | Elastic, HA built-in, multi-AZ, pay-as-you-go | Less control, per-request cost, provider-specific features |
# HAProxy — L7, round-robin with health checks, cookie stickiness, TLS termination
global
maxconn 20000
defaults
mode http
timeout connect 5s
timeout client 30s
timeout server 30s
option httpchk GET /healthz
frontend web_in
bind :80
bind :443 ssl crt /etc/haproxy/certs/site.pem # TLS termination
redirect scheme https if !{ ssl_fc }
default_backend web_pool
backend web_pool
balance roundrobin
cookie SRV insert indirect nocache # sticky sessions
server web1 10.0.1.11:8080 check inter 2s fall 3 rise 2 cookie s1 weight 100
server web2 10.0.1.12:8080 check inter 2s fall 3 rise 2 cookie s2 weight 100
server web3 10.0.1.13:8080 check inter 2s fall 3 rise 2 cookie s3 weight 50
# NGINX — least-connections upstream with health-relevant options
upstream app {
least_conn;
server 10.0.1.11:8080 max_fails=3 fail_timeout=10s;
server 10.0.1.12:8080 max_fails=3 fail_timeout=10s;
server 10.0.1.13:8080 backup; # only used if others are down
}
server {
listen 443 ssl;
ssl_certificate /etc/nginx/site.crt;
ssl_certificate_key /etc/nginx/site.key;
location / {
proxy_pass http://app;
proxy_next_upstream error timeout http_502 http_503;
}
}
Link and path redundancy
FHRP protects the gateway; the paths between devices need their own redundancy.
Link aggregation (LAG / EtherChannel / bond)
Bundle multiple physical links between two devices into one logical link (IEEE 802.3ad / LACP; Cisco EtherChannel; Linux bonding). Benefits: higher aggregate bandwidth and link-level redundancy — if one member fails, traffic redistributes over the survivors without STP reconvergence. Traffic is hashed across members (by MAC/IP/port), so a single flow does not exceed one member’s speed. See ./08-switching.md.
! Cisco LACP EtherChannel
interface range GigabitEthernet0/1 - 2
channel-group 1 mode active ! LACP
interface Port-channel1
switchport mode trunk
Redundant uplinks and STP
Two uplinks from an access switch to two distribution switches create a physical loop. STP/RSTP blocks the redundant path to prevent broadcast storms, then unblocks it in ~seconds if the primary fails (RSTP converges much faster than classic STP). Modern designs prefer MLAG / vPC / stacking so both uplinks are active (no STP-blocked port) while still being loop-free — active-active uplinks instead of active-passive. Details in ./08-switching.md.
ECMP (Equal-Cost Multi-Path)
At Layer 3, when the routing table has multiple equal-cost paths to a destination, the router installs all of them and hashes flows across them — active-active load sharing and redundancy at the routing layer. If one next-hop fails, its share redistributes over the remaining paths on the next routing update (or immediately with BFD). ECMP underpins spine-leaf data-center fabrics and anycast. Flows are hashed per-5-tuple so a single flow stays on one path (avoids packet reordering).
Comparing the redundancy layers
| Mechanism | Layer | Redundancy for | Active-active? |
|---|---|---|---|
| STP/RSTP | L2 | Redundant switch links (loop-free) | No (blocks backup) |
| MLAG / vPC / stack | L2 | Dual-homed uplinks | Yes |
| LACP / EtherChannel | L2 | Member links within one bundle | Yes |
| HSRP / VRRP | L3 | Default gateway | No (standby idle) |
| GLBP | L3 | Default gateway | Yes |
| ECMP | L3 | Routed paths | Yes |
| Load balancer | L4/L7 | Server pool | Yes |
| Anycast | L3 | Geographic/service endpoints | Yes |
Designing redundant topologies
HA is layered — each tier removes SPOFs at its level:
- Component redundancy — dual power supplies (on separate PDUs/feeds), dual supervisors/route processors, dual NICs bonded, hot-swappable fans.
- Device redundancy — pairs of switches/routers/firewalls/LBs, tied together with FHRP (gateways), HA clustering (firewalls/LBs), and MLAG (switches).
- Link/path redundancy — dual uplinks (MLAG/ECMP), LACP bundles, and physically diverse fiber runs so a single backhoe cut can’t sever both.
- Site/geographic redundancy — replicate the service across data centers / cloud regions/AZs. Steer traffic with GSLB / DNS-based load balancing, anycast, or provider global LBs. This is where RTO/RPO and DR meet HA.
Canonical enterprise pattern — collapsed-core / two-tier: dual distribution/core switches (MLAG), each access switch dual-homed to both, FHRP (or an SVI on the MLAG pair) for the gateway VIP, dual firewalls in an HA pair, dual routers to two ISPs running BGP for internet redundancy.
Rules of thumb:
- Remove SPOFs pair-by-pair; after each pair, ask “what single thing failing still takes us down?”
- Ensure failures are independent: separate power feeds, separate physical paths, avoid shared fate. Two “redundant” links in the same conduit are one link.
- Test failover regularly (game days / chaos drills). Untested failover is a hypothesis, not a feature — standby configs rot, certs expire, and preemption misfires are only found by pulling cables.
- Size active-active tiers for N+1: any one node’s death must be absorbable by the rest.
- Automate and monitor: alert on the loss of redundancy (a failed-but-covered link), not just on outages — otherwise you silently run without a safety net until the second failure.
Best Practices
- Eliminate SPOFs systematically. Diagram the traffic path and, for each element, ask “if this dies, are we down?” Redundant-pair anything that answers yes — including power, cabling, and cooling, not just the boxes.
- Optimize MTTR, not just MTBF. Fast automatic failover, good monitoring, and rehearsed runbooks buy more availability per dollar than premium hardware.
- Tune FHRP timers deliberately. Defaults (~3–10 s) are conservative. Use sub-second timers or BFD for fast failover, but not so aggressive that transient load spikes cause false failovers/flapping.
- Always enable interface/object tracking on FHRP so a gateway that loses its uplink steps down instead of black-holing traffic.
- Enable preemption with a delay so a recovering device converges before reclaiming the active role.
- Guard against split-brain with dedicated redundant heartbeat links plus quorum/witness or fencing.
- Make backends stateless. Externalize sessions (Redis/token) so you can load balance freely and lose nothing on failover; treat sticky sessions as a last resort.
- Use deep (L7) health checks hitting a real dependency-aware endpoint, so “up but broken” servers are pulled from rotation.
- Terminate TLS at the LB for CPU offload and central cert management; re-encrypt to backends where compliance requires end-to-end encryption.
- Plan active-active for N+1 headroom. If losing one node overloads the rest, you have a slower outage, not redundancy.
- Ensure independent failure domains. Separate power feeds, physically diverse fiber, and beware shared components/bugs that defeat statistical independence.
- Test failover regularly. Schedule failover drills; untested redundancy fails when you need it most.
- Alert on lost redundancy, not only on outages, so you fix the safety net before the second failure lands.
- Layer HA and DR together. HA (FHRP, LB, clustering) for common component failures with near-zero RTO; DR (replication, standby sites) for catastrophes with bounded RTO/RPO.
References
- RFC 5798 — Virtual Router Redundancy Protocol (VRRP) Version 3
- RFC 2281 — Cisco Hot Standby Router Protocol (HSRP)
- Cisco — First Hop Redundancy Protocols (HSRP/VRRP/GLBP) Configuration Guide
- Cisco — GLBP Overview and Configuration
- HAProxy Documentation
- NGINX — HTTP Load Balancing Guide
- AWS — Elastic Load Balancing (ALB/NLB) Documentation
- Google SRE Book — Availability, MTBF/MTTR, and the “nines”