Khả năng mở rộng & Độ tin cậyScalability & Reliability
Thuộc bộ kiến thức Backend Roadmap.
Tổng quan
Scalability (khả năng mở rộng) là khả năng của hệ thống xử lý nhiều việc hơn — nhiều user hơn, nhiều request hơn, nhiều dữ liệu hơn — bằng cách thêm tài nguyên, lý tưởng là không phải viết lại và với chi phí tăng không nhanh hơn tải. Reliability (độ tin cậy) là khả năng hệ thống vẫn làm đúng việc một cách chính xác ngay cả khi một phần của nó fail, chậm lại, hoặc hành xử bất thường. Hai khái niệm này gắn chặt với nhau: một hệ thống không scale được thì sớm muộn cũng sập dưới tải (một dạng lỗi reliability), còn một hệ thống không được thiết kế để chịu lỗi thì không thể scale an toàn, vì mỗi component thêm vào là thêm một thứ có thể hỏng.
Với một backend engineer, đây không phải là chuyện “để sau”. Những quyết định kiến trúc giúp hệ thống scalable và reliable — statelessness, load balancing, timeout, retry, idempotency, redundancy — rẻ khi thiết kế từ đầu nhưng cực kỳ đắt khi phải lắp thêm về sau. Note này bao quát: cách tư duy về scale, các pattern giúp hệ thống đứng vững dưới áp lực, các failure mode mà những pattern đó phòng thủ, và các metric (SLA/SLO/SLI, error budget, “the nines”) giúp bạn quản lý reliability như một kỷ luật kỹ thuật thay vì một niềm hy vọng.
Một mô hình tư duy hữu ích đến từ cuốn Release It! của Michael Nygard: mỗi integration point đều là một điểm có thể fail, và lỗi thì lan truyền. Việc của bạn không phải là ngăn mọi lỗi — điều đó bất khả thi trong hệ thống distributed — mà là kiềm chế lỗi để một database chậm, một node chết, hay một đợt spike traffic không kéo sập toàn bộ.
Các note liên quan trong bộ kiến thức này:
- Database Design & Scaling — replication, sharding, read replica.
- Caching — cách rẻ nhất để mua scalability và giảm tải.
- Message Brokers & Async — decoupling và load leveling bằng queue.
- Observability & Monitoring — không đo được thì không quản lý được reliability.
- Chi tiết về load balancer và networking: tài liệu roadmap
../network-engineer.
Kiến thức nền tảng
Vertical scaling vs horizontal scaling
Có hai cách cơ bản để thêm capacity.
| Vertical scaling (scale up) | Horizontal scaling (scale out) | |
|---|---|---|
| Là gì | Máy to hơn: nhiều CPU, RAM, disk nhanh hơn | Nhiều máy hơn sau một load balancer |
| Công sức | Đơn giản — đổi instance type, reboot | Cần thiết kế stateless, LB, điều phối |
| Trần | Giới hạn vật lý cứng (instance to nhất có thể mua) | Về cơ bản không giới hạn |
| Blast radius khi lỗi | Một máy to = single point of failure | Một node chết, các node khác gánh tải |
| Đường cong chi phí | Siêu tuyến tính — phần cứng cao cấp rất đắt | Gần như tuyến tính với commodity hardware |
| Downtime khi resize | Thường phải reboot | Thêm/bớt node với zero downtime |
Quy tắc thực tế: scale up trước vì đơn giản, nhưng thiết kế sao cho có thể scale out, vì vertical scaling luôn đụng trần và không cho bạn redundancy. Hệ thống thực tế kết hợp cả hai — node cỡ vừa phải, nhưng nhiều.
Statelessness: chìa khóa mở ra horizontal scale
Một service là stateless khi bất kỳ request nào cũng có thể được phục vụ bởi bất kỳ instance nào, vì instance không giữ dữ liệu riêng của client giữa các request. State được đẩy ra các backing store dùng chung: database, cache (Redis), object storage, và bản thân request/token (ví dụ JWT).
Vì sao quan trọng: nếu instance A giữ session của một user trong bộ nhớ local, user đó bắt buộc phải quay lại A. Bạn không thể tự do thêm, bớt, hay thay thế instance — bạn đã mất toàn bộ ý nghĩa của horizontal scaling, và tạo ra một single point of failure cho mỗi user.
- Stateful (tệ cho scale)clientapp A: session in RAMluôn phải trúng A
- Stateless (scale thoải mái)clientLBapp A / app B / app CRedis (sessions)Postgres (data)
Quy tắc cho statelessness:
- Không giữ session trong bộ nhớ — dùng shared session store hoặc stateless token.
- Không ghi file local mà request sau phụ thuộc vào — dùng object storage (S3).
- Không dùng cache “dính” trong process cần đồng bộ giữa các node — dùng shared cache.
- Bất kỳ node nào cũng phải có thể bị kill an toàn bất cứ lúc nào (điều này cũng bật được autoscaling và rolling deploy).
The scaling cube (AKF Scale Cube)
AKF Scale Cube mô tả ba trục scaling trực giao:
| Trục | Tên | Kỹ thuật | Ví dụ |
|---|---|---|---|
| X | Nhân bản theo chiều ngang | Clone toàn bộ app sau một load balancer | 10 app server giống hệt nhau |
| Y | Phân rã theo chức năng | Tách theo trách nhiệm thành các service | service orders, payments, users (microservices) |
| Z | Phân vùng dữ liệu | Tách theo dữ liệu / khách hàng | Shard theo user_id, hoặc database riêng cho từng tenant |
- X là dễ nhất và là điểm khởi đầu — cần statelessness và một load balancer.
- Y khớp với ranh giới team và cho phép scale riêng các subsystem “nóng”.
- Z chính là sharding — loại bỏ nút thắt single-database (xem Database Scaling).
Hệ thống trưởng thành dùng cả ba.
Performance vs scalability
Hai khái niệm này hay bị nhầm nhưng là hai thuộc tính khác nhau:
- Performance = một request đơn lẻ được phục vụ nhanh thế nào (latency) ở tải thấp.
- Scalability = hệ thống duy trì performance chấp nhận được tốt thế nào khi tải tăng.
Một hệ thống có thể nhanh với một user nhưng sụp với một nghìn (performance tốt, scalability kém). Một hệ thống khác có thể chậm vừa phải mỗi request nhưng giữ latency đó phẳng từ 1 đến 1.000.000 user (cái scalable). Hãy tối ưu cho hình dạng đường cong, không chỉ điểm xuất phát. Luôn đo tail latency (p95/p99), không phải trung bình — trung bình che giấu những user có trải nghiệm tệ nhất, và ở quy mô lớn thì cái tail chính là rất nhiều user thật.
Little’s Law — trực giác về capacity
L = λ × W, với L = số request đồng thời trong hệ thống, λ = tốc độ đến (req/s), W = thời gian trung bình trong hệ thống (s). Nếu request đến với 500/s và mỗi cái mất 0,2s, bạn có trung bình ~100 request đang xử lý và cần đủ thread/connection/worker để giữ chúng. Khi W tăng (một dependency chậm), L tăng, tài nguyên cạn, và bạn phải queue hoặc sập — đây là cơ chế đằng sau hầu hết các cú sụp do overload.
Khái niệm chính
Load balancing
Một load balancer (LB) đứng trước các instance của bạn và phân phối request đến trên chúng. Nó là trụ cột của horizontal scaling, đồng thời là một công cụ reliability: nó điều hướng traffic tránh xa các node không khỏe.
L4 vs L7 load balancing
| L4 (transport) | L7 (application) | |
|---|---|---|
| Hoạt động ở | TCP/UDP — IP và port | HTTP — path, header, cookie, method |
| Nhìn thấy | Packet/connection, không thấy nội dung | Toàn bộ HTTP request |
| Làm được | Forward connection nhanh | Routing theo path/host, TLS termination, sửa header, WAF, sticky theo cookie |
| Overhead | Rất thấp, cực nhanh | Cao hơn (phải parse request) |
| Ví dụ | AWS NLB, HAProxy (TCP mode), IPVS | AWS ALB, NGINX, Envoy, Traefik |
Dùng L7 cho microservice HTTP (route /api/orders sang một service, /api/users sang service khác; canary theo header). Dùng L4 khi cần throughput thô, protocol không phải HTTP, hoặc muốn backend tự terminate TLS.
Các thuật toán balancing
| Thuật toán | Cách chọn backend | Phù hợp cho |
|---|---|---|
| Round robin | Server kế tiếp trong vòng xoay | Node đồng nhất, request đều nhau |
| Weighted round robin | Vòng xoay có trọng số theo capacity | Instance kích cỡ khác nhau |
| Least connections | Server có ít connection active nhất | Request kéo dài / thời lượng biến thiên |
| Least response time | Ít connection nhất + latency thấp nhất | Service nhạy latency |
| IP / consistent hashing | Hash IP client hoặc key → cùng một server | Cache affinity, sticky routing không cần cookie |
| Random (power of two choices) | Chọn ngẫu nhiên 2, lấy cái ít tải hơn | Phân tải rất tốt với state tối thiểu, scale tốt |
Power of two random choices đáng nhắc đến: chọn ngẫu nhiên hai server và gửi sang cái ít tải hơn cho kết quả balancing gần tối ưu mà không cần state toàn cục, nên rất phổ biến trong các hệ thống lớn.
Health check
LB phải biết backend nào còn sống. Có hai loại:
- Liveness / passive — đánh dấu node xấu sau N response fail; ngừng gửi traffic cho nó.
- Readiness / active — định kỳ poll một health endpoint (
GET /healthz) và chỉ route đến các node pass.
Hãy làm health check đủ sâu để có ý nghĩa nhưng đủ nông để không cascade: một /healthz mà kiểm tra database sẽ đánh dấu mọi node là unhealthy ngay khi database chớp lỗi, kéo sập một service lẽ ra vẫn phục vụ được traffic từ cache. Cách tách phổ biến: /livez (process còn chạy không?) vs /readyz (phục vụ được không, kể cả dependency quan trọng?).
Sticky sessions
Sticky sessions (session affinity) ghim một client vào một backend (qua cookie hoặc IP hash). Đây là cái nạng cho app stateful — nó phá vỡ phân tải đều, gãy khi một node chết (các user bị ghim mất session), và làm phức tạp việc deploy. Hãy ưu tiên statelessness với shared session store. Nếu buộc phải dùng stickiness, dùng affinity theo cookie với TTL ngắn và đảm bảo app degrade mượt khi bị re-pin. Xem chi tiết load balancer trong tài liệu ../network-engineer.
Resilience patterns
Đây là các pattern phòng thủ từ Release It! của Nygard và kho kiến thức reliability của cloud. Mỗi pattern chặn một loại lỗi cụ thể.
Timeout
Đừng bao giờ gọi network mà không có timeout. Mặc định trong nhiều thư viện là vô hạn, nghĩa là một dependency treo có thể làm cạn sạch thread/connection khi request dồn lại chờ mãi mãi (Little’s Law hiện hình). Đặt timeout ở mọi tầng: connect timeout, read/socket timeout, total request timeout. Hãy phân bổ ngân sách: nếu SLA của bạn là 300ms và bạn gọi ba service, chúng không thể mỗi cái có timeout 1s. Timeout biến một cú treo (chí mạng, vô hạn) thành một lỗi nhanh (có thể phục hồi).
Retry với exponential backoff + jitter
Lỗi tạm thời (một cú chớp mạng, một node đang reboot) đáng để retry. Nhưng retry ngây thơ thì nguy hiểm:
- Chỉ retry các thao tác idempotent (hoặc dùng idempotency key — xem bên dưới). Retry một
POST /chargekhông idempotent có thể tính tiền hai lần. - Giới hạn số lần retry (ví dụ tối đa 3). Retry vô hạn biến một cú chớp thành một cơn bão.
- Backoff theo cấp số nhân: chờ 100ms, 200ms, 400ms… để không dập một dependency đang khốn khổ.
- Thêm jitter (ngẫu nhiên hóa độ trễ). Không có jitter, tất cả client fail cùng một khoảnh khắc sẽ retry cùng một khoảnh khắc — một retry storm / thundering herd đồng bộ giữ cho service đang hồi phục tiếp tục chết.
import random, time
def call_with_retry(fn, max_attempts=3, base=0.1, cap=2.0):
for attempt in range(max_attempts):
try:
return fn()
except TransientError:
if attempt == max_attempts - 1:
raise
# full jitter: ngủ đều trong [0, min(cap, base * 2**attempt)]
backoff = min(cap, base * (2 ** attempt))
time.sleep(random.uniform(0, backoff))
Mặc định AWS khuyến nghị là exponential backoff with full jitter. Cũng nên cân nhắc retry budget (giới hạn retry ở mức ví dụ 10% tổng request) để retry không bao giờ khuếch đại tải quá mức.
Circuit breaker
Một circuit breaker ngừng gọi một dependency đang fail rõ ràng, cho nó không gian để hồi phục và fail nhanh thay vì dồn lại những cú gọi chậm chắc chắn thất bại. Nó là một máy trạng thái:
| Trạng thái | Hành vi | Chuyển tiếp |
|---|---|---|
| Closed | Cho request đi qua; đếm lỗi | Tỷ lệ lỗi vượt ngưỡng → Open |
| Open | Fail ngay lập tức (hoặc fallback), không gửi request | Sau một khoảng cooldown → Half-Open |
| Half-Open | Cho vài cú gọi thử | Thành công → Closed; lỗi → Open |
Không có breaker, một dependency chậm khiến mọi thread caller block trên timeout của nó, làm cạn pool và cascade lỗi ngược lên trên. Breaker biến điều đó thành một lỗi tức thì, rẻ (lý tưởng là ghép với một fallback). Thư viện: Resilience4j (Java), Polly (.NET), và các service mesh (Envoy/Istio) cung cấp ở tầng hạ tầng.
Bulkhead
Đặt theo tên các khoang tàu ngăn không cho một khoang ngập nước làm chìm cả con tàu. Cô lập tài nguyên để một dependency quá tải không nuốt hết một pool dùng chung. Ví dụ: cấp cho các cú gọi service recommendations hay lỗi một thread pool / connection pool riêng gồm 10, tách khỏi pool checkout. Nếu recommendations treo, nó chỉ làm cạn 10 thread của nó và không hơn — checkout vẫn chạy. Bulkhead biến một cú outage toàn bộ thành một sự suy giảm cục bộ, được kiềm chế.
Rate limiting & throttling
Rate limiting giới hạn một client (hoặc cả hệ thống) được gửi bao nhiêu request trong một khung thời gian, bảo vệ bạn khỏi lạm dụng, client lỗi, và spike traffic. Throttling là hành động từ chối/trì hoãn các request vượt giới hạn (thường là HTTP 429 Too Many Requests với header Retry-After).
Các thuật toán phổ biến:
| Thuật toán | Hành vi | Ghi chú |
|---|---|---|
| Token bucket | Token nạp lại với tốc độ R, mỗi request tiêu một; bucket cho phép burst tới cỡ B | Phổ biến nhất; cho phép burst có kiểm soát |
| Leaky bucket | Request chảy ra với tốc độ cố định; tràn thì bị từ chối | Làm mượt output về tốc độ ổn định |
| Fixed window | Đếm theo khung lịch (mỗi phút) | Đơn giản nhưng có vấn đề burst ở rìa cửa sổ |
| Sliding window | Đếm cuộn trong N giây gần nhất | Mượt hơn, chính xác hơn, tốn hơn chút |
Hãy rate limit theo key (API key, user, IP) và thực thi trong một shared store (Redis) để giới hạn là toàn cục trên các instance, không phải theo từng node.
Backpressure
Backpressure là việc một hệ thống báo cho upstream chậm lại khi nó không theo kịp, thay vì âm thầm buffer cho tới khi hết bộ nhớ. Thay vì nhận việc vô hạn, một component phát tín hiệu “tôi đầy rồi” — qua một bounded queue chặn/từ chối khi đầy, TCP flow control, HTTP 429, hoặc tín hiệu demand của reactive streams. Lựa chọn thay thế — queue vô hạn — chỉ dời cú lỗi sang một cú crash out-of-memory và che giấu latency đang tăng. Bounded queue là một tính năng, không phải hạn chế. Hệ thống async (xem Message Brokers) cung cấp backpressure tự nhiên: producer chậm lại khi queue dồn ứ.
Load shedding
Khi quá tải, phục vụ phần lớn request tốt vẫn hơn phục vụ tất cả request tệ (hoặc không phục vụ gì, sau khi crash). Load shedding chủ động bỏ hoặc từ chối một phần công việc đến để giữ trong ngưỡng capacity — ví dụ, từ chối các request ưu tiên thấp trước, hoặc trả 503 khi concurrency vượt một ngưỡng. Kết hợp với priority (bỏ traffic batch/analytics trước traffic checkout hướng user), nó giữ hệ thống sống và dòng công việc quan trọng vẫn chảy dưới overload. Google SRE nhấn mạnh việc shed trước khi bão hòa, dựa trên tải đo được.
Graceful degradation & fallback
Thiết kế sao cho khi một dependency không quan trọng fail, hệ thống mất một tính năng, không phải cả trang. Ví dụ:
- Service recommendations chết → hiện danh sách “sản phẩm phổ biến” tĩnh thay vì cá nhân hóa.
- Cache chết → phục vụ từ database (chậm hơn nhưng chạy), hoặc phục vụ dữ liệu hơi cũ.
- Payment provider A chết → failover sang provider B.
- Dữ liệu mới không có → phục vụ giá trị cache last-known-good (“stale-while-revalidate”).
Một fallback là phương án thay thế cụ thể được trả về khi đường chính fail (thường ghép với circuit breaker). Hãy quyết định fallback một cách có chủ đích cho từng dependency; một kết quả rỗng âm thầm có thể tệ hơn một lỗi rõ ràng.
Failure mode & thiết kế để chịu lỗi
Bạn phải hiểu các cách hệ thống distributed fail, vì mỗi resilience pattern ở trên chặn một cách cụ thể.
Cascading failure (lỗi dây chuyền)
Lỗi của một component làm quá tải component khác, cái đó fail, làm quá tải cái tiếp theo — một chuỗi domino. Đường điển hình: DB chậm → thread app block chờ → thread pool cạn → app ngừng đáp health check → LB gỡ node → các node còn lại nhận nhiều tải hơn → chúng chết nhanh hơn. Phòng thủ: timeout, circuit breaker, bulkhead, load shedding, và health check không khuếch đại.
Thundering herd
Nhiều client đập vào cùng một tài nguyên cùng một khoảnh khắc. Các nguyên nhân phổ biến:
- Một cache key phổ biến hết hạn → hàng nghìn request đồng thời miss và dẫm đạp lên database. Cách khắc phục: request coalescing / single-flight (một request tính lại, các cái khác chờ), TTL so le có jitter, hoặc
stale-while-revalidate. Xem Caching. - Một service restart và mọi client reconnect cùng lúc → reconnect backoff có jitter.
- Một cron job kích cùng một giây ở mọi nơi → rải lịch.
Retry storm
Một dependency chậm; client retry; retry thêm tải; dependency chậm hơn; retry nhiều hơn — một vòng lặp phản hồi dương giữ cho một hệ thống đang hồi phục tiếp tục chết. Phòng thủ: backoff có jitter, retry budget, circuit breaker (ngừng hẳn retry khi dependency rõ ràng đã chết).
Single point of failure (SPOF)
Bất kỳ component nào mà lỗi của nó kéo sập cả hệ thống: một database primary duy nhất, một load balancer duy nhất, một AZ duy nhất, một config service dùng chung, thậm chí một deploy pipeline duy nhất. Tìm chúng bằng cách hỏi “chuyện gì xảy ra nếu cái này chết?” cho mọi ô trong sơ đồ kiến trúc. Loại bỏ bằng redundancy (xem bên dưới).
Chaos engineering
Bạn không thực sự biết hệ thống chịu được lỗi cho tới khi bạn gây ra lỗi và quan sát. Chaos engineering là thực hành có kỷ luật việc chủ động tiêm lỗi có kiểm soát vào production (hoặc môi trường giống production) — kill instance, thêm latency, bỏ một dependency, chia cắt network — để xác minh resilience của bạn đứng vững và tìm điểm yếu trước khi chúng tìm bạn. Được Netflix tiên phong với Chaos Monkey. Hãy làm với một giả thuyết, một blast radius nhỏ, và một nút abort. Nó biến “chúng tôi nghĩ mình xử được node chết” thành “chúng tôi đã xác minh nó hôm thứ Ba”.
Reliability metrics
Bạn quản lý reliability bằng con số, không phải cảm giác.
Availability — the nines
Availability = uptime / tổng thời gian, thường phát biểu bằng “số nines”:
| Availability | Downtime / năm | Downtime / tháng | Nhãn thường gặp |
|---|---|---|---|
| 99% (“two nines”) | 3,65 ngày | 7,2 giờ | Công cụ nội bộ |
| 99,9% (“three nines”) | 8,77 giờ | 43,2 phút | SaaS điển hình |
| 99,95% | 4,38 giờ | 21,6 phút | Nghiệp vụ quan trọng |
| 99,99% (“four nines”) | 52,6 phút | 4,32 phút | High availability |
| 99,999% (“five nines”) | 5,26 phút | 25,9 giây | Telecom / rất tốn kém |
Mỗi nine thêm vào khó và tốn hơn khoảng 10 lần. Availability nhân với nhau qua các dependency: một service phụ thuộc ba component, mỗi cái 99,9%, có trần khoảng ~99,7% (0,999³) — đây là lý do các chuỗi dependency sâu thì mong manh, và vì sao redundancy và graceful degradation lại quan trọng.
SLI, SLO, SLA
Bộ từ vựng của Google SRE, từ trong ra ngoài:
| Thuật ngữ | Ý nghĩa | Ví dụ |
|---|---|---|
| SLI (Indicator) | Một metric đo được về sức khỏe service | ”% request phục vụ < 300ms với 2xx/3xx” |
| SLO (Objective) | Mục tiêu nội bộ cho một SLI | ”99,9% request đạt SLI trong 28 ngày” |
| SLA (Agreement) | Hợp đồng có hậu quả (hoàn tiền) nếu vi phạm | ”99,5% uptime hoặc chúng tôi credit 10%” |
Nguyên tắc: SLA < SLO — lời hứa công khai (SLA) nên lỏng hơn mục tiêu nội bộ (SLO) để bạn có biên độ trước khi phải đền tiền ai. Chọn SLI phản ánh trải nghiệm của user (latency, error rate, availability, freshness), không phải proxy nội bộ như CPU.
Error budget
Nếu SLO của bạn là 99,9%, bạn được phép 0,1% không đáng tin cậy — đó là error budget của bạn (~43 phút/tháng). Điều này định nghĩa lại reliability như một tài nguyên để tiêu:
- Còn budget → ship feature nhanh, dám mạo hiểm.
- Cạn budget → đóng băng các launch rủi ro, ưu tiên công việc reliability.
Nó giải quyết mâu thuẫn muôn thuở dev-vs-ops bằng một con số cả hai bên đồng ý, và thừa nhận rằng 100% là mục tiêu sai — nó đắt đến bất khả thi và user không phân biệt được 100% với 99,99% khi chính network của họ đã chập chờn.
MTBF / MTTR và RTO / RPO
| Metric | Ý nghĩa | Cải thiện bằng |
|---|---|---|
| MTBF (Mean Time Between Failures) | Bao lâu thì hỏng một lần | Redundancy, chất lượng, testing |
| MTTR (Mean Time To Recovery) | Mất bao lâu để phục hồi | Rollback nhanh, observability tốt, runbook, tự động hóa |
| RTO (Recovery Time Objective) | Downtime tối đa chấp nhận được sau thảm họa | Hạ tầng standby, automated failover |
| RPO (Recovery Point Objective) | Mất dữ liệu tối đa chấp nhận được (theo thời gian) | Tần suất backup/replication |
Tư duy reliability hiện đại ưu tiên giảm MTTR hơn là tối đa hóa MTBF — bạn không thể ngăn mọi lỗi, nên hãy đầu tư vào phục hồi nhanh. RTO/RPO định hướng chiến lược disaster-recovery và backup: RPO 5 phút nghĩa là bạn phải replicate/backup ít nhất với tần suất đó.
High availability & redundancy
High availability (HA) nghĩa là loại bỏ single point of failure thông qua redundancy để hệ thống sống sót khi mất component.
Active-active vs active-passive
| Active-active | Active-passive (failover) | |
|---|---|---|
| Cấu hình | Mọi replica phục vụ traffic đồng thời | Standby nằm chờ cho tới khi primary fail |
| Capacity dùng | 100% (mọi node làm việc) | ~50% (standby nhàn) — tốn hơn trên mỗi đơn vị phục vụ |
| Failover | Tức thì — LB chỉ ngừng dùng node chết | Mất thời gian promote standby (RTO cao hơn) |
| Độ phức tạp | Cần state không xung đột/được điều phối | Đơn giản hơn; kinh điển là DB primary + replica |
| Ví dụ | App tier stateless sau một LB | Postgres primary với một standby có thể promote |
Failover
Failover là tự động chuyển sang một component dự phòng khi cái đang active fail. Các mối quan tâm chính: phát hiện lỗi nhanh và đáng tin (health check, heartbeat), tránh split-brain (hai node cùng nghĩ mình là primary — ngăn bằng quorum/consensus như Raft hoặc fencing token), và kiểm thử rằng failover thực sự chạy (failover chưa được test thường là không chạy).
Multi-AZ và multi-region
- Multi-AZ (Availability Zone): trải instance qua các datacenter cô lập trong một region. Sống sót khi một datacenter fail; latency giữa các AZ thấp; là baseline chuẩn cho HA production.
- Multi-region: replicate qua các region địa lý. Sống sót khi cả một region outage và phục vụ user gần họ hơn (latency thấp hơn), nhưng đưa vào độ phức tạp về latency xuyên region, data-consistency, và chi phí. Thường dành cho các tầng availability cao nhất hoặc tệp user toàn cầu.
Thang trade-off: một instance → nhiều instance (một AZ) → multi-AZ → multi-region, mỗi bước mua thêm availability với thêm chi phí và độ phức tạp.
Idempotency & consistency dưới quy mô lớn
Ở quy mô lớn, exactly-once delivery về cơ bản là bất khả thi — network rớt response, client retry, và message bị gửi lại. Câu trả lời thực tế là idempotency: thiết kế thao tác sao cho làm hai lần cho cùng kết quả như làm một lần.
- Dùng idempotency key: client gửi một key duy nhất kèm request (ví dụ
Idempotency-Key: <uuid>); server ghi lại key + kết quả, và khi gặp key lặp thì trả kết quả đã lưu thay vì thực thi lại. Đây là cách Stripe làm cho retry thanh toán an toàn. - Ưu tiên thiết kế idempotent tự nhiên:
PUT(set state) hơnPOST(append), upsert, “set balance = X” hơn “add X”. - Với consumer của message queue, hãy chờ đợi delivery kiểu at-least-once và dedupe theo message ID.
Consistency dưới quy mô là một phổ (xem Database Scaling cho CAP và replication): strong consistency tốn latency và availability; eventual consistency scale tốt hơn nhưng buộc bạn xử lý các read cũ. Chọn theo từng use case — số dư ngân hàng cần strong; số đếm “like” chịu được eventual.
Các fallacy của distributed computing
Danh sách kinh điển của Peter Deutsch về những giả định sai khiến hệ thống distributed fail. Hãy thấm chúng:
- Network đáng tin cậy. (Không phải — thiết kế cho message rớt/trùng lặp.)
- Latency bằng 0. (Gọi remote chậm hơn local hàng bậc; batch và cache.)
- Bandwidth vô hạn. (Kích thước payload quan trọng ở quy mô lớn.)
- Network an toàn. (Giả định thù địch; mã hóa và xác thực.)
- Topology không đổi. (Node đến rồi đi; đừng hardcode địa chỉ — dùng discovery.)
- Chỉ có một administrator. (Nhiều team và dependency bạn không kiểm soát.)
- Chi phí transport bằng 0. (Serialization, bandwidth, và hạ tầng đều tốn.)
- Network là đồng nhất. (Phần cứng, protocol, version pha trộn.)
Gần như mọi resilience pattern trong note này tồn tại vì một trong các fallacy này là sai.
Capacity planning, autoscaling & profiling
Capacity planning
Ước lượng tài nguyên cần để đáp ứng nhu cầu ở mức reliability mục tiêu. Các bước: đo mức dùng hiện tại và chi phí mỗi request, dự báo tăng trưởng và đỉnh (kể cả spike theo mùa/launch), load test để tìm điểm gãy thật (đừng đoán), và cấp phát có headroom (thường nhắm ~60-70% utilization để có chỗ hấp thụ spike và mất một node). Lên kế hoạch cho đỉnh, không phải trung bình.
Autoscaling
Tự động thêm/bớt instance dựa trên tải. Cần statelessness (node mới phải hoán đổi được).
- Reactive / metric-based: scale theo CPU, memory, request rate, hoặc queue depth. Đơn giản nhưng trễ — scale up mất thời gian (boot + warmup), nên có thể quá chậm với spike gắt.
- Scheduled: pre-scale cho các mẫu đã biết (giờ hành chính, một đợt sale lúc trưa).
- Predictive: dùng ML dự báo nhu cầu và scale trước.
Cạm bẫy: cold start (node mới chậm cho tới khi ấm — pre-warm hoặc giữ một warm pool), flapping (scale lên xuống liên tục — dùng cooldown và hysteresis), và scale một tầng trong khi DB downstream không theo kịp (nhiều app node hơn → nhiều DB connection hơn → DB chảy). Scale-out không miễn phí ở downstream. Queue depth thường là tín hiệu autoscaling tốt hơn CPU cho các async worker.
Profiling để tìm bottleneck
Bạn không thể tối ưu cái chưa đo, và trực giác về bottleneck thường sai. Dùng profiler (CPU, memory, allocation), flame graph, distributed tracing (Observability), và load test (k6, Gatling, Locust) để tìm ràng buộc thật. Áp dụng Định luật Amdahl: độ tăng tốc từ tối ưu một phần bị giới hạn bởi phần đó chiếm bao nhiêu tổng thời gian — hãy tối ưu cái đóng góp lớn nhất vào tail latency trước. Bottleneck thường gặp: N+1 query, thiếu index, tranh chấp lock, cú gọi đồng bộ lẽ ra nên async, connection pool quá nhỏ, và overhead serialization.
Bảng tổng hợp các pattern
| Pattern | Vấn đề nó giải quyết | Trade-off / chi phí chính |
|---|---|---|
| Horizontal scaling | Trần capacity của một máy | Cần statelessness + LB; nhiều thành phần hơn |
| Statelessness | Không thể thêm/bớt node tự do | Phải đẩy state ra ngoài (shared cache/DB) |
| Load balancing | Phân tải; route vòng qua node chết | Bản thân LB phải được làm redundant |
| Timeout | Dependency treo làm cạn tài nguyên | Quá ngắn → lỗi giả; phải phân bổ ngân sách |
| Retry + backoff + jitter | Lỗi tạm thời | Chỉ an toàn cho op idempotent; có thể khuếch đại tải |
| Circuit breaker | Cascading failure từ một dependency ốm | Tuning ngưỡng; cần một fallback |
| Bulkhead | Một dependency ngốn hết tài nguyên chung | Peak utilization thấp hơn (pool cô lập) |
| Rate limiting / throttling | Lạm dụng, spike, overload | Từ chối traffic hợp lệ ở biên; phải tuning |
| Backpressure | Buffer vô hạn → OOM | Đẩy tải/latency ngược về caller |
| Load shedding | Sụp toàn bộ dưới overload | Một số request bị bỏ có chủ đích |
| Graceful degradation / fallback | Dependency không quan trọng fail | Giảm chức năng; thêm nhánh code |
| Idempotency key | Op trùng lặp do retry | Cần lưu trữ + tra cứu key |
| Caching | Tải lặp lại lên store chậm | Dữ liệu cũ, phức tạp invalidation |
| Async / message queue | Coupling, spike tải | Eventual consistency; overhead vận hành |
| Redundancy (HA) | Single point of failure | Gấp đôi chi phí hạ tầng |
| Autoscaling | Vừa cỡ với nhu cầu dao động | Cold start, flapping, áp lực downstream |
| Chaos engineering | Giả định về lỗi chưa được test | Rủi ro gây incident; cần kỷ luật |
Best Practices
- Thiết kế stateless từ ngày đầu. Đẩy session và state ra ngoài. Đây là quyết định đòn bẩy lớn nhất cho scale, và đau đớn khi phải sửa về sau.
- Đặt timeout cho mọi cú gọi network, và phân bổ ngân sách sao cho tổng vừa với latency SLO. Timeout vô hạn là nguyên nhân phổ biến nhất của cascading failure.
- Chỉ retry thao tác idempotent, với exponential backoff + full jitter có giới hạn và một retry budget. Đừng bao giờ retry mù quáng.
- Bọc các dependency hay lỗi trong một circuit breaker với fallback được định nghĩa rõ. Fail nhanh, đừng dồn ứ.
- Cô lập tài nguyên bằng bulkhead để một dependency xấu không làm cạn cả thread/connection pool.
- Giới hạn queue và buffer để có backpressure thay vì crash out-of-memory.
- Shed load một cách thông minh dưới overload — từ chối việc ưu tiên thấp trước, giữ đường core sống; phục vụ tốt phần lớn user hơn là crash cho tất cả.
- Làm health check có ý nghĩa nhưng không cascade — tách liveness khỏi readiness; đừng để một cú chớp DB đánh dấu mọi node là chết.
- Diệt single point of failure: LB redundant, multi-AZ, DB failover. Hỏi “nếu cái ô này chết thì sao?” cho mọi component.
- Thiết kế API idempotent (idempotency key, upsert) — giả định mọi request có thể được gửi hai lần.
- Cache mạnh tay nhưng lên kế hoạch invalidation và chống stampede (single-flight, TTL có jitter). Xem Caching.
- Định nghĩa SLI/SLO theo góc nhìn của user và quản lý một error budget. Đừng đuổi theo 100%. Đo p95/p99, không bao giờ chỉ dùng trung bình.
- Tối ưu MTTR, không chỉ MTBF — rollback nhanh, runbook, observability tốt, và phục hồi tự động hơn là cố không bao giờ fail.
- Load test để tìm điểm gãy thật và profile để tìm bottleneck thật — đừng bao giờ đoán. Cấp phát có headroom.
- Thực hành chaos engineering và test failover — redundancy chưa được test thì không phải redundancy.
- Nhớ rằng availability nhân với nhau — giảm thiểu các chuỗi dependency đồng bộ sâu; ưu tiên async và graceful degradation.
- Autoscale theo tín hiệu đúng (thường queue depth hơn CPU), phòng cold start và flapping, và để mắt tới capacity downstream.
Tài liệu tham khảo
- AWS Well-Architected Framework — Reliability Pillar
- Google SRE Book — Site Reliability Engineering (xem các chương về SLO, Handling Overload, Addressing Cascading Failures)
- Google SRE Workbook — Implementing SLOs
- Michael T. Nygard — Release It! Design and Deploy Production-Ready Software (ấn bản 2, Pragmatic Bookshelf)
- Timeouts, Retries, and Backoff with Jitter — AWS Builders’ Library
- Using Load Shedding to Avoid Overload — AWS Builders’ Library
- Martin Fowler — CircuitBreaker
- The AKF Scale Cube và Fallacies of Distributed Computing (Wikipedia)
Part of the Backend Roadmap knowledge base.
Overview
Scalability is a system’s ability to handle more work — more users, more requests, more data — by adding resources, ideally without a rewrite and ideally at a cost that grows no faster than the load. Reliability is a system’s ability to keep doing the right thing correctly even when parts of it fail, slow down, or behave unexpectedly. The two are deeply linked: a system that cannot scale eventually falls over under load (a reliability failure), and a system that is not designed for failure cannot scale safely because every added component is one more thing that can break.
For a backend engineer these are not “later” concerns. The architectural decisions that make a system scalable and reliable — statelessness, load balancing, timeouts, retries, idempotency, redundancy — are cheap to design in early and brutally expensive to retrofit. This note covers how to reason about scale, the patterns that keep systems standing under stress, the failure modes those patterns defend against, and the metrics (SLA/SLO/SLI, error budgets, the “nines”) that let you manage reliability as an engineering discipline rather than a hope.
A useful mental model comes from Michael Nygard’s Release It!: every integration point is a potential failure, and failures propagate. Your job is not to prevent all failure — that is impossible in a distributed system — but to contain it so that one slow database, one dead node, or one traffic spike does not take down everything.
Related notes in this knowledge base:
- Database Design & Scaling — replication, sharding, read replicas.
- Caching — the cheapest way to buy scalability and reduce load.
- Message Brokers & Async — decoupling and load leveling with queues.
- Observability & Monitoring — you cannot manage reliability you cannot measure.
- Load balancer internals and networking:
../network-engineerroadmap material.
Fundamentals
Vertical vs horizontal scaling
There are two fundamental ways to add capacity.
| Vertical scaling (scale up) | Horizontal scaling (scale out) | |
|---|---|---|
| What | Bigger machine: more CPU, RAM, faster disk | More machines behind a load balancer |
| Effort | Trivial — change instance type, reboot | Requires stateless design, LB, coordination |
| Ceiling | Hard physical limit (biggest instance you can buy) | Effectively unbounded |
| Failure blast radius | One big machine = single point of failure | One node dies, others absorb the load |
| Cost curve | Superlinear — top-end hardware is priced steeply | Roughly linear with commodity hardware |
| Downtime to resize | Usually a reboot | Add/remove nodes with zero downtime |
The practical rule: scale up first because it is simple, but design so you can scale out, because vertical scaling always hits a wall and gives you no redundancy. Real systems combine both — moderately sized nodes, many of them.
Statelessness: the key that unlocks horizontal scale
A service is stateless when any request can be served by any instance, because the instance holds no client-specific data between requests. State is pushed to shared backing stores: databases, caches (Redis), object storage, and the request/token itself (e.g. a JWT).
Why it matters: if instance A holds a user’s session in local memory, that user must return to A. You cannot freely add, remove, or replace instances — you have lost the whole point of horizontal scaling, and you have created a single point of failure per user.
- Stateful (bad for scale)clientapp A: session in RAMmust always hit A
- Stateless (scales freely)clientLBapp A / app B / app CRedis (sessions)Postgres (data)
Rules for statelessness:
- No in-memory sessions — use a shared session store or stateless tokens.
- No local file writes that later requests depend on — use object storage (S3).
- No “sticky” in-process caches that must be consistent across nodes — use a shared cache.
- Any node must be safely killable at any time (this also enables autoscaling and rolling deploys).
The scaling cube (AKF Scale Cube)
The AKF Scale Cube describes three orthogonal axes of scaling:
| Axis | Name | Technique | Example |
|---|---|---|---|
| X | Horizontal duplication | Clone the whole app behind a load balancer | 10 identical app servers |
| Y | Functional decomposition | Split by responsibility into services | orders, payments, users services (microservices) |
| Z | Data partitioning | Split by data / customer | Shard by user_id, or per-tenant databases |
- X is the easiest and where you start — it needs statelessness and a load balancer.
- Y aligns with team boundaries and lets you scale hot subsystems independently.
- Z is sharding — it removes the single-database bottleneck (see Database Scaling).
Mature systems use all three.
Performance vs scalability
These are often confused but are different properties:
- Performance = how fast a single request is served (latency) at low load.
- Scalability = how well the system maintains acceptable performance as load grows.
A system can be fast for one user and collapse at a thousand (great performance, poor scalability). Another can be moderately fast per request but hold that latency flat from 1 to 1,000,000 users (the scalable one). Optimize for the shape of the curve, not just the starting point. Always measure tail latency (p95/p99), not the average — the average hides the users having the worst experience, and at scale the tail is many real users.
Little’s Law — the capacity intuition
L = λ × W, where L = concurrent requests in the system, λ = arrival rate (req/s), W = average time in system (s). If requests arrive at 500/s and each takes 0.2s, you have ~100 requests in flight on average and need enough threads/connections/workers to hold them. When W rises (a slow dependency), L rises, resources exhaust, and you queue or fall over — this is the mechanism behind most overload collapses.
Key Concepts
Load balancing
A load balancer (LB) sits in front of your instances and distributes incoming requests across them. It is the linchpin of horizontal scaling and also a reliability tool: it routes traffic away from unhealthy nodes.
L4 vs L7 load balancing
| L4 (transport) | L7 (application) | |
|---|---|---|
| Operates on | TCP/UDP — IPs and ports | HTTP — paths, headers, cookies, methods |
| Sees | Packets/connections, not content | Full HTTP request |
| Can do | Fast connection forwarding | Path/host routing, TLS termination, header rewrite, WAF, sticky by cookie |
| Overhead | Very low, extremely fast | Higher (parses requests) |
| Examples | AWS NLB, HAProxy (TCP mode), IPVS | AWS ALB, NGINX, Envoy, Traefik |
Use L7 for HTTP microservices (routing /api/orders to one service, /api/users to another; canary by header). Use L4 when you need raw throughput, non-HTTP protocols, or want the backend to terminate TLS itself.
Balancing algorithms
| Algorithm | How it picks a backend | Best for |
|---|---|---|
| Round robin | Next server in rotation | Homogeneous nodes, uniform requests |
| Weighted round robin | Rotation biased by capacity weights | Mixed instance sizes |
| Least connections | Server with fewest active connections | Long-lived / variable-duration requests |
| Least response time | Fewest connections + lowest latency | Latency-sensitive services |
| IP / consistent hashing | Hash of client IP or key → same server | Cache affinity, sticky routing without cookies |
| Random (power of two choices) | Pick 2 at random, choose the less loaded | Great load spread with minimal state, scales well |
Power of two random choices deserves a mention: picking two servers at random and sending to the less loaded one gives near-optimal balancing without any global state, which is why it is popular in large systems.
Health checks
The LB must know which backends are alive. Two kinds:
- Liveness / passive — mark a node bad after N failed responses; stop sending it traffic.
- Readiness / active — periodically poll a health endpoint (
GET /healthz) and only route to nodes that pass.
Make the health check deep enough to be meaningful but shallow enough not to cascade: a /healthz that checks the database will mark every node unhealthy the moment the database blips, taking down a service that could have served cached traffic. A common split: /livez (is the process up?) vs /readyz (can it serve, incl. critical deps?).
Sticky sessions
Sticky sessions (session affinity) pin a client to one backend (via cookie or IP hash). They are a crutch for stateful apps — they undermine even load distribution, break when a node dies (the pinned users lose their session), and complicate deploys. Prefer statelessness with a shared session store. If you must use stickiness, use cookie-based affinity with a short TTL and ensure your app degrades gracefully when re-pinned. See load balancer internals in the ../network-engineer material.
Resilience patterns
These are the defensive patterns from Nygard’s Release It! and the cloud reliability canon. Each contains a specific failure.
Timeouts
Never make a network call without a timeout. The default in many libraries is infinite, which means one hung dependency can exhaust all your threads/connections as requests pile up waiting forever (Little’s Law in action). Set timeouts at every layer: connect timeout, read/socket timeout, total request timeout. Budget them: if your SLA is 300ms and you call three services, they cannot each have a 1s timeout. Timeouts turn a hang (fatal, unbounded) into a fast error (recoverable).
Retries with exponential backoff + jitter
Transient failures (a brief network blip, a rebooting node) are worth retrying. But naive retries are dangerous:
- Retry only idempotent operations (or use idempotency keys — see below). Retrying a non-idempotent
POST /chargecan double-charge. - Bound the retries (e.g. max 3). Infinite retries turn a blip into a storm.
- Back off exponentially: wait 100ms, 200ms, 400ms… so you do not hammer a struggling dependency.
- Add jitter (randomize the delay). Without jitter, all clients that failed at the same instant retry at the same instant — a synchronized retry storm / thundering herd that keeps the recovering service down.
import random, time
def call_with_retry(fn, max_attempts=3, base=0.1, cap=2.0):
for attempt in range(max_attempts):
try:
return fn()
except TransientError:
if attempt == max_attempts - 1:
raise
# full jitter: sleep uniformly in [0, min(cap, base * 2**attempt)]
backoff = min(cap, base * (2 ** attempt))
time.sleep(random.uniform(0, backoff))
AWS’s recommended default is exponential backoff with full jitter. Also consider retry budgets (cap retries to e.g. 10% of total requests) so retries can never more than modestly amplify load.
Circuit breaker
A circuit breaker stops calling a dependency that is clearly failing, giving it room to recover and failing fast instead of piling up doomed, slow calls. It is a state machine:
| State | Behavior | Transition |
|---|---|---|
| Closed | Calls pass through; count failures | Failure rate exceeds threshold → Open |
| Open | Calls fail immediately (or fall back), no request sent | After a cooldown → Half-Open |
| Half-Open | Allow a few trial calls | Success → Closed; failure → Open |
Without a breaker, a slow dependency causes every caller thread to block on its timeout, exhausting the pool and cascading the failure upstream. The breaker converts that into an instant, cheap failure (paired ideally with a fallback). Libraries: Resilience4j (Java), Polly (.NET), and service meshes (Envoy/Istio) provide it at the infra layer.
Bulkhead
Named after ship compartments that stop one flooded section from sinking the vessel. Isolate resources so one overloaded dependency cannot consume all of a shared pool. Example: give calls to the flaky recommendations service their own thread pool / connection pool of 10, separate from the checkout pool. If recommendations hangs, it exhausts its 10 threads and no more — checkout keeps working. Bulkheads turn a total outage into a partial, contained degradation.
Rate limiting & throttling
Rate limiting caps how many requests a client (or the whole system) may make in a window, protecting you from abuse, buggy clients, and traffic spikes. Throttling is the act of rejecting/delaying requests over the limit (usually HTTP 429 Too Many Requests with a Retry-After header).
Common algorithms:
| Algorithm | Behavior | Notes |
|---|---|---|
| Token bucket | Tokens refill at rate R, each request spends one; bucket allows bursts up to size B | Most popular; allows controlled bursts |
| Leaky bucket | Requests drain at a fixed rate; overflow rejected | Smooths output to a steady rate |
| Fixed window | Count per calendar window (per minute) | Simple but has edge-of-window burst problem |
| Sliding window | Rolling count over the last N seconds | Smoother, more accurate, slightly costlier |
Rate limit per key (API key, user, IP) and enforce it in a shared store (Redis) so the limit is global across instances, not per-node.
Backpressure
Backpressure is a system telling its upstream to slow down when it cannot keep up, instead of silently buffering until it runs out of memory. Instead of accepting unbounded work, a component signals “I’m full” — via a bounded queue that blocks/rejects when full, TCP flow control, HTTP 429, or reactive-streams demand signals. The alternative — unbounded queues — just moves the failure to an out-of-memory crash and hides growing latency. Bounded queues are a feature, not a limitation. Async systems (see Message Brokers) provide natural backpressure: producers slow when the queue backs up.
Load shedding
When overloaded, it is better to serve most requests well than all requests badly (or none, after a crash). Load shedding deliberately drops or rejects a fraction of incoming work to stay within capacity — for example, rejecting low-priority requests first, or returning 503 once concurrency exceeds a threshold. Combined with priority (shed batch/analytics traffic before user-facing checkout), it keeps the system alive and the important work flowing under overload. Google SRE emphasizes shedding before you saturate, based on measured load.
Graceful degradation & fallbacks
Design so that when a non-critical dependency fails, the system loses a feature, not the whole page. Examples:
- Recommendations service down → show a static “popular items” list instead of personalized ones.
- Cache down → serve from the database (slower but working), or serve slightly stale data.
- Payment provider A down → failover to provider B.
- Fresh data unavailable → serve last-known-good cached value (“stale-while-revalidate”).
A fallback is the specific alternative returned when the primary path fails (often paired with a circuit breaker). Decide fallbacks deliberately per dependency; a silent empty result can be worse than a clear error.
Failure modes & designing for failure
You must understand the ways distributed systems fail, because the resilience patterns above each defend against a specific one.
Cascading failures
One component’s failure overloads another, which fails, which overloads the next — a domino chain. Classic path: DB slows → app threads block waiting → thread pool exhausts → app stops responding to health checks → LB removes nodes → remaining nodes get more load → they die faster. Defenses: timeouts, circuit breakers, bulkheads, load shedding, and health checks that do not amplify.
Thundering herd
Many clients hit the same resource at the same instant. Common triggers:
- A popular cache key expires → thousands of requests simultaneously miss and stampede the database. Fix: request coalescing / single-flight (one request recomputes, others wait), staggered TTLs with jitter, or
stale-while-revalidate. See Caching. - A service restarts and every client reconnects at once → jittered reconnect backoff.
- A cron job fires the same second everywhere → spread the schedule.
Retry storm
A dependency slows; clients retry; retries add load; the dependency slows more; more retries — a positive feedback loop that keeps a recovering system down. Defenses: backoff with jitter, retry budgets, circuit breakers (which stop retries entirely when a dependency is clearly down).
Single point of failure (SPOF)
Any component whose failure takes down the whole system: a single database primary, a single load balancer, a single AZ, a shared config service, even a single deploy pipeline. Find them by asking “what happens if this dies?” for every box in your architecture diagram. Remove with redundancy (see below).
Chaos engineering
You do not truly know your system tolerates failure until you cause failure and watch. Chaos engineering is the disciplined practice of injecting controlled failures in production (or production-like environments) — killing instances, adding latency, dropping a dependency, partitioning the network — to verify your resilience holds and to find weaknesses before they find you. Pioneered by Netflix’s Chaos Monkey. Do it with a hypothesis, a small blast radius, and an abort switch. It turns “we think we handle a node death” into “we verified it Tuesday.”
Reliability metrics
You manage reliability numerically, not by feel.
Availability — the nines
Availability = uptime / total time, usually stated as “nines”:
| Availability | Downtime / year | Downtime / month | Common label |
|---|---|---|---|
| 99% (“two nines”) | 3.65 days | 7.2 hours | Internal tools |
| 99.9% (“three nines”) | 8.77 hours | 43.2 min | Typical SaaS |
| 99.95% | 4.38 hours | 21.6 min | Business-critical |
| 99.99% (“four nines”) | 52.6 min | 4.32 min | High availability |
| 99.999% (“five nines”) | 5.26 min | 25.9 sec | Telecom / very costly |
Each extra nine is roughly 10× harder and more expensive. Availability multiplies across dependencies: a service that depends on three components each at 99.9% has a ceiling of ~99.7% (0.999³) — this is why deep dependency chains are fragile and why redundancy and graceful degradation matter.
SLI, SLO, SLA
The Google SRE vocabulary, from the inside out:
| Term | Meaning | Example |
|---|---|---|
| SLI (Indicator) | A measured metric of service health | ”% of requests served < 300ms with 2xx/3xx” |
| SLO (Objective) | The internal target for an SLI | ”99.9% of requests meet the SLI over 28 days” |
| SLA (Agreement) | A contract with consequences (refunds) if breached | ”99.5% uptime or we credit 10%” |
Rule of thumb: SLA < SLO — your public promise (SLA) should be looser than your internal target (SLO) so you have margin before you owe anyone money. Pick SLIs that reflect the user’s experience (latency, error rate, availability, freshness), not internal proxies like CPU.
Error budgets
If your SLO is 99.9%, you are permitted 0.1% unreliability — that is your error budget (~43 min/month). This reframes reliability as a resource to spend:
- Budget remaining → ship features fast, take risks.
- Budget exhausted → freeze risky launches, prioritize reliability work.
It resolves the eternal dev-vs-ops tension with a number both sides agree on, and it acknowledges that 100% is the wrong target — it is impossibly expensive and users cannot tell the difference between 100% and 99.99% given their own network flakiness.
MTBF / MTTR and RTO / RPO
| Metric | Meaning | Improve by |
|---|---|---|
| MTBF (Mean Time Between Failures) | How often it breaks | Redundancy, quality, testing |
| MTTR (Mean Time To Recovery) | How long to recover | Fast rollback, good observability, runbooks, automation |
| RTO (Recovery Time Objective) | Max acceptable downtime after disaster | Standby infra, automated failover |
| RPO (Recovery Point Objective) | Max acceptable data loss (time) | Backup/replication frequency |
Modern reliability thinking prioritizes lowering MTTR over maximizing MTBF — you cannot prevent all failures, so invest in recovering fast. RTO/RPO drive your disaster-recovery and backup strategy: an RPO of 5 minutes means you must replicate/back up at least that often.
High availability & redundancy
High availability (HA) means eliminating single points of failure through redundancy so the system survives component loss.
Active-active vs active-passive
| Active-active | Active-passive (failover) | |
|---|---|---|
| Setup | All replicas serve traffic simultaneously | Standby idle until primary fails |
| Capacity used | 100% (all nodes working) | ~50% (standby idle) — costlier per unit served |
| Failover | Instant — LB just stops using the dead node | Takes time to promote standby (higher RTO) |
| Complexity | Needs conflict-free/coordinated state | Simpler; classic DB primary + replica |
| Example | Stateless app tier behind an LB | Postgres primary with a promotable standby |
Failover
Failover is automatically switching to a redundant component when the active one fails. Key concerns: fast and reliable failure detection (health checks, heartbeats), avoiding split-brain (two nodes both think they are primary — prevent with quorum/consensus like Raft or a fencing token), and testing that failover actually works (untested failover usually does not).
Multi-AZ and multi-region
- Multi-AZ (Availability Zone): spread instances across isolated datacenters within one region. Survives a datacenter failure; low latency between AZs; the standard baseline for production HA.
- Multi-region: replicate across geographic regions. Survives a whole-region outage and serves users closer to them (lower latency), but introduces cross-region latency, data-consistency, and cost complexity. Usually reserved for the highest availability tiers or global user bases.
The trade-off ladder: single instance → multi-instance (one AZ) → multi-AZ → multi-region, each step buying more availability at more cost and complexity.
Idempotency & consistency under scale
At scale, exactly-once delivery is essentially impossible — networks drop responses, clients retry, and messages get redelivered. The practical answer is idempotency: design operations so that doing them twice has the same effect as doing them once.
- Use idempotency keys: the client sends a unique key with a request (e.g.
Idempotency-Key: <uuid>); the server records the key + result, and on a repeat key returns the stored result instead of re-executing. This is how Stripe makes payment retries safe. - Prefer naturally idempotent designs:
PUT(set state) overPOST(append), upserts, “set balance to X” over “add X”. - For consumers of a message queue, expect at-least-once delivery and dedupe on a message ID.
Consistency under scale is a spectrum (see Database Scaling for CAP and replication): strong consistency costs latency and availability; eventual consistency scales better but forces you to handle stale reads. Choose per use case — a bank balance needs strong; a “like” count tolerates eventual.
The fallacies of distributed computing
Peter Deutsch’s classic list of false assumptions that cause distributed systems to fail. Internalize them:
- The network is reliable. (It is not — design for dropped/duplicated messages.)
- Latency is zero. (Remote calls are orders of magnitude slower than local; batch and cache.)
- Bandwidth is infinite. (Payload size matters at scale.)
- The network is secure. (Assume hostile; encrypt and authenticate.)
- Topology doesn’t change. (Nodes come and go; do not hardcode addresses — use discovery.)
- There is one administrator. (Many teams and dependencies you don’t control.)
- Transport cost is zero. (Serialization, bandwidth, and infra all cost.)
- The network is homogeneous. (Mixed hardware, protocols, versions.)
Nearly every resilience pattern in this note exists because one of these fallacies is false.
Capacity planning, autoscaling & profiling
Capacity planning
Estimate the resources needed to meet demand at target reliability. Steps: measure current usage and per-request cost, forecast growth and peaks (including seasonal/launch spikes), load test to find the real breaking point (do not guess), and provision with headroom (commonly target ~60-70% utilization so you have room to absorb spikes and lose a node). Plan for peak, not average.
Autoscaling
Automatically add/remove instances based on load. Requires statelessness (new nodes must be interchangeable).
- Reactive / metric-based: scale on CPU, memory, request rate, or queue depth. Simple but lags — scaling up takes time (boot + warmup), so it can be too slow for sharp spikes.
- Scheduled: pre-scale for known patterns (business hours, a sale at noon).
- Predictive: ML-forecast demand and scale ahead of it.
Pitfalls: cold starts (new nodes are slow until warm — pre-warm or keep a warm pool), flapping (scaling up and down rapidly — use cooldowns and hysteresis), and scaling a tier while its downstream DB cannot keep up (more app nodes → more DB connections → DB melts). Scale-out is not free downstream. Queue depth is often a better autoscaling signal than CPU for async workers.
Profiling to find bottlenecks
You cannot optimize what you have not measured, and intuition about bottlenecks is usually wrong. Use profilers (CPU, memory, allocation), flame graphs, distributed tracing (Observability), and load tests (k6, Gatling, Locust) to find the actual constraint. Apply Amdahl’s Law: the speedup from optimizing one part is limited by how much of the total time that part represents — optimize the biggest contributor to tail latency first. Common bottlenecks: N+1 queries, missing indexes, lock contention, synchronous calls that should be async, undersized connection pools, and serialization overhead.
Patterns summary table
| Pattern | Problem it solves | Key trade-off / cost |
|---|---|---|
| Horizontal scaling | Capacity ceiling of one machine | Requires statelessness + LB; more moving parts |
| Statelessness | Can’t add/remove nodes freely | Must externalize state (shared cache/DB) |
| Load balancing | Distribute load; route around dead nodes | LB itself must be made redundant |
| Timeout | Hung dependency exhausts resources | Too short → false failures; must budget |
| Retry + backoff + jitter | Transient failures | Only safe for idempotent ops; can amplify load |
| Circuit breaker | Cascading failure from a sick dependency | Tuning thresholds; needs a fallback |
| Bulkhead | One dependency starving shared resources | Lower peak utilization (isolated pools) |
| Rate limiting / throttling | Abuse, spikes, overload | Rejects legitimate traffic at the edge; tuning |
| Backpressure | Unbounded buffering → OOM | Pushes load/latency back to caller |
| Load shedding | Total collapse under overload | Some requests dropped deliberately |
| Graceful degradation / fallback | Non-critical dependency failure | Reduced functionality; extra code paths |
| Idempotency keys | Duplicate ops from retries | Storage + lookup for keys |
| Caching | Repeated load on slow stores | Staleness, invalidation complexity |
| Async / message queue | Coupling, load spikes | Eventual consistency; operational overhead |
| Redundancy (HA) | Single points of failure | Double the infra cost |
| Autoscaling | Right-size to fluctuating demand | Cold starts, flapping, downstream pressure |
| Chaos engineering | Untested failure assumptions | Risk of injected incidents; needs discipline |
Best Practices
- Design stateless from day one. Externalize sessions and state. It is the single highest-leverage decision for scale, and painful to retrofit.
- Put a timeout on every network call, and budget them so the sum fits your latency SLO. Infinite timeouts are the most common cause of cascading failure.
- Retry only idempotent operations, with bounded exponential backoff + full jitter and a retry budget. Never retry blindly.
- Wrap flaky dependencies in a circuit breaker with a defined fallback. Fail fast, do not pile up.
- Isolate resources with bulkheads so one bad dependency cannot drain the whole thread/connection pool.
- Bound your queues and buffers to get backpressure instead of out-of-memory crashes.
- Shed load intelligently under overload — reject low-priority work first, keep the core path alive; serving most users well beats crashing for all.
- Make health checks meaningful but non-cascading — separate liveness from readiness; do not let a DB blip mark every node dead.
- Kill single points of failure: redundant LBs, multi-AZ, DB failover. Ask “what if this one box dies?” for every component.
- Design idempotent APIs (idempotency keys, upserts) — assume every request may be delivered twice.
- Cache aggressively but plan invalidation and stampede protection (single-flight, jittered TTLs). See Caching.
- Define SLIs/SLOs from the user’s perspective and manage an error budget. Do not chase 100%. Measure p95/p99, never just the average.
- Optimize MTTR, not just MTBF — fast rollback, runbooks, good observability, and automated recovery beat trying to never fail.
- Load test to find your real breaking point and profile to find real bottlenecks — never guess. Provision with headroom.
- Practice chaos engineering and test your failover — untested redundancy is not redundancy.
- Remember availability multiplies — minimize deep synchronous dependency chains; prefer async and graceful degradation.
- Autoscale on the right signal (often queue depth over CPU), guard against cold starts and flapping, and watch downstream capacity.
References
- AWS Well-Architected Framework — Reliability Pillar
- Google SRE Book — Site Reliability Engineering (see chapters on SLOs, Handling Overload, Addressing Cascading Failures)
- Google SRE Workbook — Implementing SLOs
- Michael T. Nygard — Release It! Design and Deploy Production-Ready Software (2nd ed., Pragmatic Bookshelf)
- Timeouts, Retries, and Backoff with Jitter — AWS Builders’ Library
- Using Load Shedding to Avoid Overload — AWS Builders’ Library
- Martin Fowler — CircuitBreaker
- The AKF Scale Cube and Fallacies of Distributed Computing (Wikipedia)