← Backend← Backend
BackendBackend19 Th7, 2026Jul 19, 202631 phút đọc26 min read

Khả năng mở rộng & Độ tin cậyScalability & Reliability

Thuộc bộ kiến thức Backend Roadmap.

Tổng quan

Scalability (khả năng mở rộng) là khả năng của hệ thống xử lý nhiều việc hơn — nhiều user hơn, nhiều request hơn, nhiều dữ liệu hơn — bằng cách thêm tài nguyên, lý tưởng là không phải viết lại và với chi phí tăng không nhanh hơn tải. Reliability (độ tin cậy) là khả năng hệ thống vẫn làm đúng việc một cách chính xác ngay cả khi một phần của nó fail, chậm lại, hoặc hành xử bất thường. Hai khái niệm này gắn chặt với nhau: một hệ thống không scale được thì sớm muộn cũng sập dưới tải (một dạng lỗi reliability), còn một hệ thống không được thiết kế để chịu lỗi thì không thể scale an toàn, vì mỗi component thêm vào là thêm một thứ có thể hỏng.

Với một backend engineer, đây không phải là chuyện “để sau”. Những quyết định kiến trúc giúp hệ thống scalable và reliable — statelessness, load balancing, timeout, retry, idempotency, redundancy — rẻ khi thiết kế từ đầu nhưng cực kỳ đắt khi phải lắp thêm về sau. Note này bao quát: cách tư duy về scale, các pattern giúp hệ thống đứng vững dưới áp lực, các failure mode mà những pattern đó phòng thủ, và các metric (SLA/SLO/SLI, error budget, “the nines”) giúp bạn quản lý reliability như một kỷ luật kỹ thuật thay vì một niềm hy vọng.

Một mô hình tư duy hữu ích đến từ cuốn Release It! của Michael Nygard: mỗi integration point đều là một điểm có thể fail, và lỗi thì lan truyền. Việc của bạn không phải là ngăn mọi lỗi — điều đó bất khả thi trong hệ thống distributed — mà là kiềm chế lỗi để một database chậm, một node chết, hay một đợt spike traffic không kéo sập toàn bộ.

Các note liên quan trong bộ kiến thức này:

Kiến thức nền tảng

Vertical scaling vs horizontal scaling

Có hai cách cơ bản để thêm capacity.

Vertical scaling (scale up)Horizontal scaling (scale out)
Là gìMáy to hơn: nhiều CPU, RAM, disk nhanh hơnNhiều máy hơn sau một load balancer
Công sứcĐơn giản — đổi instance type, rebootCần thiết kế stateless, LB, điều phối
TrầnGiới hạn vật lý cứng (instance to nhất có thể mua)Về cơ bản không giới hạn
Blast radius khi lỗiMột máy to = single point of failureMột node chết, các node khác gánh tải
Đường cong chi phíSiêu tuyến tính — phần cứng cao cấp rất đắtGần như tuyến tính với commodity hardware
Downtime khi resizeThường phải rebootThêm/bớt node với zero downtime

Quy tắc thực tế: scale up trước vì đơn giản, nhưng thiết kế sao cho có thể scale out, vì vertical scaling luôn đụng trần và không cho bạn redundancy. Hệ thống thực tế kết hợp cả hai — node cỡ vừa phải, nhưng nhiều.

Statelessness: chìa khóa mở ra horizontal scale

Một service là stateless khi bất kỳ request nào cũng có thể được phục vụ bởi bất kỳ instance nào, vì instance không giữ dữ liệu riêng của client giữa các request. State được đẩy ra các backing store dùng chung: database, cache (Redis), object storage, và bản thân request/token (ví dụ JWT).

Vì sao quan trọng: nếu instance A giữ session của một user trong bộ nhớ local, user đó bắt buộc phải quay lại A. Bạn không thể tự do thêm, bớt, hay thay thế instance — bạn đã mất toàn bộ ý nghĩa của horizontal scaling, và tạo ra một single point of failure cho mỗi user.

Stateful vs stateless
  1. Stateful (tệ cho scale)
    client
    app A: session in RAMluôn phải trúng A
  2. Stateless (scale thoải mái)
    client
    LB
    app A / app B / app C
    Redis (sessions)
    Postgres (data)

Quy tắc cho statelessness:

The scaling cube (AKF Scale Cube)

AKF Scale Cube mô tả ba trục scaling trực giao:

TrụcTênKỹ thuậtVí dụ
XNhân bản theo chiều ngangClone toàn bộ app sau một load balancer10 app server giống hệt nhau
YPhân rã theo chức năngTách theo trách nhiệm thành các serviceservice orders, payments, users (microservices)
ZPhân vùng dữ liệuTách theo dữ liệu / khách hàngShard theo user_id, hoặc database riêng cho từng tenant

Hệ thống trưởng thành dùng cả ba.

Performance vs scalability

Hai khái niệm này hay bị nhầm nhưng là hai thuộc tính khác nhau:

Một hệ thống có thể nhanh với một user nhưng sụp với một nghìn (performance tốt, scalability kém). Một hệ thống khác có thể chậm vừa phải mỗi request nhưng giữ latency đó phẳng từ 1 đến 1.000.000 user (cái scalable). Hãy tối ưu cho hình dạng đường cong, không chỉ điểm xuất phát. Luôn đo tail latency (p95/p99), không phải trung bình — trung bình che giấu những user có trải nghiệm tệ nhất, và ở quy mô lớn thì cái tail chính là rất nhiều user thật.

Little’s Law — trực giác về capacity

L = λ × W, với L = số request đồng thời trong hệ thống, λ = tốc độ đến (req/s), W = thời gian trung bình trong hệ thống (s). Nếu request đến với 500/s và mỗi cái mất 0,2s, bạn có trung bình ~100 request đang xử lý và cần đủ thread/connection/worker để giữ chúng. Khi W tăng (một dependency chậm), L tăng, tài nguyên cạn, và bạn phải queue hoặc sập — đây là cơ chế đằng sau hầu hết các cú sụp do overload.

Khái niệm chính

Load balancing

Một load balancer (LB) đứng trước các instance của bạn và phân phối request đến trên chúng. Nó là trụ cột của horizontal scaling, đồng thời là một công cụ reliability: nó điều hướng traffic tránh xa các node không khỏe.

L4 vs L7 load balancing

L4 (transport)L7 (application)
Hoạt động ởTCP/UDP — IP và portHTTP — path, header, cookie, method
Nhìn thấyPacket/connection, không thấy nội dungToàn bộ HTTP request
Làm đượcForward connection nhanhRouting theo path/host, TLS termination, sửa header, WAF, sticky theo cookie
OverheadRất thấp, cực nhanhCao hơn (phải parse request)
Ví dụAWS NLB, HAProxy (TCP mode), IPVSAWS ALB, NGINX, Envoy, Traefik

Dùng L7 cho microservice HTTP (route /api/orders sang một service, /api/users sang service khác; canary theo header). Dùng L4 khi cần throughput thô, protocol không phải HTTP, hoặc muốn backend tự terminate TLS.

Các thuật toán balancing

Thuật toánCách chọn backendPhù hợp cho
Round robinServer kế tiếp trong vòng xoayNode đồng nhất, request đều nhau
Weighted round robinVòng xoay có trọng số theo capacityInstance kích cỡ khác nhau
Least connectionsServer có ít connection active nhấtRequest kéo dài / thời lượng biến thiên
Least response timeÍt connection nhất + latency thấp nhấtService nhạy latency
IP / consistent hashingHash IP client hoặc key → cùng một serverCache affinity, sticky routing không cần cookie
Random (power of two choices)Chọn ngẫu nhiên 2, lấy cái ít tải hơnPhân tải rất tốt với state tối thiểu, scale tốt

Power of two random choices đáng nhắc đến: chọn ngẫu nhiên hai server và gửi sang cái ít tải hơn cho kết quả balancing gần tối ưu mà không cần state toàn cục, nên rất phổ biến trong các hệ thống lớn.

Health check

LB phải biết backend nào còn sống. Có hai loại:

Hãy làm health check đủ sâu để có ý nghĩa nhưng đủ nông để không cascade: một /healthz mà kiểm tra database sẽ đánh dấu mọi node là unhealthy ngay khi database chớp lỗi, kéo sập một service lẽ ra vẫn phục vụ được traffic từ cache. Cách tách phổ biến: /livez (process còn chạy không?) vs /readyz (phục vụ được không, kể cả dependency quan trọng?).

Sticky sessions

Sticky sessions (session affinity) ghim một client vào một backend (qua cookie hoặc IP hash). Đây là cái nạng cho app stateful — nó phá vỡ phân tải đều, gãy khi một node chết (các user bị ghim mất session), và làm phức tạp việc deploy. Hãy ưu tiên statelessness với shared session store. Nếu buộc phải dùng stickiness, dùng affinity theo cookie với TTL ngắn và đảm bảo app degrade mượt khi bị re-pin. Xem chi tiết load balancer trong tài liệu ../network-engineer.

Resilience patterns

Đây là các pattern phòng thủ từ Release It! của Nygard và kho kiến thức reliability của cloud. Mỗi pattern chặn một loại lỗi cụ thể.

Timeout

Đừng bao giờ gọi network mà không có timeout. Mặc định trong nhiều thư viện là vô hạn, nghĩa là một dependency treo có thể làm cạn sạch thread/connection khi request dồn lại chờ mãi mãi (Little’s Law hiện hình). Đặt timeout ở mọi tầng: connect timeout, read/socket timeout, total request timeout. Hãy phân bổ ngân sách: nếu SLA của bạn là 300ms và bạn gọi ba service, chúng không thể mỗi cái có timeout 1s. Timeout biến một cú treo (chí mạng, vô hạn) thành một lỗi nhanh (có thể phục hồi).

Retry với exponential backoff + jitter

Lỗi tạm thời (một cú chớp mạng, một node đang reboot) đáng để retry. Nhưng retry ngây thơ thì nguy hiểm:

import random, time

def call_with_retry(fn, max_attempts=3, base=0.1, cap=2.0):
    for attempt in range(max_attempts):
        try:
            return fn()
        except TransientError:
            if attempt == max_attempts - 1:
                raise
            # full jitter: ngủ đều trong [0, min(cap, base * 2**attempt)]
            backoff = min(cap, base * (2 ** attempt))
            time.sleep(random.uniform(0, backoff))

Mặc định AWS khuyến nghị là exponential backoff with full jitter. Cũng nên cân nhắc retry budget (giới hạn retry ở mức ví dụ 10% tổng request) để retry không bao giờ khuếch đại tải quá mức.

Circuit breaker

Một circuit breaker ngừng gọi một dependency đang fail rõ ràng, cho nó không gian để hồi phục và fail nhanh thay vì dồn lại những cú gọi chậm chắc chắn thất bại. Nó là một máy trạng thái:

Trạng tháiHành viChuyển tiếp
ClosedCho request đi qua; đếm lỗiTỷ lệ lỗi vượt ngưỡng → Open
OpenFail ngay lập tức (hoặc fallback), không gửi requestSau một khoảng cooldown → Half-Open
Half-OpenCho vài cú gọi thửThành công → Closed; lỗi → Open

Không có breaker, một dependency chậm khiến mọi thread caller block trên timeout của nó, làm cạn pool và cascade lỗi ngược lên trên. Breaker biến điều đó thành một lỗi tức thì, rẻ (lý tưởng là ghép với một fallback). Thư viện: Resilience4j (Java), Polly (.NET), và các service mesh (Envoy/Istio) cung cấp ở tầng hạ tầng.

Bulkhead

Đặt theo tên các khoang tàu ngăn không cho một khoang ngập nước làm chìm cả con tàu. Cô lập tài nguyên để một dependency quá tải không nuốt hết một pool dùng chung. Ví dụ: cấp cho các cú gọi service recommendations hay lỗi một thread pool / connection pool riêng gồm 10, tách khỏi pool checkout. Nếu recommendations treo, nó chỉ làm cạn 10 thread của nó và không hơn — checkout vẫn chạy. Bulkhead biến một cú outage toàn bộ thành một sự suy giảm cục bộ, được kiềm chế.

Rate limiting & throttling

Rate limiting giới hạn một client (hoặc cả hệ thống) được gửi bao nhiêu request trong một khung thời gian, bảo vệ bạn khỏi lạm dụng, client lỗi, và spike traffic. Throttling là hành động từ chối/trì hoãn các request vượt giới hạn (thường là HTTP 429 Too Many Requests với header Retry-After).

Các thuật toán phổ biến:

Thuật toánHành viGhi chú
Token bucketToken nạp lại với tốc độ R, mỗi request tiêu một; bucket cho phép burst tới cỡ BPhổ biến nhất; cho phép burst có kiểm soát
Leaky bucketRequest chảy ra với tốc độ cố định; tràn thì bị từ chốiLàm mượt output về tốc độ ổn định
Fixed windowĐếm theo khung lịch (mỗi phút)Đơn giản nhưng có vấn đề burst ở rìa cửa sổ
Sliding windowĐếm cuộn trong N giây gần nhấtMượt hơn, chính xác hơn, tốn hơn chút

Hãy rate limit theo key (API key, user, IP) và thực thi trong một shared store (Redis) để giới hạn là toàn cục trên các instance, không phải theo từng node.

Backpressure

Backpressure là việc một hệ thống báo cho upstream chậm lại khi nó không theo kịp, thay vì âm thầm buffer cho tới khi hết bộ nhớ. Thay vì nhận việc vô hạn, một component phát tín hiệu “tôi đầy rồi” — qua một bounded queue chặn/từ chối khi đầy, TCP flow control, HTTP 429, hoặc tín hiệu demand của reactive streams. Lựa chọn thay thế — queue vô hạn — chỉ dời cú lỗi sang một cú crash out-of-memory và che giấu latency đang tăng. Bounded queue là một tính năng, không phải hạn chế. Hệ thống async (xem Message Brokers) cung cấp backpressure tự nhiên: producer chậm lại khi queue dồn ứ.

Load shedding

Khi quá tải, phục vụ phần lớn request tốt vẫn hơn phục vụ tất cả request tệ (hoặc không phục vụ gì, sau khi crash). Load shedding chủ động bỏ hoặc từ chối một phần công việc đến để giữ trong ngưỡng capacity — ví dụ, từ chối các request ưu tiên thấp trước, hoặc trả 503 khi concurrency vượt một ngưỡng. Kết hợp với priority (bỏ traffic batch/analytics trước traffic checkout hướng user), nó giữ hệ thống sống và dòng công việc quan trọng vẫn chảy dưới overload. Google SRE nhấn mạnh việc shed trước khi bão hòa, dựa trên tải đo được.

Graceful degradation & fallback

Thiết kế sao cho khi một dependency không quan trọng fail, hệ thống mất một tính năng, không phải cả trang. Ví dụ:

Một fallback là phương án thay thế cụ thể được trả về khi đường chính fail (thường ghép với circuit breaker). Hãy quyết định fallback một cách có chủ đích cho từng dependency; một kết quả rỗng âm thầm có thể tệ hơn một lỗi rõ ràng.

Failure mode & thiết kế để chịu lỗi

Bạn phải hiểu các cách hệ thống distributed fail, vì mỗi resilience pattern ở trên chặn một cách cụ thể.

Cascading failure (lỗi dây chuyền)

Lỗi của một component làm quá tải component khác, cái đó fail, làm quá tải cái tiếp theo — một chuỗi domino. Đường điển hình: DB chậm → thread app block chờ → thread pool cạn → app ngừng đáp health check → LB gỡ node → các node còn lại nhận nhiều tải hơn → chúng chết nhanh hơn. Phòng thủ: timeout, circuit breaker, bulkhead, load shedding, và health check không khuếch đại.

Thundering herd

Nhiều client đập vào cùng một tài nguyên cùng một khoảnh khắc. Các nguyên nhân phổ biến:

Retry storm

Một dependency chậm; client retry; retry thêm tải; dependency chậm hơn; retry nhiều hơn — một vòng lặp phản hồi dương giữ cho một hệ thống đang hồi phục tiếp tục chết. Phòng thủ: backoff có jitter, retry budget, circuit breaker (ngừng hẳn retry khi dependency rõ ràng đã chết).

Single point of failure (SPOF)

Bất kỳ component nào mà lỗi của nó kéo sập cả hệ thống: một database primary duy nhất, một load balancer duy nhất, một AZ duy nhất, một config service dùng chung, thậm chí một deploy pipeline duy nhất. Tìm chúng bằng cách hỏi “chuyện gì xảy ra nếu cái này chết?” cho mọi ô trong sơ đồ kiến trúc. Loại bỏ bằng redundancy (xem bên dưới).

Chaos engineering

Bạn không thực sự biết hệ thống chịu được lỗi cho tới khi bạn gây ra lỗi và quan sát. Chaos engineering là thực hành có kỷ luật việc chủ động tiêm lỗi có kiểm soát vào production (hoặc môi trường giống production) — kill instance, thêm latency, bỏ một dependency, chia cắt network — để xác minh resilience của bạn đứng vững và tìm điểm yếu trước khi chúng tìm bạn. Được Netflix tiên phong với Chaos Monkey. Hãy làm với một giả thuyết, một blast radius nhỏ, và một nút abort. Nó biến “chúng tôi nghĩ mình xử được node chết” thành “chúng tôi đã xác minh nó hôm thứ Ba”.

Reliability metrics

Bạn quản lý reliability bằng con số, không phải cảm giác.

Availability — the nines

Availability = uptime / tổng thời gian, thường phát biểu bằng “số nines”:

AvailabilityDowntime / nămDowntime / thángNhãn thường gặp
99% (“two nines”)3,65 ngày7,2 giờCông cụ nội bộ
99,9% (“three nines”)8,77 giờ43,2 phútSaaS điển hình
99,95%4,38 giờ21,6 phútNghiệp vụ quan trọng
99,99% (“four nines”)52,6 phút4,32 phútHigh availability
99,999% (“five nines”)5,26 phút25,9 giâyTelecom / rất tốn kém

Mỗi nine thêm vào khó và tốn hơn khoảng 10 lần. Availability nhân với nhau qua các dependency: một service phụ thuộc ba component, mỗi cái 99,9%, có trần khoảng ~99,7% (0,999³) — đây là lý do các chuỗi dependency sâu thì mong manh, và vì sao redundancy và graceful degradation lại quan trọng.

SLI, SLO, SLA

Bộ từ vựng của Google SRE, từ trong ra ngoài:

Thuật ngữÝ nghĩaVí dụ
SLI (Indicator)Một metric đo được về sức khỏe service”% request phục vụ < 300ms với 2xx/3xx”
SLO (Objective)Mục tiêu nội bộ cho một SLI”99,9% request đạt SLI trong 28 ngày”
SLA (Agreement)Hợp đồng có hậu quả (hoàn tiền) nếu vi phạm”99,5% uptime hoặc chúng tôi credit 10%”

Nguyên tắc: SLA < SLO — lời hứa công khai (SLA) nên lỏng hơn mục tiêu nội bộ (SLO) để bạn có biên độ trước khi phải đền tiền ai. Chọn SLI phản ánh trải nghiệm của user (latency, error rate, availability, freshness), không phải proxy nội bộ như CPU.

Error budget

Nếu SLO của bạn là 99,9%, bạn được phép 0,1% không đáng tin cậy — đó là error budget của bạn (~43 phút/tháng). Điều này định nghĩa lại reliability như một tài nguyên để tiêu:

Nó giải quyết mâu thuẫn muôn thuở dev-vs-ops bằng một con số cả hai bên đồng ý, và thừa nhận rằng 100% là mục tiêu sai — nó đắt đến bất khả thi và user không phân biệt được 100% với 99,99% khi chính network của họ đã chập chờn.

MTBF / MTTR và RTO / RPO

MetricÝ nghĩaCải thiện bằng
MTBF (Mean Time Between Failures)Bao lâu thì hỏng một lầnRedundancy, chất lượng, testing
MTTR (Mean Time To Recovery)Mất bao lâu để phục hồiRollback nhanh, observability tốt, runbook, tự động hóa
RTO (Recovery Time Objective)Downtime tối đa chấp nhận được sau thảm họaHạ tầng standby, automated failover
RPO (Recovery Point Objective)Mất dữ liệu tối đa chấp nhận được (theo thời gian)Tần suất backup/replication

Tư duy reliability hiện đại ưu tiên giảm MTTR hơn là tối đa hóa MTBF — bạn không thể ngăn mọi lỗi, nên hãy đầu tư vào phục hồi nhanh. RTO/RPO định hướng chiến lược disaster-recovery và backup: RPO 5 phút nghĩa là bạn phải replicate/backup ít nhất với tần suất đó.

High availability & redundancy

High availability (HA) nghĩa là loại bỏ single point of failure thông qua redundancy để hệ thống sống sót khi mất component.

Active-active vs active-passive

Active-activeActive-passive (failover)
Cấu hìnhMọi replica phục vụ traffic đồng thờiStandby nằm chờ cho tới khi primary fail
Capacity dùng100% (mọi node làm việc)~50% (standby nhàn) — tốn hơn trên mỗi đơn vị phục vụ
FailoverTức thì — LB chỉ ngừng dùng node chếtMất thời gian promote standby (RTO cao hơn)
Độ phức tạpCần state không xung đột/được điều phốiĐơn giản hơn; kinh điển là DB primary + replica
Ví dụApp tier stateless sau một LBPostgres primary với một standby có thể promote

Failover

Failover là tự động chuyển sang một component dự phòng khi cái đang active fail. Các mối quan tâm chính: phát hiện lỗi nhanh và đáng tin (health check, heartbeat), tránh split-brain (hai node cùng nghĩ mình là primary — ngăn bằng quorum/consensus như Raft hoặc fencing token), và kiểm thử rằng failover thực sự chạy (failover chưa được test thường là không chạy).

Multi-AZ và multi-region

Thang trade-off: một instance → nhiều instance (một AZ) → multi-AZ → multi-region, mỗi bước mua thêm availability với thêm chi phí và độ phức tạp.

Idempotency & consistency dưới quy mô lớn

Ở quy mô lớn, exactly-once delivery về cơ bản là bất khả thi — network rớt response, client retry, và message bị gửi lại. Câu trả lời thực tế là idempotency: thiết kế thao tác sao cho làm hai lần cho cùng kết quả như làm một lần.

Consistency dưới quy mô là một phổ (xem Database Scaling cho CAP và replication): strong consistency tốn latency và availability; eventual consistency scale tốt hơn nhưng buộc bạn xử lý các read cũ. Chọn theo từng use case — số dư ngân hàng cần strong; số đếm “like” chịu được eventual.

Các fallacy của distributed computing

Danh sách kinh điển của Peter Deutsch về những giả định sai khiến hệ thống distributed fail. Hãy thấm chúng:

  1. Network đáng tin cậy. (Không phải — thiết kế cho message rớt/trùng lặp.)
  2. Latency bằng 0. (Gọi remote chậm hơn local hàng bậc; batch và cache.)
  3. Bandwidth vô hạn. (Kích thước payload quan trọng ở quy mô lớn.)
  4. Network an toàn. (Giả định thù địch; mã hóa và xác thực.)
  5. Topology không đổi. (Node đến rồi đi; đừng hardcode địa chỉ — dùng discovery.)
  6. Chỉ có một administrator. (Nhiều team và dependency bạn không kiểm soát.)
  7. Chi phí transport bằng 0. (Serialization, bandwidth, và hạ tầng đều tốn.)
  8. Network là đồng nhất. (Phần cứng, protocol, version pha trộn.)

Gần như mọi resilience pattern trong note này tồn tại vì một trong các fallacy này là sai.

Capacity planning, autoscaling & profiling

Capacity planning

Ước lượng tài nguyên cần để đáp ứng nhu cầu ở mức reliability mục tiêu. Các bước: đo mức dùng hiện tại và chi phí mỗi request, dự báo tăng trưởng và đỉnh (kể cả spike theo mùa/launch), load test để tìm điểm gãy thật (đừng đoán), và cấp phát có headroom (thường nhắm ~60-70% utilization để có chỗ hấp thụ spike và mất một node). Lên kế hoạch cho đỉnh, không phải trung bình.

Autoscaling

Tự động thêm/bớt instance dựa trên tải. Cần statelessness (node mới phải hoán đổi được).

Cạm bẫy: cold start (node mới chậm cho tới khi ấm — pre-warm hoặc giữ một warm pool), flapping (scale lên xuống liên tục — dùng cooldown và hysteresis), và scale một tầng trong khi DB downstream không theo kịp (nhiều app node hơn → nhiều DB connection hơn → DB chảy). Scale-out không miễn phí ở downstream. Queue depth thường là tín hiệu autoscaling tốt hơn CPU cho các async worker.

Profiling để tìm bottleneck

Bạn không thể tối ưu cái chưa đo, và trực giác về bottleneck thường sai. Dùng profiler (CPU, memory, allocation), flame graph, distributed tracing (Observability), và load test (k6, Gatling, Locust) để tìm ràng buộc thật. Áp dụng Định luật Amdahl: độ tăng tốc từ tối ưu một phần bị giới hạn bởi phần đó chiếm bao nhiêu tổng thời gian — hãy tối ưu cái đóng góp lớn nhất vào tail latency trước. Bottleneck thường gặp: N+1 query, thiếu index, tranh chấp lock, cú gọi đồng bộ lẽ ra nên async, connection pool quá nhỏ, và overhead serialization.

Bảng tổng hợp các pattern

PatternVấn đề nó giải quyếtTrade-off / chi phí chính
Horizontal scalingTrần capacity của một máyCần statelessness + LB; nhiều thành phần hơn
StatelessnessKhông thể thêm/bớt node tự doPhải đẩy state ra ngoài (shared cache/DB)
Load balancingPhân tải; route vòng qua node chếtBản thân LB phải được làm redundant
TimeoutDependency treo làm cạn tài nguyênQuá ngắn → lỗi giả; phải phân bổ ngân sách
Retry + backoff + jitterLỗi tạm thờiChỉ an toàn cho op idempotent; có thể khuếch đại tải
Circuit breakerCascading failure từ một dependency ốmTuning ngưỡng; cần một fallback
BulkheadMột dependency ngốn hết tài nguyên chungPeak utilization thấp hơn (pool cô lập)
Rate limiting / throttlingLạm dụng, spike, overloadTừ chối traffic hợp lệ ở biên; phải tuning
BackpressureBuffer vô hạn → OOMĐẩy tải/latency ngược về caller
Load sheddingSụp toàn bộ dưới overloadMột số request bị bỏ có chủ đích
Graceful degradation / fallbackDependency không quan trọng failGiảm chức năng; thêm nhánh code
Idempotency keyOp trùng lặp do retryCần lưu trữ + tra cứu key
CachingTải lặp lại lên store chậmDữ liệu cũ, phức tạp invalidation
Async / message queueCoupling, spike tảiEventual consistency; overhead vận hành
Redundancy (HA)Single point of failureGấp đôi chi phí hạ tầng
AutoscalingVừa cỡ với nhu cầu dao độngCold start, flapping, áp lực downstream
Chaos engineeringGiả định về lỗi chưa được testRủi ro gây incident; cần kỷ luật

Best Practices

Tài liệu tham khảo

Part of the Backend Roadmap knowledge base.

Overview

Scalability is a system’s ability to handle more work — more users, more requests, more data — by adding resources, ideally without a rewrite and ideally at a cost that grows no faster than the load. Reliability is a system’s ability to keep doing the right thing correctly even when parts of it fail, slow down, or behave unexpectedly. The two are deeply linked: a system that cannot scale eventually falls over under load (a reliability failure), and a system that is not designed for failure cannot scale safely because every added component is one more thing that can break.

For a backend engineer these are not “later” concerns. The architectural decisions that make a system scalable and reliable — statelessness, load balancing, timeouts, retries, idempotency, redundancy — are cheap to design in early and brutally expensive to retrofit. This note covers how to reason about scale, the patterns that keep systems standing under stress, the failure modes those patterns defend against, and the metrics (SLA/SLO/SLI, error budgets, the “nines”) that let you manage reliability as an engineering discipline rather than a hope.

A useful mental model comes from Michael Nygard’s Release It!: every integration point is a potential failure, and failures propagate. Your job is not to prevent all failure — that is impossible in a distributed system — but to contain it so that one slow database, one dead node, or one traffic spike does not take down everything.

Related notes in this knowledge base:

Fundamentals

Vertical vs horizontal scaling

There are two fundamental ways to add capacity.

Vertical scaling (scale up)Horizontal scaling (scale out)
WhatBigger machine: more CPU, RAM, faster diskMore machines behind a load balancer
EffortTrivial — change instance type, rebootRequires stateless design, LB, coordination
CeilingHard physical limit (biggest instance you can buy)Effectively unbounded
Failure blast radiusOne big machine = single point of failureOne node dies, others absorb the load
Cost curveSuperlinear — top-end hardware is priced steeplyRoughly linear with commodity hardware
Downtime to resizeUsually a rebootAdd/remove nodes with zero downtime

The practical rule: scale up first because it is simple, but design so you can scale out, because vertical scaling always hits a wall and gives you no redundancy. Real systems combine both — moderately sized nodes, many of them.

Statelessness: the key that unlocks horizontal scale

A service is stateless when any request can be served by any instance, because the instance holds no client-specific data between requests. State is pushed to shared backing stores: databases, caches (Redis), object storage, and the request/token itself (e.g. a JWT).

Why it matters: if instance A holds a user’s session in local memory, that user must return to A. You cannot freely add, remove, or replace instances — you have lost the whole point of horizontal scaling, and you have created a single point of failure per user.

Stateful vs stateless
  1. Stateful (bad for scale)
    client
    app A: session in RAMmust always hit A
  2. Stateless (scales freely)
    client
    LB
    app A / app B / app C
    Redis (sessions)
    Postgres (data)

Rules for statelessness:

The scaling cube (AKF Scale Cube)

The AKF Scale Cube describes three orthogonal axes of scaling:

AxisNameTechniqueExample
XHorizontal duplicationClone the whole app behind a load balancer10 identical app servers
YFunctional decompositionSplit by responsibility into servicesorders, payments, users services (microservices)
ZData partitioningSplit by data / customerShard by user_id, or per-tenant databases

Mature systems use all three.

Performance vs scalability

These are often confused but are different properties:

A system can be fast for one user and collapse at a thousand (great performance, poor scalability). Another can be moderately fast per request but hold that latency flat from 1 to 1,000,000 users (the scalable one). Optimize for the shape of the curve, not just the starting point. Always measure tail latency (p95/p99), not the average — the average hides the users having the worst experience, and at scale the tail is many real users.

Little’s Law — the capacity intuition

L = λ × W, where L = concurrent requests in the system, λ = arrival rate (req/s), W = average time in system (s). If requests arrive at 500/s and each takes 0.2s, you have ~100 requests in flight on average and need enough threads/connections/workers to hold them. When W rises (a slow dependency), L rises, resources exhaust, and you queue or fall over — this is the mechanism behind most overload collapses.

Key Concepts

Load balancing

A load balancer (LB) sits in front of your instances and distributes incoming requests across them. It is the linchpin of horizontal scaling and also a reliability tool: it routes traffic away from unhealthy nodes.

L4 vs L7 load balancing

L4 (transport)L7 (application)
Operates onTCP/UDP — IPs and portsHTTP — paths, headers, cookies, methods
SeesPackets/connections, not contentFull HTTP request
Can doFast connection forwardingPath/host routing, TLS termination, header rewrite, WAF, sticky by cookie
OverheadVery low, extremely fastHigher (parses requests)
ExamplesAWS NLB, HAProxy (TCP mode), IPVSAWS ALB, NGINX, Envoy, Traefik

Use L7 for HTTP microservices (routing /api/orders to one service, /api/users to another; canary by header). Use L4 when you need raw throughput, non-HTTP protocols, or want the backend to terminate TLS itself.

Balancing algorithms

AlgorithmHow it picks a backendBest for
Round robinNext server in rotationHomogeneous nodes, uniform requests
Weighted round robinRotation biased by capacity weightsMixed instance sizes
Least connectionsServer with fewest active connectionsLong-lived / variable-duration requests
Least response timeFewest connections + lowest latencyLatency-sensitive services
IP / consistent hashingHash of client IP or key → same serverCache affinity, sticky routing without cookies
Random (power of two choices)Pick 2 at random, choose the less loadedGreat load spread with minimal state, scales well

Power of two random choices deserves a mention: picking two servers at random and sending to the less loaded one gives near-optimal balancing without any global state, which is why it is popular in large systems.

Health checks

The LB must know which backends are alive. Two kinds:

Make the health check deep enough to be meaningful but shallow enough not to cascade: a /healthz that checks the database will mark every node unhealthy the moment the database blips, taking down a service that could have served cached traffic. A common split: /livez (is the process up?) vs /readyz (can it serve, incl. critical deps?).

Sticky sessions

Sticky sessions (session affinity) pin a client to one backend (via cookie or IP hash). They are a crutch for stateful apps — they undermine even load distribution, break when a node dies (the pinned users lose their session), and complicate deploys. Prefer statelessness with a shared session store. If you must use stickiness, use cookie-based affinity with a short TTL and ensure your app degrades gracefully when re-pinned. See load balancer internals in the ../network-engineer material.

Resilience patterns

These are the defensive patterns from Nygard’s Release It! and the cloud reliability canon. Each contains a specific failure.

Timeouts

Never make a network call without a timeout. The default in many libraries is infinite, which means one hung dependency can exhaust all your threads/connections as requests pile up waiting forever (Little’s Law in action). Set timeouts at every layer: connect timeout, read/socket timeout, total request timeout. Budget them: if your SLA is 300ms and you call three services, they cannot each have a 1s timeout. Timeouts turn a hang (fatal, unbounded) into a fast error (recoverable).

Retries with exponential backoff + jitter

Transient failures (a brief network blip, a rebooting node) are worth retrying. But naive retries are dangerous:

import random, time

def call_with_retry(fn, max_attempts=3, base=0.1, cap=2.0):
    for attempt in range(max_attempts):
        try:
            return fn()
        except TransientError:
            if attempt == max_attempts - 1:
                raise
            # full jitter: sleep uniformly in [0, min(cap, base * 2**attempt)]
            backoff = min(cap, base * (2 ** attempt))
            time.sleep(random.uniform(0, backoff))

AWS’s recommended default is exponential backoff with full jitter. Also consider retry budgets (cap retries to e.g. 10% of total requests) so retries can never more than modestly amplify load.

Circuit breaker

A circuit breaker stops calling a dependency that is clearly failing, giving it room to recover and failing fast instead of piling up doomed, slow calls. It is a state machine:

StateBehaviorTransition
ClosedCalls pass through; count failuresFailure rate exceeds threshold → Open
OpenCalls fail immediately (or fall back), no request sentAfter a cooldown → Half-Open
Half-OpenAllow a few trial callsSuccess → Closed; failure → Open

Without a breaker, a slow dependency causes every caller thread to block on its timeout, exhausting the pool and cascading the failure upstream. The breaker converts that into an instant, cheap failure (paired ideally with a fallback). Libraries: Resilience4j (Java), Polly (.NET), and service meshes (Envoy/Istio) provide it at the infra layer.

Bulkhead

Named after ship compartments that stop one flooded section from sinking the vessel. Isolate resources so one overloaded dependency cannot consume all of a shared pool. Example: give calls to the flaky recommendations service their own thread pool / connection pool of 10, separate from the checkout pool. If recommendations hangs, it exhausts its 10 threads and no more — checkout keeps working. Bulkheads turn a total outage into a partial, contained degradation.

Rate limiting & throttling

Rate limiting caps how many requests a client (or the whole system) may make in a window, protecting you from abuse, buggy clients, and traffic spikes. Throttling is the act of rejecting/delaying requests over the limit (usually HTTP 429 Too Many Requests with a Retry-After header).

Common algorithms:

AlgorithmBehaviorNotes
Token bucketTokens refill at rate R, each request spends one; bucket allows bursts up to size BMost popular; allows controlled bursts
Leaky bucketRequests drain at a fixed rate; overflow rejectedSmooths output to a steady rate
Fixed windowCount per calendar window (per minute)Simple but has edge-of-window burst problem
Sliding windowRolling count over the last N secondsSmoother, more accurate, slightly costlier

Rate limit per key (API key, user, IP) and enforce it in a shared store (Redis) so the limit is global across instances, not per-node.

Backpressure

Backpressure is a system telling its upstream to slow down when it cannot keep up, instead of silently buffering until it runs out of memory. Instead of accepting unbounded work, a component signals “I’m full” — via a bounded queue that blocks/rejects when full, TCP flow control, HTTP 429, or reactive-streams demand signals. The alternative — unbounded queues — just moves the failure to an out-of-memory crash and hides growing latency. Bounded queues are a feature, not a limitation. Async systems (see Message Brokers) provide natural backpressure: producers slow when the queue backs up.

Load shedding

When overloaded, it is better to serve most requests well than all requests badly (or none, after a crash). Load shedding deliberately drops or rejects a fraction of incoming work to stay within capacity — for example, rejecting low-priority requests first, or returning 503 once concurrency exceeds a threshold. Combined with priority (shed batch/analytics traffic before user-facing checkout), it keeps the system alive and the important work flowing under overload. Google SRE emphasizes shedding before you saturate, based on measured load.

Graceful degradation & fallbacks

Design so that when a non-critical dependency fails, the system loses a feature, not the whole page. Examples:

A fallback is the specific alternative returned when the primary path fails (often paired with a circuit breaker). Decide fallbacks deliberately per dependency; a silent empty result can be worse than a clear error.

Failure modes & designing for failure

You must understand the ways distributed systems fail, because the resilience patterns above each defend against a specific one.

Cascading failures

One component’s failure overloads another, which fails, which overloads the next — a domino chain. Classic path: DB slows → app threads block waiting → thread pool exhausts → app stops responding to health checks → LB removes nodes → remaining nodes get more load → they die faster. Defenses: timeouts, circuit breakers, bulkheads, load shedding, and health checks that do not amplify.

Thundering herd

Many clients hit the same resource at the same instant. Common triggers:

Retry storm

A dependency slows; clients retry; retries add load; the dependency slows more; more retries — a positive feedback loop that keeps a recovering system down. Defenses: backoff with jitter, retry budgets, circuit breakers (which stop retries entirely when a dependency is clearly down).

Single point of failure (SPOF)

Any component whose failure takes down the whole system: a single database primary, a single load balancer, a single AZ, a shared config service, even a single deploy pipeline. Find them by asking “what happens if this dies?” for every box in your architecture diagram. Remove with redundancy (see below).

Chaos engineering

You do not truly know your system tolerates failure until you cause failure and watch. Chaos engineering is the disciplined practice of injecting controlled failures in production (or production-like environments) — killing instances, adding latency, dropping a dependency, partitioning the network — to verify your resilience holds and to find weaknesses before they find you. Pioneered by Netflix’s Chaos Monkey. Do it with a hypothesis, a small blast radius, and an abort switch. It turns “we think we handle a node death” into “we verified it Tuesday.”

Reliability metrics

You manage reliability numerically, not by feel.

Availability — the nines

Availability = uptime / total time, usually stated as “nines”:

AvailabilityDowntime / yearDowntime / monthCommon label
99% (“two nines”)3.65 days7.2 hoursInternal tools
99.9% (“three nines”)8.77 hours43.2 minTypical SaaS
99.95%4.38 hours21.6 minBusiness-critical
99.99% (“four nines”)52.6 min4.32 minHigh availability
99.999% (“five nines”)5.26 min25.9 secTelecom / very costly

Each extra nine is roughly 10× harder and more expensive. Availability multiplies across dependencies: a service that depends on three components each at 99.9% has a ceiling of ~99.7% (0.999³) — this is why deep dependency chains are fragile and why redundancy and graceful degradation matter.

SLI, SLO, SLA

The Google SRE vocabulary, from the inside out:

TermMeaningExample
SLI (Indicator)A measured metric of service health”% of requests served < 300ms with 2xx/3xx”
SLO (Objective)The internal target for an SLI”99.9% of requests meet the SLI over 28 days”
SLA (Agreement)A contract with consequences (refunds) if breached”99.5% uptime or we credit 10%”

Rule of thumb: SLA < SLO — your public promise (SLA) should be looser than your internal target (SLO) so you have margin before you owe anyone money. Pick SLIs that reflect the user’s experience (latency, error rate, availability, freshness), not internal proxies like CPU.

Error budgets

If your SLO is 99.9%, you are permitted 0.1% unreliability — that is your error budget (~43 min/month). This reframes reliability as a resource to spend:

It resolves the eternal dev-vs-ops tension with a number both sides agree on, and it acknowledges that 100% is the wrong target — it is impossibly expensive and users cannot tell the difference between 100% and 99.99% given their own network flakiness.

MTBF / MTTR and RTO / RPO

MetricMeaningImprove by
MTBF (Mean Time Between Failures)How often it breaksRedundancy, quality, testing
MTTR (Mean Time To Recovery)How long to recoverFast rollback, good observability, runbooks, automation
RTO (Recovery Time Objective)Max acceptable downtime after disasterStandby infra, automated failover
RPO (Recovery Point Objective)Max acceptable data loss (time)Backup/replication frequency

Modern reliability thinking prioritizes lowering MTTR over maximizing MTBF — you cannot prevent all failures, so invest in recovering fast. RTO/RPO drive your disaster-recovery and backup strategy: an RPO of 5 minutes means you must replicate/back up at least that often.

High availability & redundancy

High availability (HA) means eliminating single points of failure through redundancy so the system survives component loss.

Active-active vs active-passive

Active-activeActive-passive (failover)
SetupAll replicas serve traffic simultaneouslyStandby idle until primary fails
Capacity used100% (all nodes working)~50% (standby idle) — costlier per unit served
FailoverInstant — LB just stops using the dead nodeTakes time to promote standby (higher RTO)
ComplexityNeeds conflict-free/coordinated stateSimpler; classic DB primary + replica
ExampleStateless app tier behind an LBPostgres primary with a promotable standby

Failover

Failover is automatically switching to a redundant component when the active one fails. Key concerns: fast and reliable failure detection (health checks, heartbeats), avoiding split-brain (two nodes both think they are primary — prevent with quorum/consensus like Raft or a fencing token), and testing that failover actually works (untested failover usually does not).

Multi-AZ and multi-region

The trade-off ladder: single instance → multi-instance (one AZ) → multi-AZ → multi-region, each step buying more availability at more cost and complexity.

Idempotency & consistency under scale

At scale, exactly-once delivery is essentially impossible — networks drop responses, clients retry, and messages get redelivered. The practical answer is idempotency: design operations so that doing them twice has the same effect as doing them once.

Consistency under scale is a spectrum (see Database Scaling for CAP and replication): strong consistency costs latency and availability; eventual consistency scales better but forces you to handle stale reads. Choose per use case — a bank balance needs strong; a “like” count tolerates eventual.

The fallacies of distributed computing

Peter Deutsch’s classic list of false assumptions that cause distributed systems to fail. Internalize them:

  1. The network is reliable. (It is not — design for dropped/duplicated messages.)
  2. Latency is zero. (Remote calls are orders of magnitude slower than local; batch and cache.)
  3. Bandwidth is infinite. (Payload size matters at scale.)
  4. The network is secure. (Assume hostile; encrypt and authenticate.)
  5. Topology doesn’t change. (Nodes come and go; do not hardcode addresses — use discovery.)
  6. There is one administrator. (Many teams and dependencies you don’t control.)
  7. Transport cost is zero. (Serialization, bandwidth, and infra all cost.)
  8. The network is homogeneous. (Mixed hardware, protocols, versions.)

Nearly every resilience pattern in this note exists because one of these fallacies is false.

Capacity planning, autoscaling & profiling

Capacity planning

Estimate the resources needed to meet demand at target reliability. Steps: measure current usage and per-request cost, forecast growth and peaks (including seasonal/launch spikes), load test to find the real breaking point (do not guess), and provision with headroom (commonly target ~60-70% utilization so you have room to absorb spikes and lose a node). Plan for peak, not average.

Autoscaling

Automatically add/remove instances based on load. Requires statelessness (new nodes must be interchangeable).

Pitfalls: cold starts (new nodes are slow until warm — pre-warm or keep a warm pool), flapping (scaling up and down rapidly — use cooldowns and hysteresis), and scaling a tier while its downstream DB cannot keep up (more app nodes → more DB connections → DB melts). Scale-out is not free downstream. Queue depth is often a better autoscaling signal than CPU for async workers.

Profiling to find bottlenecks

You cannot optimize what you have not measured, and intuition about bottlenecks is usually wrong. Use profilers (CPU, memory, allocation), flame graphs, distributed tracing (Observability), and load tests (k6, Gatling, Locust) to find the actual constraint. Apply Amdahl’s Law: the speedup from optimizing one part is limited by how much of the total time that part represents — optimize the biggest contributor to tail latency first. Common bottlenecks: N+1 queries, missing indexes, lock contention, synchronous calls that should be async, undersized connection pools, and serialization overhead.

Patterns summary table

PatternProblem it solvesKey trade-off / cost
Horizontal scalingCapacity ceiling of one machineRequires statelessness + LB; more moving parts
StatelessnessCan’t add/remove nodes freelyMust externalize state (shared cache/DB)
Load balancingDistribute load; route around dead nodesLB itself must be made redundant
TimeoutHung dependency exhausts resourcesToo short → false failures; must budget
Retry + backoff + jitterTransient failuresOnly safe for idempotent ops; can amplify load
Circuit breakerCascading failure from a sick dependencyTuning thresholds; needs a fallback
BulkheadOne dependency starving shared resourcesLower peak utilization (isolated pools)
Rate limiting / throttlingAbuse, spikes, overloadRejects legitimate traffic at the edge; tuning
BackpressureUnbounded buffering → OOMPushes load/latency back to caller
Load sheddingTotal collapse under overloadSome requests dropped deliberately
Graceful degradation / fallbackNon-critical dependency failureReduced functionality; extra code paths
Idempotency keysDuplicate ops from retriesStorage + lookup for keys
CachingRepeated load on slow storesStaleness, invalidation complexity
Async / message queueCoupling, load spikesEventual consistency; operational overhead
Redundancy (HA)Single points of failureDouble the infra cost
AutoscalingRight-size to fluctuating demandCold starts, flapping, downstream pressure
Chaos engineeringUntested failure assumptionsRisk of injected incidents; needs discipline

Best Practices

References