Cloud Design Patterns & Độ tin cậyCloud Design Patterns & Reliability
Thuộc kho kiến thức theo DevOps Roadmap.
Tổng quan
Cloud design patterns là những giải pháp đã được kiểm chứng, tái sử dụng được cho các vấn đề lặp đi lặp lại khi xây dựng hệ thống phân tán: lỗi cục bộ (partial failure), tải biến động, mạng không đáng tin cậy và tính nhất quán dữ liệu ở quy mô lớn. Trên cloud, lỗi không phải là ngoại lệ—nó là trạng thái vận hành bình thường. Instance bị thu hồi, availability zone suy giảm, dependency throttle bạn. Hệ thống sống tốt là hệ thống được thiết kế cho lỗi, chứ không phải thiết kế với hy vọng lỗi sẽ không xảy ra.
Với kỹ sư DevOps, các pattern này là ngôn ngữ chung giữa development và operations. Những thuật ngữ như RTO/RPO, circuit breaker, bulkhead, warm standby xuất hiện trong architecture review, postmortem sự cố, đàm phán SLO và runbook disaster recovery. Nắm được catalog pattern (AWS Well-Architected, Azure Architecture Center, Google SRE) giúp bạn đánh giá thiết kế nhanh chóng: điều gì xảy ra khi dependency này chậm lại? Blast radius là gì? Hệ thống scale ra sao, và chi phí khôi phục là bao nhiêu?
Có hai ý tưởng làm nền cho tất cả những gì bên dưới. Thứ nhất, các ngộ nhận của distributed computing (fallacies of distributed computing): mạng không đáng tin, latency không bằng 0, bandwidth không vô hạn, topology có thay đổi. Mọi pattern ở đây là một biện pháp phòng vệ trước một trong những giả định sai này. Thứ hai, lỗi mang tính xác suất và tích lũy: với hàng nghìn thành phần, luôn có cái gì đó đang hỏng ở đâu đó; một hệ thống chống chịu tốt sẽ khoanh vùng và hấp thụ những lỗi đó thay vì để chúng lan thành sự cố người dùng thấy được.
Chủ đề này gồm bốn nhóm: tính sẵn sàng và dự phòng (availability, redundancy), khả năng mở rộng (scalability), các pattern quản lý dữ liệu, và các pattern thiết kế/triển khai cho khả năng chống chịu—cộng thêm các chiến lược disaster recovery và phần giới thiệu chaos engineering, kỷ luật kiểm chứng rằng tất cả những thứ trên thực sự hoạt động.
Kiến thức nền tảng
Tính sẵn sàng và dự phòng
- High availability (HA): loại bỏ điểm lỗi đơn (single point of failure) bằng dự phòng. Availability thường biểu diễn theo “số 9”. Mỗi số 9 tăng thêm gần như nhân đôi chi phí và độ phức tạp, nên hãy “mua” số 9 ở nơi cần, không phải khắp nơi.
| Availability | Downtime / năm | Downtime / tháng | Dùng điển hình |
|---|---|---|---|
| 99% (“two nines”) | 3.65 ngày | 7.2 giờ | Công cụ nội bộ, best-effort |
| 99.9% (“three nines”) | 8.77 giờ | 43.8 phút | Baseline SaaS tiêu chuẩn |
| 99.95% | 4.38 giờ | 21.9 phút | Dịch vụ nghiệp vụ trả phí |
| 99.99% (“four nines”) | 52.6 phút | 4.4 phút | Nền tảng trọng yếu |
| 99.999% (“five nines”) | 5.26 phút | 26 giây | Telco / hạ tầng lõi (rất đắt) |
- Các kiểu dự phòng: active-active (mọi replica đều phục vụ traffic; lỗi chỉ làm giảm công suất) so với active-passive (một bản standby tiếp quản khi lỗi, có độ trễ failover). N+1 là một bản dự phòng ngoài mức tối thiểu; 2N là nhân đôi hoàn toàn.
- Multi-AZ: chạy các instance dự phòng trên nhiều availability zone (nguồn điện/mạng/làm mát độc lập trong một region, nhưng đủ gần để replication đồng bộ độ trễ thấp). Bảo vệ trước sự cố data center với chi phí latency thấp; là baseline mặc định cho production.
- Multi-region: bảo vệ trước sự cố toàn region và giảm latency cho người dùng toàn cầu, nhưng kéo theo độ trễ replication dữ liệu, các vấn đề nhất quán, latency cao hơn cho ghi đồng bộ và chi phí đáng kể. Hãy quyết định dựa trên yêu cầu nghiệp vụ, không phải theo trào lưu.
- RTO (Recovery Time Objective): thời gian tối đa chấp nhận được để khôi phục dịch vụ sau sự cố—con số “chúng ta được down bao lâu”.
- RPO (Recovery Point Objective): lượng dữ liệu tối đa chấp nhận mất, đo theo thời gian (ví dụ RPO 5 phút = mất tối đa dữ liệu của 5 phút cuối)—con số “chúng ta được mất bao nhiêu dữ liệu”. RTO/RPO quyết định chiến lược DR và chi phí của nó.
Availability tổng hợp: các dependency nối tiếp nhân với nhau. Hai service 99.9% nối tiếp cho 0.999 × 0.999 ≈ 99.8%—tệ hơn từng cái riêng lẻ, vì lỗi ở bất kỳ cái nào cũng làm đứt chuỗi. Các đường dự phòng song song cải thiện availability: hai đường 99% song song cho 1 − (0.01 × 0.01) = 99.99%. Bài học thực tế: giảm số hard dependency thường rẻ hơn thêm số 9 cho từng cái. Hãy hỏi với mỗi dependency, “đây là hard dependency (nó lỗi thì request của tôi lỗi) hay soft (tôi có thể suy giảm cấp)?” và chuyển hard thành soft ở mọi nơi có thể.
Khả năng mở rộng
- Vertical scaling (scale up): máy to hơn. Đơn giản, không cần sửa code, và đôi khi là nước đi đầu tiên đúng đắn, nhưng có trần cứng (instance lớn nhất), thường phải downtime khi resize, và cỗ máy lớn vẫn là điểm lỗi đơn.
- Horizontal scaling (scale out): thêm máy phía sau load balancer. Gần như không giới hạn, cho phép rolling update và chịu lỗi—nhưng đòi hỏi ứng dụng phải “hợp tác” (không state cục bộ, không session trong bộ nhớ).
- Thiết kế stateless là điều kiện tiên quyết để scale out: đưa session state vào Redis/database, file vào object storage, để bất kỳ instance nào cũng phục vụ được bất kỳ request nào. Instance trở nên dùng-xong-bỏ (“cattle, not pets”) nên bạn có thể thêm, bớt, thay thoải mái.
- Load balancing: phân phối request qua các instance. Chú ý thuật toán (round-robin so với least-connections so với consistent hashing) và health check—một instance không khỏe bị bỏ lại trong vòng xoay là một sự cố cháy chậm.
- Autoscaling: tự động điều chỉnh công suất theo metrics (CPU, request per target, độ sâu queue) hoặc theo lịch. Kubernetes có HPA (scale số pod), VPA (right-size request của pod), Cluster Autoscaler/Karpenter (thêm/bớt node) và KEDA (scale theo event/queue, gồm cả scale về 0). Autoscaling phản ứng luôn trễ hơn tải, nên hãy kết hợp với một khoảng headroom dự phòng, và dùng scaling dự đoán/theo lịch cho các đỉnh đã biết (giờ hành chính, sự kiện sale).
# Kubernetes HPA: scale theo CPU, scale-down thận trọng
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: api
spec:
scaleTargetRef: { apiVersion: apps/v1, kind: Deployment, name: api }
minReplicas: 3
maxReplicas: 30
metrics:
- type: Resource
resource:
name: cpu
target: { type: Utilization, averageUtilization: 65 }
behavior:
scaleDown:
stabilizationWindowSeconds: 300 # chờ 5 phút trước khi scale down
stabilizationWindowSeconds khi scale-down rất quan trọng: nó ngăn flapping, khi một cú giảm tải ngắn làm bạn scale xuống, tải quay lại làm bạn scale lên, và bạn dao động liên tục—trả giá bằng cold start và bất ổn. Scale lên nhanh, scale xuống chậm.
Các pattern quản lý dữ liệu
- Cache-aside (lazy loading): ứng dụng kiểm tra cache trước; nếu miss thì đọc database và ghi vào cache kèm TTL. Đơn giản và phổ biến nhất. Đề phòng hai kiểu lỗi: dữ liệu cũ (stale, giảm thiểu bằng TTL hợp lý và invalidate tường minh khi ghi) và cache stampede (khi một key nóng hết hạn, hàng nghìn request đập vào DB cùng lúc—giảm thiểu bằng TTL có jitter, gộp request/single-flight, hoặc refresh-ahead).
- Read-through / write-through / write-behind: các phương án mà chính cache sở hữu quyền truy cập DB. Write-through giữ cache và DB nhất quán, đổi lại là latency ghi; write-behind đệm ghi để tăng throughput, đổi lại là nguy cơ mất dữ liệu khi crash.
- Replication: nhân bản dữ liệu để scale read và tăng availability. Replication đồng bộ bảo vệ RPO (ghi chưa được ack cho đến khi replica có dữ liệu) nhưng tăng latency ghi và có thể nghẽn khi một replica chậm; bất đồng bộ nhanh hơn nhưng có thể mất phần đuôi các ghi chưa replicate khi failover. Leader-follower là topology phổ biến; cần biết hành vi failover của database và nguy cơ split-brain.
- Sharding (phân mảnh): chia dữ liệu ra nhiều node theo partition key khi một node không gánh nổi tải ghi hoặc kích thước dataset. Chọn key phân bố đều để tránh hot shard; lưu ý truy vấn cross-shard và resharding (chia lại khi lớn lên) là phần khó nhất. Consistent hashing giảm thiểu dịch chuyển dữ liệu khi thêm/bớt node.
- CQRS (Command Query Responsibility Segregation): tách write model khỏi read model, cho phép mỗi bên scale và tối ưu độc lập (ví dụ: ghi chuẩn hóa, view đọc phi chuẩn hóa tinh chỉnh theo pattern truy vấn). Đổi lại là eventual consistency giữa hai bên và độ phức tạp vận hành—dùng cho domain đọc nhiều với hình dạng đọc/ghi khác nhau, không phải khắp nơi.
- Event sourcing: lưu chuỗi sự kiện làm nguồn sự thật thay vì trạng thái hiện tại; trạng thái được suy ra bằng cách replay sự kiện. Kết hợp tự nhiên với CQRS và cho audit trail đầy đủ cùng khả năng “du hành thời gian” khi debug, nhưng trả giá bằng độ phức tạp đáng kể: versioning schema sự kiện, rebuild projection, và thực tế là bạn không thể đơn giản “xóa một hàng”. Dùng có chọn lọc, ở domain nơi lịch sử chính là sản phẩm (sổ cái, audit, workflow).
- Materialized view: tính trước và lưu kết quả một truy vấn tốn kém, refresh theo lịch hoặc khi ghi, để phục vụ read rẻ. Một điểm trung gian thực dụng trước khi làm CQRS đầy đủ.
- Idempotency & outbox pattern: để ghi an toàn dưới retry, cho các thao tác một idempotency key; để publish sự kiện đáng tin cùng với một lần ghi DB, dùng transactional outbox (ghi sự kiện vào bảng outbox trong cùng transaction rồi relay) thay vì dual-write vào DB và một broker (có thể lỗi một phần).
Khái niệm chính
Các pattern chống chịu lỗi
| Pattern | Vấn đề giải quyết | Ý tưởng chính |
|---|---|---|
| Retry với backoff + jitter | Lỗi tạm thời (chập chờn, throttle ngắn) | Retry có giới hạn số lần, delay tăng theo cấp số nhân và ngẫu nhiên hóa |
| Circuit breaker | Dependency lỗi kéo caller sập theo | Sau N lần lỗi thì fail fast (open); định kỳ thăm dò (half-open) trước khi đóng lại |
| Bulkhead | Một dependency “ồn ào” chiếm hết tài nguyên chung | Chia pool (thread, connection, pod) theo từng dependency/tenant để khoanh vùng thiệt hại |
| Throttling / rate limiting | Quá tải và client lạm dụng/mất kiểm soát | Từ chối hoặc xếp hàng phần tải dư từ sớm (token bucket, sliding window); trả về 429 |
| Queue-based load leveling | Producer bùng nổ đè bẹp consumer | Đặt queue ở giữa; consumer xử lý theo tốc độ bền vững |
| Strangler fig | Viết lại kiểu big-bang đầy rủi ro | Route traffic qua một facade; di trú chức năng từng phần sang hệ thống mới |
| Timeout | Dependency chậm giữ tài nguyên vô thời hạn | Chặn mọi lần chờ; giải phóng tài nguyên và fail một cách xác định |
| Fallback / suy giảm cấp | Một dependency không sẵn sàng | Phục vụ kết quả từ cache/mặc định/giảm cấp thay vì lỗi cứng |
| Health check + đẩy khỏi load balancer | Một instance ốm vẫn nhận traffic | Thăm dò liveness/readiness; loại instance lỗi khỏi vòng xoay |
Retry với exponential backoff và full jitter:
import random, time
def call_with_retry(fn, max_attempts=5, base=0.2, cap=10.0):
for attempt in range(max_attempts):
try:
return fn()
except TransientError:
if attempt == max_attempts - 1:
raise
# full jitter: sleep trong [0, min(cap, base * 2^attempt)]
time.sleep(random.uniform(0, min(cap, base * 2 ** attempt)))
Jitter rất quan trọng: không có nó, tất cả client cùng lỗi sẽ retry theo từng đợt đồng bộ (“thundering herd”) và đè chết luôn service đang hồi phục. Chỉ retry các thao tác idempotent, tôn trọng header Retry-After khi server gửi về, và kết hợp retry với một retry budget (giới hạn retry theo tỷ lệ tổng request, ví dụ ≤10%) để retry không thể làm hơn gấp đôi tải khi có sự cố. Anti-pattern kinh điển là retry ở mọi tầng: nếu ba service lồng nhau mỗi cái retry ba lần, một request người dùng thành tới 27 backend call đúng lúc hệ thống ít gánh nổi nhất.
Các trạng thái của circuit breaker: closed (bình thường, đếm lỗi) → open (fail fast, không gọi, trong một khoảng cool-down) → half-open (cho gọi thăm dò giới hạn) → trở về closed khi thành công, hoặc về open nếu tiếp tục lỗi. Giá trị nằm ở hai điểm: nó fail fast (caller nhận lỗi ngay thay vì chờ timeout, giải phóng thread của nó) và cho dependency khoảng thở để hồi phục (ngừng đập vào một service đang chật vật). Kết hợp với fallback để fail fast vẫn trả được câu trả lời dùng được—response từ cache, giá trị mặc định, hoặc tính năng giảm cấp—thay vì một lỗi trống rỗng.
Bulkhead lấy tên từ khoang tàu thủy: cách ly các pool tài nguyên để nước tràn vào một khoang không nhấn chìm cả con tàu. Trong thực tế: cho mỗi dependency phía dưới (hoặc mỗi tenant) một connection pool / thread pool / giới hạn concurrency riêng, để khi dependency X treo, chỉ pool-X cạn còn call tới Y vẫn thành công. Không có bulkhead, một dependency chậm có thể ngốn hết thread trong một pool dùng chung và kéo sập cả service.
Queue-based load leveling tách producer khỏi consumer bằng một queue bền. Ingress bùng nổ, gồ ghề trở thành workload đều đặn theo nhịp consumer; queue hấp thụ đỉnh để consumer chạy ở tốc độ bền vững của nó. Giám sát độ sâu queue và tuổi message như SLI hạng nhất—backlog tăng dần là cảnh báo sớm nhất rằng consumer không theo kịp, từ rất lâu trước khi người dùng nhận ra.
Throttling / rate limiting bảo vệ một service khỏi quá tải và khỏi client lạm dụng hoặc lỗi. Thuật toán phổ biến: token bucket (cho phép burst tới kích thước bucket, nạp lại ở tốc độ đều) và sliding window (làm mượt tốc độ trên một cửa sổ trượt). Trả về HTTP 429 kèm header Retry-After để client ngoan back off. Rate-limit cả ở edge (tự bảo vệ khỏi thế giới) và giữa các service nội bộ (bảo vệ dependency khỏi chính bạn).
Strangler fig (đặt theo loài dây leo mọc quanh một cái cây và dần thay thế nó) di trú một hệ thống cũ theo từng bước: đặt một facade routing phía trước, rồi bóc từng chức năng một sang bản triển khai mới, chuyển traffic khi mỗi phần được kiểm chứng. Hệ thống luôn ship được và đảo ngược được trong suốt quá trình—trái ngược với viết lại big-bang, thứ không ship được hàng tháng và mạo hiểm ở một lần cutover thảm khốc.
Các chiến lược disaster recovery
| Chiến lược | RTO | RPO | Chi phí tương đối | Cách hoạt động |
|---|---|---|---|---|
| Backup & restore | Vài giờ–vài ngày | Vài giờ | $ | Backup dữ liệu/cấu hình; dựng lại hạ tầng từ IaC khi cần |
| Pilot light | Vài chục phút–vài giờ | Vài phút | $$ | Dữ liệu cốt lõi replicate liên tục; chạy tối thiểu vài service; scale lên khi thảm họa |
| Warm standby | Vài phút | Vài giây–vài phút | $$$ | Bản sao đầy đủ chức năng nhưng thu nhỏ luôn chạy; scale lên full khi failover |
| Multi-site active-active | Gần bằng 0 | Gần bằng 0 | $$$$ | Đủ công suất ở nhiều region phục vụ đồng thời; failover = chuyển hướng traffic |
Cách tư duy về cái thang này: mỗi bậc mua cho bạn RTO/RPO thấp hơn với nhiều tiền hơn và phức tạp vận hành hơn. Backup & restore chỉ giữ dữ liệu “ấm” và dựng lại mọi thứ khác—rẻ, nhưng chậm và phụ thuộc vào việc IaC của bạn thực sự chạy được. Pilot light giữ phần “khó replicate” (database, replicate liên tục) luôn bật, còn mọi thứ khác tắt cho đến khi cần. Warm standby chạy một bản sao nhỏ nhưng hoàn chỉnh của cả stack, scale lên khi failover—không có nguy cơ cold start, chỉ là vấn đề công suất. Active-active chạy đủ công suất ở mọi region đồng thời, nên “failover” chỉ là gỡ một region khỏi pool traffic; nó cũng đóng vai trò như chiến lược scaling và giảm latency, nhưng buộc bạn phải giải bài toán nhất quán dữ liệu và giải quyết xung đột đa region.
Cách chọn: xác định RTO/RPO cho từng workload dựa trên tác động nghiệp vụ, rồi mua chiến lược rẻ nhất đáp ứng được—đừng “mạ vàng” một dashboard báo cáo lên active-active. Quan trọng nhất, IaC (Terraform/CloudFormation/Pulumi) và runbook tự động hóa, được diễn tập định kỳ là thứ khiến các tầng rẻ hơn trở nên đáng tin: kế hoạch DR chưa được test chỉ là một giả thuyết, không phải kế hoạch. Cũng cần quyết định failover so với failback (quay về primary khi nó hồi phục) một cách tường minh—nhiều team diễn tập failover mà quên failback, rồi chật vật khi cần quay lại.
Chaos engineering (giới thiệu)
Chaos engineering là thực hành chạy các thí nghiệm có kiểm soát, inject lỗi thật (kill instance, thêm latency, chặn một dependency, đánh sập một AZ, làm đầy đĩa) để kiểm chứng hệ thống hoạt động đúng như thiết kế—trước khi một sự cố chứng minh rằng nó không. Phương pháp:
- Định nghĩa “steady state” bằng output đo được (success rate, latency SLI, số đơn/phút)—một tín hiệu bên ngoài, liên quan nghiệp vụ, không phải metric nội bộ.
- Đặt giả thuyết: “steady state vẫn giữ nếu kill 20% số pod” / “nếu DB payments bị +200 ms latency”.
- Inject lỗi trong blast radius giới hạn: bắt đầu ở staging, rồi production với một phần trăm nhỏ và một cơ chế abort tự động (“nút dừng”) gắn với metric steady-state của bạn.
- Đo đạc, so với giả thuyết, sửa điểm yếu tìm được, rồi tự động hóa thí nghiệm để nó chạy liên tục và bắt được hồi quy.
Công cụ: Chaos Mesh và LitmusChaos (Kubernetes-native, CNCF), Gremlin (thương mại, có guardrail an toàn), AWS Fault Injection Service, và Chaos Monkey / Simian Army của Netflix (khởi nguồn của kỷ luật này). Hãy bắt đầu với GameDay—các buổi diễn tập lỗi có lịch, do con người chạy—trước khi tự động hóa hoàn toàn. Điểm cốt lõi về văn hóa chính là lý do tồn tại của cả kỷ luật này: các cơ chế reliability không bao giờ được diễn tập (failover, circuit breaker, runbook DR, autoscaling) sẽ không hoạt động khi bạn thực sự cần chúng. Chaos engineering biến “chúng ta nghĩ nó chống chịu được” thành “chúng ta đã kiểm chứng rằng nó chống chịu được”.
Best Practices
- Thiết kế cho lỗi ở mọi tầng. Mặc định instance, AZ và dependency sẽ hỏng; đưa xử lý lỗi thành câu hỏi bắt buộc trong design review bằng cách luôn hỏi “điều gì xảy ra khi X chậm, hoặc sập, hoặc trả về rác?”
- Chốt RTO/RPO cho từng workload trước khi chọn kiến trúc. Chi tiền cho DR mà không có mục tiêu rõ ràng thì hoặc là lãng phí (over-provisioning một service ít giá trị), hoặc là tự tin ảo (bảo vệ thiếu một service trọng yếu). Hãy để tác động nghiệp vụ đặt con số.
- Mặc định multi-AZ; multi-region phải có lý do. Multi-AZ là bảo hiểm rẻ và nên là baseline production; multi-region là cả một chương trình lớn với chi phí nhất quán dữ liệu thật—chỉ làm khi có yêu cầu nghiệp vụ hoặc latency thực sự.
- Giảm hard dependency; chuyển hard thành soft. Mỗi dependency đồng bộ trên đường request đều làm giảm availability tổng hợp. Hãy tự hỏi liệu có thể suy giảm cấp êm ái (cache, mặc định, async) thay vì fail khi một dependency sập không.
- Giữ service stateless; đưa state ra ngoài. Session vào Redis, file vào object storage, state vào database—đây chính là thứ mở khóa autoscaling, rolling deploy và thay instance không đau đớn.
- Autoscale theo metric phản ánh đúng tải, và đặt trần. Request-per-target hoặc độ sâu queue thường tốt hơn CPU làm tín hiệu tải; đặt
maxReplicas, scale lên nhanh, và scale xuống chậm (stabilization window) để tránh flapping và chi phí mất kiểm soát. - Mọi lời gọi qua mạng phải có timeout, retry giới hạn với backoff + jitter, và idempotency. Thiếu timeout và retry vô hạn là hai “bộ khuếch đại” hàng đầu biến một trục trặc nhỏ thành sự cố lớn; một retry budget giữ retry không thể làm hơn gấp đôi tải.
- Fail fast bằng circuit breaker và luôn có fallback. Một câu trả lời giảm cấp nhưng nhanh tốt hơn một timeout chậm giữ chặt thread. Gắn mỗi breaker với một fallback đã định (cache, mặc định, tính năng giảm cấp) để lỗi diễn ra êm ái, không trống rỗng.
- Cách ly tài nguyên bằng bulkhead. Cho mỗi dependency (hoặc tenant) một connection/thread pool và giới hạn concurrency riêng để một dependency treo không ngốn hết tài nguyên mà mọi request khác cần.
- Rate-limit ở edge và cả giữa các service nội bộ. Tự bảo vệ mình khỏi client—và bảo vệ dependency khỏi chính bạn. Trả về 429 kèm
Retry-After, và thiết kế client tôn trọng nó. - Tách các workload bùng nổ bằng queue. Queue-based load leveling biến đỉnh tải thành công việc đều đặn theo nhịp consumer; giám sát độ sâu và tuổi message của queue như SLI hạng nhất để backlog tăng dần cảnh báo bạn trước khi người dùng bị ảnh hưởng.
- Làm cho ghi idempotent và dùng outbox pattern cho sự kiện. Idempotency key làm retry an toàn; transactional outbox tránh bài toán dual-write khi bạn phải vừa lưu state vừa publish một sự kiện.
- Test backup bằng cách restore thật, theo lịch. Backup chưa từng được restore không phải là backup, mà là một niềm hy vọng. Tự động hóa việc kiểm chứng restore và tổ chức game day DR đầy đủ ít nhất mỗi năm một lần.
- Ưu tiên strangler fig thay vì viết lại big-bang. Di trú từng bước phía sau một facade routing giữ cho hệ thống luôn ship được và đảo ngược được trong suốt quá trình, và cho phép hủy hoặc tạm dừng với thiệt hại giới hạn.
- Thực hành chaos engineering với blast radius có kiểm soát. Bắt đầu nhỏ ở staging, định nghĩa metric steady-state và điều kiện abort tự động, chạy GameDay do con người trước, rồi đưa thí nghiệm lên production khi đã thành thục.
- Đo bằng SLI/SLO và error budget. Bạn không thể quản lý reliability mà bạn không đo; SLO biến “nó đã đủ tin cậy chưa?” thành một câu hỏi dữ liệu, và error budget cho một cách có nguyên tắc để cân bằng tốc độ với rủi ro.
- Tự động hóa hạ tầng và khôi phục bằng IaC. Các tầng DR rẻ hơn và việc dựng lại nhanh chỉ đáng tin nếu hạ tầng của bạn được viết thành code, đánh version và diễn tập định kỳ—khôi phục kiểu click-ops dưới áp lực sự cố sẽ thất bại.
- Thêm observability cho suy giảm cấp, không chỉ cho sự cố. Alert theo các chỉ báo dẫn trước (latency tăng, tuổi queue tăng, error rate leo thang, breaker trip) để bạn hành động khi hệ thống đang suy giảm, chứ không phải sau khi nó đã sập.
Tài liệu tham khảo
- Azure Architecture Center — Cloud Design Patterns — https://learn.microsoft.com/en-us/azure/architecture/patterns/
- AWS Well-Architected Framework — Reliability Pillar — https://docs.aws.amazon.com/wellarchitected/latest/reliability-pillar/welcome.html
- AWS Disaster Recovery Whitepaper — https://docs.aws.amazon.com/whitepapers/latest/disaster-recovery-workloads-on-aws/disaster-recovery-workloads-on-aws.html
- AWS Builders’ Library — Timeouts, Retries and Backoff with Jitter — https://aws.amazon.com/builders-library/timeouts-retries-and-backoff-with-jitter/
- AWS Builders’ Library — Avoiding fallback in distributed systems — https://aws.amazon.com/builders-library/avoiding-fallback-in-distributed-systems/
- Google SRE Book (đọc miễn phí) — https://sre.google/sre-book/table-of-contents/
- Google SRE Workbook (đọc miễn phí) — https://sre.google/workbook/table-of-contents/
- Martin Fowler — CQRS — https://martinfowler.com/bliki/CQRS.html
- Martin Fowler — Event Sourcing — https://martinfowler.com/eaaDev/EventSourcing.html
- Martin Fowler — Strangler Fig Application — https://martinfowler.com/bliki/StranglerFigApplication.html
- Microservices.io — Transactional Outbox pattern — https://microservices.io/patterns/data/transactional-outbox.html
- Principles of Chaos Engineering — https://principlesofchaos.org/
- Chaos Mesh — https://chaos-mesh.org/
- LitmusChaos — https://litmuschaos.io/
- AWS Fault Injection Service — https://docs.aws.amazon.com/fis/
- roadmap.sh DevOps roadmap — https://roadmap.sh/devops
- Sách: Designing Data-Intensive Applications (Martin Kleppmann, O’Reilly)
- Sách: Release It! ấn bản 2 (Michael Nygard, Pragmatic Bookshelf)
Part of the DevOps Roadmap knowledge base.
Overview
Cloud design patterns are proven, reusable solutions to the recurring problems of building distributed systems: partial failure, variable load, network unreliability, and data consistency at scale. In the cloud, failure is not an exception—it is the normal operating condition. Instances get recycled, availability zones degrade, dependencies throttle you. Systems that thrive are designed for failure rather than designed hoping failure won’t happen.
For DevOps engineers these patterns are the shared vocabulary between development and operations. Terms like RTO/RPO, circuit breaker, bulkhead, and warm standby show up in architecture reviews, incident postmortems, SLO negotiations, and disaster-recovery runbooks. Knowing the pattern catalog (AWS Well-Architected, Azure Architecture Center, Google SRE) lets you evaluate designs quickly: what happens when this dependency slows down? What is the blast radius? How does it scale, and what does recovery cost?
Two ideas underpin everything below. First, the fallacies of distributed computing: the network is not reliable, latency is not zero, bandwidth is not infinite, topology does change. Every pattern here is a defense against one of these false assumptions. Second, failure is probabilistic and compounding: with thousands of components, something is always failing somewhere; a resilient system localizes and absorbs those failures instead of letting them cascade into a user-visible outage.
This topic covers four groups: availability and redundancy, scalability, data management patterns, and design/implementation patterns for resilience—plus disaster recovery strategies and an introduction to chaos engineering, the discipline of verifying that all of the above actually works.
Fundamentals
Availability and redundancy
- High availability (HA): eliminating single points of failure through redundancy. Availability is usually expressed in “nines”. Each extra nine roughly multiplies cost and complexity, so buy nines where they matter and not everywhere.
| Availability | Downtime / year | Downtime / month | Typical use |
|---|---|---|---|
| 99% (“two nines”) | 3.65 days | 7.2 h | Internal tools, best-effort |
| 99.9% (“three nines”) | 8.77 h | 43.8 min | Standard SaaS baseline |
| 99.95% | 4.38 h | 21.9 min | Paid business services |
| 99.99% (“four nines”) | 52.6 min | 4.4 min | Critical platforms |
| 99.999% (“five nines”) | 5.26 min | 26 s | Telco / core infra (very expensive) |
- Redundancy types: active-active (all replicas serve traffic; failure just removes capacity) vs active-passive (a standby takes over on failure, with a failover delay). N+1 means one spare beyond the minimum; 2N means a full duplicate.
- Multi-AZ: run redundant instances across availability zones (independent power/network/cooling within a region, but close enough for low-latency synchronous replication). Protects against data-center failure at low latency cost; the default baseline for production.
- Multi-region: protects against whole-region failure and reduces latency for global users, but introduces data-replication lag, consistency questions, higher latency for synchronous writes, and significant cost. Justify with business requirements, not fashion.
- RTO (Recovery Time Objective): maximum acceptable time to restore service after failure—the “how long can we be down” number.
- RPO (Recovery Point Objective): maximum acceptable data loss, measured as time (e.g., RPO of 5 minutes = lose at most the last 5 minutes of data)—the “how much data can we lose” number. RTO/RPO drive the DR strategy and its cost.
Composite availability: serial dependencies multiply. Two 99.9% services in series give 0.999 × 0.999 ≈ 99.8%—worse than either alone, because a failure in either breaks the chain. Redundant parallel paths improve availability: two 99% paths in parallel give 1 − (0.01 × 0.01) = 99.99%. The practical lesson: reducing the number of hard dependencies is often cheaper than adding nines to each one. Ask of every dependency, “is this a hard dependency (its failure fails my request) or a soft one (I can degrade)?” and convert hard to soft wherever possible.
Scalability
- Vertical scaling (scale up): bigger machine. Simple, no code changes, and sometimes the right first move, but it has a hard ceiling (the largest instance), usually requires downtime to resize, and the big box remains a single point of failure.
- Horizontal scaling (scale out): more machines behind a load balancer. Near-unlimited, enables rolling updates and failure tolerance—but requires the application to cooperate (no local state, no in-memory sessions).
- Stateless design is the prerequisite for scaling out: keep session state in Redis/database, files in object storage, and make any instance able to serve any request. Instances become disposable (“cattle, not pets”) so you can add, remove, and replace them freely.
- Load balancing: distributes requests across instances. Watch the algorithm (round-robin vs least-connections vs consistent hashing) and health checks—an unhealthy instance left in rotation is a slow-burning outage.
- Autoscaling: adjust capacity automatically from metrics (CPU, requests per target, queue depth) or schedules. Kubernetes offers HPA (scale pod count), VPA (right-size pod requests), Cluster Autoscaler/Karpenter (add/remove nodes), and KEDA (event/queue-driven scaling, including scale-to-zero). Reactive autoscaling always lags load, so combine it with a buffer of headroom, and use predictive/scheduled scaling for known spikes (business hours, sales events).
# Kubernetes HPA: scale on CPU, conservative scale-down
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: api
spec:
scaleTargetRef: { apiVersion: apps/v1, kind: Deployment, name: api }
minReplicas: 3
maxReplicas: 30
metrics:
- type: Resource
resource:
name: cpu
target: { type: Utilization, averageUtilization: 65 }
behavior:
scaleDown:
stabilizationWindowSeconds: 300 # wait 5 min before scaling down
The stabilizationWindowSeconds on scale-down is important: it prevents flapping, where a brief dip in load scales you down, the returning load scales you back up, and you thrash—paying in cold starts and instability. Scale up fast, scale down slowly.
Data management patterns
- Cache-aside (lazy loading): the app checks the cache first; on miss, it reads the database and populates the cache with a TTL. Simple and dominant. Watch for two failure modes: stale data (mitigate with sensible TTLs and explicit invalidation on write) and cache stampede (when a hot key expires, thousands of requests hit the DB at once—mitigate with jittered TTLs, request coalescing/single-flight, or refresh-ahead).
- Read-through / write-through / write-behind: alternatives where the cache itself owns DB access. Write-through keeps cache and DB consistent at the cost of write latency; write-behind buffers writes for throughput at the risk of loss on crash.
- Replication: copies of data for read scaling and availability. Synchronous replication protects RPO (the write isn’t acknowledged until replicas have it) but adds write latency and can stall on a slow replica; asynchronous is faster but can lose the tail of un-replicated writes on failover. Leader-follower is the common topology; know your database’s failover behavior and whether it risks split-brain.
- Sharding (partitioning): split data across nodes by a partition key when one node can’t hold the write load or the dataset. Choose keys with even distribution to avoid hot shards; beware that cross-shard queries and resharding (re-splitting as you grow) are the hard parts. Consistent hashing minimizes data movement when nodes are added/removed.
- CQRS (Command Query Responsibility Segregation): separate the write model from the read model(s), letting each scale and be optimized independently (e.g., normalized writes, denormalized read views tuned for query patterns). It adds eventual consistency between the two and operational complexity—use it for read-heavy domains with divergent read/write shapes, not everywhere.
- Event sourcing: store the sequence of events as the source of truth instead of current state; state is derived by replaying events. Pairs naturally with CQRS and gives a full audit trail and time-travel debugging, at the price of significant complexity: event schema versioning, rebuilding projections, and the fact that you can never simply “delete a row”. Use sparingly, in domains where the history is the product (ledgers, audit, workflow).
- Materialized view: precompute and store an expensive query result, refreshed on a schedule or on write, to serve reads cheaply. A pragmatic middle ground before full CQRS.
- Idempotency & the outbox pattern: to make writes safe under retries, give operations idempotency keys; to reliably publish events alongside a DB write, use the transactional outbox (write the event to an outbox table in the same transaction, then relay it) rather than dual-writing to DB and a broker (which can partially fail).
Key Concepts
Resilience patterns
| Pattern | Problem it solves | Key idea |
|---|---|---|
| Retry with backoff + jitter | Transient failures (blips, brief throttling) | Retry a bounded number of times with exponentially growing, randomized delays |
| Circuit breaker | A failing dependency drags callers down | After N failures, fail fast (open); periodically probe (half-open) before closing |
| Bulkhead | One noisy dependency exhausts shared resources | Partition pools (threads, connections, pods) per dependency/tenant to contain damage |
| Throttling / rate limiting | Overload and abusive/runaway clients | Reject or queue excess load early (token bucket, sliding window); return 429 |
| Queue-based load leveling | Spiky producers overwhelm consumers | Put a queue between them; consumers drain at a sustainable rate |
| Strangler fig | Risky big-bang rewrites | Route traffic through a facade; migrate functionality piece by piece to the new system |
| Timeout | A slow dependency ties up resources indefinitely | Cap every wait; free the resource and fail deterministically |
| Fallback / graceful degradation | A dependency is unavailable | Serve a cached/default/reduced result instead of a hard error |
| Health check + load-balancer eviction | A sick instance keeps receiving traffic | Probe liveness/readiness; remove failing instances from rotation |
Retry with exponential backoff and full jitter:
import random, time
def call_with_retry(fn, max_attempts=5, base=0.2, cap=10.0):
for attempt in range(max_attempts):
try:
return fn()
except TransientError:
if attempt == max_attempts - 1:
raise
# full jitter: sleep in [0, min(cap, base * 2^attempt)]
time.sleep(random.uniform(0, min(cap, base * 2 ** attempt)))
Jitter matters: without it, all clients that failed together retry in synchronized waves (“thundering herd”) and re-kill the recovering service. Only retry idempotent operations, honor Retry-After when the server sends it, and combine retries with a retry budget (cap retries as a fraction of total requests, e.g. ≤10%) so retries can’t more than double your load during an incident. The classic anti-pattern is retries at every layer: if three nested services each retry three times, one user request becomes up to 27 backend calls exactly when the system can least afford it.
Circuit breaker states: closed (normal, counting failures) → open (fail fast, no calls, for a cool-down period) → half-open (limited probe calls) → back to closed on success, or back to open on continued failure. The value is twofold: it fails fast (a caller gets an immediate error instead of waiting for a timeout, freeing its threads) and it gives the dependency room to recover (it stops hammering a struggling service). Combine it with fallbacks so failing fast still yields a usable answer—a cached response, a default value, or a degraded feature—rather than a blank error.
Bulkhead takes its name from ship compartments: isolate resource pools so a flood in one compartment doesn’t sink the ship. In practice: give each downstream dependency (or each tenant) its own connection pool / thread pool / concurrency limit, so when dependency X hangs, only the X-pool is exhausted and calls to Y still succeed. Without bulkheads, one slow dependency can consume every thread in a shared pool and take the whole service down.
Queue-based load leveling decouples producers from consumers with a durable queue. Spiky, bursty ingress becomes a steady, consumer-paced workload; the queue absorbs the spike so the consumer runs at its sustainable rate. Monitor queue depth and message age as first-class SLIs—a growing backlog is your earliest warning that consumers can’t keep up, well before users notice.
Throttling / rate limiting protects a service from overload and from abusive or buggy clients. Common algorithms: token bucket (allows bursts up to a bucket size, refills at a steady rate) and sliding window (smooths the rate over a moving interval). Return HTTP 429 with a Retry-After header so well-behaved clients back off. Rate-limit both at the edge (protect yourself from the world) and between internal services (protect your dependencies from you).
Strangler fig (named after the vine that grows around a tree and gradually replaces it) migrates a legacy system incrementally: put a routing facade in front, then peel off one capability at a time to a new implementation, shifting traffic as each piece is proven. The system stays shippable and reversible throughout—contrast the big-bang rewrite, which is unshippable for months and risks a catastrophic single cutover.
Disaster recovery strategies
| Strategy | RTO | RPO | Relative cost | How it works |
|---|---|---|---|---|
| Backup & restore | Hours–days | Hours | $ | Back up data/config; rebuild infrastructure from IaC on demand |
| Pilot light | Tens of minutes–hours | Minutes | $$ | Core data replicated continuously; minimal services running; scale up on disaster |
| Warm standby | Minutes | Seconds–minutes | $$$ | Scaled-down but fully functional copy always running; scale to full on failover |
| Multi-site active-active | Near zero | Near zero | $$$$ | Full capacity in multiple regions serving simultaneously; failover = shift traffic |
How to think about the ladder: each rung buys you lower RTO/RPO for more money and more operational complexity. Backup & restore keeps only data warm and rebuilds everything else—cheap, but slow and dependent on your IaC actually working. Pilot light keeps the “hard to replicate” part (the database, replicated continuously) always on, and everything else off until needed. Warm standby runs a small but complete copy of the whole stack that you scale up on failover—no cold-start risk, just capacity. Active-active runs full capacity in every region simultaneously, so “failover” is just removing a region from the traffic pool; it also doubles as a scaling and latency strategy, but forces you to solve multi-region data consistency and conflict resolution.
Choosing: derive RTO/RPO per workload from business impact, then buy the cheapest strategy that meets them—do not gold-plate a reporting dashboard to active-active. Crucially, IaC (Terraform/CloudFormation/Pulumi) and automated, regularly tested runbooks are what make the cheaper tiers credible: an untested DR plan is a hypothesis, not a plan. Also decide failover vs failback (returning to the primary once it recovers) explicitly—many teams rehearse failover and forget failback, then struggle to return.
Chaos engineering (intro)
Chaos engineering is the practice of running controlled experiments that inject real failures (kill instances, add latency, block a dependency, fail an AZ, exhaust a disk) to verify the system behaves as designed—before an incident proves it doesn’t. Method:
- Define “steady state” via measurable output (success rate, latency SLI, orders/minute)—an external, business-relevant signal, not an internal metric.
- Hypothesize: “steady state persists if we kill 20% of pods” / “if the payments DB gets +200 ms latency”.
- Inject the failure in a limited blast radius: start in staging, then production with a small percentage and an automatic abort (“stop button”) tied to your steady-state metric.
- Measure, compare with the hypothesis, fix the weaknesses you find, then automate the experiment so it runs continuously and catches regressions.
Tools: Chaos Mesh and LitmusChaos (Kubernetes-native, CNCF), Gremlin (commercial, with safety guardrails), AWS Fault Injection Service, and Netflix’s Chaos Monkey / Simian Army (the origin of the discipline). Start with GameDays—scheduled, human-run failure exercises—before fully automating. The cultural point is the whole reason the discipline exists: reliability mechanisms you never exercise (failover, circuit breakers, DR runbooks, autoscaling) will not work when you finally need them. Chaos engineering turns “we think it’s resilient” into “we’ve verified it’s resilient”.
Best Practices
- Design for failure at every layer. Assume instances, AZs, and dependencies will fail; make failure handling explicit in design reviews by always asking “what happens when X is slow, or down, or returns garbage?”
- Set RTO/RPO per workload before choosing architecture. DR spend without stated objectives is either waste (over-provisioning a low-value service) or false confidence (under-protecting a critical one). Let the business impact set the number.
- Default to multi-AZ; justify multi-region. Multi-AZ is cheap insurance and should be your production baseline; multi-region is a major program with real data-consistency costs—adopt it only for genuine business or latency requirements.
- Reduce hard dependencies; convert hard to soft. Every synchronous dependency in your request path lowers composite availability. Ask whether you can degrade gracefully (cache, default, async) instead of failing when a dependency is down.
- Keep services stateless; externalize state. Sessions in Redis, files in object storage, state in databases—this is what unlocks autoscaling, rolling deploys, and painless instance replacement.
- Autoscale on the metric that reflects load, and cap it. Requests-per-target or queue depth usually beat CPU as a load signal; set a
maxReplicas, scale up quickly, and scale down slowly (stabilization window) to avoid flapping and runaway cost. - Every network call gets a timeout, bounded retries with jittered backoff, and idempotency. Missing timeouts and unbounded retries are the top two amplifiers that turn a brief blip into a full outage; a retry budget keeps retries from more than doubling load.
- Fail fast with circuit breakers and always provide a fallback. A fast degraded answer beats a slow timeout that ties up threads. Pair every breaker with a defined fallback (cache, default, reduced feature) so the failure is graceful, not blank.
- Isolate resources with bulkheads. Give each dependency (or tenant) its own connection/thread pool and concurrency limit so one hung dependency can’t exhaust the resources every other request needs.
- Rate-limit at the edge and between internal services. Protect yourself from clients—and your dependencies from you. Return 429 with
Retry-After, and design clients to honor it. - Decouple spiky workloads with queues. Queue-based load leveling turns traffic spikes into steady, consumer-paced work; monitor queue depth and message age as first-class SLIs so a growing backlog alerts you before users are affected.
- Make writes idempotent and use the outbox pattern for events. Idempotency keys make retries safe; a transactional outbox avoids the dual-write problem when you must both persist state and publish an event.
- Test backups by restoring them, on a schedule. A backup that has never been restored is not a backup, it’s a hope. Automate restore verification and include full DR game days at least annually.
- Prefer strangler fig over big-bang rewrites. Incremental migration behind a routing facade keeps the system shippable and reversible throughout, and lets you abort or pause with limited damage.
- Practice chaos engineering with a controlled blast radius. Start small in staging, define steady-state metrics and automatic abort conditions, run human GameDays first, and graduate experiments to production once they’re routine.
- Instrument with SLIs/SLOs and error budgets. You cannot manage reliability you don’t measure; SLOs turn “is it reliable enough?” into a data question, and error budgets give a principled way to balance velocity against risk.
- Automate infrastructure and recovery with IaC. Cheaper DR tiers and fast rebuilds are only credible if your infrastructure is codified, versioned, and regularly exercised—click-ops recovery under incident pressure fails.
- Add observability for degradation, not just outages. Alert on the leading indicators (rising latency, growing queue age, climbing error rate, breaker trips) so you act while the system is degrading, not after it’s down.
References
- Azure Architecture Center — Cloud Design Patterns — https://learn.microsoft.com/en-us/azure/architecture/patterns/
- AWS Well-Architected Framework — Reliability Pillar — https://docs.aws.amazon.com/wellarchitected/latest/reliability-pillar/welcome.html
- AWS Disaster Recovery Whitepaper (workloads on AWS) — https://docs.aws.amazon.com/whitepapers/latest/disaster-recovery-workloads-on-aws/disaster-recovery-workloads-on-aws.html
- AWS Builders’ Library — Timeouts, Retries and Backoff with Jitter — https://aws.amazon.com/builders-library/timeouts-retries-and-backoff-with-jitter/
- AWS Builders’ Library — Avoiding fallback in distributed systems — https://aws.amazon.com/builders-library/avoiding-fallback-in-distributed-systems/
- Google SRE Book (free online) — https://sre.google/sre-book/table-of-contents/
- Google SRE Workbook (free online) — https://sre.google/workbook/table-of-contents/
- Martin Fowler — CQRS — https://martinfowler.com/bliki/CQRS.html
- Martin Fowler — Event Sourcing — https://martinfowler.com/eaaDev/EventSourcing.html
- Martin Fowler — Strangler Fig Application — https://martinfowler.com/bliki/StranglerFigApplication.html
- Microservices.io — Transactional Outbox pattern — https://microservices.io/patterns/data/transactional-outbox.html
- Principles of Chaos Engineering — https://principlesofchaos.org/
- Chaos Mesh — https://chaos-mesh.org/
- LitmusChaos — https://litmuschaos.io/
- AWS Fault Injection Service — https://docs.aws.amazon.com/fis/
- roadmap.sh DevOps roadmap — https://roadmap.sh/devops
- Book: Designing Data-Intensive Applications (Martin Kleppmann, O’Reilly)
- Book: Release It! 2nd ed. (Michael Nygard, Pragmatic Bookshelf)