Nhà cung cấp Cloud (Cloud Providers)Cloud Providers
Thuộc bộ tài liệu DevOps Roadmap.
Tổng quan
Cloud provider cung cấp tài nguyên tính toán — server, storage, networking, database và các managed service cấp cao hơn — qua internet theo mô hình trả tiền theo mức sử dụng (pay-as-you-go). Thay vì mua và lắp đặt phần cứng vật lý, team có thể provision hạ tầng trong vài phút qua API hoặc web console. Ba nhà cung cấp lớn nhất là Amazon Web Services (AWS), Microsoft Azure và Google Cloud Platform (GCP); ngoài ra còn có các đơn vị nhỏ hơn như DigitalOcean, Hetzner, Alibaba Cloud phục vụ các ngách riêng (đơn giản, giá rẻ, hoặc thị trường khu vực).
Với DevOps engineer, cloud là môi trường triển khai mặc định của hầu hết hệ thống hiện đại. Gần như mọi thực hành DevOps — Infrastructure as Code, autoscaling, immutable deployment, managed Kubernetes, serverless — đều xây dựng trên các dịch vụ nền tảng của cloud. Hiểu cách các provider tổ chức dịch vụ, cách tính tiền, và ranh giới trách nhiệm giữa provider và bạn là kiến thức bắt buộc.
Vì ba ông lớn cung cấp các dịch vụ cốt lõi gần như tương đương nhau dưới những cái tên khác nhau, kỹ năng thực tế không phải là thuộc lòng catalog của một provider, mà là nắm vững khái niệm nền (compute, block/object storage, VPC networking, IAM) để có thể “ánh xạ” giữa các cloud. Một mô hình tư duy hữu ích: mọi cloud đều là một kho chứa cùng những primitive giống nhau — một cỗ máy chạy code, một disk, một bucket, một mạng riêng, một hệ thống identity, một queue, một database — chỉ khác thương hiệu. Học primitive một lần, còn tên trên console chỉ là bài toán dịch thuật.
Sự thay đổi về kinh tế mà cloud mang lại
Thay đổi sâu sắc nhất cloud đem lại là về tài chính, không phải kỹ thuật. Phần cứng on-premises là chi phí đầu tư (CapEx): bạn trả một khoản lớn từ đầu, khấu hao nhiều năm, và trả cho công suất đỉnh ngay cả khi không dùng. Cloud là chi phí vận hành (OpEx): trả cho những gì bạn dùng, khi bạn dùng, và có thể tắt đi. Điều này biến hạ tầng thành chi phí biến đổi co giãn theo quy mô kinh doanh — mạnh mẽ khi tăng trưởng, nguy hiểm khi buông lỏng, vì chính sự co giãn cho phép bạn scale lên một triệu user cũng cho phép một vòng lặp cấu hình sai đẩy hóa đơn lên năm chữ số chỉ sau một đêm. Cloud financial management (FinOps) ra đời chính vì công tơ không bao giờ ngừng chạy.
Kiến thức nền tảng
Mô hình dịch vụ: IaaS vs PaaS vs SaaS
| Mô hình | Bạn quản lý | Provider quản lý | Ví dụ |
|---|---|---|---|
| IaaS (Infrastructure as a Service) | OS, runtime, app, dữ liệu | Phần cứng, ảo hóa, network | EC2, Azure VMs, Compute Engine |
| PaaS (Platform as a Service) | App, dữ liệu | OS, runtime, scaling, vá lỗi | Elastic Beanstalk, App Service, Cloud Run |
| SaaS (Software as a Service) | Cấu hình, cách dùng | Toàn bộ phần còn lại | Gmail, Office 365, Datadog |
Càng xuống thấp (IaaS) bạn càng có nhiều quyền kiểm soát nhưng gánh nặng vận hành càng lớn. Càng lên cao (PaaS/SaaS) bạn triển khai càng nhanh nhưng phải chấp nhận ràng buộc của provider.
Một nhóm thứ tư đáng gọi tên là CaaS (Containers as a Service) và FaaS (Functions as a Service), nằm giữa IaaS và PaaS: managed Kubernetes (EKS/AKS/GKE) và serverless function (Lambda/Cloud Functions) cho phép bạn ship container hoặc code mà không phải quản lý VM bên dưới. Ranh giới khá mờ — Cloud Run vừa là “serverless container”, vừa là CaaS, vừa là PaaS tùy góc nhìn. Đừng sa đà vào phân loại; hãy dùng nó để đặt câu hỏi thật sự: tôi muốn chịu trách nhiệm bao nhiêu phần của stack này lúc 3 giờ sáng khi có sự cố?
“Managed” là một dải liên tục, không phải công tắc bật/tắt. RDS quản lý database engine, backup và failover — nhưng bạn vẫn chọn kích thước instance, tinh chỉnh tham số và thiết kế schema. Một database serverless hoàn toàn như DynamoDB hay Aurora Serverless loại bỏ luôn cả quyết định sizing. Mỗi nấc leo lên là đánh đổi một cái núm bạn có thể vặn để bớt đi một thứ có thể hỏng.
Region và Availability Zone
- Region là một khu vực địa lý (ví dụ
us-east-1,westeurope,asia-southeast1) chứa nhiều data center. - Availability Zone (AZ) là một hoặc nhiều data center độc lập trong region, có nguồn điện, làm mát và network riêng. Các AZ trong cùng region đủ gần để replication đồng bộ với latency thấp (vài mili-giây) nhưng đủ xa để cháy nổ, ngập lụt hay mất điện ở một AZ không kéo sập AZ khác.
- Triển khai trên nhiều AZ giúp chống lỗi cấp data center; triển khai nhiều region giúp chống sự cố toàn region và giảm latency cho người dùng toàn cầu.
- Traffic trong một AZ thường miễn phí; traffic giữa các AZ và giữa các region thì tính tiền — đây là khoản chi phí hay bị bỏ quên.
Một số khái niệm địa lý khác quan trọng trong thực tế:
- Edge location / PoP (Point of Presence) — hàng trăm điểm nhỏ dùng bởi CDN (CloudFront, Azure Front Door, Cloud CDN) và DNS để phục vụ nội dung cache gần người dùng. Chúng không phải region đầy đủ; bạn không chạy database ở đó được.
- Local Zones / Wavelength (AWS), Edge Zones (Azure) — phần mở rộng của region đặt tại một đô thị hoặc mạng viễn thông cho use case cần latency cực thấp (game, live media, 5G).
- Region pair (Azure) — Azure ghép cặp các region để replication và bảo trì so le; một số service tự động replicate sang region ghép cặp.
- Tiêu chí chọn region — cân nhắc trên bốn trục: (1) latency tới người dùng, (2) data residency / chủ quyền dữ liệu (GDPR, quy định bảo vệ dữ liệu nội địa), (3) độ sẵn có của service — không phải service hay instance type nào cũng có ở mọi region, và region mới thường bị chậm cập nhật, và (4) giá — cùng một VM có thể đắt hơn 20-50% ở region này so với region khác (
us-east-1vàus-west-2thường thuộc nhóm rẻ nhất của AWS).
Mô hình tư duy về blast radius: một instance có thể chết (deploy ≥2); một AZ có thể chết (trải trên nhiều AZ); một region có thể chết (multi-region cho hệ thống tier-0); và hiếm khi, một control plane toàn cầu có thể chết (ví dụ sự cố IAM hay DNS global kéo theo nhiều region — lý do vì sao dự phòng toàn cầu thật sự đôi khi nghĩa là multi-cloud, không chỉ multi-region).
Shared responsibility model (mô hình trách nhiệm chia sẻ)
Provider bảo mật hạ tầng cloud; bạn bảo mật những gì bạn đặt lên đó.
| Trách nhiệm của provider | Trách nhiệm của bạn |
|---|---|
| An ninh vật lý, phần cứng | IAM user, role, policy |
| Hypervisor và host OS | Vá guest OS (với IaaS) |
| Độ sẵn sàng của managed service | Mã hóa dữ liệu, quản lý key |
| Hạ tầng mạng | Security group, firewall rule, phân đoạn mạng |
| Control plane toàn cầu | Code ứng dụng và secrets |
Phần lớn sự cố bảo mật trên cloud đến từ cấu hình sai phía khách hàng (S3 bucket public, IAM role quá quyền), không phải lỗi của provider.
Ranh giới này dịch chuyển theo mô hình dịch vụ. Đây là điểm tinh tế quan trọng nhất: service càng được quản lý nhiều thì provider gánh càng nhiều, nhưng dữ liệu, identity và cấu hình quyền truy cập luôn là của bạn.
- IaaS (EC2): bạn vá OS, cấu hình firewall, mã hóa disk, quản lý app. Provider chỉ đảm bảo phần cứng và hypervisor.
- PaaS (App Service, RDS): provider vá OS và runtime; bạn vẫn sở hữu việc phân loại dữ liệu, kiểm soát truy cập và bảo mật tầng ứng dụng.
- SaaS (Office 365): provider chạy gần như mọi thứ; bạn vẫn quyết định ai có quyền truy cập, identity được quản trị ra sao, và đưa dữ liệu gì vào.
Một câu tóm tắt sống sót qua mọi mô hình dịch vụ: “Của, Trong, và Quyền truy cập.” Provider bảo mật an ninh của cloud; bạn bảo mật những gì đặt trong cloud và ai có quyền truy cập vào đó. Cấu hình sai — bucket public, key không rotate, role quá rộng — luôn nằm ở phía khách hàng.
Mô hình tính giá
| Mô hình | Cách hoạt động | Phù hợp với |
|---|---|---|
| On-demand | Trả theo giây/giờ, không cam kết | Workload khó dự đoán, thử nghiệm |
| Reserved / Savings Plans / CUD | Cam kết 1–3 năm, giảm 30–70% | Tải nền ổn định (database, core service) |
| Spot / Preemptible | Dùng capacity dư, rẻ tới 90%, có thể bị thu hồi đột ngột | Batch job, CI runner, workload stateless chịu được gián đoạn |
| Free tier | Mức dùng miễn phí cho tài khoản mới hoặc vĩnh viễn | Học tập, prototype |
So sánh các lựa chọn cam kết:
| Lựa chọn | Tên theo provider | Mức giảm | Độ linh hoạt |
|---|---|---|---|
| Standard Reserved Instances | AWS RI | Tới ~72% | Khóa theo instance family/region |
| Convertible RIs | AWS | Tới ~54% | Đổi được family, cam kết dài hơn |
| Savings Plans (Compute) | AWS | Tới ~66% | Áp cho EC2/Fargate/Lambda theo mức chi $/giờ |
| Reserved VM Instances | Azure | Tới ~72% | Theo VM series/region, đổi được |
| Committed Use Discounts | GCP CUD | Tới ~57% (resource) / ~70% (flexible) | Theo mức chi hoặc theo resource |
| Sustained Use Discounts | GCP | Tự động, tới ~30% | Không cam kết — tự áp cho VM chạy lâu |
Cơ chế spot/preemptible rất quan trọng. Capacity spot có thể bị thu hồi với cảnh báo ngắn (AWS báo trước 2 phút; GCP Spot VM báo 30 giây; Azure Spot tùy trường hợp). Hãy thiết kế để chịu được điều đó: checkpoint công việc, drain nhẹ nhàng, trải trên nhiều instance type và AZ để một đợt khan hiếm capacity không thu hồi tất cả cùng lúc, và không bao giờ đặt database primary có state lên spot. Spot lý tưởng cho CI runner, batch/ETL, rendering, ML training có checkpoint, và Kubernetes worker node sau một cluster autoscaler.
Egress là cái bẫy giá làm mọi người bất ngờ. Ingress (dữ liệu vào cloud) gần như luôn miễn phí. Egress (dữ liệu ra internet, và thường cả giữa region, giữa AZ) bị tính tiền và có thể chiếm phần lớn hóa đơn với workload nặng dữ liệu hay media. Mô hình tư duy: check-in thì rẻ, check-out thì đắt. Đây cũng là đòn bẩy lock-in có chủ đích — chuyển hàng terabyte sang cloud khác tốn tiền thật — nên các nhà quản lý (và EU Data Act) đã thúc ép provider miễn phí egress cho khách hàng rời đi.
Khái niệm chính
Bảng ánh xạ dịch vụ cốt lõi: AWS vs Azure vs GCP
| Nhóm | AWS | Azure | GCP |
|---|---|---|---|
| Máy ảo | EC2 | Virtual Machines | Compute Engine |
| Autoscaling group | Auto Scaling Group | VM Scale Sets | Managed Instance Groups |
| Object storage | S3 | Blob Storage | Cloud Storage |
| Block storage | EBS | Managed Disks | Persistent Disk / Hyperdisk |
| File storage | EFS / FSx | Azure Files | Filestore |
| Mạng ảo | VPC | VNet | VPC |
| Load balancer | ELB/ALB/NLB | Load Balancer / App Gateway | Cloud Load Balancing |
| CDN | CloudFront | Front Door / CDN | Cloud CDN |
| DNS | Route 53 | Azure DNS | Cloud DNS |
| IAM | IAM | Entra ID (Azure AD) + RBAC | Cloud IAM |
| SQL managed | RDS / Aurora | Azure SQL / Database for PostgreSQL | Cloud SQL / AlloyDB |
| NoSQL | DynamoDB | Cosmos DB | Firestore / Bigtable |
| Data warehouse | Redshift | Synapse / Fabric | BigQuery |
| Managed Kubernetes | EKS | AKS | GKE |
| Serverless container | Fargate / App Runner | Container Apps | Cloud Run |
| Serverless function | Lambda | Azure Functions | Cloud Functions / Cloud Run functions |
| Container registry | ECR | ACR | Artifact Registry |
| Secrets | Secrets Manager | Key Vault | Secret Manager |
| Key management | KMS | Key Vault | Cloud KMS |
| Monitoring | CloudWatch | Azure Monitor | Cloud Monitoring (Ops Suite) |
| Message queue | SQS / SNS | Service Bus / Event Grid | Pub/Sub |
| Streaming managed | Kinesis / MSK | Event Hubs | Pub/Sub / Managed Kafka |
| IaC (native) | CloudFormation / CDK | ARM / Bicep | Deployment Manager / Config Connector |
| Redis/cache managed | ElastiCache | Azure Cache for Redis | Memorystore |
Hãy coi bảng này là “phiến đá Rosetta”, không phải chân lý — service chồng lấn và hay bị đổi tên. Giá trị nằm ở chỗ khi ai đó nói “chúng tôi dùng Pub/Sub”, bạn lập tức biết vai trò của nó (vùng đất của SQS/SNS + Kinesis) dù chưa từng chạm vào GCP.
Các lựa chọn compute trong cùng một provider
Ngay trong một cloud, “compute” cũng là một cái thang từ kiểm-soát-nhiều-nhất tới quản-lý-ít-nhất, và chọn sai nấc là lỗi phổ biến:
- Bare metal / dedicated host — cô lập vật lý hoàn toàn cho licensing hoặc compliance; hiếm gặp.
- Virtual machine (EC2, Azure VMs, Compute Engine) — đơn vị IaaS cổ điển; bạn sở hữu OS.
- Managed Kubernetes (EKS/AKS/GKE) — container được điều phối; bạn sở hữu workload và (thường cả) worker node.
- Serverless container (Fargate, Cloud Run, Container Apps) — mang container image tới, quên node đi.
- Function (Lambda, Cloud Functions, Azure Functions) — mang function tới, quên container đi.
- PaaS app hosting (Elastic Beanstalk, App Service, App Engine) — mang code tới, quên gần như mọi thứ.
Quy tắc chung: bắt đầu ở nấc cao nhất mà yêu cầu cho phép, chỉ tụt xuống khi có một ràng buộc cụ thể (latency, chi phí ở quy mô lớn, phần cứng đặc biệt, licensing) buộc bạn. Nhiều team với tay tới Kubernetes trong khi Cloud Run là đủ, để rồi trả giá cho độ phức tạp của cluster mà họ không cần.
Instance family và sizing
VM có các family tối ưu cho từng loại workload — tên khác nhau nhưng nhóm thì phổ quát:
| Workload | AWS family | Azure series | GCP family |
|---|---|---|---|
| General purpose | M, T (burstable) | D, B (burstable) | E2, N2 |
| Compute optimized | C | F | C2 |
| Memory optimized | R, X | E, M | M, N2-highmem |
| Storage optimized | I, D | L | — |
| GPU / accelerated | P, G, Inf/Trn | N-series | A (GPU), TPU VMs |
Instance burstable (T-series, B-series) tích lũy CPU credit khi rảnh và tiêu chúng khi có đột biến — rẻ cho workload bùng phát nhẹ, trung bình thấp như web server nhỏ, nhưng bị throttle mạnh nếu chạy CPU cao liên tục và cạn credit. Đừng bao giờ đặt một service bận liên tục lên instance burstable rồi thắc mắc sao nó chậm.
Right-sizing là kỷ luật liên tục, không phải lựa chọn một lần: đo mức dùng CPU/memory/IO thật và thu nhỏ instance dư thừa. Khoản lãng phí phổ biến nhất trong mọi hóa đơn cloud là VM quá cỡ, dùng không hết, mua “cho chắc”.
Các provider nhỏ đáng biết
- DigitalOcean — giao diện đơn giản, giá phẳng dễ đoán, Droplet (VM), managed Kubernetes và database, App Platform (PaaS), Spaces (object storage tương thích S3). Rất hợp với startup, side project và lập trình viên muốn tránh độ phức tạp của ba ông lớn. Giá egress hào phóng hơn nhiều so với hyperscaler.
- Hetzner — server dedicated và cloud giá cực cạnh tranh tại châu Âu (và nay cả Mỹ); phổ biến cho self-hosted Kubernetes, production quy mô homelab và workload nhạy chi phí. Bạn đánh đổi catalog managed service khổng lồ của hyperscaler lấy giá mỗi core và mỗi TB thấp hơn hẳn.
- Alibaba Cloud — provider thống trị tại Trung Quốc và phần lớn Đông Nam Á; gần như bắt buộc nếu phục vụ người dùng đại lục (giấy phép ICP, region nội địa, tích hợp hệ sinh thái thanh toán và CDN địa phương).
- Oracle Cloud (OCI) — free tier rộng rãi và egress cạnh tranh; mạnh cho workload Oracle Database.
- Cloudflare — không phải IaaS đầy đủ, nhưng nền tảng edge (Workers, R2 object storage với egress bằng 0, D1, KV) ngày càng là lựa chọn serverless-first cho đúng loại workload.
Mô thức chung: hyperscaler thắng ở bề rộng managed service và compliance doanh nghiệp; các đơn vị chuyên biệt thắng ở giá, sự đơn giản, hoặc một region/hệ sinh thái cụ thể. Kiến trúc phổ biến là “hyperscaler cho managed database và compliance, đơn vị chuyên biệt cho CDN hoặc object storage nơi egress chi phối”.
Nếm thử CLI
# AWS: tạo VM
aws ec2 run-instances --image-id ami-0abcdef1234567890 \
--instance-type t3.micro --count 1
# Azure: tạo VM
az vm create --resource-group myRG --name myVM \
--image Ubuntu2204 --size Standard_B1s
# GCP: tạo VM
gcloud compute instances create my-vm \
--machine-type=e2-micro --zone=asia-southeast1-a
Các dòng lệnh này chỉ để học. Thực tế bạn không bao giờ tạo hạ tầng production theo kiểu mệnh lệnh (imperative) — bạn khai báo nó trong Terraform/OpenTofu hoặc IaC native của provider để review được, quản lý version và tái tạo được. CLI dùng để kiểm tra, sửa khẩn cấp (break-glass) và scripting, không phải để dựng stack bằng tay.
Cấu trúc account và organization
Mọi provider đều có một phân cấp để cô lập môi trường và áp policy tập trung:
| Khái niệm | AWS | Azure | GCP |
|---|---|---|---|
| Container gốc | Organization | Tenant (Entra ID) | Organization |
| Nhóm | Organizational Units (OUs) | Management Groups | Folders |
| Biên cô lập | Account | Subscription | Project |
| Guardrail | Service Control Policies (SCPs) | Azure Policy | Organization Policies |
Mô thức chuẩn là landing zone: tách account/subscription/project theo môi trường (dev, staging, prod) và thường theo team, với một account “management” trung tâm cho billing và logging, một account “security” cho audit trail. Nó mang lại cô lập blast-radius (sai ở dev không chạm được prod), phân bổ chi phí sạch, và một nơi để cưỡng chế guardrail toàn org. AWS Control Tower, Azure Landing Zones và blueprint landing zone của GCP tự động hóa việc thiết lập.
Cách chọn provider
- Hệ sinh thái sẵn có — công ty dùng nhiều Microsoft (AD, Office 365)? Azure tích hợp tự nhiên và thường đi kèm trong hợp đồng doanh nghiệp. Thiên về data/ML? BigQuery và Vertex AI của GCP rất mạnh. Cần catalog rộng nhất, thị phần lớn nhất (nguồn nhân lực dồi dào nhất)? AWS.
- Kỹ năng của team — kinh nghiệm sẵn có và khả năng tuyển dụng thường quan trọng hơn khác biệt kỹ thuật; một team thành thạo cloud này ship nhanh hơn chính team đó vật lộn với cloud “tốt hơn” nhưng lạ lẫm.
- Region khả dụng — kiểm tra latency và yêu cầu data residency; xác nhận đúng service và instance type bạn cần thực sự tồn tại ở region đích. Với Việt Nam thường dùng region Singapore (
ap-southeast-1,asia-southeast1). - Chi phí cho workload của bạn — mô phỏng usage thực tế; phí egress và giá managed service chênh lệch đáng kể. Dùng pricing calculator của từng provider với số liệu thực, không phải giá niêm yết.
- Compliance — chứng chỉ bắt buộc (PCI DSS, HIPAA, SOC 2, FedRAMP, ISO 27001, quy định địa phương) theo từng region và từng service.
- Managed service bạn thật sự dựa vào — cloud “tốt nhất” thường là cloud có đúng managed service gỡ bỏ nhiều toil nhất cho workload của bạn (BigQuery cho analytics, DynamoDB cho key-value quy mô lớn, Cosmos DB cho multi-model toàn cầu).
Cân nhắc multi-cloud
- Ưu điểm: giảm đòn bẩy vendor lock-in, đáp ứng yêu cầu pháp lý/data residency, chọn được dịch vụ tốt nhất từng mảng (ví dụ BigQuery trên GCP nhưng mọi thứ khác trên AWS), và cho khả năng chịu lỗi thật sự trước sự cố sập cả một provider với hệ thống tier-0.
- Nhược điểm: nhân đôi bề mặt vận hành (hai mô hình IAM, hai stack network, hai hệ thống billing, hai bộ runbook on-call), phức tạp hóa thế trận bảo mật, làm loãng chiều sâu chuyên môn của team, và egress giữa các cloud rất đắt và chậm.
- Portability ≠ multi-cloud. Giữ workload có tính di động (container, Kubernetes, Terraform, data store open-source) là bảo hiểm rẻ và đáng làm. Còn chạy cùng một workload live trên nhiều cloud thì đắt và hiếm khi đáng với team nhỏ.
- Quan điểm thực dụng: phần lớn team nên đi sâu vào một provider và giữ portability như một lớp phòng hờ, thay vì dàn trải công sức trên hai. Hãy để dành active/active multi-cloud thật sự cho những tổ chức đủ quy mô, nhân sự và động lực pháp lý. Multi-cloud “để dự phòng” thường là giải pháp đi tìm vấn đề với team dưới vài trăm kỹ sư.
- Hybrid cloud (on-prem + cloud) là một mô thức riêng, phổ biến ở doanh nghiệp đã có data center, chi phí phần cứng chìm, hoặc dữ liệu về mặt pháp lý không được rời khỏi cơ sở; các công cụ như AWS Outposts, Azure Arc/Stack và Anthos mở rộng control plane cloud xuống on-prem.
Best Practices
- Dùng Infrastructure as Code ngay từ đầu (Terraform/OpenTofu, Pulumi, CloudFormation, Bicep) — click tay trên console tạo ra hạ tầng “snowflake” không tài liệu, không tái tạo được, không review được mà không ai dựng lại nổi sau sự cố.
- Bật MFA và không bao giờ dùng tài khoản root/owner cho công việc hằng ngày — tạo IAM user/role có phạm vi giới hạn; khóa kỹ tài khoản root/global-admin sau MFA và chỉ dùng cho các thao tác break-glass thật sự cần.
- Áp dụng least privilege cho IAM — bắt đầu với quyền hẹp rồi mở dần; ưu tiên role và credential ngắn hạn thay vì access key sống lâu; audit bằng AWS IAM Access Analyzer, Azure PIM, GCP Policy Analyzer.
- Đặt billing alert và budget ngay lập tức — hóa đơn “bất ngờ” là nỗi đau phổ biến nhất của người mới; cảnh báo ở mức 50/80/100% chi tiêu dự kiến, và bật anomaly detection để một vòng lặp chạy loạn báo động trong vài giờ, không phải cuối tháng.
- Triển khai production trên ít nhất hai AZ — chạy một AZ nghĩa là bảo trì data center định kỳ trở thành sự cố của bạn; multi-AZ là lớp chịu lỗi rẻ nhất bạn từng mua.
- Gắn tag cho mọi resource (team, environment, service, cost-center, owner) — tag phục vụ phân bổ chi phí, tự động dọn dẹp, truy vết ownership và policy bảo mật; cưỡng chế tagging bằng policy-as-code để resource thiếu tag bị chặn hoặc gắn cờ.
- Ưu tiên managed service cho các việc “nặng mà không tạo khác biệt” — tự vận hành PostgreSQL, Kafka hay Redis trên VM hiếm khi hơn bản managed khi tính cả backup, vá lỗi, failover và gánh nặng on-call. Chỉ tự quản lý khi có lý do chi phí hoặc kiểm soát cụ thể.
- Để ý chi phí egress và cross-AZ — dữ liệu vào cloud miễn phí, ra thì không; giữ traffic trao đổi nhiều trong cùng AZ/region, dùng private endpoint thay vì đi qua internet công cộng, và đặt CDN trước nội dung tĩnh.
- Dùng spot/preemptible cho workload chịu được gián đoạn — tiết kiệm 60–90% cho CI, batch, ETL và service stateless chịu lỗi; thiết kế để checkpoint và drain nhẹ nhàng khi có cảnh báo thu hồi.
- Chỉ mua reservation/savings plan sau khi usage đã ổn định — đo 2–3 tháng tải ổn định trước, rồi cam kết cho phần tải nền; để phần đỉnh đột biến trên on-demand hoặc spot.
- Right-size liên tục — khoản lớn nhất trong hầu hết hóa đơn là compute quá cỡ, dùng không hết, mua “cho chắc”; xem lại mức dùng và thu nhỏ hằng tháng.
- Tổ chức landing zone rõ ràng — tách account/project/subscription theo môi trường để cô lập blast radius, phân bổ chi phí sạch và cưỡng chế guardrail tập trung (SCP/Azure Policy/Org Policy).
- Mã hóa mặc định và quản lý key có chủ đích — bật mã hóa at-rest và in-transit ở mọi nơi (thường miễn phí và mặc định bật); dùng KMS của provider, và dùng customer-managed key khi compliance yêu cầu kiểm soát vòng đời key.
- Tập trung logging và audit trail — bật CloudTrail / Azure Activity Log / Cloud Audit Logs trong một account chuyên dụng, chỉ-ghi để có bản ghi bất biến về ai làm gì lúc nào — không thể thiếu khi có sự cố và khi audit.
- Tự động dọn dẹp resource tạm — môi trường dev/test, volume mồ côi, IP không gắn, snapshot cũ âm thầm ngốn tiền; lên lịch teardown và chạy một “cost janitor” (cloud-nuke, gợi ý của provider) định kỳ.
- Học khái niệm, đừng học console — VPC, IAM, object storage, load balancing và identity federation dùng được trên mọi provider; bố cục console và tên nút bấm thì không, và chúng thay đổi liên tục.
Tài liệu tham khảo
- roadmap.sh — DevOps Roadmap
- AWS Documentation
- Azure Documentation
- Google Cloud Documentation
- AWS Well-Architected Framework
- Google Cloud Architecture Framework
- Azure Well-Architected Framework
- AWS Shared Responsibility Model
- Microsoft — Shared responsibility in the cloud
- Google Cloud — Shared responsibilities and shared fate
- So sánh dịch vụ AWS và Azure (Microsoft)
- GCP cho người quen AWS — so sánh dịch vụ (Google)
- AWS Global Infrastructure — Regions và AZs
- AWS Pricing Calculator
- Azure Pricing Calculator
- Google Cloud Pricing Calculator
- FinOps Foundation
- DigitalOcean Documentation
- Hetzner Cloud Documentation
- Sách: Cloud Strategy — Gregor Hohpe
- Sách: The Cloud Adoption Playbook — Moe Abdula và cộng sự
Part of the DevOps Roadmap knowledge base.
Overview
Cloud providers deliver computing resources — servers, storage, networking, databases, and higher-level managed services — over the internet on a pay-as-you-go basis. Instead of buying and racking physical hardware, teams provision infrastructure in minutes through an API or web console. The three dominant providers are Amazon Web Services (AWS), Microsoft Azure, and Google Cloud Platform (GCP), with smaller players like DigitalOcean, Hetzner, and Alibaba Cloud serving specific niches (simplicity, price, or regional presence).
For DevOps engineers, the cloud is the default deployment target for most modern systems. Nearly every practice in the DevOps toolbox — infrastructure as code, autoscaling, immutable deployments, managed Kubernetes, serverless — is built on cloud primitives. Understanding how providers structure their services, how they charge, and where their responsibility ends and yours begins is foundational knowledge.
Because the big three offer largely equivalent core services under different names, the practical skill is less about memorizing one provider’s catalog and more about understanding the underlying concepts (compute, block/object storage, VPC networking, IAM) so you can map them across clouds. A useful mental model: every cloud is a warehouse of the same primitives — a machine that runs code, a disk, a bucket, a private network, an identity system, a queue, a database — dressed in different branding. Learn the primitives once and the console names become a translation exercise.
The economic shift the cloud represents
The deeper change the cloud introduced is financial, not technical. On-premises hardware is a capital expense (CapEx): you pay a large sum up front, depreciate it over years, and pay for peak capacity even when idle. Cloud is an operating expense (OpEx): you pay for what you use, when you use it, and you can turn it off. This turns infrastructure into a variable cost that scales with the business — powerful when growing, dangerous when unmanaged, because the same elasticity that lets you scale to a million users also lets a misconfigured loop scale your bill to five figures overnight. Cloud financial management (FinOps) exists precisely because the meter never stops.
Fundamentals
Service models: IaaS vs PaaS vs SaaS
| Model | You manage | Provider manages | Examples |
|---|---|---|---|
| IaaS (Infrastructure as a Service) | OS, runtime, app, data | Hardware, virtualization, networking | EC2, Azure VMs, Compute Engine |
| PaaS (Platform as a Service) | App, data | OS, runtime, scaling, patching | Elastic Beanstalk, App Service, Cloud Run |
| SaaS (Software as a Service) | Configuration, usage | Everything else | Gmail, Office 365, Datadog |
The lower you go in the stack (IaaS), the more control and operational burden you have. The higher you go (PaaS/SaaS), the faster you move but the more you accept the provider’s constraints.
A common fourth category worth naming is CaaS (Containers as a Service) and FaaS (Functions as a Service), which sit between IaaS and PaaS: managed Kubernetes (EKS/AKS/GKE) and serverless functions (Lambda/Cloud Functions) let you ship containers or code without managing the underlying VMs. The boundaries are fuzzy — Cloud Run is simultaneously “serverless containers,” a CaaS, and a PaaS depending on who you ask. Don’t over-index on the taxonomy; use it to ask the real question: how much of this stack do I want to be responsible for at 3 a.m. during an incident?
“Managed” is a spectrum, not a switch. RDS manages the database engine, backups, and failover — but you still choose instance size, tune parameters, and design your schema. A fully serverless database like DynamoDB or Aurora Serverless removes even the sizing decision. Each rung up the ladder trades a knob you could turn for one less thing that can break.
Regions and Availability Zones
- A region is a geographic area (e.g.,
us-east-1,westeurope,asia-southeast1) containing multiple data centers. - An Availability Zone (AZ) is one or more isolated data centers within a region, with independent power, cooling, and networking. AZs within a region are close enough for low-latency synchronous replication (single-digit milliseconds) but far enough apart that a fire, flood, or power failure in one does not take down another.
- Deploying across multiple AZs protects against data-center failure; deploying across multiple regions protects against regional outages and reduces latency for global users.
- Data transfer within an AZ is usually free; cross-AZ and cross-region traffic costs money — an often-overlooked cost driver.
Additional geographic concepts that matter in practice:
- Edge locations / PoPs (Points of Presence) — hundreds of small sites used by CDNs (CloudFront, Azure Front Door, Cloud CDN) and DNS to serve cached content close to users. They are not full regions; you cannot run a database there.
- Local Zones / Wavelength (AWS), Edge Zones (Azure) — extensions of a region placed in a metro area or telco network for ultra-low-latency use cases (gaming, live media, 5G).
- Region pairs (Azure) — Azure pairs regions for replication and staggered maintenance; some services replicate to the paired region by default.
- Region selection criteria — pick regions on four axes: (1) latency to your users, (2) data residency / sovereignty laws (GDPR, local data-protection rules), (3) service availability — not every service or instance type exists in every region, and new regions lag, and (4) price — the same VM can cost 20-50% more in one region than another (
us-east-1andus-west-2are typically among the cheapest AWS regions).
Mental model for blast radius: a single instance can fail (deploy ≥2); an AZ can fail (spread across AZs); a region can fail (multi-region for tier-0 systems); and rarely, a global control plane can fail (e.g., a global IAM or DNS outage takes multiple regions with it — the reason true global redundancy sometimes means multi-cloud, not just multi-region).
Shared responsibility model
The provider secures the cloud; you secure what you put in it.
| Provider’s responsibility | Your responsibility |
|---|---|
| Physical security, hardware | IAM users, roles, and policies |
| Hypervisor and host OS | Guest OS patching (IaaS) |
| Managed service availability | Data encryption choices, key management |
| Network infrastructure | Security groups, firewall rules, network segmentation |
| Global control plane | Application code and secrets |
Most cloud breaches are caused by customer-side misconfiguration (public S3 buckets, over-privileged IAM roles), not provider failures.
The dividing line shifts with the service model. This is the single most important nuance: the more managed the service, the more the provider takes on, but data, identity, and access configuration are always yours.
- IaaS (EC2): you patch the OS, configure the firewall, encrypt the disk, manage the app. The provider only guarantees the hardware and hypervisor.
- PaaS (App Service, RDS): the provider patches the OS and runtime; you still own data classification, access control, and app-layer security.
- SaaS (Office 365): the provider runs almost everything; you still own who has access, how identities are governed, and what data you put in.
A one-line heuristic that survives every service model: “Of, In, and Access.” The provider secures the security of the cloud; you secure what you put in the cloud and who has access to it. Misconfiguration — a public bucket, an unrotated key, an over-broad role — is always on the customer side of the line.
Pricing models
| Model | How it works | Best for |
|---|---|---|
| On-demand | Pay per second/hour, no commitment | Unpredictable or spiky workloads, experiments |
| Reserved / Savings Plans / CUDs | Commit to 1–3 years for 30–70% discount | Steady baseline load (databases, core services) |
| Spot / Preemptible | Bid on spare capacity, up to 90% off, can be reclaimed with short notice | Batch jobs, CI runners, stateless fault-tolerant workloads |
| Free tier | Limited free usage for new accounts or forever-free tiers | Learning and prototyping |
Commitment options compared:
| Option | Provider term | Discount | Flexibility |
|---|---|---|---|
| Standard Reserved Instances | AWS RI | Up to ~72% | Locked to instance family/region |
| Convertible RIs | AWS | Up to ~54% | Can change family, longer commitment |
| Savings Plans (Compute) | AWS | Up to ~66% | Applies across EC2/Fargate/Lambda by $/hour spend |
| Reserved VM Instances | Azure | Up to ~72% | Per VM series/region, exchangeable |
| Committed Use Discounts | GCP CUDs | Up to ~57% (resource) / ~70% (flexible) | Spend- or resource-based |
| Sustained Use Discounts | GCP | Automatic, up to ~30% | No commitment — applied automatically for long-running VMs |
Spot/preemptible mechanics matter. Spot capacity can be reclaimed with a short warning (AWS gives a 2-minute notice; GCP Spot VMs a 30-second notice; Azure Spot varies). Architect for it: checkpoint work, drain gracefully, spread across instance types and AZs so a single capacity crunch doesn’t reclaim everything at once, and never put a stateful primary database on spot. Spot is ideal for CI runners, batch/ETL, rendering, ML training with checkpointing, and Kubernetes worker nodes behind a cluster autoscaler.
Egress is the pricing trap that surprises everyone. Ingress (data into the cloud) is almost always free. Egress (data out to the internet, and often cross-region and cross-AZ) is metered and can dominate a bill for data-heavy or media-heavy workloads. The mental model: it’s cheap to check in, expensive to check out. This is also a deliberate lock-in lever — moving terabytes to another cloud costs real money — which is why regulators (and the EU Data Act) have pushed providers toward waiving egress fees for customers leaving.
Key Concepts
Core service mapping: AWS vs Azure vs GCP
| Category | AWS | Azure | GCP |
|---|---|---|---|
| Virtual machines | EC2 | Virtual Machines | Compute Engine |
| Autoscaling group | Auto Scaling Group | VM Scale Sets | Managed Instance Groups |
| Object storage | S3 | Blob Storage | Cloud Storage |
| Block storage | EBS | Managed Disks | Persistent Disk / Hyperdisk |
| File storage | EFS / FSx | Azure Files | Filestore |
| Virtual network | VPC | VNet | VPC |
| Load balancer | ELB/ALB/NLB | Load Balancer / App Gateway | Cloud Load Balancing |
| CDN | CloudFront | Front Door / CDN | Cloud CDN |
| DNS | Route 53 | Azure DNS | Cloud DNS |
| IAM | IAM | Entra ID (Azure AD) + RBAC | Cloud IAM |
| Managed SQL | RDS / Aurora | Azure SQL / Database for PostgreSQL | Cloud SQL / AlloyDB |
| NoSQL | DynamoDB | Cosmos DB | Firestore / Bigtable |
| Data warehouse | Redshift | Synapse / Fabric | BigQuery |
| Managed Kubernetes | EKS | AKS | GKE |
| Serverless containers | Fargate / App Runner | Container Apps | Cloud Run |
| Serverless functions | Lambda | Azure Functions | Cloud Functions / Cloud Run functions |
| Container registry | ECR | ACR | Artifact Registry |
| Secrets | Secrets Manager | Key Vault | Secret Manager |
| Key management | KMS | Key Vault | Cloud KMS |
| Monitoring | CloudWatch | Azure Monitor | Cloud Monitoring (Ops Suite) |
| Message queue | SQS / SNS | Service Bus / Event Grid | Pub/Sub |
| Managed streaming | Kinesis / MSK | Event Hubs | Pub/Sub / Managed Kafka |
| IaC (native) | CloudFormation / CDK | ARM / Bicep | Deployment Manager / Config Connector |
| Managed Redis/cache | ElastiCache | Azure Cache for Redis | Memorystore |
Treat this table as a Rosetta Stone, not gospel — services overlap and get renamed. The value is in recognizing that when someone says “we use Pub/Sub,” you immediately know its role (SQS/SNS + Kinesis territory) even if you’ve never touched GCP.
Compute options within a single provider
Even inside one cloud, “compute” is a ladder from most-control to least-management, and picking the wrong rung is a common mistake:
- Bare metal / dedicated hosts — full physical isolation for licensing or compliance; rare.
- Virtual machines (EC2, Azure VMs, Compute Engine) — the classic IaaS unit; you own the OS.
- Managed Kubernetes (EKS/AKS/GKE) — orchestrated containers; you own workloads and (often) worker nodes.
- Serverless containers (Fargate, Cloud Run, Container Apps) — bring a container image, forget the nodes.
- Functions (Lambda, Cloud Functions, Azure Functions) — bring a function, forget the container.
- PaaS app hosting (Elastic Beanstalk, App Service, App Engine) — bring code, forget almost everything.
The general rule: start as high on the ladder as your requirements allow, and drop down only when a concrete constraint (latency, cost at scale, special hardware, licensing) forces you to. Teams routinely reach for Kubernetes when Cloud Run would have done, paying for cluster complexity they don’t need.
Instance families and sizing
VMs come in families tuned for different workloads — the naming differs but the categories are universal:
| Workload | AWS family | Azure series | GCP family |
|---|---|---|---|
| General purpose | M, T (burstable) | D, B (burstable) | E2, N2 |
| Compute optimized | C | F | C2 |
| Memory optimized | R, X | E, M | M, N2-highmem |
| Storage optimized | I, D | L | — |
| GPU / accelerated | P, G, Inf/Trn | N-series | A (GPU), TPU VMs |
Burstable instances (T-series, B-series) accumulate CPU credits while idle and spend them during spikes — cheap for bursty, low-average workloads like small web servers, but they throttle hard if you sustain high CPU and run out of credits. Never put a steadily-busy service on a burstable instance and then wonder why it’s slow.
Right-sizing is a continuous discipline, not a one-time choice: measure real CPU/memory/IO utilization and shrink over-provisioned instances. The most common waste in any cloud bill is oversized, under-utilized VMs bought “to be safe.”
Smaller providers worth knowing
- DigitalOcean — simple UX, predictable flat pricing, Droplets (VMs), managed Kubernetes and databases, App Platform (PaaS), Spaces (S3-compatible object storage). Great for startups, side projects, and developers who want to avoid the complexity of the big three. Its egress pricing is far more generous than hyperscalers.
- Hetzner — extremely price-competitive dedicated and cloud servers in Europe (and now the US); popular for self-hosted Kubernetes, homelab-grade production, and cost-sensitive workloads. You trade the vast managed-service catalog of the hyperscalers for dramatically lower per-core and per-TB prices.
- Alibaba Cloud — the dominant provider in China and much of Southeast Asia; essential if you serve mainland Chinese users (ICP licensing, local regions, integration with local payment and CDN ecosystems).
- Oracle Cloud (OCI) — aggressive free tier and competitive egress; strong for Oracle Database workloads.
- Cloudflare — not a full IaaS, but its edge platform (Workers, R2 object storage with zero egress fees, D1, KV) is increasingly a serverless-first alternative for the right workloads.
The pattern: hyperscalers win on breadth of managed services and enterprise compliance; specialists win on price, simplicity, or a specific region/ecosystem. A common architecture is “hyperscaler for the managed database and compliance, specialist for the CDN or object storage where egress dominates.”
CLI quick taste
# AWS: launch a VM
aws ec2 run-instances --image-id ami-0abcdef1234567890 \
--instance-type t3.micro --count 1
# Azure: create a VM
az vm create --resource-group myRG --name myVM \
--image Ubuntu2204 --size Standard_B1s
# GCP: create a VM
gcloud compute instances create my-vm \
--machine-type=e2-micro --zone=asia-southeast1-a
These one-liners are for learning only. In practice you never create production infrastructure imperatively — you declare it in Terraform/OpenTofu or the provider’s native IaC so it’s reviewable, version-controlled, and reproducible. The CLI is for inspection, break-glass fixes, and scripting, not for standing up your stack by hand.
Account and organization structure
Every provider offers a hierarchy for isolating environments and applying policy centrally:
| Concept | AWS | Azure | GCP |
|---|---|---|---|
| Top container | Organization | Tenant (Entra ID) | Organization |
| Grouping | Organizational Units (OUs) | Management Groups | Folders |
| Isolation boundary | Account | Subscription | Project |
| Guardrails | Service Control Policies (SCPs) | Azure Policy | Organization Policies |
The standard pattern is a landing zone: separate accounts/subscriptions/projects per environment (dev, staging, prod) and often per team, with a central “management” account for billing and logging and a “security” account for audit trails. This gives blast-radius isolation (a mistake in dev can’t touch prod), clean cost attribution, and a place to enforce org-wide guardrails. AWS Control Tower, Azure Landing Zones, and GCP’s landing zone blueprints automate the setup.
How to choose a provider
- Existing ecosystem — heavy Microsoft shop (AD, Office 365)? Azure integrates naturally and often comes bundled in enterprise agreements. Data/ML focused? GCP’s BigQuery and Vertex AI are strong. Broadest catalog and market share (largest talent pool)? AWS.
- Team expertise — hiring and existing knowledge often outweigh technical differences; a team fluent in one cloud ships faster than the same team fighting an unfamiliar “better” one.
- Region availability — check latency and data-residency requirements for your users; confirm the specific services and instance types you need actually exist in your target region.
- Pricing for your workload — model your actual usage; egress fees and managed-service premiums differ significantly. Use each provider’s pricing calculator with realistic numbers, not list prices.
- Compliance — required certifications (PCI DSS, HIPAA, SOC 2, FedRAMP, ISO 27001, local regulations) per region and per service.
- Managed services you’ll actually lean on — the “best” cloud is often the one with the specific managed service that removes the most toil for your workload (BigQuery for analytics, DynamoDB for key-value at scale, Cosmos DB for global multi-model).
Multi-cloud considerations
- Pros: reduces vendor lock-in leverage, meets regulatory/data-residency demands, lets you pick best-of-breed services (e.g., BigQuery on GCP but everything else on AWS), and provides genuine resilience against a whole-provider outage for tier-0 systems.
- Cons: doubles operational surface (two IAM models, two networking stacks, two billing systems, two sets of on-call runbooks), complicates security posture, fractures your team’s depth of expertise, and cross-cloud egress is expensive and slow.
- Portability ≠ multi-cloud. Keeping workloads portable (containers, Kubernetes, Terraform, open-source data stores) is cheap insurance and worth doing. Actively running the same workload live on multiple clouds is expensive and rarely worth it for small teams.
- Pragmatic stance: most teams are best served going deep on one provider and keeping portability as a hedge, rather than diluting effort across two. Reserve true active/active multi-cloud for organizations with the scale, staff, and regulatory drivers to justify it. Multi-cloud “for redundancy” is usually a solution in search of a problem for teams under a few hundred engineers.
- Hybrid cloud (on-prem + cloud) is a distinct pattern, common in enterprises with existing data centers, sunk hardware costs, or data that legally cannot leave the premises; tools like AWS Outposts, Azure Arc/Stack, and Anthos extend cloud control planes on-prem.
Best Practices
- Use Infrastructure as Code from day one (Terraform/OpenTofu, Pulumi, CloudFormation, Bicep) — clicking in the console creates undocumented, unreproducible, un-reviewable “snowflake” infrastructure that nobody can rebuild after an outage.
- Enable MFA and never use the root/owner account for daily work — create scoped IAM users/roles; lock away the root/global-admin account, put it behind MFA, and use it only for break-glass tasks that genuinely require it.
- Apply least privilege in IAM — start with narrow permissions and widen as needed; prefer roles and short-lived credentials over long-lived access keys; audit with tools like AWS IAM Access Analyzer, Azure PIM, and GCP Policy Analyzer.
- Set billing alerts and budgets immediately — surprise bills are the most common newcomer pain; alert at 50/80/100% of expected spend, and enable anomaly detection so a runaway loop pages you in hours, not at month-end.
- Deploy across at least two AZs for anything production — single-AZ deployments turn routine data-center maintenance into your outage; multi-AZ is the cheapest resilience you will ever buy.
- Tag every resource (team, environment, service, cost-center, owner) — tags power cost allocation, cleanup automation, ownership tracking, and security policy; enforce tagging with policy-as-code so untagged resources are blocked or flagged.
- Prefer managed services for undifferentiated heavy lifting — running your own PostgreSQL, Kafka, or Redis on VMs rarely beats the managed equivalent once you factor in backups, patching, failover, and the on-call burden. Reserve self-management for cases with a concrete cost or control justification.
- Watch egress and cross-AZ costs — data into the cloud is free, data out is not; keep chatty traffic within one AZ/region, use private endpoints instead of routing through the public internet, and put a CDN in front of static content.
- Use spot/preemptible instances for interruptible workloads — 60–90% savings for CI, batch, ETL, and stateless fault-tolerant services; design them to checkpoint and drain gracefully on the reclamation warning.
- Buy reservations/savings plans only after usage stabilizes — measure 2–3 months of steady-state usage first, then commit to your baseline; leave the spiky top of the curve on on-demand or spot.
- Right-size continuously — the biggest line item in most bills is oversized, under-utilized compute bought “to be safe”; review utilization and shrink monthly.
- Keep an organization/landing-zone structure — separate accounts/projects/subscriptions per environment for blast-radius isolation, clean cost attribution, and centralized guardrails (SCPs/Azure Policy/Org Policies).
- Encrypt by default and manage keys deliberately — enable encryption at rest and in transit everywhere (it’s usually free and on by default); use the provider KMS, and use customer-managed keys where compliance requires control over the key lifecycle.
- Centralize logging and audit trails — turn on CloudTrail / Azure Activity Log / Cloud Audit Logs in a dedicated, write-once account so you have an immutable record of who did what when — indispensable during incidents and audits.
- Automate cleanup of ephemeral resources — dev/test environments, orphaned volumes, unattached IPs, and old snapshots quietly accrue cost; schedule teardown and run a “cost janitor” (e.g., cloud-nuke, provider recommendations) regularly.
- Learn concepts, not consoles — VPCs, IAM, object storage, load balancing, and identity federation transfer across all providers; console layouts and exact button names do not, and they change constantly.
References
- roadmap.sh — DevOps Roadmap
- AWS Documentation
- Azure Documentation
- Google Cloud Documentation
- AWS Well-Architected Framework
- Google Cloud Architecture Framework
- Azure Well-Architected Framework
- AWS Shared Responsibility Model
- Microsoft — Shared responsibility in the cloud
- Google Cloud — Shared responsibilities and shared fate
- Compare AWS and Azure services (Microsoft)
- GCP for AWS professionals — service comparison (Google)
- AWS Global Infrastructure — Regions and AZs
- AWS Pricing Calculator
- Azure Pricing Calculator
- Google Cloud Pricing Calculator
- FinOps Foundation
- DigitalOcean Documentation
- Hetzner Cloud Documentation
- Book: Cloud Strategy — Gregor Hohpe
- Book: The Cloud Adoption Playbook — Moe Abdula et al.