← DevOps← DevOps
DevOpsDevOps18 Th7, 2026Jul 18, 202624 phút đọc20 min read

Nhà cung cấp Cloud (Cloud Providers)Cloud Providers

Thuộc bộ tài liệu DevOps Roadmap.

Tổng quan

Cloud provider cung cấp tài nguyên tính toán — server, storage, networking, database và các managed service cấp cao hơn — qua internet theo mô hình trả tiền theo mức sử dụng (pay-as-you-go). Thay vì mua và lắp đặt phần cứng vật lý, team có thể provision hạ tầng trong vài phút qua API hoặc web console. Ba nhà cung cấp lớn nhất là Amazon Web Services (AWS), Microsoft AzureGoogle Cloud Platform (GCP); ngoài ra còn có các đơn vị nhỏ hơn như DigitalOcean, Hetzner, Alibaba Cloud phục vụ các ngách riêng (đơn giản, giá rẻ, hoặc thị trường khu vực).

Với DevOps engineer, cloud là môi trường triển khai mặc định của hầu hết hệ thống hiện đại. Gần như mọi thực hành DevOps — Infrastructure as Code, autoscaling, immutable deployment, managed Kubernetes, serverless — đều xây dựng trên các dịch vụ nền tảng của cloud. Hiểu cách các provider tổ chức dịch vụ, cách tính tiền, và ranh giới trách nhiệm giữa provider và bạn là kiến thức bắt buộc.

Vì ba ông lớn cung cấp các dịch vụ cốt lõi gần như tương đương nhau dưới những cái tên khác nhau, kỹ năng thực tế không phải là thuộc lòng catalog của một provider, mà là nắm vững khái niệm nền (compute, block/object storage, VPC networking, IAM) để có thể “ánh xạ” giữa các cloud. Một mô hình tư duy hữu ích: mọi cloud đều là một kho chứa cùng những primitive giống nhau — một cỗ máy chạy code, một disk, một bucket, một mạng riêng, một hệ thống identity, một queue, một database — chỉ khác thương hiệu. Học primitive một lần, còn tên trên console chỉ là bài toán dịch thuật.

Sự thay đổi về kinh tế mà cloud mang lại

Thay đổi sâu sắc nhất cloud đem lại là về tài chính, không phải kỹ thuật. Phần cứng on-premises là chi phí đầu tư (CapEx): bạn trả một khoản lớn từ đầu, khấu hao nhiều năm, và trả cho công suất đỉnh ngay cả khi không dùng. Cloud là chi phí vận hành (OpEx): trả cho những gì bạn dùng, khi bạn dùng, và có thể tắt đi. Điều này biến hạ tầng thành chi phí biến đổi co giãn theo quy mô kinh doanh — mạnh mẽ khi tăng trưởng, nguy hiểm khi buông lỏng, vì chính sự co giãn cho phép bạn scale lên một triệu user cũng cho phép một vòng lặp cấu hình sai đẩy hóa đơn lên năm chữ số chỉ sau một đêm. Cloud financial management (FinOps) ra đời chính vì công tơ không bao giờ ngừng chạy.

Kiến thức nền tảng

Mô hình dịch vụ: IaaS vs PaaS vs SaaS

Mô hìnhBạn quản lýProvider quản lýVí dụ
IaaS (Infrastructure as a Service)OS, runtime, app, dữ liệuPhần cứng, ảo hóa, networkEC2, Azure VMs, Compute Engine
PaaS (Platform as a Service)App, dữ liệuOS, runtime, scaling, vá lỗiElastic Beanstalk, App Service, Cloud Run
SaaS (Software as a Service)Cấu hình, cách dùngToàn bộ phần còn lạiGmail, Office 365, Datadog

Càng xuống thấp (IaaS) bạn càng có nhiều quyền kiểm soát nhưng gánh nặng vận hành càng lớn. Càng lên cao (PaaS/SaaS) bạn triển khai càng nhanh nhưng phải chấp nhận ràng buộc của provider.

Một nhóm thứ tư đáng gọi tên là CaaS (Containers as a Service)FaaS (Functions as a Service), nằm giữa IaaS và PaaS: managed Kubernetes (EKS/AKS/GKE) và serverless function (Lambda/Cloud Functions) cho phép bạn ship container hoặc code mà không phải quản lý VM bên dưới. Ranh giới khá mờ — Cloud Run vừa là “serverless container”, vừa là CaaS, vừa là PaaS tùy góc nhìn. Đừng sa đà vào phân loại; hãy dùng nó để đặt câu hỏi thật sự: tôi muốn chịu trách nhiệm bao nhiêu phần của stack này lúc 3 giờ sáng khi có sự cố?

“Managed” là một dải liên tục, không phải công tắc bật/tắt. RDS quản lý database engine, backup và failover — nhưng bạn vẫn chọn kích thước instance, tinh chỉnh tham số và thiết kế schema. Một database serverless hoàn toàn như DynamoDB hay Aurora Serverless loại bỏ luôn cả quyết định sizing. Mỗi nấc leo lên là đánh đổi một cái núm bạn có thể vặn để bớt đi một thứ có thể hỏng.

Region và Availability Zone

Một số khái niệm địa lý khác quan trọng trong thực tế:

Mô hình tư duy về blast radius: một instance có thể chết (deploy ≥2); một AZ có thể chết (trải trên nhiều AZ); một region có thể chết (multi-region cho hệ thống tier-0); và hiếm khi, một control plane toàn cầu có thể chết (ví dụ sự cố IAM hay DNS global kéo theo nhiều region — lý do vì sao dự phòng toàn cầu thật sự đôi khi nghĩa là multi-cloud, không chỉ multi-region).

Shared responsibility model (mô hình trách nhiệm chia sẻ)

Provider bảo mật hạ tầng cloud; bạn bảo mật những gì bạn đặt lên đó.

Trách nhiệm của providerTrách nhiệm của bạn
An ninh vật lý, phần cứngIAM user, role, policy
Hypervisor và host OSVá guest OS (với IaaS)
Độ sẵn sàng của managed serviceMã hóa dữ liệu, quản lý key
Hạ tầng mạngSecurity group, firewall rule, phân đoạn mạng
Control plane toàn cầuCode ứng dụng và secrets

Phần lớn sự cố bảo mật trên cloud đến từ cấu hình sai phía khách hàng (S3 bucket public, IAM role quá quyền), không phải lỗi của provider.

Ranh giới này dịch chuyển theo mô hình dịch vụ. Đây là điểm tinh tế quan trọng nhất: service càng được quản lý nhiều thì provider gánh càng nhiều, nhưng dữ liệu, identity và cấu hình quyền truy cập luôn là của bạn.

Một câu tóm tắt sống sót qua mọi mô hình dịch vụ: “Của, Trong, và Quyền truy cập.” Provider bảo mật an ninh của cloud; bạn bảo mật những gì đặt trong cloud và ai có quyền truy cập vào đó. Cấu hình sai — bucket public, key không rotate, role quá rộng — luôn nằm ở phía khách hàng.

Mô hình tính giá

Mô hìnhCách hoạt độngPhù hợp với
On-demandTrả theo giây/giờ, không cam kếtWorkload khó dự đoán, thử nghiệm
Reserved / Savings Plans / CUDCam kết 1–3 năm, giảm 30–70%Tải nền ổn định (database, core service)
Spot / PreemptibleDùng capacity dư, rẻ tới 90%, có thể bị thu hồi đột ngộtBatch job, CI runner, workload stateless chịu được gián đoạn
Free tierMức dùng miễn phí cho tài khoản mới hoặc vĩnh viễnHọc tập, prototype

So sánh các lựa chọn cam kết:

Lựa chọnTên theo providerMức giảmĐộ linh hoạt
Standard Reserved InstancesAWS RITới ~72%Khóa theo instance family/region
Convertible RIsAWSTới ~54%Đổi được family, cam kết dài hơn
Savings Plans (Compute)AWSTới ~66%Áp cho EC2/Fargate/Lambda theo mức chi $/giờ
Reserved VM InstancesAzureTới ~72%Theo VM series/region, đổi được
Committed Use DiscountsGCP CUDTới ~57% (resource) / ~70% (flexible)Theo mức chi hoặc theo resource
Sustained Use DiscountsGCPTự động, tới ~30%Không cam kết — tự áp cho VM chạy lâu

Cơ chế spot/preemptible rất quan trọng. Capacity spot có thể bị thu hồi với cảnh báo ngắn (AWS báo trước 2 phút; GCP Spot VM báo 30 giây; Azure Spot tùy trường hợp). Hãy thiết kế để chịu được điều đó: checkpoint công việc, drain nhẹ nhàng, trải trên nhiều instance type và AZ để một đợt khan hiếm capacity không thu hồi tất cả cùng lúc, và không bao giờ đặt database primary có state lên spot. Spot lý tưởng cho CI runner, batch/ETL, rendering, ML training có checkpoint, và Kubernetes worker node sau một cluster autoscaler.

Egress là cái bẫy giá làm mọi người bất ngờ. Ingress (dữ liệu vào cloud) gần như luôn miễn phí. Egress (dữ liệu ra internet, và thường cả giữa region, giữa AZ) bị tính tiền và có thể chiếm phần lớn hóa đơn với workload nặng dữ liệu hay media. Mô hình tư duy: check-in thì rẻ, check-out thì đắt. Đây cũng là đòn bẩy lock-in có chủ đích — chuyển hàng terabyte sang cloud khác tốn tiền thật — nên các nhà quản lý (và EU Data Act) đã thúc ép provider miễn phí egress cho khách hàng rời đi.

Khái niệm chính

Bảng ánh xạ dịch vụ cốt lõi: AWS vs Azure vs GCP

NhómAWSAzureGCP
Máy ảoEC2Virtual MachinesCompute Engine
Autoscaling groupAuto Scaling GroupVM Scale SetsManaged Instance Groups
Object storageS3Blob StorageCloud Storage
Block storageEBSManaged DisksPersistent Disk / Hyperdisk
File storageEFS / FSxAzure FilesFilestore
Mạng ảoVPCVNetVPC
Load balancerELB/ALB/NLBLoad Balancer / App GatewayCloud Load Balancing
CDNCloudFrontFront Door / CDNCloud CDN
DNSRoute 53Azure DNSCloud DNS
IAMIAMEntra ID (Azure AD) + RBACCloud IAM
SQL managedRDS / AuroraAzure SQL / Database for PostgreSQLCloud SQL / AlloyDB
NoSQLDynamoDBCosmos DBFirestore / Bigtable
Data warehouseRedshiftSynapse / FabricBigQuery
Managed KubernetesEKSAKSGKE
Serverless containerFargate / App RunnerContainer AppsCloud Run
Serverless functionLambdaAzure FunctionsCloud Functions / Cloud Run functions
Container registryECRACRArtifact Registry
SecretsSecrets ManagerKey VaultSecret Manager
Key managementKMSKey VaultCloud KMS
MonitoringCloudWatchAzure MonitorCloud Monitoring (Ops Suite)
Message queueSQS / SNSService Bus / Event GridPub/Sub
Streaming managedKinesis / MSKEvent HubsPub/Sub / Managed Kafka
IaC (native)CloudFormation / CDKARM / BicepDeployment Manager / Config Connector
Redis/cache managedElastiCacheAzure Cache for RedisMemorystore

Hãy coi bảng này là “phiến đá Rosetta”, không phải chân lý — service chồng lấn và hay bị đổi tên. Giá trị nằm ở chỗ khi ai đó nói “chúng tôi dùng Pub/Sub”, bạn lập tức biết vai trò của nó (vùng đất của SQS/SNS + Kinesis) dù chưa từng chạm vào GCP.

Các lựa chọn compute trong cùng một provider

Ngay trong một cloud, “compute” cũng là một cái thang từ kiểm-soát-nhiều-nhất tới quản-lý-ít-nhất, và chọn sai nấc là lỗi phổ biến:

  1. Bare metal / dedicated host — cô lập vật lý hoàn toàn cho licensing hoặc compliance; hiếm gặp.
  2. Virtual machine (EC2, Azure VMs, Compute Engine) — đơn vị IaaS cổ điển; bạn sở hữu OS.
  3. Managed Kubernetes (EKS/AKS/GKE) — container được điều phối; bạn sở hữu workload và (thường cả) worker node.
  4. Serverless container (Fargate, Cloud Run, Container Apps) — mang container image tới, quên node đi.
  5. Function (Lambda, Cloud Functions, Azure Functions) — mang function tới, quên container đi.
  6. PaaS app hosting (Elastic Beanstalk, App Service, App Engine) — mang code tới, quên gần như mọi thứ.

Quy tắc chung: bắt đầu ở nấc cao nhất mà yêu cầu cho phép, chỉ tụt xuống khi có một ràng buộc cụ thể (latency, chi phí ở quy mô lớn, phần cứng đặc biệt, licensing) buộc bạn. Nhiều team với tay tới Kubernetes trong khi Cloud Run là đủ, để rồi trả giá cho độ phức tạp của cluster mà họ không cần.

Instance family và sizing

VM có các family tối ưu cho từng loại workload — tên khác nhau nhưng nhóm thì phổ quát:

WorkloadAWS familyAzure seriesGCP family
General purposeM, T (burstable)D, B (burstable)E2, N2
Compute optimizedCFC2
Memory optimizedR, XE, MM, N2-highmem
Storage optimizedI, DL
GPU / acceleratedP, G, Inf/TrnN-seriesA (GPU), TPU VMs

Instance burstable (T-series, B-series) tích lũy CPU credit khi rảnh và tiêu chúng khi có đột biến — rẻ cho workload bùng phát nhẹ, trung bình thấp như web server nhỏ, nhưng bị throttle mạnh nếu chạy CPU cao liên tục và cạn credit. Đừng bao giờ đặt một service bận liên tục lên instance burstable rồi thắc mắc sao nó chậm.

Right-sizing là kỷ luật liên tục, không phải lựa chọn một lần: đo mức dùng CPU/memory/IO thật và thu nhỏ instance dư thừa. Khoản lãng phí phổ biến nhất trong mọi hóa đơn cloud là VM quá cỡ, dùng không hết, mua “cho chắc”.

Các provider nhỏ đáng biết

Mô thức chung: hyperscaler thắng ở bề rộng managed service và compliance doanh nghiệp; các đơn vị chuyên biệt thắng ở giá, sự đơn giản, hoặc một region/hệ sinh thái cụ thể. Kiến trúc phổ biến là “hyperscaler cho managed database và compliance, đơn vị chuyên biệt cho CDN hoặc object storage nơi egress chi phối”.

Nếm thử CLI

# AWS: tạo VM
aws ec2 run-instances --image-id ami-0abcdef1234567890 \
  --instance-type t3.micro --count 1

# Azure: tạo VM
az vm create --resource-group myRG --name myVM \
  --image Ubuntu2204 --size Standard_B1s

# GCP: tạo VM
gcloud compute instances create my-vm \
  --machine-type=e2-micro --zone=asia-southeast1-a

Các dòng lệnh này chỉ để học. Thực tế bạn không bao giờ tạo hạ tầng production theo kiểu mệnh lệnh (imperative) — bạn khai báo nó trong Terraform/OpenTofu hoặc IaC native của provider để review được, quản lý version và tái tạo được. CLI dùng để kiểm tra, sửa khẩn cấp (break-glass) và scripting, không phải để dựng stack bằng tay.

Cấu trúc account và organization

Mọi provider đều có một phân cấp để cô lập môi trường và áp policy tập trung:

Khái niệmAWSAzureGCP
Container gốcOrganizationTenant (Entra ID)Organization
NhómOrganizational Units (OUs)Management GroupsFolders
Biên cô lậpAccountSubscriptionProject
GuardrailService Control Policies (SCPs)Azure PolicyOrganization Policies

Mô thức chuẩn là landing zone: tách account/subscription/project theo môi trường (dev, staging, prod) và thường theo team, với một account “management” trung tâm cho billing và logging, một account “security” cho audit trail. Nó mang lại cô lập blast-radius (sai ở dev không chạm được prod), phân bổ chi phí sạch, và một nơi để cưỡng chế guardrail toàn org. AWS Control Tower, Azure Landing Zones và blueprint landing zone của GCP tự động hóa việc thiết lập.

Cách chọn provider

  1. Hệ sinh thái sẵn có — công ty dùng nhiều Microsoft (AD, Office 365)? Azure tích hợp tự nhiên và thường đi kèm trong hợp đồng doanh nghiệp. Thiên về data/ML? BigQuery và Vertex AI của GCP rất mạnh. Cần catalog rộng nhất, thị phần lớn nhất (nguồn nhân lực dồi dào nhất)? AWS.
  2. Kỹ năng của team — kinh nghiệm sẵn có và khả năng tuyển dụng thường quan trọng hơn khác biệt kỹ thuật; một team thành thạo cloud này ship nhanh hơn chính team đó vật lộn với cloud “tốt hơn” nhưng lạ lẫm.
  3. Region khả dụng — kiểm tra latency và yêu cầu data residency; xác nhận đúng service và instance type bạn cần thực sự tồn tại ở region đích. Với Việt Nam thường dùng region Singapore (ap-southeast-1, asia-southeast1).
  4. Chi phí cho workload của bạn — mô phỏng usage thực tế; phí egress và giá managed service chênh lệch đáng kể. Dùng pricing calculator của từng provider với số liệu thực, không phải giá niêm yết.
  5. Compliance — chứng chỉ bắt buộc (PCI DSS, HIPAA, SOC 2, FedRAMP, ISO 27001, quy định địa phương) theo từng region và từng service.
  6. Managed service bạn thật sự dựa vào — cloud “tốt nhất” thường là cloud có đúng managed service gỡ bỏ nhiều toil nhất cho workload của bạn (BigQuery cho analytics, DynamoDB cho key-value quy mô lớn, Cosmos DB cho multi-model toàn cầu).

Cân nhắc multi-cloud

Best Practices

  1. Dùng Infrastructure as Code ngay từ đầu (Terraform/OpenTofu, Pulumi, CloudFormation, Bicep) — click tay trên console tạo ra hạ tầng “snowflake” không tài liệu, không tái tạo được, không review được mà không ai dựng lại nổi sau sự cố.
  2. Bật MFA và không bao giờ dùng tài khoản root/owner cho công việc hằng ngày — tạo IAM user/role có phạm vi giới hạn; khóa kỹ tài khoản root/global-admin sau MFA và chỉ dùng cho các thao tác break-glass thật sự cần.
  3. Áp dụng least privilege cho IAM — bắt đầu với quyền hẹp rồi mở dần; ưu tiên role và credential ngắn hạn thay vì access key sống lâu; audit bằng AWS IAM Access Analyzer, Azure PIM, GCP Policy Analyzer.
  4. Đặt billing alert và budget ngay lập tức — hóa đơn “bất ngờ” là nỗi đau phổ biến nhất của người mới; cảnh báo ở mức 50/80/100% chi tiêu dự kiến, và bật anomaly detection để một vòng lặp chạy loạn báo động trong vài giờ, không phải cuối tháng.
  5. Triển khai production trên ít nhất hai AZ — chạy một AZ nghĩa là bảo trì data center định kỳ trở thành sự cố của bạn; multi-AZ là lớp chịu lỗi rẻ nhất bạn từng mua.
  6. Gắn tag cho mọi resource (team, environment, service, cost-center, owner) — tag phục vụ phân bổ chi phí, tự động dọn dẹp, truy vết ownership và policy bảo mật; cưỡng chế tagging bằng policy-as-code để resource thiếu tag bị chặn hoặc gắn cờ.
  7. Ưu tiên managed service cho các việc “nặng mà không tạo khác biệt” — tự vận hành PostgreSQL, Kafka hay Redis trên VM hiếm khi hơn bản managed khi tính cả backup, vá lỗi, failover và gánh nặng on-call. Chỉ tự quản lý khi có lý do chi phí hoặc kiểm soát cụ thể.
  8. Để ý chi phí egress và cross-AZ — dữ liệu vào cloud miễn phí, ra thì không; giữ traffic trao đổi nhiều trong cùng AZ/region, dùng private endpoint thay vì đi qua internet công cộng, và đặt CDN trước nội dung tĩnh.
  9. Dùng spot/preemptible cho workload chịu được gián đoạn — tiết kiệm 60–90% cho CI, batch, ETL và service stateless chịu lỗi; thiết kế để checkpoint và drain nhẹ nhàng khi có cảnh báo thu hồi.
  10. Chỉ mua reservation/savings plan sau khi usage đã ổn định — đo 2–3 tháng tải ổn định trước, rồi cam kết cho phần tải nền; để phần đỉnh đột biến trên on-demand hoặc spot.
  11. Right-size liên tục — khoản lớn nhất trong hầu hết hóa đơn là compute quá cỡ, dùng không hết, mua “cho chắc”; xem lại mức dùng và thu nhỏ hằng tháng.
  12. Tổ chức landing zone rõ ràng — tách account/project/subscription theo môi trường để cô lập blast radius, phân bổ chi phí sạch và cưỡng chế guardrail tập trung (SCP/Azure Policy/Org Policy).
  13. Mã hóa mặc định và quản lý key có chủ đích — bật mã hóa at-rest và in-transit ở mọi nơi (thường miễn phí và mặc định bật); dùng KMS của provider, và dùng customer-managed key khi compliance yêu cầu kiểm soát vòng đời key.
  14. Tập trung logging và audit trail — bật CloudTrail / Azure Activity Log / Cloud Audit Logs trong một account chuyên dụng, chỉ-ghi để có bản ghi bất biến về ai làm gì lúc nào — không thể thiếu khi có sự cố và khi audit.
  15. Tự động dọn dẹp resource tạm — môi trường dev/test, volume mồ côi, IP không gắn, snapshot cũ âm thầm ngốn tiền; lên lịch teardown và chạy một “cost janitor” (cloud-nuke, gợi ý của provider) định kỳ.
  16. Học khái niệm, đừng học console — VPC, IAM, object storage, load balancing và identity federation dùng được trên mọi provider; bố cục console và tên nút bấm thì không, và chúng thay đổi liên tục.

Tài liệu tham khảo

Part of the DevOps Roadmap knowledge base.

Overview

Cloud providers deliver computing resources — servers, storage, networking, databases, and higher-level managed services — over the internet on a pay-as-you-go basis. Instead of buying and racking physical hardware, teams provision infrastructure in minutes through an API or web console. The three dominant providers are Amazon Web Services (AWS), Microsoft Azure, and Google Cloud Platform (GCP), with smaller players like DigitalOcean, Hetzner, and Alibaba Cloud serving specific niches (simplicity, price, or regional presence).

For DevOps engineers, the cloud is the default deployment target for most modern systems. Nearly every practice in the DevOps toolbox — infrastructure as code, autoscaling, immutable deployments, managed Kubernetes, serverless — is built on cloud primitives. Understanding how providers structure their services, how they charge, and where their responsibility ends and yours begins is foundational knowledge.

Because the big three offer largely equivalent core services under different names, the practical skill is less about memorizing one provider’s catalog and more about understanding the underlying concepts (compute, block/object storage, VPC networking, IAM) so you can map them across clouds. A useful mental model: every cloud is a warehouse of the same primitives — a machine that runs code, a disk, a bucket, a private network, an identity system, a queue, a database — dressed in different branding. Learn the primitives once and the console names become a translation exercise.

The economic shift the cloud represents

The deeper change the cloud introduced is financial, not technical. On-premises hardware is a capital expense (CapEx): you pay a large sum up front, depreciate it over years, and pay for peak capacity even when idle. Cloud is an operating expense (OpEx): you pay for what you use, when you use it, and you can turn it off. This turns infrastructure into a variable cost that scales with the business — powerful when growing, dangerous when unmanaged, because the same elasticity that lets you scale to a million users also lets a misconfigured loop scale your bill to five figures overnight. Cloud financial management (FinOps) exists precisely because the meter never stops.

Fundamentals

Service models: IaaS vs PaaS vs SaaS

ModelYou manageProvider managesExamples
IaaS (Infrastructure as a Service)OS, runtime, app, dataHardware, virtualization, networkingEC2, Azure VMs, Compute Engine
PaaS (Platform as a Service)App, dataOS, runtime, scaling, patchingElastic Beanstalk, App Service, Cloud Run
SaaS (Software as a Service)Configuration, usageEverything elseGmail, Office 365, Datadog

The lower you go in the stack (IaaS), the more control and operational burden you have. The higher you go (PaaS/SaaS), the faster you move but the more you accept the provider’s constraints.

A common fourth category worth naming is CaaS (Containers as a Service) and FaaS (Functions as a Service), which sit between IaaS and PaaS: managed Kubernetes (EKS/AKS/GKE) and serverless functions (Lambda/Cloud Functions) let you ship containers or code without managing the underlying VMs. The boundaries are fuzzy — Cloud Run is simultaneously “serverless containers,” a CaaS, and a PaaS depending on who you ask. Don’t over-index on the taxonomy; use it to ask the real question: how much of this stack do I want to be responsible for at 3 a.m. during an incident?

“Managed” is a spectrum, not a switch. RDS manages the database engine, backups, and failover — but you still choose instance size, tune parameters, and design your schema. A fully serverless database like DynamoDB or Aurora Serverless removes even the sizing decision. Each rung up the ladder trades a knob you could turn for one less thing that can break.

Regions and Availability Zones

Additional geographic concepts that matter in practice:

Mental model for blast radius: a single instance can fail (deploy ≥2); an AZ can fail (spread across AZs); a region can fail (multi-region for tier-0 systems); and rarely, a global control plane can fail (e.g., a global IAM or DNS outage takes multiple regions with it — the reason true global redundancy sometimes means multi-cloud, not just multi-region).

Shared responsibility model

The provider secures the cloud; you secure what you put in it.

Provider’s responsibilityYour responsibility
Physical security, hardwareIAM users, roles, and policies
Hypervisor and host OSGuest OS patching (IaaS)
Managed service availabilityData encryption choices, key management
Network infrastructureSecurity groups, firewall rules, network segmentation
Global control planeApplication code and secrets

Most cloud breaches are caused by customer-side misconfiguration (public S3 buckets, over-privileged IAM roles), not provider failures.

The dividing line shifts with the service model. This is the single most important nuance: the more managed the service, the more the provider takes on, but data, identity, and access configuration are always yours.

A one-line heuristic that survives every service model: “Of, In, and Access.” The provider secures the security of the cloud; you secure what you put in the cloud and who has access to it. Misconfiguration — a public bucket, an unrotated key, an over-broad role — is always on the customer side of the line.

Pricing models

ModelHow it worksBest for
On-demandPay per second/hour, no commitmentUnpredictable or spiky workloads, experiments
Reserved / Savings Plans / CUDsCommit to 1–3 years for 30–70% discountSteady baseline load (databases, core services)
Spot / PreemptibleBid on spare capacity, up to 90% off, can be reclaimed with short noticeBatch jobs, CI runners, stateless fault-tolerant workloads
Free tierLimited free usage for new accounts or forever-free tiersLearning and prototyping

Commitment options compared:

OptionProvider termDiscountFlexibility
Standard Reserved InstancesAWS RIUp to ~72%Locked to instance family/region
Convertible RIsAWSUp to ~54%Can change family, longer commitment
Savings Plans (Compute)AWSUp to ~66%Applies across EC2/Fargate/Lambda by $/hour spend
Reserved VM InstancesAzureUp to ~72%Per VM series/region, exchangeable
Committed Use DiscountsGCP CUDsUp to ~57% (resource) / ~70% (flexible)Spend- or resource-based
Sustained Use DiscountsGCPAutomatic, up to ~30%No commitment — applied automatically for long-running VMs

Spot/preemptible mechanics matter. Spot capacity can be reclaimed with a short warning (AWS gives a 2-minute notice; GCP Spot VMs a 30-second notice; Azure Spot varies). Architect for it: checkpoint work, drain gracefully, spread across instance types and AZs so a single capacity crunch doesn’t reclaim everything at once, and never put a stateful primary database on spot. Spot is ideal for CI runners, batch/ETL, rendering, ML training with checkpointing, and Kubernetes worker nodes behind a cluster autoscaler.

Egress is the pricing trap that surprises everyone. Ingress (data into the cloud) is almost always free. Egress (data out to the internet, and often cross-region and cross-AZ) is metered and can dominate a bill for data-heavy or media-heavy workloads. The mental model: it’s cheap to check in, expensive to check out. This is also a deliberate lock-in lever — moving terabytes to another cloud costs real money — which is why regulators (and the EU Data Act) have pushed providers toward waiving egress fees for customers leaving.

Key Concepts

Core service mapping: AWS vs Azure vs GCP

CategoryAWSAzureGCP
Virtual machinesEC2Virtual MachinesCompute Engine
Autoscaling groupAuto Scaling GroupVM Scale SetsManaged Instance Groups
Object storageS3Blob StorageCloud Storage
Block storageEBSManaged DisksPersistent Disk / Hyperdisk
File storageEFS / FSxAzure FilesFilestore
Virtual networkVPCVNetVPC
Load balancerELB/ALB/NLBLoad Balancer / App GatewayCloud Load Balancing
CDNCloudFrontFront Door / CDNCloud CDN
DNSRoute 53Azure DNSCloud DNS
IAMIAMEntra ID (Azure AD) + RBACCloud IAM
Managed SQLRDS / AuroraAzure SQL / Database for PostgreSQLCloud SQL / AlloyDB
NoSQLDynamoDBCosmos DBFirestore / Bigtable
Data warehouseRedshiftSynapse / FabricBigQuery
Managed KubernetesEKSAKSGKE
Serverless containersFargate / App RunnerContainer AppsCloud Run
Serverless functionsLambdaAzure FunctionsCloud Functions / Cloud Run functions
Container registryECRACRArtifact Registry
SecretsSecrets ManagerKey VaultSecret Manager
Key managementKMSKey VaultCloud KMS
MonitoringCloudWatchAzure MonitorCloud Monitoring (Ops Suite)
Message queueSQS / SNSService Bus / Event GridPub/Sub
Managed streamingKinesis / MSKEvent HubsPub/Sub / Managed Kafka
IaC (native)CloudFormation / CDKARM / BicepDeployment Manager / Config Connector
Managed Redis/cacheElastiCacheAzure Cache for RedisMemorystore

Treat this table as a Rosetta Stone, not gospel — services overlap and get renamed. The value is in recognizing that when someone says “we use Pub/Sub,” you immediately know its role (SQS/SNS + Kinesis territory) even if you’ve never touched GCP.

Compute options within a single provider

Even inside one cloud, “compute” is a ladder from most-control to least-management, and picking the wrong rung is a common mistake:

  1. Bare metal / dedicated hosts — full physical isolation for licensing or compliance; rare.
  2. Virtual machines (EC2, Azure VMs, Compute Engine) — the classic IaaS unit; you own the OS.
  3. Managed Kubernetes (EKS/AKS/GKE) — orchestrated containers; you own workloads and (often) worker nodes.
  4. Serverless containers (Fargate, Cloud Run, Container Apps) — bring a container image, forget the nodes.
  5. Functions (Lambda, Cloud Functions, Azure Functions) — bring a function, forget the container.
  6. PaaS app hosting (Elastic Beanstalk, App Service, App Engine) — bring code, forget almost everything.

The general rule: start as high on the ladder as your requirements allow, and drop down only when a concrete constraint (latency, cost at scale, special hardware, licensing) forces you to. Teams routinely reach for Kubernetes when Cloud Run would have done, paying for cluster complexity they don’t need.

Instance families and sizing

VMs come in families tuned for different workloads — the naming differs but the categories are universal:

WorkloadAWS familyAzure seriesGCP family
General purposeM, T (burstable)D, B (burstable)E2, N2
Compute optimizedCFC2
Memory optimizedR, XE, MM, N2-highmem
Storage optimizedI, DL
GPU / acceleratedP, G, Inf/TrnN-seriesA (GPU), TPU VMs

Burstable instances (T-series, B-series) accumulate CPU credits while idle and spend them during spikes — cheap for bursty, low-average workloads like small web servers, but they throttle hard if you sustain high CPU and run out of credits. Never put a steadily-busy service on a burstable instance and then wonder why it’s slow.

Right-sizing is a continuous discipline, not a one-time choice: measure real CPU/memory/IO utilization and shrink over-provisioned instances. The most common waste in any cloud bill is oversized, under-utilized VMs bought “to be safe.”

Smaller providers worth knowing

The pattern: hyperscalers win on breadth of managed services and enterprise compliance; specialists win on price, simplicity, or a specific region/ecosystem. A common architecture is “hyperscaler for the managed database and compliance, specialist for the CDN or object storage where egress dominates.”

CLI quick taste

# AWS: launch a VM
aws ec2 run-instances --image-id ami-0abcdef1234567890 \
  --instance-type t3.micro --count 1

# Azure: create a VM
az vm create --resource-group myRG --name myVM \
  --image Ubuntu2204 --size Standard_B1s

# GCP: create a VM
gcloud compute instances create my-vm \
  --machine-type=e2-micro --zone=asia-southeast1-a

These one-liners are for learning only. In practice you never create production infrastructure imperatively — you declare it in Terraform/OpenTofu or the provider’s native IaC so it’s reviewable, version-controlled, and reproducible. The CLI is for inspection, break-glass fixes, and scripting, not for standing up your stack by hand.

Account and organization structure

Every provider offers a hierarchy for isolating environments and applying policy centrally:

ConceptAWSAzureGCP
Top containerOrganizationTenant (Entra ID)Organization
GroupingOrganizational Units (OUs)Management GroupsFolders
Isolation boundaryAccountSubscriptionProject
GuardrailsService Control Policies (SCPs)Azure PolicyOrganization Policies

The standard pattern is a landing zone: separate accounts/subscriptions/projects per environment (dev, staging, prod) and often per team, with a central “management” account for billing and logging and a “security” account for audit trails. This gives blast-radius isolation (a mistake in dev can’t touch prod), clean cost attribution, and a place to enforce org-wide guardrails. AWS Control Tower, Azure Landing Zones, and GCP’s landing zone blueprints automate the setup.

How to choose a provider

  1. Existing ecosystem — heavy Microsoft shop (AD, Office 365)? Azure integrates naturally and often comes bundled in enterprise agreements. Data/ML focused? GCP’s BigQuery and Vertex AI are strong. Broadest catalog and market share (largest talent pool)? AWS.
  2. Team expertise — hiring and existing knowledge often outweigh technical differences; a team fluent in one cloud ships faster than the same team fighting an unfamiliar “better” one.
  3. Region availability — check latency and data-residency requirements for your users; confirm the specific services and instance types you need actually exist in your target region.
  4. Pricing for your workload — model your actual usage; egress fees and managed-service premiums differ significantly. Use each provider’s pricing calculator with realistic numbers, not list prices.
  5. Compliance — required certifications (PCI DSS, HIPAA, SOC 2, FedRAMP, ISO 27001, local regulations) per region and per service.
  6. Managed services you’ll actually lean on — the “best” cloud is often the one with the specific managed service that removes the most toil for your workload (BigQuery for analytics, DynamoDB for key-value at scale, Cosmos DB for global multi-model).

Multi-cloud considerations

Best Practices

  1. Use Infrastructure as Code from day one (Terraform/OpenTofu, Pulumi, CloudFormation, Bicep) — clicking in the console creates undocumented, unreproducible, un-reviewable “snowflake” infrastructure that nobody can rebuild after an outage.
  2. Enable MFA and never use the root/owner account for daily work — create scoped IAM users/roles; lock away the root/global-admin account, put it behind MFA, and use it only for break-glass tasks that genuinely require it.
  3. Apply least privilege in IAM — start with narrow permissions and widen as needed; prefer roles and short-lived credentials over long-lived access keys; audit with tools like AWS IAM Access Analyzer, Azure PIM, and GCP Policy Analyzer.
  4. Set billing alerts and budgets immediately — surprise bills are the most common newcomer pain; alert at 50/80/100% of expected spend, and enable anomaly detection so a runaway loop pages you in hours, not at month-end.
  5. Deploy across at least two AZs for anything production — single-AZ deployments turn routine data-center maintenance into your outage; multi-AZ is the cheapest resilience you will ever buy.
  6. Tag every resource (team, environment, service, cost-center, owner) — tags power cost allocation, cleanup automation, ownership tracking, and security policy; enforce tagging with policy-as-code so untagged resources are blocked or flagged.
  7. Prefer managed services for undifferentiated heavy lifting — running your own PostgreSQL, Kafka, or Redis on VMs rarely beats the managed equivalent once you factor in backups, patching, failover, and the on-call burden. Reserve self-management for cases with a concrete cost or control justification.
  8. Watch egress and cross-AZ costs — data into the cloud is free, data out is not; keep chatty traffic within one AZ/region, use private endpoints instead of routing through the public internet, and put a CDN in front of static content.
  9. Use spot/preemptible instances for interruptible workloads — 60–90% savings for CI, batch, ETL, and stateless fault-tolerant services; design them to checkpoint and drain gracefully on the reclamation warning.
  10. Buy reservations/savings plans only after usage stabilizes — measure 2–3 months of steady-state usage first, then commit to your baseline; leave the spiky top of the curve on on-demand or spot.
  11. Right-size continuously — the biggest line item in most bills is oversized, under-utilized compute bought “to be safe”; review utilization and shrink monthly.
  12. Keep an organization/landing-zone structure — separate accounts/projects/subscriptions per environment for blast-radius isolation, clean cost attribution, and centralized guardrails (SCPs/Azure Policy/Org Policies).
  13. Encrypt by default and manage keys deliberately — enable encryption at rest and in transit everywhere (it’s usually free and on by default); use the provider KMS, and use customer-managed keys where compliance requires control over the key lifecycle.
  14. Centralize logging and audit trails — turn on CloudTrail / Azure Activity Log / Cloud Audit Logs in a dedicated, write-once account so you have an immutable record of who did what when — indispensable during incidents and audits.
  15. Automate cleanup of ephemeral resources — dev/test environments, orphaned volumes, unattached IPs, and old snapshots quietly accrue cost; schedule teardown and run a “cost janitor” (e.g., cloud-nuke, provider recommendations) regularly.
  16. Learn concepts, not consoles — VPCs, IAM, object storage, load balancing, and identity federation transfer across all providers; console layouts and exact button names do not, and they change constantly.

References