Tổng quan & Hạ tầng toàn cầuOverview & Global Infrastructure
Part of the Cloud knowledge base — AWS deep dive. Complements DevOps Roadmap.
Tổng quan
Amazon Web Services (AWS) là nhà cung cấp cloud lớn nhất theo thị phần, cung cấp khoảng 200 dịch vụ trải rộng từ compute, storage, networking, database, analytics, machine learning đến các công cụ cho developer. Ra mắt năm 2006 với S3 và EC2, AWS tiên phong mô hình hạ tầng pay-as-you-go và đến nay vẫn là hình mẫu tham chiếu mà hầu hết các cloud khác đem ra so sánh.
Trang này là tấm bản đồ bạn nên mở sẵn khi học phần còn lại của AWS. Nó bao quát những mảnh nằm bên dưới mọi dịch vụ: dấu chân vật lý (Region, Availability Zone, edge location), mô hình account và organization giúp cô lập workload, shared responsibility model xác định ai bảo mật cái gì, Well-Architected Framework đặt tên cho các đánh đổi, các cách bạn thực sự “nói chuyện” với API (Console, CLI, SDK, CloudShell), và các đòn bẩy giá quyết định hóa đơn của bạn. Nắm vững những thứ này thì mọi dịch vụ riêng lẻ — EC2, Lambda, RDS — chỉ còn là biến thể của những chủ đề bạn đã hiểu.
Chuyển đổi tư duy quan trọng nhất với AWS là: mọi thứ đều là một API call được IAM cấp quyền, chạy trong một Region, và tính tiền theo mức sử dụng. Console chỉ là một website thay bạn gọi các API đó. Khi điều này “thông”, nền tảng không còn giống 200 sản phẩm rời rạc mà trở thành một hệ thống nhất quán với một catalog dịch vụ rất lớn.
Vì sao hạ tầng toàn cầu phải hiểu trước tiên
Nơi tài nguyên của bạn thực sự “sống” quyết định latency, tuân thủ về data-residency, câu chuyện resilience, và một phần bất ngờ lớn trong hóa đơn. Một workload ở us-east-1 hành xử khác với chính workload đó ở ap-southeast-1 — khác giá, khác mức khả dụng dịch vụ, khác blast radius khi có sự cố. Bạn không thể thiết kế một hệ thống đáng tin cậy hay tối ưu chi phí nếu chưa hiểu địa lý nơi nó chạy — đó là lý do đây là chủ đề đầu tiên, không phải phần phụ.
Kiến thức nền tảng
Region, Availability Zone và edge
AWS tổ chức hạ tầng vật lý theo một phân cấp chặt chẽ:
| Tầng | Là gì | Số lượng (xấp xỉ) | Cô lập lỗi |
|---|---|---|---|
| Region | Một khu vực địa lý (ví dụ us-east-1 N. Virginia, eu-west-1 Ireland, ap-southeast-1 Singapore) | ~34 | Cả region có thể lỗi; độc lập với region khác |
| Availability Zone (AZ) | Một hoặc nhiều data center riêng biệt trong Region, có nguồn điện, làm mát, network độc lập | 3–6 mỗi region (~108) | Một AZ lỗi không kéo sập AZ khác |
| Edge location / PoP | Site nhỏ cho CloudFront CDN, Route 53 DNS, Global Accelerator | 600+ | Chỉ cache nội dung, không phải compute đầy đủ |
| Local Zone | Mở rộng region đặt tại đô thị để latency vài mili-giây | ~30+ đô thị | Chạy tập con dịch vụ gần người dùng |
| Outposts | Rack do AWS quản lý đặt trong data center của bạn | on-prem | Hybrid — mở rộng Region vào tòa nhà của bạn |
| Wavelength | Compute AWS nhúng trong mạng 5G nhà mạng | vài nhà mạng | Latency cực thấp cho mobile/edge |
Các quy tắc rút ra từ phân cấp này:
- Tên Region là tiền tố; tên AZ thêm một chữ cái —
us-east-1a,us-east-1b,us-east-1c. Quan trọng: chữ cái AZ được ngẫu nhiên hóa theo account:us-east-1acủa bạn có thể là data center vật lý khác vớius-east-1acủa account khác. AWS xáo trộn để khách hàng không dồn hết vào zone “đầu tiên”. Dùng AZ ID (use1-az1) khi cần tham chiếu ổn định giữa các account. - AZ là đơn vị của resilience. Triển khai trên ≥2 AZ thì lỗi cấp data center trở thành chuyện nhỏ. Đây là mức reliability rẻ nhất bạn từng mua — hầu hết managed service (RDS Multi-AZ, ELB, EKS) đều AZ-aware mặc định.
- Traffic trong cùng AZ miễn phí; traffic cross-AZ và cross-region đều tính tiền. Một mạng microservice “nói nhiều” trải bừa qua nhiều AZ có thể tạo hóa đơn data-transfer bất ngờ.
us-east-1(N. Virginia) đặc biệt. Là region cũ nhất, lớn nhất, rẻ nhất; nhiều dịch vụ global có control plane ở đó (IAM, CloudFront, Route 53, Organizations, ACM cho cert CloudFront). Sự cốus-east-1trong lịch sử có blast radius lớn bất thường vì quá nhiều bộ máy global nằm ở đó.
Global vs Regional service
Điểm này khiến người mới thường xuyên vấp. Một số dịch vụ AWS là global (một namespace dùng chung mọi region), số khác là regional (cô lập theo từng region, và bạn phải chọn nơi chúng chạy).
| Dịch vụ Global | Dịch vụ Regional |
|---|---|
| IAM (user, role, policy) | EC2, EBS, VPC |
| Route 53 (DNS) | S3 (bucket là regional dù namespace global) |
| CloudFront (CDN) | RDS, DynamoDB, Lambda |
| AWS Organizations | SQS, SNS, ECS, EKS |
| WAF (cho CloudFront), Shield | CloudWatch, KMS key |
| Billing / Cost Explorer | Secrets Manager, ELB |
Hệ quả thực tế: một IAM role hoạt động ở mọi region, nhưng một EC2 security group chỉ tồn tại ở region bạn tạo. S3 là ca tinh tế — namespace bucket là global (tên bucket duy nhất toàn cầu), nhưng dữ liệu bucket nằm ở một region bạn chọn. AMI, snapshot, key pair đều là regional và phải copy để dùng ở nơi khác.
Mô hình account và AWS Organizations
Một AWS account là đơn vị cơ bản của cô lập, billing và bảo mật trong AWS. Mọi thứ — tài nguyên, IAM, quota, billing — đều gắn phạm vi vào một account. Best practice hiện đại nhấn mạnh multi-account: bạn không chạy dev, staging, prod trong cùng một account.
AWS Organizations cho phép quản lý tập trung nhiều account:
- Organizational Unit (OU) — nhóm account theo phân cấp (ví dụ OU
Prod, OUNonProd, OUSecurity, OUSandbox). Bạn gắn policy ở cấp OU và nó lan xuống các account thành viên. - Service Control Policy (SCP) — lan can bảo vệ toàn org đặt ra mức quyền tối đa mà bất kỳ account (hay IAM user/role của nó) có thể có. SCP là bộ lọc, không phải cấp quyền: nó chỉ có thể loại bỏ quyền, không bao giờ thêm. Ngay cả root user của account cũng bị SCP ràng buộc. Dùng phổ biến: cấm tắt CloudTrail, cấm rời org, giới hạn vào các region được duyệt, chặn xóa tài nguyên bảo mật.
- Consolidated billing — mọi account thành viên gộp về một payer account với một hóa đơn duy nhất. Chiết khấu theo khối lượng và lợi ích Reserved Instance/Savings Plan được gộp và chia sẻ toàn org, thường tiết kiệm hơn so với account riêng lẻ.
- Management (payer) account — account đã tạo ra org. Giữ nó gần như rỗng: không workload, khóa chặt, chỉ dùng để quản trị org và billing. Đây là account duy nhất SCP không hạn chế được, nên là mục tiêu giá trị cao nhất.
AWS Control Tower tự động thiết lập môi trường multi-account well-architected — một landing zone — với mặc định hợp lý: một management account, một Log Archive account (tập trung log CloudTrail/Config), một Audit account (công cụ bảo mật), OU dựng sẵn, và “guardrails” (SCP + Config rule được quản lý). Đây là con đường nhanh nhất đến một org tuân theo khuyến nghị của chính AWS thay vì tự dựng bằng tay.
Shared responsibility model
AWS bảo mật cloud; bạn bảo mật những gì bạn đặt vào trong đó. Ranh giới dịch chuyển theo mức “managed” của dịch vụ.
| AWS chịu trách nhiệm (“security of the cloud”) | Bạn chịu trách nhiệm (“security in the cloud”) |
|---|---|
| Data center vật lý, phần cứng | IAM identity, role, policy |
| Hypervisor, host OS, ảo hóa | Vá lỗi guest OS (EC2), cấu hình firewall |
| Phần mềm managed service (engine RDS, runtime Lambda) | Phân loại dữ liệu, chọn mã hóa, key policy |
| Backbone mạng toàn cầu | Security Group, NACL, phân đoạn network |
| Tính khả dụng của chính dịch vụ | Code ứng dụng, secret, cấu hình truy cập |
Với IaaS (EC2) bạn vá OS và cấu hình firewall; AWS chỉ đảm bảo phần cứng và hypervisor. Với managed service (RDS, ECS) AWS vá engine/runtime, nhưng dữ liệu, kiểm soát truy cập và mức phơi bày network vẫn của bạn. Với serverless hoàn toàn (S3, Lambda, DynamoDB) AWS lo gần hết, nhưng bạn vẫn sở hữu việc ai có quyền truy cập và bạn lưu gì. Đại đa số sự cố bảo mật AWS là do khách hàng cấu hình sai — bucket S3 công khai, IAM policy quá rộng, access key rò rỉ — chứ không phải AWS lỗi.
Well-Architected Framework — 6 trụ cột
Khung của AWS để đánh giá kiến trúc theo một tập đánh đổi nhất quán. Biết các trụ cột cho bạn từ vựng để lập luận về mọi thiết kế.
| Trụ cột | Câu hỏi cốt lõi | Đòn bẩy ví dụ |
|---|---|---|
| Operational Excellence | Vận hành và cải tiến trơn tru không? | IaC, runbook, observability, thay đổi nhỏ thường xuyên |
| Security | Dữ liệu và truy cập có được bảo vệ? | IAM least privilege, mã hóa, CloudTrail, GuardDuty |
| Reliability | Có phục hồi được sau lỗi? | Multi-AZ, health check, backup, auto-recovery |
| Performance Efficiency | Dùng đúng tài nguyên hiệu quả không? | Right-sizing, serverless, caching, tầng storage phù hợp |
| Cost Optimization | Có tránh chi tiêu thừa? | Savings Plans, spot, lifecycle policy, tagging |
| Sustainability | Có giảm thiểu tác động môi trường? | Region hiệu quả, right-sizing, managed service, Graviton |
Well-Architected Tool (miễn phí, trong console) dẫn bạn qua bộ câu hỏi theo từng trụ cột và đánh dấu các lỗ hổng rủi ro cao. Sustainability được thêm năm 2021 làm trụ cột thứ sáu. Giá trị thật của khung là như một lăng kính review — trước khi ship, hãy hỏi một câu khó từ mỗi trụ cột và bạn sẽ bắt được hầu hết lỗi kiến trúc.
Khái niệm chính
Cách bạn truy cập AWS
Chỉ có một control plane (AWS API); mọi thứ bên dưới là client gọi nó.
| Cách | Là gì | Tốt cho |
|---|---|---|
| Management Console | Web UI tại console.aws.amazon.com | Học, khám phá, tác vụ một lần, dashboard |
| AWS CLI | Công cụ dòng lệnh (aws ...) | Scripting, automation, break-glass, CI |
| SDK | Thư viện theo ngôn ngữ (boto3 cho Python, aws-sdk cho JS…) | Code ứng dụng gọi AWS |
| CloudShell | Shell trên trình duyệt, đã xác thực CLI sẵn | Lệnh nhanh không cần cài đặt cục bộ |
| IaC (CloudFormation/CDK/Terraform) | Hạ tầng khai báo | Mọi thứ production — review được, tái lập được |
# Cấu hình credential một lần (cho CLI/SDK). Ưu tiên SSO hơn key dài hạn.
aws configure sso # hiện đại: IAM Identity Center
# hoặc, key dài hạn kiểu cũ (tránh dùng cho người):
aws configure # nhập Access Key, Secret, region
# Tôi là ai? — lệnh debug hữu ích nhất trong AWS.
aws sts get-caller-identity
# {
# "UserId": "AROA...:alice",
# "Account": "123456789012",
# "Arn": "arn:aws:sts::123456789012:assumed-role/Admin/alice"
# }
# Liệt kê bucket S3, instance EC2 trong một region
aws s3 ls
aws ec2 describe-instances --region ap-southeast-1 \
--query 'Reservations[].Instances[].[InstanceId,State.Name,InstanceType]' \
--output table
aws sts get-caller-identity trả lời “tôi đang thực sự vận hành với account và identity nào ngay lúc này?” — điều đầu tiên cần kiểm tra khi một lệnh bí ẩn trả về Access Denied hoặc chạm nhầm môi trường.
ARN — địa chỉ tài nguyên phổ quát
Mọi tài nguyên AWS có một Amazon Resource Name (ARN), định danh duy nhất toàn cầu với cú pháp cố định:
arn:partition:service:region:account-id:resource-type/resource-id
arn:aws:s3:::my-bucket/data/file.txt # S3: ARN không có region/account
arn:aws:iam::123456789012:role/DeployRole # IAM: global, nên region để trống
arn:aws:ec2:ap-southeast-1:123456789012:instance/i-0abc123
arn:aws:lambda:us-east-1:123456789012:function:my-fn
Bạn tham chiếu tài nguyên bằng ARN xuyên suốt IAM policy, nên đọc ARN trôi chảy — nhận ra service, region, account — là kỹ năng cốt lõi. Chú ý dịch vụ global (IAM, S3) để trống trường region.
Kiến thức nền về giá
Giá AWS là trả theo mức dùng, nhưng cách bạn cam kết thay đổi mức giá đáng kể:
| Mô hình | Cách hoạt động | Chiết khấu vs on-demand | Tốt cho |
|---|---|---|---|
| On-Demand | Trả theo giây/giờ, không cam kết | gốc (0%) | Workload dao động, khó đoán, ngắn hạn, hoặc mới |
| Savings Plans | Cam kết $/giờ chi tiêu compute 1–3 năm | tới ~72% | Compute baseline ổn định (linh hoạt EC2/Fargate/Lambda) |
| Reserved Instances (RI) | Cam kết một instance family/region 1–3 năm | tới ~72% | Instance ổn định, đã biết rõ (nhất là RDS, ElastiCache) |
| Spot Instances | Đấu giá năng lực dư, thu hồi với báo trước 2 phút | tới ~90% | Batch chịu lỗi, CI, stateless, ML có checkpoint |
| Free Tier | Ưu đãi 12 tháng, always-free, và dùng thử | 100% (giới hạn) | Học và prototype |
Hướng dẫn giúp tiết kiệm tiền thật:
- Compute Savings Plans thường là cam kết mặc định tốt nhất — áp dụng tự động cho EC2, Fargate và Lambda theo đô-la chi tiêu, nên bạn không bị khóa vào một instance family như RI chuẩn.
- Chỉ mua cam kết sau khi usage ổn định (2–3 tháng steady state). Cam kết cho baseline, để phần đỉnh dao động cho on-demand hoặc spot.
- Spot cần kiến trúc, không chỉ một checkbox — xử lý thông báo gián đoạn 2 phút, trải qua nhiều instance type và AZ, không bao giờ chạy database primary stateful trên spot.
- Data egress là cú sốc hóa đơn kinh điển. Ingress miễn phí; egress ra internet và traffic cross-region đều tính tiền. “Check in thì rẻ, check out thì đắt.”
- Free Tier có ba loại: miễn phí 12 tháng (750 giờ/tháng EC2 t2/t3.micro, 5 GB S3), always-free (1M request Lambda/tháng, 25 GB DynamoDB), và dùng thử ngắn. Đặt billing alarm ngay ngày mở account — free tier có giới hạn, vượt qua trong im lặng là bất ngờ số 1 của người mới.
Mô hình tư duy về blast radius
- Một instance có thể lỗi → triển khai ≥2 sau một load balancer.
- Một AZ có thể lỗi → trải qua ≥2 AZ (mặc định cho prod).
- Một Region có thể lỗi → multi-region cho hệ thống tier-0 (đắt, phức tạp).
- Một global control plane (IAM, Route 53) hiếm khi lỗi → đây là lý do redundancy toàn cầu thật sự đôi khi nghĩa là multi-cloud, và vì sao các phụ thuộc vào
us-east-1đáng được soi kỹ.
Best Practices
-
Không bao giờ dùng root account cho công việc hằng ngày. Root user có thể đóng account và bỏ qua hầu hết kiểm soát. Bật MFA phần cứng hoặc ảo, xóa hoàn toàn access key của nó, cất credential trong két vật lý, và chỉ dùng cho vài tác vụ thực sự cần root (đổi phương thức thanh toán, đóng account). Mọi thứ khác qua IAM Identity Center hoặc IAM role.
-
Áp dụng chiến lược multi-account từ đầu. Account riêng theo môi trường (dev/staging/prod) và thường theo team cho bạn cô lập blast radius cứng, quy chi phí sạch, và ranh giới bảo mật độc lập. Một account cho tất cả là “foot-gun” biến lỗi ở dev thành sự cố ở prod.
-
Dùng AWS Organizations với SCP làm guardrail. SCP thực thi các bất biến toàn org mà không ai — kể cả admin account — vi phạm được: cấm rời org, cấm tắt CloudTrail/Config, giới hạn vào region được duyệt. Guardrail ngăn cả lớp lỗi thay vì bắt sau khi đã xảy ra.
-
Khởi tạo bằng Control Tower hoặc blueprint landing-zone. Tự dựng thiết lập multi-account tuân thủ dễ sai; Control Tower cho bạn Log Archive account, Audit account, OU nền, và guardrail được quản lý ngay từ đầu, khớp với khuyến nghị của chính AWS.
-
Đặt billing alarm và budget ngay ngày đầu. Hóa đơn bất ngờ là nỗi đau phổ biến nhất của người mới. Cấu hình AWS Budgets cảnh báo ở 50/80/100% mức chi dự kiến và bật Cost Anomaly Detection để tài nguyên chạy lồng gọi bạn trong vài giờ, không phải cuối tháng.
-
Tag mọi thứ, và thực thi tagging bằng policy. Một schema tag nhất quán (
Environment,Owner,CostCenter,Service) hỗ trợ phân bổ chi phí, dọn dẹp tự động và security policy. Dùng Tag Policy và SCP để tài nguyên không tag bị đánh dấu hoặc chặn — tag hồi tố không bao giờ thực sự xảy ra. -
Ưu tiên IAM Identity Center (SSO) hơn IAM user dài hạn. Truy cập của con người nên ngắn hạn, federated và quản trị tập trung. Access key dài hạn rò rỉ, không hết hạn, và là nguyên nhân gốc của phần lớn breach. Chỉ để dành IAM user cho các ca legacy hoặc lập trình thực sự không thể federate.
-
Chọn region có chủ đích theo bốn trục: latency đến người dùng, luật data-residency, mức khả dụng dịch vụ/tính năng (tính năng và dịch vụ mới về
us-east-1trước và trễ ở nơi khác), và giá (region chênh 20–50% cho cùng tài nguyên). Đừng mặc địnhus-east-1chỉ vì tutorial làm thế. -
Thiết kế chịu lỗi Availability Zone theo mặc định. Trải mọi workload production qua ≥2 AZ — đây là khoản đầu tư reliability rẻ nhất trong AWS. Để dành chi phí và độ phức tạp lớn hơn nhiều của multi-region cho hệ thống thực sự tier-0.
-
Chạy mọi thứ qua Infrastructure as Code. Tài nguyên “snowflake” tạo bằng console không có tài liệu, không tái lập, không review được. Khai báo hạ tầng trong CloudFormation, CDK hoặc Terraform để nó được version-control, peer-review, và dựng lại được sau thảm họa. Dùng console và CLI để inspect và break-glass, không phải để dựng production.
-
Tập trung log vào một account chuyên dụng, khóa chặt. Gộp CloudTrail, Config và VPC Flow Logs vào một Log Archive account ghi-một-lần mà ngay cả admin ở account khác cũng không sửa được. Dấu vết audit bất biến này là không thể thiếu khi có sự cố và khi audit tuân thủ.
-
Mã hóa theo mặc định và quản lý key bằng KMS. Bật mã hóa at-rest và in-transit ở mọi nơi (thường miễn phí và bật sẵn). Dùng AWS KMS, và áp dụng customer-managed key ở nơi tuân thủ đòi hỏi kiểm soát vòng đời key và dấu vết audit mọi lần dùng.
-
Right-size và mua cam kết sau khi đo. Khoản lớn nhất trong hầu hết hóa đơn AWS là compute quá cỡ, kém dùng, mua “cho chắc”. Xem xét mức dùng bằng Compute Optimizer và Cost Explorer, thu nhỏ hằng tháng, rồi cam kết Savings Plans cho baseline ổn định bạn đã đo.
-
Review kiến trúc theo các trụ cột Well-Architected trước khi ship. Hỏi một câu khó từ mỗi trong sáu trụ cột — operational excellence, security, reliability, performance, cost, sustainability. Dùng Well-Architected Tool miễn phí để lộ sớm các lỗ hổng rủi ro cao, khi còn rẻ để sửa.
-
Bật GuardDuty, Security Hub và Config toàn org. Các managed service này liên tục phát hiện mối đe dọa, cấu hình sai và trôi dạt khỏi best practice trên mọi account. Bật sớm rẻ hơn nhiều so với phát hiện một key bị xâm phạm hay một database công khai sau khi sự đã rồi.
Tài liệu tham khảo
- AWS Documentation home
- AWS Global Infrastructure — Regions and AZs
- AWS Global Infrastructure map
- Regional and global AWS services (fault isolation whitepaper)
- AWS Organizations User Guide
- Service Control Policies (SCPs)
- AWS Control Tower User Guide
- Organizing your AWS environment using multiple accounts (whitepaper)
- AWS Shared Responsibility Model
- AWS Well-Architected Framework
- Well-Architected Framework — the six pillars
- AWS CLI documentation
- AWS CloudShell
- Amazon Resource Names (ARNs)
- AWS Pricing overview · Savings Plans · Spot Instances
- AWS Free Tier
- AWS Pricing Calculator
- AWS Architecture Center
Part of the Cloud knowledge base — AWS deep dive. Complements DevOps Roadmap.
Overview
Amazon Web Services (AWS) is the largest cloud provider by market share, offering ~200 services spanning compute, storage, networking, databases, analytics, machine learning, and developer tooling. Launched in 2006 with S3 and EC2, AWS pioneered the pay-as-you-go infrastructure model and remains the reference implementation most other clouds are measured against.
This page is the map you keep open while learning the rest of AWS. It covers the pieces that sit underneath every service: the physical footprint (Regions, Availability Zones, edge locations), the account and organization model that isolates your workloads, the shared responsibility model that defines who secures what, the Well-Architected Framework that names the trade-offs, the ways you actually talk to the API (Console, CLI, SDK, CloudShell), and the pricing levers that decide your bill. Master these and every individual service — EC2, Lambda, RDS — becomes a variation on themes you already understand.
The single most important mental shift for AWS is that everything is an API call authorized by IAM, running in a Region, billed by usage. The console is just a website that makes those API calls for you. Once that clicks, the platform stops feeling like 200 unrelated products and starts feeling like one consistent system with a very large service catalog.
Why the global infrastructure matters before anything else
Where your resources physically live determines your latency, your data-residency compliance, your resilience story, and a surprisingly large share of your bill. A workload in us-east-1 behaves differently from the same workload in ap-southeast-1 — different price, different service availability, different failure blast radius. You cannot design a reliable or cost-effective system without first understanding the geography it runs on, which is why this is the first topic, not an afterthought.
Fundamentals
Regions, Availability Zones, and edge
AWS organizes its physical infrastructure into a strict hierarchy:
| Layer | What it is | Count (approx.) | Failure isolation |
|---|---|---|---|
| Region | A geographic area (e.g., us-east-1 N. Virginia, eu-west-1 Ireland, ap-southeast-1 Singapore) | ~34 | A whole region can fail; independent from other regions |
| Availability Zone (AZ) | One or more discrete data centers in a Region with independent power, cooling, networking | 3–6 per region (~108 total) | An AZ can fail without taking down siblings |
| Edge location / PoP | Small sites for CloudFront CDN, Route 53 DNS, and Global Accelerator | 600+ | Content-caching only, not full compute |
| Local Zone | Region extension in a metro for single-digit-ms latency | ~30+ metros | Runs a subset of services close to users |
| Outposts | AWS-managed racks in your data center | on-prem | Hybrid — extends a Region into your building |
| Wavelength | AWS compute embedded in telco 5G networks | select carriers | Ultra-low-latency mobile/edge |
Key rules that flow from this hierarchy:
- A Region name is a prefix; AZ names append a letter —
us-east-1a,us-east-1b,us-east-1c. Critically, AZ letters are randomized per account: yourus-east-1amay be a different physical data center than another account’sus-east-1a. AWS scrambles them so customers don’t all pile into the “first” zone. Use AZ IDs (use1-az1) when you need a stable cross-account reference. - AZs are the unit of resilience. Deploy across ≥2 AZs and a data-center failure is a non-event. This is the cheapest reliability you will ever buy — most AWS managed services (RDS Multi-AZ, ELB, EKS) are AZ-aware by default.
- Data transfer within an AZ is free; cross-AZ and cross-region traffic is metered. A chatty microservice mesh spread carelessly across AZs can generate a surprising data-transfer bill.
us-east-1(N. Virginia) is special. It is the oldest, largest, and cheapest region; many global services have their control plane there (IAM, CloudFront, Route 53, Organizations, ACM for CloudFront certs). Aus-east-1outage has historically had outsized blast radius because so much global machinery lives there.
Global vs Regional services
This distinction trips up newcomers constantly. Some AWS services are global (one namespace shared across all regions), others are regional (isolated per region, and you must pick where they run).
| Global services | Regional services |
|---|---|
| IAM (users, roles, policies) | EC2, EBS, VPC |
| Route 53 (DNS) | S3 (buckets are regional despite a global namespace) |
| CloudFront (CDN) | RDS, DynamoDB, Lambda |
| AWS Organizations | SQS, SNS, ECS, EKS |
| WAF (for CloudFront), Shield | CloudWatch, KMS keys |
| Billing / Cost Explorer | Secrets Manager, ELB |
Practical consequences: an IAM role works in every region, but an EC2 security group only exists in the region you created it. S3 is a subtle case — the bucket namespace is global (bucket names are globally unique), but the bucket data lives in one region you choose. AMIs, snapshots, and key pairs are regional and must be copied to use them elsewhere.
The account model and AWS Organizations
An AWS account is the fundamental unit of isolation, billing, and security in AWS. Everything — resources, IAM, quotas, billing — is scoped to an account. Modern best practice is emphatically multi-account: you do not run dev, staging, and prod in one account.
AWS Organizations lets you centrally manage many accounts:
- Organizational Units (OUs) — hierarchical groupings of accounts (e.g.,
ProdOU,NonProdOU,SecurityOU,SandboxOU). You attach policies at the OU level and they cascade to member accounts. - Service Control Policies (SCPs) — organization-wide guardrails that set the maximum permissions any account (or its IAM users/roles) can ever have. An SCP is a filter, not a grant: it can only remove permissions, never add them. Even the account root user is bound by SCPs. Classic uses: deny disabling CloudTrail, deny leaving the org, restrict to approved regions, block deletion of security resources.
- Consolidated billing — all member accounts roll up to one payer account with a single bill. Volume discounts and Reserved Instance/Savings Plan benefits are pooled and shared across the whole org, which often saves money versus separate accounts.
- Management (payer) account — the account that created the org. Keep it nearly empty: no workloads, tightly locked down, used only for org administration and billing. It is the one account SCPs cannot restrict, so it is your highest-value target.
AWS Control Tower automates setting up a well-architected multi-account environment — a landing zone — with sensible defaults: a management account, a Log Archive account (centralized CloudTrail/Config logs), an Audit account (security tooling), pre-built OUs, and “guardrails” (managed SCPs + Config rules). It is the fastest path to an org that follows AWS’s own recommendations rather than hand-rolling one.
The shared responsibility model
AWS secures the cloud; you secure what you put in it. The dividing line shifts with how managed the service is.
| AWS is responsible for (“security of the cloud”) | You are responsible for (“security in the cloud”) |
|---|---|
| Physical data centers, hardware | IAM identities, roles, and policies |
| Hypervisor, host OS, virtualization | Guest OS patching (EC2), firewall config |
| Managed-service software (RDS engine, Lambda runtime) | Data classification, encryption choices, key policy |
| Global network backbone | Security Groups, NACLs, network segmentation |
| Availability of the service itself | Application code, secrets, and access configuration |
For IaaS (EC2) you patch the OS and configure the firewall; AWS only guarantees the hardware and hypervisor. For managed services (RDS, ECS) AWS patches the engine/runtime, but data, access control, and network exposure are still yours. For fully serverless (S3, Lambda, DynamoDB) AWS runs almost everything, yet you still own who has access and what you store. The vast majority of AWS security incidents are customer-side misconfigurations — a public S3 bucket, an over-broad IAM policy, a leaked access key — not AWS failures.
The Well-Architected Framework — 6 pillars
AWS’s framework for evaluating architectures against a consistent set of trade-offs. Knowing the pillars gives you the vocabulary to reason about any design.
| Pillar | Core question | Example levers |
|---|---|---|
| Operational Excellence | Can we run and improve this smoothly? | IaC, runbooks, observability, small frequent changes |
| Security | Is data and access protected? | Least privilege IAM, encryption, CloudTrail, GuardDuty |
| Reliability | Does it recover from failure? | Multi-AZ, health checks, backups, auto-recovery |
| Performance Efficiency | Are we using the right resources efficiently? | Right-sizing, serverless, caching, appropriate storage tiers |
| Cost Optimization | Are we avoiding unnecessary spend? | Savings Plans, spot, lifecycle policies, tagging |
| Sustainability | Are we minimizing environmental impact? | Efficient regions, right-sizing, managed services, Graviton |
The Well-Architected Tool (free, in the console) walks you through a questionnaire per pillar and flags high-risk gaps. Sustainability was added in 2021 as the sixth pillar. The framework’s real value is as a review lens — before shipping, ask one hard question from each pillar and you will catch most architectural sins.
Key Concepts
How you access AWS
There is one control plane (the AWS API); everything below is a client that calls it.
| Method | What it is | Best for |
|---|---|---|
| Management Console | The web UI at console.aws.amazon.com | Learning, exploration, one-off tasks, dashboards |
| AWS CLI | Command-line tool (aws ...) | Scripting, automation, break-glass, CI |
| SDKs | Language libraries (boto3 for Python, aws-sdk for JS, etc.) | Application code that calls AWS |
| CloudShell | Browser-based shell with the CLI pre-authenticated | Quick commands without local setup |
| IaC (CloudFormation/CDK/Terraform) | Declarative infrastructure | Everything production — reviewable, reproducible |
# Configure credentials once (for CLI/SDK). Prefer SSO over long-lived keys.
aws configure sso # modern: IAM Identity Center
# or, legacy long-lived keys (avoid for humans):
aws configure # prompts for Access Key, Secret, region
# Who am I? — the single most useful debugging command in AWS.
aws sts get-caller-identity
# {
# "UserId": "AROA...:alice",
# "Account": "123456789012",
# "Arn": "arn:aws:sts::123456789012:assumed-role/Admin/alice"
# }
# List S3 buckets, EC2 instances in a region
aws s3 ls
aws ec2 describe-instances --region ap-southeast-1 \
--query 'Reservations[].Instances[].[InstanceId,State.Name,InstanceType]' \
--output table
aws sts get-caller-identity answers “which account and identity am I actually operating as right now?” — the first thing to check when a command mysteriously returns Access Denied or touches the wrong environment.
ARNs — the universal resource address
Every AWS resource has an Amazon Resource Name (ARN), a globally unique identifier with a fixed grammar:
arn:partition:service:region:account-id:resource-type/resource-id
arn:aws:s3:::my-bucket/data/file.txt # S3: no region/account in ARN
arn:aws:iam::123456789012:role/DeployRole # IAM: global, so region is empty
arn:aws:ec2:ap-southeast-1:123456789012:instance/i-0abc123
arn:aws:lambda:us-east-1:123456789012:function:my-fn
You reference resources by ARN throughout IAM policies, so reading an ARN fluently — spotting the service, region, and account — is a core skill. Note how global services (IAM, S3) leave the region field empty.
Pricing fundamentals
AWS pricing is pay-per-use, but how you commit changes the rate dramatically:
| Model | How it works | Discount vs on-demand | Best for |
|---|---|---|---|
| On-Demand | Pay per second/hour, no commitment | baseline (0%) | Spiky, unpredictable, short-lived, or new workloads |
| Savings Plans | Commit to $/hour of compute spend for 1–3 yrs | up to ~72% | Steady baseline compute (flexible across EC2/Fargate/Lambda) |
| Reserved Instances (RIs) | Commit to a specific instance family/region 1–3 yrs | up to ~72% | Steady, well-known instance shapes (esp. RDS, ElastiCache) |
| Spot Instances | Bid on spare capacity, reclaimed with 2-min notice | up to ~90% | Fault-tolerant batch, CI, stateless, checkpointed ML |
| Free Tier | 12-month, always-free, and trial offers | 100% (limited) | Learning and prototyping |
Guidance that saves real money:
- Compute Savings Plans are usually the best default commitment — they apply automatically across EC2, Fargate, and Lambda by dollar-spend, so you don’t lock yourself to one instance family the way standard RIs do.
- Buy commitments only after usage stabilizes (2–3 months of steady state). Commit to your baseline, leave the spiky top on on-demand or spot.
- Spot needs architecture, not just a checkbox — handle the 2-minute interruption notice, spread across instance types and AZs, never run a stateful primary on spot.
- Data egress is the classic bill shock. Ingress is free; egress to the internet and cross-region traffic is metered. “Cheap to check in, expensive to check out.”
- The Free Tier has three flavors: 12-month free (750 hrs/mo of t2/t3.micro EC2, 5 GB S3), always-free (1M Lambda requests/mo, 25 GB DynamoDB), and short trials. Set a billing alarm the day you open an account — the free tier has limits, and exceeding them silently is the #1 newcomer surprise.
The mental model for blast radius
- A single instance can fail → deploy ≥2 behind a load balancer.
- An AZ can fail → spread across ≥2 AZs (default for prod).
- A Region can fail → multi-region for tier-0 systems (expensive, complex).
- A global control plane (IAM, Route 53) can rarely fail → this is why true global redundancy sometimes means multi-cloud, and why
us-east-1dependencies deserve scrutiny.
Best Practices
-
Never use the root account for daily work. The root user can close the account and bypass most controls. Enable a hardware or virtual MFA on it, remove its access keys entirely, store its credentials in a physical safe, and use it only for the handful of tasks that genuinely require root (changing the account’s payment method, closing the account). Everything else goes through IAM Identity Center or IAM roles.
-
Adopt a multi-account strategy from the start. Separate accounts per environment (dev/staging/prod) and often per team give you hard blast-radius isolation, clean cost attribution, and independent security boundaries. One account for everything is a foot-gun that turns a dev mistake into a prod incident.
-
Use AWS Organizations with SCPs as guardrails. SCPs enforce org-wide invariants nobody — not even an account admin — can violate: deny leaving the org, deny disabling CloudTrail/Config, restrict to approved regions. Guardrails prevent whole classes of mistakes rather than catching them after the fact.
-
Bootstrap with Control Tower or a landing-zone blueprint. Hand-rolling a compliant multi-account setup is error-prone; Control Tower gives you a Log Archive account, an Audit account, baseline OUs, and managed guardrails out of the box that match AWS’s own recommendations.
-
Set billing alarms and budgets on day one. Surprise bills are the most common newcomer pain. Configure AWS Budgets with alerts at 50/80/100% of expected spend and enable Cost Anomaly Detection so a runaway resource pages you in hours, not at month-end.
-
Tag everything, and enforce tagging with policy. A consistent tag schema (
Environment,Owner,CostCenter,Service) powers cost allocation, automated cleanup, and security policy. Use Tag Policies and SCPs so untagged resources are flagged or blocked — retroactive tagging never happens. -
Prefer IAM Identity Center (SSO) over long-lived IAM users. Human access should be short-lived, federated, and centrally governed. Long-lived access keys leak, don’t expire, and are the root cause of a huge share of breaches. Reserve IAM users for legacy or programmatic edge cases that genuinely can’t federate.
-
Choose regions deliberately on four axes: latency to users, data-residency law, service/feature availability (new features and services land in
us-east-1first and lag elsewhere), and price (regions vary 20–50% for the same resource). Don’t default tous-east-1just because tutorials do. -
Design for Availability Zone failure by default. Spread every production workload across ≥2 AZs — it is the single cheapest reliability investment in AWS. Reserve the far greater cost and complexity of multi-region for genuinely tier-0 systems.
-
Run everything through Infrastructure as Code. Console-created “snowflake” resources are undocumented, unreproducible, and un-reviewable. Declare infrastructure in CloudFormation, CDK, or Terraform so it is version-controlled, peer-reviewed, and rebuildable after a disaster. Use the console and CLI for inspection and break-glass, not for standing up production.
-
Centralize logs in a dedicated, locked-down account. Aggregate CloudTrail, Config, and VPC Flow Logs into a write-once Log Archive account that even admins in other accounts cannot tamper with. This immutable audit trail is indispensable during incidents and compliance audits.
-
Encrypt by default and manage keys with KMS. Enable encryption at rest and in transit everywhere (it’s usually free and often on by default). Use AWS KMS, and adopt customer-managed keys where compliance requires control over the key lifecycle and an audit trail of every use.
-
Right-size and buy commitments after measuring. The biggest line item in most AWS bills is oversized, under-utilized compute bought “to be safe.” Review utilization with Compute Optimizer and Cost Explorer, shrink monthly, then commit Savings Plans to the stable baseline you’ve measured.
-
Review architectures against the Well-Architected pillars before shipping. Ask one hard question from each of the six pillars — operational excellence, security, reliability, performance, cost, sustainability. Use the free Well-Architected Tool to surface high-risk gaps early, when they’re cheap to fix.
-
Enable GuardDuty, Security Hub, and Config org-wide. These managed services continuously detect threats, misconfigurations, and drift from best practice across all accounts. Turning them on early is far cheaper than discovering a compromised key or a public database after the fact.
References
- AWS Documentation home
- AWS Global Infrastructure — Regions and AZs
- AWS Global Infrastructure map
- Regional and global AWS services (fault isolation whitepaper)
- AWS Organizations User Guide
- Service Control Policies (SCPs)
- AWS Control Tower User Guide
- Organizing your AWS environment using multiple accounts (whitepaper)
- AWS Shared Responsibility Model
- AWS Well-Architected Framework
- Well-Architected Framework — the six pillars
- AWS CLI documentation
- AWS CloudShell
- Amazon Resource Names (ARNs)
- AWS Pricing overview · Savings Plans · Spot Instances
- AWS Free Tier
- AWS Pricing Calculator
- AWS Architecture Center