IaC & Cost ManagementIaC & Cost Management
Part of the Cloud knowledge base — AWS deep dive. Complements DevOps Roadmap.
Tổng quan
Trang này bàn về hai lĩnh vực mà một kỹ sư DevOps sống chung mỗi ngày: Infrastructure as Code (IaC) — định nghĩa hạ tầng cloud của bạn bằng các file được quản lý phiên bản, có thể review, và lặp lại được thay vì click quanh console — và cost management (quản lý chi phí) — giữ cho khoản chi tiêu phát sinh luôn rõ ràng, được quy trách nhiệm, và được tối ưu. Chúng đi cùng nhau vì chính những tính chất khiến hạ tầng an toàn (declarative, được gắn tag, được review, được tự động hóa) cũng chính là những gì khiến nó rẻ để vận hành và dễ để quy trách nhiệm.
Về phía IaC, AWS cung cấp một bộ công cụ native — CloudFormation (engine provisioning nền tảng), CDK (ngôn ngữ lập trình thật sự, tổng hợp ra CloudFormation), và SAM (một phần mở rộng của CloudFormation tập trung cho serverless) — trong khi thế giới rộng hơn chạy trên Terraform, tiêu chuẩn thực tế cloud-agnostic. Không có một câu trả lời đúng duy nhất: CloudFormation/CDK cho tích hợp AWS sâu nhất và drift detection mà không cần quản lý state file; Terraform cho khả năng chuyển đổi đa cloud, một hệ sinh thái module khổng lồ, và một workflow plan/apply tuyệt vời. Hầu hết các team chọn một công cụ chính và giữ tính nhất quán.
Về phía chi phí, thực tế chi phối là hóa đơn cloud là sản phẩm của các quyết định kỹ thuật, và nó chỉ kiểm soát được nếu nó rõ ràng và có thể quy trách nhiệm. Hạ tầng không gắn tag, không giám sát tạo ra một hóa đơn mà không ai giải thích được và không ai làm chủ. Nghề này là xây dựng một kỷ luật gắn tag, theo dõi chi tiêu bằng Cost Explorer, Budgets, và Anomaly Detection, và khớp các cam kết mua (Savings Plans, Reserved, Spot) với hình dạng workload. Làm tốt, cost management không phải là thắt lưng buộc bụng — mà là chi tiêu có chủ đích vào những gì quan trọng và loại bỏ lãng phí một cách tự động.
Kiến thức nền tảng
Tại sao lại cần IaC
| Manual console | Infrastructure as Code |
|---|---|
| Không tái tạo được; các môi trường snowflake | Dev/stage/prod giống hệt nhau từ một nguồn duy nhất |
| Không có lịch sử; “ai đã đổi cái này?” | Lịch sử Git, review PR, blame |
| Không có dry-run; thay đổi áp dụng trực tiếp | plan / change set xem trước khi apply |
| Drift tích tụ âm thầm | Drift detection và re-apply |
| Onboarding = kiến thức truyền miệng | Repo chính là tài liệu |
IaC biến hạ tầng thành phần mềm: được peer-review, được test, được version, và rollback như bất kỳ đoạn code nào khác. Mọi môi trường AWS nghiêm túc trong năm 2026 đều được định nghĩa theo cách này.
Họ IaC native của AWS
- CloudFormation — dịch vụ nền tảng. Bạn khai báo resource trong các template YAML/JSON; CloudFormation tạo chúng thành một stack mà nó theo dõi và quản lý, xử lý thứ tự phụ thuộc, rollback khi thất bại, và xóa. Không có state file để lưu trữ — AWS giữ state. Nó là engine mà CDK và SAM biên dịch xuống.
- CDK (Cloud Development Kit) — viết hạ tầng bằng TypeScript, Python, Java, Go, hoặc C#. Các construct cấp cao đóng gói các mặc định theo best-practice (một dịch vụ Fargate có load-balancer trong hơn chục dòng), và
cdk synthphát ra một template CloudFormation. Lý tưởng khi bạn muốn vòng lặp, điều kiện, unit test, và hỗ trợ IDE thay vì YAML. - SAM (Serverless Application Model) — một macro/phần mở rộng của CloudFormation với cú pháp ngắn gọn cho Lambda, API Gateway, DynamoDB, và Step Functions, cùng một CLI để test cục bộ (
sam local invoke) và deploy nhanh. Con đường nhẹ nhàng nhất cho các ứng dụng serverless.
Cơ chế cốt lõi của CloudFormation
- Stack — một instance đã deploy của một template; đơn vị của create/update/delete. Xóa một stack và (theo mặc định) mọi resource nó tạo ra đều bị hủy.
- Change set — một bản xem trước những gì một cập nhật sẽ làm (add/modify/replace/remove) trước khi bạn thực thi nó. Cờ “replace” rất quan trọng: một số thay đổi property buộc phải tạo lại resource (và gây downtime / mất dữ liệu).
- StackSets — deploy cùng một stack trên nhiều account và region từ một thao tác duy nhất, được điều khiển bởi AWS Organizations. Đây là cách bạn triển khai các guardrail toàn tổ chức (một bucket logging, một Config conformance pack, một IAM baseline) tới mọi account một cách nhất quán.
- Drift detection — so sánh các resource đang chạy với template và báo cáo các thay đổi ngoài luồng mà ai đó đã thực hiện bằng tay.
- Nested stacks / modules — phân rã các template lớn thành các thành phần tái sử dụng được.
Khái niệm chính
CDK vs. CloudFormation vs. Terraform
| Tiêu chí | CloudFormation | CDK | Terraform |
|---|---|---|---|
| Ngôn ngữ | YAML/JSON | TS/Python/Java/Go/C# | HCL |
| Phạm vi cloud | Chỉ AWS | Chỉ AWS (chủ yếu) | Đa cloud + 3000 provider |
| State | Do AWS quản lý | Do AWS quản lý (qua CFN) | Bạn quản lý (S3 backend) |
| Xem trước | Change sets | cdk diff → change sets | terraform plan |
| Hệ sinh thái | Do AWS xuất bản | Construct Hub | Terraform Registry (khổng lồ) |
| Drift | Detection tích hợp sẵn | Tích hợp sẵn (qua CFN) | plan hiển thị drift khi refresh |
| Phù hợp nhất cho | Thuần AWS, không state ngoài | Thuần AWS với code thật | Đa cloud, tái sử dụng module lớn nhất |
Cách chọn: dốc toàn lực vào AWS với một team không thích quản lý state → CloudFormation hoặc CDK (CDK nếu bạn muốn tính năng ngôn ngữ thật sự). Đa cloud, hoặc bạn coi trọng workflow plan và hệ sinh thái module của Terraform, hoặc bạn đã có sẵn chuyên môn Terraform → Terraform. Ứng dụng serverless-first → SAM. Không có câu trả lời sai; tính nhất quán trong một team thắng những lợi thế biên nhỏ của bất kỳ công cụ nào.
Terraform trên AWS: provider và remote state
Terraform mô tả AWS bằng provider aws và lưu ánh xạ giữa config của bạn với các resource thật trong một state file. Nơi state đó nằm là quyết định vận hành quan trọng nhất: state cục bộ trên một laptop là công thức cho thảm họa (mất file = hạ tầng mồ côi, không có locking = hai kỹ sư làm hỏng apply của nhau). Câu trả lời tiêu chuẩn là một remote backend trong S3 với một bảng lock DynamoDB.
Ví dụ thực tế — Terraform S3 backend với DynamoDB locking:
# backend.tf
terraform {
required_version = ">= 1.6"
required_providers {
aws = {
source = "hashicorp/aws"
version = "~> 5.0"
}
}
backend "s3" {
bucket = "acme-terraform-state-prod" # versioned, encrypted, private
key = "networking/terraform.tfstate" # path per component
region = "us-east-1"
dynamodb_table = "terraform-locks" # state-lock table (LockID hash key)
encrypt = true # SSE on the state object
}
}
provider "aws" {
region = "us-east-1"
default_tags {
tags = {
Environment = "prod"
ManagedBy = "terraform"
CostCenter = "platform"
Owner = "platform-team"
}
}
}
Vì sao mỗi phần lại quan trọng:
- S3 bucket — bền, versioned (để bạn có thể rollback một state bị hỏng), được mã hóa, với public access bị chặn. State có thể chứa secret ở dạng plaintext, nên hãy đối xử với bucket như tài sản nhạy cảm.
key— một state file cho mỗi thành phần logic (networking, data, app) để blast radius và thời gian apply luôn nhỏ; một state đơn khối duy nhất thì chậm và nguy hiểm.- Bảng lock DynamoDB — một bảng đơn giản với partition key
LockID(String). Terraform lấy một lock ở đây trước mỗi lần apply để hai người (hoặc hai lần chạy CI) không thể biến đổi state đồng thời. Lưu ý: các phiên bản Terraform mới hơn cũng cung cấp locking native trên S3 (use_lockfile), giảm nhu cầu dùng DynamoDB, nhưng pattern DynamoDB vẫn là cái được triển khai rộng rãi nhất. default_tags— mọi resource mà provider này tạo ra đều tự động kế thừa các tag này. Đây là xương sống của việc quy trách nhiệm chi phí (xem bên dưới) — đặt một lần và không bao giờ quên một tag nào nữa.
Lưu ý về bootstrap: bucket state và bảng lock bản thân chúng không thể nằm trong state mà chúng làm nền, nên hãy tạo chúng một lần bằng tay hoặc trong một cấu hình “bootstrap” nhỏ riêng biệt.
SAM cho serverless
SAM tỏa sáng cho các ứng dụng lấy Lambda làm trung tâm. Một template.yaml dùng các loại resource AWS::Serverless::* được mở rộng thành CloudFormation đầy đủ:
Transform: AWS::Serverless-2016-10-31
Resources:
CheckoutFn:
Type: AWS::Serverless::Function
Properties:
Handler: app.handler
Runtime: python3.13
MemorySize: 512
Timeout: 10
Events:
Api:
Type: HttpApi
Properties: { Path: /checkout, Method: post }
Policies:
- DynamoDBCrudPolicy: { TableName: !Ref OrdersTable }
sam local invoke chạy function trong một container đối với một test event, sam local start-api dựng API cục bộ, và sam deploy --guided đóng gói và ship nó. Giá trị nằm ở sự ngắn gọn (cú pháp rút gọn Policies đó tạo ra một IAM role có phạm vi hẹp) và vòng phản hồi cục bộ nhanh.
Bộ công cụ quản lý chi phí
| Công cụ | Mục đích | Nhịp độ |
|---|---|---|
| Cost Explorer | Trực quan hóa và phân tích chi tiêu theo service/tag/account; dự báo | Ad hoc / review hàng tuần |
| AWS Budgets | Đặt ngưỡng; cảnh báo (hoặc hành động) khi chi tiêu/sử dụng vượt qua chúng | Liên tục, có cảnh báo |
| Cost Anomaly Detection | ML phát hiện các đợt tăng chi tiêu bất thường và cảnh báo | Liên tục, tự động |
| Cost Allocation Tags | Quy chi phí cho team/env/product qua tag | Nền tảng, liên tục |
| Compute Optimizer | Khuyến nghị right-sizing cho EC2/EBS/Lambda/ASG | Review định kỳ |
| Cost & Usage Report (CUR) | Dữ liệu billing thô, theo từng dòng (tới S3 / Athena / QuickSight) | Cho phân tích sâu / FinOps |
Chiến lược gắn tag
Tag là nền tảng của mọi thứ liên quan đến chi phí — bạn không thể phân bổ, lập ngân sách, hay tối ưu khoản chi tiêu mà bạn không thể quy trách nhiệm. Một bộ tag bắt buộc tối thiểu:
| Tag | Ví dụ | Mục đích |
|---|---|---|
Environment | prod / staging / dev | Chia chi tiêu theo giai đoạn; tìm lãng phí ở dev |
CostCenter / Team | platform / checkout | Charge-back / show-back cho các chủ sở hữu |
Owner | platform-team | Hỏi ai; ai dọn dẹp |
Application / Service | checkout-api | Chi phí theo từng product |
ManagedBy | terraform | Phân biệt resource IaC với resource thủ công (drift) |
Áp đặt tag bằng Tag Policies (Organizations), Service Control Policies (từ chối tạo resource nếu thiếu tag bắt buộc), hoặc default_tags của IaC. Sau đó kích hoạt chúng thành cost allocation tags trong Billing console để chúng xuất hiện trong Cost Explorer và CUR — một tag chỉ trở thành một chiều chi phí sau khi bạn kích hoạt nó, và chỉ từ thời điểm đó trở đi. Lưu ý giá trị tag phân biệt hoa thường (Prod ≠ prod), nên hãy chuẩn hóa cách viết hoa.
Các mô hình mua: Savings Plans vs. Reserved vs. Spot
| Mô hình | Chiết khấu | Cam kết | Linh hoạt | Có thể gián đoạn | Dùng cho |
|---|---|---|---|---|---|
| On-Demand | baseline | không | toàn phần | không | Đột biến, ngắn hạn, không đoán trước |
| Spot | tới ~90% | không | bất kỳ capacity dư nào | Có (báo trước 2 phút) | Batch chịu lỗi, CI, fleet stateless |
| Reserved Instances | tới ~72% | 1 hoặc 3 năm | khóa theo family/region (Standard) | không | Cũ; các instance ổn định cụ thể |
| Compute Savings Plans | tới ~66% | 1 hoặc 3 năm ($/giờ) | bất kỳ region/family/EC2/Fargate/Lambda | không | Chi tiêu baseline ổn định (mặc định hiện đại) |
| EC2 Instance Savings Plans | tới ~72% | 1 hoặc 3 năm ($/giờ) | trong một family/region | không | Chi tiêu ổn định trên một family đã biết |
Luồng quyết định: đo lường baseline ổn định của bạn (lượng compute bạn chạy 24/7 mỗi tháng) và phủ nó bằng Compute Savings Plans (linh hoạt) hoặc EC2 Instance Savings Plans (chiết khấu sâu hơn nếu family ổn định). Chạy công việc có thể gián đoạn, stateless (CI runner, batch, big-data, dev fleet) trên Spot để giảm tới 90%. Để phần đột biến, không đoán trước ở đỉnh trên On-Demand. Cách tiếp cận phân tầng này — Savings Plans cho phần sàn, Spot cho phần đàn hồi ở giữa, On-Demand cho các đỉnh — là cấu trúc chi phí FinOps tiêu chuẩn. Reserved Instances phần lớn đã bị Savings Plans thay thế cho compute, dù RI vẫn áp dụng cho RDS, ElastiCache, Redshift, và OpenSearch.
Budgets và anomaly detection trong thực tế
Một Budget đặt một giới hạn chi phí/sử dụng hàng tháng (hoặc tùy chỉnh) và thông báo qua SNS tại các ngưỡng có thể cấu hình (ví dụ 80% thực tế, 100% thực tế, 100% dự báo). Budget Actions có thể đi xa hơn và áp dụng một SCP hoặc IAM policy để tự động kìm hãm chi tiêu khi một ngưỡng bị chạm. Cost Anomaly Detection bổ sung cho budget: thay vì một ngưỡng cố định, nó học chi tiêu bình thường của mỗi service và cảnh báo về các đợt tăng bất thường về mặt thống kê — bắt được, chẳng hạn, một NAT gateway chạy mất kiểm soát hoặc một S3 bucket vô tình public đang tích lũy request, nhiều ngày trước khi cảnh báo budget hàng tháng kịp báo.
Một lưu ý thực tế về budget + tagging: tạo budget theo từng chiều cost-allocation (theo tag team, theo môi trường) thay vì một budget toàn account, để một cảnh báo rơi vào team đã gây ra nó. Một budget toàn account nói cho bạn biết “chúng ta vượt rồi” nhưng không biết “ai” — các budget theo tag làm cho cảnh báo có thể hành động được và tự động định tuyến trách nhiệm.
Best Practices
- Định nghĩa toàn bộ hạ tầng bằng code; cấm resource production tạo từ console. Lý do: resource thủ công là các snowflake không tái tạo được, không lịch sử hay review, và chúng drift; repo phải là nguồn sự thật duy nhất.
- Luôn xem trước trước khi apply (
terraform plan/ change sets). Lý do: thay đổi hạ tầng có thể thay thế resource và gây downtime hoặc mất dữ liệu; bản xem trước bắt các thay đổi phá hủy trước khi chúng thực thi. - Lưu trữ Terraform state từ xa trong S3 versioned, được mã hóa với locking. Lý do: state cục bộ dễ mất hoặc hỏng, và không có locking thì các apply đồng thời phá hủy state của nhau; S3 versioning cho phép bạn rollback một state xấu.
- Chia state theo thành phần, không phải một khối đơn khối. Lý do: một state file khổng lồ duy nhất khiến mọi apply chậm và biến bất kỳ sai lầm nào thành blast radius toàn môi trường; state theo từng thành phần giữ thay đổi nhỏ và an toàn.
- Đặt default tags ở cấp provider (hoặc Tag Policies) để không gì bị thiếu tag. Lý do: resource không có tag là vô hình đối với cost allocation và tự động dọn dẹp; áp đặt tag lúc tạo là cách duy nhất giữ cho việc quy trách nhiệm được đầy đủ.
- Kích hoạt cost allocation tags sớm và chuẩn hóa cách viết hoa của chúng. Lý do: tag chỉ trở thành chiều chi phí sau khi kích hoạt và chỉ áp dụng về sau, và
Prod≠prodlàm phân mảnh báo cáo của bạn — hãy đặt quy ước trước khi chi tiêu tích lũy. - Phủ baseline ổn định bằng Savings Plans, không phải On-Demand. Lý do: trả On-Demand cho workload 24/7 để lại tới ~66% trên bàn; một Compute Savings Plan 1 năm rủi ro thấp và linh hoạt qua family và region.
- Chạy workload có thể gián đoạn, stateless trên Spot. Lý do: CI, batch, và fleet stateless chịu được việc thu hồi trong 2 phút và được giảm tới 90% — đòn bẩy lớn nhất trên chi phí compute biến đổi.
- Tạo budget theo tag cho từng team/môi trường, không phải một budget toàn account. Lý do: một cảnh báo theo phạm vi định tuyến tới chủ sở hữu đã gây ra khoản vượt và có thể hành động được; một cảnh báo toàn account chỉ nói “chúng ta vượt rồi” mà không có trách nhiệm cụ thể.
- Bật Cost Anomaly Detection trên nền của budget. Lý do: budget bắt một ngưỡng vào cuối tháng; anomaly detection bắt một resource chạy mất kiểm soát trong vòng một ngày, trước khi nó trở thành một hóa đơn bất ngờ lớn.
- Review Compute Optimizer và right-size thường xuyên. Lý do: hầu hết các fleet đều được cấp phát dư; hành động theo khuyến nghị right-sizing và Graviton đáng tin cậy cắt giảm hàng chục phần trăm mà không tốn chi phí về độ tin cậy.
- Giữ secret ra khỏi state và template; tham chiếu Secrets Manager/SSM lúc deploy. Lý do: Terraform state và các parameter CloudFormation có thể để lộ secret plaintext; lấy chúng lúc runtime giữ chúng ra khỏi bucket state và version control.
- Dùng StackSets / Terraform module để áp đặt guardrail toàn tổ chức một cách nhất quán. Lý do: triển khai cùng một baseline (logging, Config pack, IAM) tới mọi account bằng tay đảm bảo drift; triển khai đa account tập trung giữ trạng thái đồng nhất.
- Chạy IaC qua CI với plan-on-PR và apply-on-merge. Lý do: các plan được peer-review bắt sai lầm trước khi chúng đến production và tạo lịch sử thay đổi có thể audit, hệt như review code ứng dụng.
- Ghim phiên bản provider và module. Lý do: phiên bản không ghim âm thầm kéo về các thay đổi phá vỡ trong lần apply tiếp theo; ghim làm cho build tái tạo được và nâng cấp có chủ đích.
- Gắn tag cho resource để tự động hóa vòng đời (TTL/hết hạn trên dev) và xóa hạ tầng nhàn rỗi. Lý do: các môi trường dev bị lãng quên và volume/snapshot/IP mồ côi là lãng phí thuần túy; tự động hóa dọn dẹp theo tag thu hồi chi tiêu liên tục mà không cần con người giám sát.
Tài liệu tham khảo
- AWS CloudFormation User Guide
- CloudFormation StackSets
- AWS CDK Developer Guide
- AWS SAM Developer Guide
- Terraform AWS Provider
- Terraform S3 backend
- AWS Cost Explorer
- AWS Budgets
- AWS Cost Anomaly Detection
- Cost allocation tags
- AWS Savings Plans
- Amazon EC2 Spot Instances
- AWS Compute Optimizer
- AWS Well-Architected — Cost Optimization Pillar
Part of the Cloud knowledge base — AWS deep dive. Complements DevOps Roadmap.
Overview
This page covers two disciplines that a DevOps engineer lives inside every day: Infrastructure as Code (IaC) — defining your cloud in version-controlled, reviewable, repeatable files instead of clicking around the console — and cost management — keeping the resulting spend visible, attributed, and optimized. They belong together because the same properties that make infrastructure safe (declarative, tagged, reviewed, automated) are exactly what make it cheap to operate and easy to attribute.
On the IaC side AWS offers a native stack — CloudFormation (the underlying provisioning engine), CDK (real programming languages that synthesize to CloudFormation), and SAM (a serverless-focused CloudFormation extension) — while the wider world runs on Terraform, the cloud-agnostic de-facto standard. There is no single right answer: CloudFormation/CDK give the deepest AWS integration and drift detection with no state file to manage; Terraform gives multi-cloud portability, a huge module ecosystem, and a superb plan/apply workflow. Most teams pick one primary tool and stay consistent.
On the cost side the governing reality is that the cloud bill is a product of engineering decisions, and it is only controllable if it is visible and attributable. Un-tagged, unmonitored infrastructure produces a bill nobody can explain and nobody owns. The craft is building a tagging discipline, watching spend with Cost Explorer, Budgets, and Anomaly Detection, and matching purchase commitments (Savings Plans, Reserved, Spot) to workload shape. Done well, cost management is not austerity — it is spending deliberately on what matters and eliminating waste automatically.
Fundamentals
Why IaC at all
| Manual console | Infrastructure as Code |
|---|---|
| Not reproducible; snowflake environments | Identical dev/stage/prod from one source |
| No history; “who changed this?” | Git history, PR review, blame |
| No dry-run; changes are live | plan / change sets preview before apply |
| Drift accumulates silently | Drift detection & re-apply |
| Onboarding = tribal knowledge | The repo is the documentation |
IaC turns infrastructure into software: peer-reviewed, tested, versioned, and rolled back like any other code. Every serious AWS environment in 2026 is defined this way.
The AWS-native IaC family
- CloudFormation — the foundational service. You declare resources in YAML/JSON templates; CloudFormation creates them as a stack it tracks and manages, handling dependency ordering, rollback on failure, and deletion. No state file to store — AWS holds the state. It is the engine CDK and SAM compile down to.
- CDK (Cloud Development Kit) — write infrastructure in TypeScript, Python, Java, Go, or C#. High-level constructs encapsulate best-practice defaults (a load-balanced Fargate service in a dozen lines), and
cdk synthemits a CloudFormation template. Ideal when you want loops, conditionals, unit tests, and IDE support over YAML. - SAM (Serverless Application Model) — a CloudFormation macro/extension with terse syntax for Lambda, API Gateway, DynamoDB, and Step Functions, plus a CLI for local testing (
sam local invoke) and fast deploys. The lightest path for serverless apps.
CloudFormation core mechanics
- Stack — a deployed instance of a template; the unit of create/update/delete. Delete a stack and (by default) every resource it made is destroyed.
- Change set — a preview of what an update will do (add/modify/replace/remove) before you execute it. The “replace” flag is critical: some property changes force resource re-creation (and downtime / data loss).
- StackSets — deploy the same stack across many accounts and regions from a single operation, driven by AWS Organizations. This is how you roll out org-wide guardrails (a logging bucket, a Config conformance pack, an IAM baseline) to every account consistently.
- Drift detection — compares the live resources to the template and reports out-of-band changes someone made by hand.
- Nested stacks / modules — decompose large templates into reusable components.
Key Concepts
CDK vs. CloudFormation vs. Terraform
| Dimension | CloudFormation | CDK | Terraform |
|---|---|---|---|
| Language | YAML/JSON | TS/Python/Java/Go/C# | HCL |
| Cloud scope | AWS only | AWS only (primarily) | Multi-cloud + 3000 providers |
| State | Managed by AWS | Managed by AWS (via CFN) | You manage (S3 backend) |
| Preview | Change sets | cdk diff → change sets | terraform plan |
| Ecosystem | AWS-published | Construct Hub | Terraform Registry (huge) |
| Drift | Built-in detection | Built-in (via CFN) | plan shows drift on refresh |
| Best for | Pure-AWS, no external state | Pure-AWS with real code | Multi-cloud, largest module reuse |
How to choose: all-in on AWS with a team that dislikes managing state → CloudFormation or CDK (CDK if you want real language features). Multi-cloud, or you value Terraform’s plan workflow and module ecosystem, or you already have Terraform expertise → Terraform. Serverless-first app → SAM. There is no wrong answer; consistency within a team beats the marginal advantages of any one tool.
Terraform on AWS: provider and remote state
Terraform describes AWS with the aws provider and stores the mapping between your config and real resources in a state file. Where that state lives is the single most important operational decision: local state on a laptop is a recipe for disaster (lost file = orphaned infra, no locking = two engineers corrupting each other’s applies). The standard answer is a remote backend in S3 with a DynamoDB lock table.
Worked example — Terraform S3 backend with DynamoDB locking:
# backend.tf
terraform {
required_version = ">= 1.6"
required_providers {
aws = {
source = "hashicorp/aws"
version = "~> 5.0"
}
}
backend "s3" {
bucket = "acme-terraform-state-prod" # versioned, encrypted, private
key = "networking/terraform.tfstate" # path per component
region = "us-east-1"
dynamodb_table = "terraform-locks" # state-lock table (LockID hash key)
encrypt = true # SSE on the state object
}
}
provider "aws" {
region = "us-east-1"
default_tags {
tags = {
Environment = "prod"
ManagedBy = "terraform"
CostCenter = "platform"
Owner = "platform-team"
}
}
}
Why each piece matters:
- S3 bucket — durable, versioned (so you can roll back a corrupted state), encrypted, with public access blocked. State can contain secrets in plaintext, so treat the bucket as sensitive.
key— one state file per logical component (networking, data, app) so blast radius and apply times stay small; a monolithic single state is slow and dangerous.- DynamoDB lock table — a simple table with a
LockID(String) partition key. Terraform takes a lock here before every apply so two people (or two CI runs) cannot mutate state simultaneously. Note: newer Terraform versions also offer S3-native locking (use_lockfile), reducing the need for DynamoDB, but the DynamoDB pattern remains the most widely deployed. default_tags— every resource this provider creates inherits these tags automatically. This is the backbone of cost attribution (see below) — set it once and never forget a tag again.
Bootstrapping note: the state bucket and lock table cannot themselves live in the state they back, so create them once by hand or in a tiny separate “bootstrap” configuration.
SAM for serverless
SAM shines for Lambda-centric apps. A template.yaml uses AWS::Serverless::* resource types that expand into full CloudFormation:
Transform: AWS::Serverless-2016-10-31
Resources:
CheckoutFn:
Type: AWS::Serverless::Function
Properties:
Handler: app.handler
Runtime: python3.13
MemorySize: 512
Timeout: 10
Events:
Api:
Type: HttpApi
Properties: { Path: /checkout, Method: post }
Policies:
- DynamoDBCrudPolicy: { TableName: !Ref OrdersTable }
sam local invoke runs the function in a container against a test event, sam local start-api stands up the API locally, and sam deploy --guided packages and ships it. The value is the terseness (that Policies shorthand generates a scoped IAM role) and the fast local feedback loop.
Cost management toolchain
| Tool | Purpose | Cadence |
|---|---|---|
| Cost Explorer | Visualize and analyze spend by service/tag/account; forecast | Ad hoc / weekly review |
| AWS Budgets | Set thresholds; alert (or act) when spend/usage crosses them | Continuous, alerting |
| Cost Anomaly Detection | ML detects unusual spend spikes and alerts | Continuous, automatic |
| Cost Allocation Tags | Attribute cost to team/env/product via tags | Foundational, ongoing |
| Compute Optimizer | Right-sizing recommendations for EC2/EBS/Lambda/ASG | Periodic review |
| Cost & Usage Report (CUR) | The raw, line-item billing data (to S3 / Athena / QuickSight) | For deep analysis / FinOps |
Tagging strategy
Tags are the foundation of everything cost-related — you cannot allocate, budget, or optimize spend you cannot attribute. A minimal mandatory tag set:
| Tag | Example | Purpose |
|---|---|---|
Environment | prod / staging / dev | Split spend by stage; find dev waste |
CostCenter / Team | platform / checkout | Charge-back / show-back to owners |
Owner | platform-team | Who to ask; who cleans up |
Application / Service | checkout-api | Per-product cost |
ManagedBy | terraform | Distinguish IaC from manual (drift) resources |
Enforce tags with Tag Policies (Organizations), Service Control Policies (deny resource creation without required tags), or IaC default_tags. Then activate them as cost allocation tags in the Billing console so they appear in Cost Explorer and the CUR — a tag only becomes a cost dimension after you activate it, and only from that point forward. Note tag values are case-sensitive (Prod ≠ prod), so standardize casing.
Purchase models: Savings Plans vs. Reserved vs. Spot
| Model | Discount | Commitment | Flexibility | Interruptible | Use for |
|---|---|---|---|---|---|
| On-Demand | baseline | none | total | no | Spiky, short-lived, unpredictable |
| Spot | up to ~90% | none | any spare capacity | Yes (2-min notice) | Fault-tolerant batch, CI, stateless fleets |
| Reserved Instances | up to ~72% | 1 or 3 yr | family/region-locked (Standard) | no | Legacy; specific steady instances |
| Compute Savings Plans | up to ~66% | 1 or 3 yr ($/hr) | any region/family/EC2/Fargate/Lambda | no | Steady baseline spend (modern default) |
| EC2 Instance Savings Plans | up to ~72% | 1 or 3 yr ($/hr) | within one family/region | no | Steady spend on a known family |
Decision flow: measure your steady baseline (the compute you run 24/7 every month) and cover it with Compute Savings Plans (flexible) or EC2 Instance Savings Plans (deeper discount if the family is stable). Run interruptible, stateless work (CI runners, batch, big-data, dev fleets) on Spot for up to 90% off. Leave the spiky, unpredictable top on On-Demand. This layered approach — Savings Plans for the floor, Spot for the elastic middle, On-Demand for the peaks — is the standard FinOps cost structure. Reserved Instances are largely superseded by Savings Plans for compute, though RIs still apply to RDS, ElastiCache, Redshift, and OpenSearch.
Budgets and anomaly detection in practice
A Budget sets a monthly (or custom) cost/usage limit and notifies via SNS at configurable thresholds (e.g. 80% actual, 100% actual, 100% forecast). Budget Actions can go further and apply an SCP or IAM policy to throttle spend automatically when a threshold is hit. Cost Anomaly Detection complements budgets: instead of a fixed threshold, it learns each service’s normal spend and alerts on statistically unusual spikes — catching, say, a runaway NAT gateway or an accidentally-public S3 bucket racking up requests, days before the monthly budget alarm would.
A practical budget + tagging note: create budgets per cost-allocation dimension (per team tag, per environment) rather than one account-wide budget, so an alert lands on the team that caused it. An account-wide budget tells you “we’re over” but not “who” — tag-scoped budgets make the alert actionable and route ownership automatically.
Best Practices
- Define all infrastructure as code; ban console-created production resources. Rationale: manual resources are unreproducible snowflakes with no history or review, and they drift; the repo must be the single source of truth.
- Always preview before applying (
terraform plan/ change sets). Rationale: infrastructure changes can replace resources and cause downtime or data loss; the preview catches destructive changes before they execute. - Store Terraform state remotely in versioned, encrypted S3 with locking. Rationale: local state is lost or corrupted easily, and without locking concurrent applies destroy each other’s state; S3 versioning lets you roll back a bad state.
- Split state by component, not one monolith. Rationale: a single giant state file makes every apply slow and turns any mistake into a whole-environment blast radius; per-component state keeps changes small and safe.
- Set provider-level default tags (or Tag Policies) so nothing is untagged. Rationale: untagged resources are invisible to cost allocation and cleanup automation; enforcing tags at creation is the only way to keep attribution complete.
- Activate cost allocation tags early and standardize their casing. Rationale: tags only become cost dimensions after activation and only prospectively, and
Prod≠prodfragments your reports — set the convention before spend accumulates. - Cover the steady baseline with Savings Plans, not On-Demand. Rationale: paying On-Demand for 24/7 workloads leaves up to ~66% on the table; a 1-year Compute Savings Plan is low-risk and flexible across family and region.
- Run interruptible, stateless workloads on Spot. Rationale: CI, batch, and stateless fleets tolerate the 2-minute reclamation and get up to 90% off — the single biggest lever on variable compute cost.
- Create tag-scoped budgets per team/environment, not one account-wide budget. Rationale: a scoped alert routes to the owner who caused the overage and is actionable; an account-wide alert only says “we’re over” without accountability.
- Enable Cost Anomaly Detection on top of budgets. Rationale: budgets catch a threshold at month-end; anomaly detection catches a runaway resource within a day, before it becomes a large surprise bill.
- Review Compute Optimizer and right-size regularly. Rationale: most fleets are over-provisioned; acting on right-sizing and Graviton recommendations reliably cuts double-digit percentages with no reliability cost.
- Keep secrets out of state and templates; reference Secrets Manager/SSM at deploy time. Rationale: Terraform state and CloudFormation parameters can expose plaintext secrets; sourcing them at runtime keeps them out of the state bucket and version control.
- Use StackSets / Terraform modules to enforce org-wide guardrails consistently. Rationale: rolling the same baseline (logging, Config packs, IAM) to every account by hand guarantees drift; centralized multi-account deployment keeps posture uniform.
- Run IaC through CI with plan-on-PR and apply-on-merge. Rationale: peer-reviewed plans catch mistakes before they reach production and create an auditable change history, exactly like application code review.
- Pin provider and module versions. Rationale: unpinned versions silently pull breaking changes on the next apply; pinning makes builds reproducible and upgrades deliberate.
- Tag resources for lifecycle automation (TTL/expiry on dev) and delete idle infra. Rationale: forgotten dev environments and orphaned volumes/snapshots/IPs are pure waste; tag-driven cleanup automation recovers spend continuously without human policing.
References
- AWS CloudFormation User Guide
- CloudFormation StackSets
- AWS CDK Developer Guide
- AWS SAM Developer Guide
- Terraform AWS Provider
- Terraform S3 backend
- AWS Cost Explorer
- AWS Budgets
- AWS Cost Anomaly Detection
- Cost allocation tags
- AWS Savings Plans
- Amazon EC2 Spot Instances
- AWS Compute Optimizer
- AWS Well-Architected — Cost Optimization Pillar