← Cloud · AWS← Cloud · AWS
Cloud · AWSCloud · AWS19 Th7, 2026Jul 19, 202618 phút đọc13 min read

IaC & Cost ManagementIaC & Cost Management

Part of the Cloud knowledge base — AWS deep dive. Complements DevOps Roadmap.

Tổng quan

Trang này bàn về hai lĩnh vực mà một kỹ sư DevOps sống chung mỗi ngày: Infrastructure as Code (IaC) — định nghĩa hạ tầng cloud của bạn bằng các file được quản lý phiên bản, có thể review, và lặp lại được thay vì click quanh console — và cost management (quản lý chi phí) — giữ cho khoản chi tiêu phát sinh luôn rõ ràng, được quy trách nhiệm, và được tối ưu. Chúng đi cùng nhau vì chính những tính chất khiến hạ tầng an toàn (declarative, được gắn tag, được review, được tự động hóa) cũng chính là những gì khiến nó rẻ để vận hành và dễ để quy trách nhiệm.

Về phía IaC, AWS cung cấp một bộ công cụ native — CloudFormation (engine provisioning nền tảng), CDK (ngôn ngữ lập trình thật sự, tổng hợp ra CloudFormation), và SAM (một phần mở rộng của CloudFormation tập trung cho serverless) — trong khi thế giới rộng hơn chạy trên Terraform, tiêu chuẩn thực tế cloud-agnostic. Không có một câu trả lời đúng duy nhất: CloudFormation/CDK cho tích hợp AWS sâu nhất và drift detection mà không cần quản lý state file; Terraform cho khả năng chuyển đổi đa cloud, một hệ sinh thái module khổng lồ, và một workflow plan/apply tuyệt vời. Hầu hết các team chọn một công cụ chính và giữ tính nhất quán.

Về phía chi phí, thực tế chi phối là hóa đơn cloud là sản phẩm của các quyết định kỹ thuật, và nó chỉ kiểm soát được nếu nó rõ ràng và có thể quy trách nhiệm. Hạ tầng không gắn tag, không giám sát tạo ra một hóa đơn mà không ai giải thích được và không ai làm chủ. Nghề này là xây dựng một kỷ luật gắn tag, theo dõi chi tiêu bằng Cost Explorer, Budgets, và Anomaly Detection, và khớp các cam kết mua (Savings Plans, Reserved, Spot) với hình dạng workload. Làm tốt, cost management không phải là thắt lưng buộc bụng — mà là chi tiêu có chủ đích vào những gì quan trọng và loại bỏ lãng phí một cách tự động.

Kiến thức nền tảng

Tại sao lại cần IaC

Manual consoleInfrastructure as Code
Không tái tạo được; các môi trường snowflakeDev/stage/prod giống hệt nhau từ một nguồn duy nhất
Không có lịch sử; “ai đã đổi cái này?”Lịch sử Git, review PR, blame
Không có dry-run; thay đổi áp dụng trực tiếpplan / change set xem trước khi apply
Drift tích tụ âm thầmDrift detection và re-apply
Onboarding = kiến thức truyền miệngRepo chính là tài liệu

IaC biến hạ tầng thành phần mềm: được peer-review, được test, được version, và rollback như bất kỳ đoạn code nào khác. Mọi môi trường AWS nghiêm túc trong năm 2026 đều được định nghĩa theo cách này.

Họ IaC native của AWS

Cơ chế cốt lõi của CloudFormation

Khái niệm chính

CDK vs. CloudFormation vs. Terraform

Tiêu chíCloudFormationCDKTerraform
Ngôn ngữYAML/JSONTS/Python/Java/Go/C#HCL
Phạm vi cloudChỉ AWSChỉ AWS (chủ yếu)Đa cloud + 3000 provider
StateDo AWS quản lýDo AWS quản lý (qua CFN)Bạn quản lý (S3 backend)
Xem trướcChange setscdk diff → change setsterraform plan
Hệ sinh tháiDo AWS xuất bảnConstruct HubTerraform Registry (khổng lồ)
DriftDetection tích hợp sẵnTích hợp sẵn (qua CFN)plan hiển thị drift khi refresh
Phù hợp nhất choThuần AWS, không state ngoàiThuần AWS với code thậtĐa cloud, tái sử dụng module lớn nhất

Cách chọn: dốc toàn lực vào AWS với một team không thích quản lý state → CloudFormation hoặc CDK (CDK nếu bạn muốn tính năng ngôn ngữ thật sự). Đa cloud, hoặc bạn coi trọng workflow plan và hệ sinh thái module của Terraform, hoặc bạn đã có sẵn chuyên môn Terraform → Terraform. Ứng dụng serverless-first → SAM. Không có câu trả lời sai; tính nhất quán trong một team thắng những lợi thế biên nhỏ của bất kỳ công cụ nào.

Terraform trên AWS: provider và remote state

Terraform mô tả AWS bằng provider aws và lưu ánh xạ giữa config của bạn với các resource thật trong một state file. Nơi state đó nằm là quyết định vận hành quan trọng nhất: state cục bộ trên một laptop là công thức cho thảm họa (mất file = hạ tầng mồ côi, không có locking = hai kỹ sư làm hỏng apply của nhau). Câu trả lời tiêu chuẩn là một remote backend trong S3 với một bảng lock DynamoDB.

Ví dụ thực tế — Terraform S3 backend với DynamoDB locking:

# backend.tf
terraform {
  required_version = ">= 1.6"

  required_providers {
    aws = {
      source  = "hashicorp/aws"
      version = "~> 5.0"
    }
  }

  backend "s3" {
    bucket         = "acme-terraform-state-prod"   # versioned, encrypted, private
    key            = "networking/terraform.tfstate" # path per component
    region         = "us-east-1"
    dynamodb_table = "terraform-locks"              # state-lock table (LockID hash key)
    encrypt        = true                            # SSE on the state object
  }
}

provider "aws" {
  region = "us-east-1"
  default_tags {
    tags = {
      Environment = "prod"
      ManagedBy   = "terraform"
      CostCenter  = "platform"
      Owner       = "platform-team"
    }
  }
}

Vì sao mỗi phần lại quan trọng:

Lưu ý về bootstrap: bucket state và bảng lock bản thân chúng không thể nằm trong state mà chúng làm nền, nên hãy tạo chúng một lần bằng tay hoặc trong một cấu hình “bootstrap” nhỏ riêng biệt.

SAM cho serverless

SAM tỏa sáng cho các ứng dụng lấy Lambda làm trung tâm. Một template.yaml dùng các loại resource AWS::Serverless::* được mở rộng thành CloudFormation đầy đủ:

Transform: AWS::Serverless-2016-10-31
Resources:
  CheckoutFn:
    Type: AWS::Serverless::Function
    Properties:
      Handler: app.handler
      Runtime: python3.13
      MemorySize: 512
      Timeout: 10
      Events:
        Api:
          Type: HttpApi
          Properties: { Path: /checkout, Method: post }
      Policies:
        - DynamoDBCrudPolicy: { TableName: !Ref OrdersTable }

sam local invoke chạy function trong một container đối với một test event, sam local start-api dựng API cục bộ, và sam deploy --guided đóng gói và ship nó. Giá trị nằm ở sự ngắn gọn (cú pháp rút gọn Policies đó tạo ra một IAM role có phạm vi hẹp) và vòng phản hồi cục bộ nhanh.

Bộ công cụ quản lý chi phí

Công cụMục đíchNhịp độ
Cost ExplorerTrực quan hóa và phân tích chi tiêu theo service/tag/account; dự báoAd hoc / review hàng tuần
AWS BudgetsĐặt ngưỡng; cảnh báo (hoặc hành động) khi chi tiêu/sử dụng vượt qua chúngLiên tục, có cảnh báo
Cost Anomaly DetectionML phát hiện các đợt tăng chi tiêu bất thường và cảnh báoLiên tục, tự động
Cost Allocation TagsQuy chi phí cho team/env/product qua tagNền tảng, liên tục
Compute OptimizerKhuyến nghị right-sizing cho EC2/EBS/Lambda/ASGReview định kỳ
Cost & Usage Report (CUR)Dữ liệu billing thô, theo từng dòng (tới S3 / Athena / QuickSight)Cho phân tích sâu / FinOps

Chiến lược gắn tag

Tag là nền tảng của mọi thứ liên quan đến chi phí — bạn không thể phân bổ, lập ngân sách, hay tối ưu khoản chi tiêu mà bạn không thể quy trách nhiệm. Một bộ tag bắt buộc tối thiểu:

TagVí dụMục đích
Environmentprod / staging / devChia chi tiêu theo giai đoạn; tìm lãng phí ở dev
CostCenter / Teamplatform / checkoutCharge-back / show-back cho các chủ sở hữu
Ownerplatform-teamHỏi ai; ai dọn dẹp
Application / Servicecheckout-apiChi phí theo từng product
ManagedByterraformPhân biệt resource IaC với resource thủ công (drift)

Áp đặt tag bằng Tag Policies (Organizations), Service Control Policies (từ chối tạo resource nếu thiếu tag bắt buộc), hoặc default_tags của IaC. Sau đó kích hoạt chúng thành cost allocation tags trong Billing console để chúng xuất hiện trong Cost Explorer và CUR — một tag chỉ trở thành một chiều chi phí sau khi bạn kích hoạt nó, và chỉ từ thời điểm đó trở đi. Lưu ý giá trị tag phân biệt hoa thường (Prodprod), nên hãy chuẩn hóa cách viết hoa.

Các mô hình mua: Savings Plans vs. Reserved vs. Spot

Mô hìnhChiết khấuCam kếtLinh hoạtCó thể gián đoạnDùng cho
On-Demandbaselinekhôngtoàn phầnkhôngĐột biến, ngắn hạn, không đoán trước
Spottới ~90%khôngbất kỳ capacity dư nàoCó (báo trước 2 phút)Batch chịu lỗi, CI, fleet stateless
Reserved Instancestới ~72%1 hoặc 3 nămkhóa theo family/region (Standard)khôngCũ; các instance ổn định cụ thể
Compute Savings Planstới ~66%1 hoặc 3 năm ($/giờ)bất kỳ region/family/EC2/Fargate/LambdakhôngChi tiêu baseline ổn định (mặc định hiện đại)
EC2 Instance Savings Planstới ~72%1 hoặc 3 năm ($/giờ)trong một family/regionkhôngChi tiêu ổn định trên một family đã biết

Luồng quyết định: đo lường baseline ổn định của bạn (lượng compute bạn chạy 24/7 mỗi tháng) và phủ nó bằng Compute Savings Plans (linh hoạt) hoặc EC2 Instance Savings Plans (chiết khấu sâu hơn nếu family ổn định). Chạy công việc có thể gián đoạn, stateless (CI runner, batch, big-data, dev fleet) trên Spot để giảm tới 90%. Để phần đột biến, không đoán trước ở đỉnh trên On-Demand. Cách tiếp cận phân tầng này — Savings Plans cho phần sàn, Spot cho phần đàn hồi ở giữa, On-Demand cho các đỉnh — là cấu trúc chi phí FinOps tiêu chuẩn. Reserved Instances phần lớn đã bị Savings Plans thay thế cho compute, dù RI vẫn áp dụng cho RDS, ElastiCache, Redshift, và OpenSearch.

Budgets và anomaly detection trong thực tế

Một Budget đặt một giới hạn chi phí/sử dụng hàng tháng (hoặc tùy chỉnh) và thông báo qua SNS tại các ngưỡng có thể cấu hình (ví dụ 80% thực tế, 100% thực tế, 100% dự báo). Budget Actions có thể đi xa hơn và áp dụng một SCP hoặc IAM policy để tự động kìm hãm chi tiêu khi một ngưỡng bị chạm. Cost Anomaly Detection bổ sung cho budget: thay vì một ngưỡng cố định, nó học chi tiêu bình thường của mỗi service và cảnh báo về các đợt tăng bất thường về mặt thống kê — bắt được, chẳng hạn, một NAT gateway chạy mất kiểm soát hoặc một S3 bucket vô tình public đang tích lũy request, nhiều ngày trước khi cảnh báo budget hàng tháng kịp báo.

Một lưu ý thực tế về budget + tagging: tạo budget theo từng chiều cost-allocation (theo tag team, theo môi trường) thay vì một budget toàn account, để một cảnh báo rơi vào team đã gây ra nó. Một budget toàn account nói cho bạn biết “chúng ta vượt rồi” nhưng không biết “ai” — các budget theo tag làm cho cảnh báo có thể hành động được và tự động định tuyến trách nhiệm.

Best Practices

  1. Định nghĩa toàn bộ hạ tầng bằng code; cấm resource production tạo từ console. Lý do: resource thủ công là các snowflake không tái tạo được, không lịch sử hay review, và chúng drift; repo phải là nguồn sự thật duy nhất.
  2. Luôn xem trước trước khi apply (terraform plan / change sets). Lý do: thay đổi hạ tầng có thể thay thế resource và gây downtime hoặc mất dữ liệu; bản xem trước bắt các thay đổi phá hủy trước khi chúng thực thi.
  3. Lưu trữ Terraform state từ xa trong S3 versioned, được mã hóa với locking. Lý do: state cục bộ dễ mất hoặc hỏng, và không có locking thì các apply đồng thời phá hủy state của nhau; S3 versioning cho phép bạn rollback một state xấu.
  4. Chia state theo thành phần, không phải một khối đơn khối. Lý do: một state file khổng lồ duy nhất khiến mọi apply chậm và biến bất kỳ sai lầm nào thành blast radius toàn môi trường; state theo từng thành phần giữ thay đổi nhỏ và an toàn.
  5. Đặt default tags ở cấp provider (hoặc Tag Policies) để không gì bị thiếu tag. Lý do: resource không có tag là vô hình đối với cost allocation và tự động dọn dẹp; áp đặt tag lúc tạo là cách duy nhất giữ cho việc quy trách nhiệm được đầy đủ.
  6. Kích hoạt cost allocation tags sớm và chuẩn hóa cách viết hoa của chúng. Lý do: tag chỉ trở thành chiều chi phí sau khi kích hoạt và chỉ áp dụng về sau, và Prodprod làm phân mảnh báo cáo của bạn — hãy đặt quy ước trước khi chi tiêu tích lũy.
  7. Phủ baseline ổn định bằng Savings Plans, không phải On-Demand. Lý do: trả On-Demand cho workload 24/7 để lại tới ~66% trên bàn; một Compute Savings Plan 1 năm rủi ro thấp và linh hoạt qua family và region.
  8. Chạy workload có thể gián đoạn, stateless trên Spot. Lý do: CI, batch, và fleet stateless chịu được việc thu hồi trong 2 phút và được giảm tới 90% — đòn bẩy lớn nhất trên chi phí compute biến đổi.
  9. Tạo budget theo tag cho từng team/môi trường, không phải một budget toàn account. Lý do: một cảnh báo theo phạm vi định tuyến tới chủ sở hữu đã gây ra khoản vượt và có thể hành động được; một cảnh báo toàn account chỉ nói “chúng ta vượt rồi” mà không có trách nhiệm cụ thể.
  10. Bật Cost Anomaly Detection trên nền của budget. Lý do: budget bắt một ngưỡng vào cuối tháng; anomaly detection bắt một resource chạy mất kiểm soát trong vòng một ngày, trước khi nó trở thành một hóa đơn bất ngờ lớn.
  11. Review Compute Optimizer và right-size thường xuyên. Lý do: hầu hết các fleet đều được cấp phát dư; hành động theo khuyến nghị right-sizing và Graviton đáng tin cậy cắt giảm hàng chục phần trăm mà không tốn chi phí về độ tin cậy.
  12. Giữ secret ra khỏi state và template; tham chiếu Secrets Manager/SSM lúc deploy. Lý do: Terraform state và các parameter CloudFormation có thể để lộ secret plaintext; lấy chúng lúc runtime giữ chúng ra khỏi bucket state và version control.
  13. Dùng StackSets / Terraform module để áp đặt guardrail toàn tổ chức một cách nhất quán. Lý do: triển khai cùng một baseline (logging, Config pack, IAM) tới mọi account bằng tay đảm bảo drift; triển khai đa account tập trung giữ trạng thái đồng nhất.
  14. Chạy IaC qua CI với plan-on-PR và apply-on-merge. Lý do: các plan được peer-review bắt sai lầm trước khi chúng đến production và tạo lịch sử thay đổi có thể audit, hệt như review code ứng dụng.
  15. Ghim phiên bản provider và module. Lý do: phiên bản không ghim âm thầm kéo về các thay đổi phá vỡ trong lần apply tiếp theo; ghim làm cho build tái tạo được và nâng cấp có chủ đích.
  16. Gắn tag cho resource để tự động hóa vòng đời (TTL/hết hạn trên dev) và xóa hạ tầng nhàn rỗi. Lý do: các môi trường dev bị lãng quên và volume/snapshot/IP mồ côi là lãng phí thuần túy; tự động hóa dọn dẹp theo tag thu hồi chi tiêu liên tục mà không cần con người giám sát.

Tài liệu tham khảo

Part of the Cloud knowledge base — AWS deep dive. Complements DevOps Roadmap.

Overview

This page covers two disciplines that a DevOps engineer lives inside every day: Infrastructure as Code (IaC) — defining your cloud in version-controlled, reviewable, repeatable files instead of clicking around the console — and cost management — keeping the resulting spend visible, attributed, and optimized. They belong together because the same properties that make infrastructure safe (declarative, tagged, reviewed, automated) are exactly what make it cheap to operate and easy to attribute.

On the IaC side AWS offers a native stack — CloudFormation (the underlying provisioning engine), CDK (real programming languages that synthesize to CloudFormation), and SAM (a serverless-focused CloudFormation extension) — while the wider world runs on Terraform, the cloud-agnostic de-facto standard. There is no single right answer: CloudFormation/CDK give the deepest AWS integration and drift detection with no state file to manage; Terraform gives multi-cloud portability, a huge module ecosystem, and a superb plan/apply workflow. Most teams pick one primary tool and stay consistent.

On the cost side the governing reality is that the cloud bill is a product of engineering decisions, and it is only controllable if it is visible and attributable. Un-tagged, unmonitored infrastructure produces a bill nobody can explain and nobody owns. The craft is building a tagging discipline, watching spend with Cost Explorer, Budgets, and Anomaly Detection, and matching purchase commitments (Savings Plans, Reserved, Spot) to workload shape. Done well, cost management is not austerity — it is spending deliberately on what matters and eliminating waste automatically.

Fundamentals

Why IaC at all

Manual consoleInfrastructure as Code
Not reproducible; snowflake environmentsIdentical dev/stage/prod from one source
No history; “who changed this?”Git history, PR review, blame
No dry-run; changes are liveplan / change sets preview before apply
Drift accumulates silentlyDrift detection & re-apply
Onboarding = tribal knowledgeThe repo is the documentation

IaC turns infrastructure into software: peer-reviewed, tested, versioned, and rolled back like any other code. Every serious AWS environment in 2026 is defined this way.

The AWS-native IaC family

CloudFormation core mechanics

Key Concepts

CDK vs. CloudFormation vs. Terraform

DimensionCloudFormationCDKTerraform
LanguageYAML/JSONTS/Python/Java/Go/C#HCL
Cloud scopeAWS onlyAWS only (primarily)Multi-cloud + 3000 providers
StateManaged by AWSManaged by AWS (via CFN)You manage (S3 backend)
PreviewChange setscdk diff → change setsterraform plan
EcosystemAWS-publishedConstruct HubTerraform Registry (huge)
DriftBuilt-in detectionBuilt-in (via CFN)plan shows drift on refresh
Best forPure-AWS, no external statePure-AWS with real codeMulti-cloud, largest module reuse

How to choose: all-in on AWS with a team that dislikes managing state → CloudFormation or CDK (CDK if you want real language features). Multi-cloud, or you value Terraform’s plan workflow and module ecosystem, or you already have Terraform expertise → Terraform. Serverless-first app → SAM. There is no wrong answer; consistency within a team beats the marginal advantages of any one tool.

Terraform on AWS: provider and remote state

Terraform describes AWS with the aws provider and stores the mapping between your config and real resources in a state file. Where that state lives is the single most important operational decision: local state on a laptop is a recipe for disaster (lost file = orphaned infra, no locking = two engineers corrupting each other’s applies). The standard answer is a remote backend in S3 with a DynamoDB lock table.

Worked example — Terraform S3 backend with DynamoDB locking:

# backend.tf
terraform {
  required_version = ">= 1.6"

  required_providers {
    aws = {
      source  = "hashicorp/aws"
      version = "~> 5.0"
    }
  }

  backend "s3" {
    bucket         = "acme-terraform-state-prod"   # versioned, encrypted, private
    key            = "networking/terraform.tfstate" # path per component
    region         = "us-east-1"
    dynamodb_table = "terraform-locks"              # state-lock table (LockID hash key)
    encrypt        = true                            # SSE on the state object
  }
}

provider "aws" {
  region = "us-east-1"
  default_tags {
    tags = {
      Environment = "prod"
      ManagedBy   = "terraform"
      CostCenter  = "platform"
      Owner       = "platform-team"
    }
  }
}

Why each piece matters:

Bootstrapping note: the state bucket and lock table cannot themselves live in the state they back, so create them once by hand or in a tiny separate “bootstrap” configuration.

SAM for serverless

SAM shines for Lambda-centric apps. A template.yaml uses AWS::Serverless::* resource types that expand into full CloudFormation:

Transform: AWS::Serverless-2016-10-31
Resources:
  CheckoutFn:
    Type: AWS::Serverless::Function
    Properties:
      Handler: app.handler
      Runtime: python3.13
      MemorySize: 512
      Timeout: 10
      Events:
        Api:
          Type: HttpApi
          Properties: { Path: /checkout, Method: post }
      Policies:
        - DynamoDBCrudPolicy: { TableName: !Ref OrdersTable }

sam local invoke runs the function in a container against a test event, sam local start-api stands up the API locally, and sam deploy --guided packages and ships it. The value is the terseness (that Policies shorthand generates a scoped IAM role) and the fast local feedback loop.

Cost management toolchain

ToolPurposeCadence
Cost ExplorerVisualize and analyze spend by service/tag/account; forecastAd hoc / weekly review
AWS BudgetsSet thresholds; alert (or act) when spend/usage crosses themContinuous, alerting
Cost Anomaly DetectionML detects unusual spend spikes and alertsContinuous, automatic
Cost Allocation TagsAttribute cost to team/env/product via tagsFoundational, ongoing
Compute OptimizerRight-sizing recommendations for EC2/EBS/Lambda/ASGPeriodic review
Cost & Usage Report (CUR)The raw, line-item billing data (to S3 / Athena / QuickSight)For deep analysis / FinOps

Tagging strategy

Tags are the foundation of everything cost-related — you cannot allocate, budget, or optimize spend you cannot attribute. A minimal mandatory tag set:

TagExamplePurpose
Environmentprod / staging / devSplit spend by stage; find dev waste
CostCenter / Teamplatform / checkoutCharge-back / show-back to owners
Ownerplatform-teamWho to ask; who cleans up
Application / Servicecheckout-apiPer-product cost
ManagedByterraformDistinguish IaC from manual (drift) resources

Enforce tags with Tag Policies (Organizations), Service Control Policies (deny resource creation without required tags), or IaC default_tags. Then activate them as cost allocation tags in the Billing console so they appear in Cost Explorer and the CUR — a tag only becomes a cost dimension after you activate it, and only from that point forward. Note tag values are case-sensitive (Prodprod), so standardize casing.

Purchase models: Savings Plans vs. Reserved vs. Spot

ModelDiscountCommitmentFlexibilityInterruptibleUse for
On-DemandbaselinenonetotalnoSpiky, short-lived, unpredictable
Spotup to ~90%noneany spare capacityYes (2-min notice)Fault-tolerant batch, CI, stateless fleets
Reserved Instancesup to ~72%1 or 3 yrfamily/region-locked (Standard)noLegacy; specific steady instances
Compute Savings Plansup to ~66%1 or 3 yr ($/hr)any region/family/EC2/Fargate/LambdanoSteady baseline spend (modern default)
EC2 Instance Savings Plansup to ~72%1 or 3 yr ($/hr)within one family/regionnoSteady spend on a known family

Decision flow: measure your steady baseline (the compute you run 24/7 every month) and cover it with Compute Savings Plans (flexible) or EC2 Instance Savings Plans (deeper discount if the family is stable). Run interruptible, stateless work (CI runners, batch, big-data, dev fleets) on Spot for up to 90% off. Leave the spiky, unpredictable top on On-Demand. This layered approach — Savings Plans for the floor, Spot for the elastic middle, On-Demand for the peaks — is the standard FinOps cost structure. Reserved Instances are largely superseded by Savings Plans for compute, though RIs still apply to RDS, ElastiCache, Redshift, and OpenSearch.

Budgets and anomaly detection in practice

A Budget sets a monthly (or custom) cost/usage limit and notifies via SNS at configurable thresholds (e.g. 80% actual, 100% actual, 100% forecast). Budget Actions can go further and apply an SCP or IAM policy to throttle spend automatically when a threshold is hit. Cost Anomaly Detection complements budgets: instead of a fixed threshold, it learns each service’s normal spend and alerts on statistically unusual spikes — catching, say, a runaway NAT gateway or an accidentally-public S3 bucket racking up requests, days before the monthly budget alarm would.

A practical budget + tagging note: create budgets per cost-allocation dimension (per team tag, per environment) rather than one account-wide budget, so an alert lands on the team that caused it. An account-wide budget tells you “we’re over” but not “who” — tag-scoped budgets make the alert actionable and route ownership automatically.

Best Practices

  1. Define all infrastructure as code; ban console-created production resources. Rationale: manual resources are unreproducible snowflakes with no history or review, and they drift; the repo must be the single source of truth.
  2. Always preview before applying (terraform plan / change sets). Rationale: infrastructure changes can replace resources and cause downtime or data loss; the preview catches destructive changes before they execute.
  3. Store Terraform state remotely in versioned, encrypted S3 with locking. Rationale: local state is lost or corrupted easily, and without locking concurrent applies destroy each other’s state; S3 versioning lets you roll back a bad state.
  4. Split state by component, not one monolith. Rationale: a single giant state file makes every apply slow and turns any mistake into a whole-environment blast radius; per-component state keeps changes small and safe.
  5. Set provider-level default tags (or Tag Policies) so nothing is untagged. Rationale: untagged resources are invisible to cost allocation and cleanup automation; enforcing tags at creation is the only way to keep attribution complete.
  6. Activate cost allocation tags early and standardize their casing. Rationale: tags only become cost dimensions after activation and only prospectively, and Prodprod fragments your reports — set the convention before spend accumulates.
  7. Cover the steady baseline with Savings Plans, not On-Demand. Rationale: paying On-Demand for 24/7 workloads leaves up to ~66% on the table; a 1-year Compute Savings Plan is low-risk and flexible across family and region.
  8. Run interruptible, stateless workloads on Spot. Rationale: CI, batch, and stateless fleets tolerate the 2-minute reclamation and get up to 90% off — the single biggest lever on variable compute cost.
  9. Create tag-scoped budgets per team/environment, not one account-wide budget. Rationale: a scoped alert routes to the owner who caused the overage and is actionable; an account-wide alert only says “we’re over” without accountability.
  10. Enable Cost Anomaly Detection on top of budgets. Rationale: budgets catch a threshold at month-end; anomaly detection catches a runaway resource within a day, before it becomes a large surprise bill.
  11. Review Compute Optimizer and right-size regularly. Rationale: most fleets are over-provisioned; acting on right-sizing and Graviton recommendations reliably cuts double-digit percentages with no reliability cost.
  12. Keep secrets out of state and templates; reference Secrets Manager/SSM at deploy time. Rationale: Terraform state and CloudFormation parameters can expose plaintext secrets; sourcing them at runtime keeps them out of the state bucket and version control.
  13. Use StackSets / Terraform modules to enforce org-wide guardrails consistently. Rationale: rolling the same baseline (logging, Config packs, IAM) to every account by hand guarantees drift; centralized multi-account deployment keeps posture uniform.
  14. Run IaC through CI with plan-on-PR and apply-on-merge. Rationale: peer-reviewed plans catch mistakes before they reach production and create an auditable change history, exactly like application code review.
  15. Pin provider and module versions. Rationale: unpinned versions silently pull breaking changes on the next apply; pinning makes builds reproducible and upgrades deliberate.
  16. Tag resources for lifecycle automation (TTL/expiry on dev) and delete idle infra. Rationale: forgotten dev environments and orphaned volumes/snapshots/IPs are pure waste; tag-driven cleanup automation recovers spend continuously without human policing.

References