← Cloud · AWS← Cloud · AWS
Cloud · AWSCloud · AWS19 Th7, 2026Jul 19, 202621 phút đọc16 min read

StorageStorage

Part of the Cloud knowledge base — AWS deep dive. Complements DevOps Roadmap.

Tổng quan

Storage trên AWS chia thành ba hình dạng vật lý, và chọn sai hình dạng là một sai lầm bạn cảm nhận mỗi ngày. Object storage (S3) giữ các blob phi cấu trúc được định địa chỉ bằng key qua HTTP — có thể scale vô hạn, rẻ, và là nơi mặc định cho gần như mọi thứ không phải là một đĩa đang sống hay một database. Block storage (EBS) trình bày các đĩa thô cho một EC2 instance duy nhất — volume mà OS và database của bạn nằm trên đó. File storage (EFS, FSx) cung cấp một hệ thống file POSIX hoặc SMB dùng chung mà nhiều máy mount cùng lúc. Quanh những thứ này là backup (AWS Backup), các cầu nối hybrid (Storage Gateway), và câu hỏi luôn hiện diện về chi phí truyền dữ liệu, thứ âm thầm chi phối nhiều hóa đơn.

Trang này đi sâu vào S3 — viên ngọc quý — bao gồm bucket, storage class, lifecycle, versioning, mã hóa, sự phân biệt giữa bucket-policy và IAM, static hosting, và event notification. Sau đó nó bao gồm các loại volume EBS và snapshot, EFS và FSx cho hệ thống file dùng chung, AWS Backup cho bảo vệ dữ liệu tập trung, Storage Gateway cho hybrid, và các bẫy chi phí egress bắt gặp mọi người. Nó kết thúc bằng một bucket policy S3 và cấu hình lifecycle thực hành.

Một cách nhìn hữu ích: chọn storage theo mẫu truy cập, không theo kích thước. Một máy cần một đĩa? EBS. Nhiều máy cần cùng các file? EFS/FSx. Thứ gì đó đọc và ghi blob theo tên qua mạng và không bao giờ cần một hệ thống file? S3 — và gần như luôn là S3. Làm đúng điều này ngay từ đầu tránh được những lần tái kiến trúc đau đớn đến từ việc gắn một hệ thống file lên thứ lẽ ra phải là object storage.

Tại sao các quyết định về storage tích lũy dồn

Storage là nơi trọng lực dữ liệu tích tụ. Một petabyte ở sai storage class tốn gấp 10 lần mức đáng lẽ; một bucket không có lifecycle rule phình mãi mãi; dữ liệu không được mã hóa mặc định trở thành một phát hiện khi audit; và egress mà bạn chưa từng đo biến một data pipeline thành một bất ngờ năm-con-số. Không như compute, bạn không thể chỉ đơn giản chấm dứt storage — dữ liệu chính là mục đích. Các quyết định về class, lifecycle, mã hóa và đường truyền được đưa ra sớm sẽ tiết kiệm tiền thật và các sự cố thật về sau, bởi vì di chuyển dữ liệu ở quy mô lớn thì chậm và đắt.

Kiến thức nền tảng

S3 — Simple Storage Service

S3 là object storage: bạn đặt các object (dữ liệu + metadata, tới 5 TB mỗi cái) vào các bucket (container có tên duy nhất toàn cầu, sống trong một Region), được định địa chỉ bằng một key (tên đầy đủ giống đường dẫn). Không có thư mục — dấu / trong một key chỉ là một ký tự, và khung nhìn “thư mục” là một tiện lợi của console. S3 cung cấp độ bền mười một số chín (99,999999999%) bằng cách lưu dư thừa trên nhiều AZ, và nó scale tới dung lượng gần như không giới hạn mà không cần provisioning.

Storage class đánh đổi chi phí/độ trễ truy xuất với giá lưu trữ. Chọn đúng class cho mỗi object là đòn bẩy chi phí lớn nhất của S3.

ClassDùng choTruy xuấtChi phí tương đối
S3 StandardDữ liệu nóng, truy cập thường xuyênTức thìMốc cơ sở (giá lưu trữ cao nhất)
S3 Intelligent-TieringMẫu truy cập chưa biết/thay đổiTức thìTự chuyển giữa các tier; phí giám sát nhỏ
S3 Standard-IA (Infrequent Access)Dữ liệu ấm truy cập tầm hằng thángTức thìLưu trữ rẻ hơn ~40%; phí truy xuất theo GB
S3 One Zone-IADữ liệu ấm có thể tạo lạiTức thìRẻ hơn nữa; một AZ (không chịu lỗi AZ)
S3 Glacier Instant RetrievalLưu trữ cần truy cập tức thìTức thì (ms)Lưu trữ rẻ; phí truy xuất cao hơn
S3 Glacier Flexible RetrievalLưu trữ, chấp nhận phút-đến-giờPhút–12hRất rẻ
S3 Glacier Deep ArchiveTuân thủ/lạnh, hiếm khi đọc12–48hLưu trữ rẻ nhất trên AWS

Intelligent-Tiering là mặc định an toàn khi bạn không biết mẫu truy cập — nó tự động chuyển object giữa các tier thường xuyên, không thường xuyên, và lưu trữ dựa trên mức sử dụng thực tế, không có phí truy xuất cho các tier tức thì. Hãy dùng nó và ngừng đoán.

Lifecycle policy tự động hóa việc chuyển class và hết hạn theo rule (prefix, tag, tuổi). Một policy điển hình: giữ log ở Standard trong 30 ngày, chuyển sang Standard-IA ở mốc 30, sang Glacier ở mốc 90, và xóa ở mốc 365. Lifecycle rule là cách bạn giữ chi phí lưu trữ không phình vô hạn — thiếu chúng, dữ liệu lạnh nằm mãi trong Standard đắt đỏ.

Versioning giữ mọi phiên bản của một object (bao gồm cả delete marker) để việc ghi đè và xóa có thể phục hồi. Đây là nền tảng để bảo vệ chống xóa nhầm và ransomware. Kết hợp với MFA Delete (yêu cầu MFA để xóa vĩnh viễn các phiên bản) và Object Lock (WORM — write-once-read-many, cho lưu giữ tuân thủ). Lưu ý: versioning nhân bội storage, nên hãy ghép nó với lifecycle rule cho hết hạn các phiên bản noncurrent cũ.

Mã hóa — S3 mã hóa mọi object mới lúc nghỉ theo mặc định (SSE-S3) kể từ 2023. Các tùy chọn:

Bucket policy vs IAM policy

Sự phân biệt này làm mọi người bối rối, nên hãy ghim chặt nó: cả hai đều là tài liệu IAM policy, nhưng chúng gắn vào những thứ khác nhau và trả lời những câu hỏi khác nhau.

IAM policyBucket policy
Gắn vàoMột user, group hoặc role (danh tính)Bucket (tài nguyên)
Trả lời”Danh tính này có thể làm gì?""Ai được truy cập bucket này, và như thế nào?”
Cross-accountChỉ cấp trong tài khoản của bạnCó thể cấp quyền cho tài khoản khác / công khai
Tốt nhất choChuẩn “principal nào của tôi chạm được bucket nào”Chia sẻ cross-account, truy cập công khai/ẩn danh, ép buộc TLS/mã hóa không phụ thuộc tổ chức

Quyền được cấp nếu hoặc một identity policy áp dụng hoặc resource policy cho phép (và không có Deny tường minh hoặc SCP nào chặn). Dùng IAM policy cho trường hợp thông thường (các principal của chính bạn), và bucket policy khi rule là về bản thân bucket — truy cập cross-account, ép buộc TLS toàn diện, hoặc (cẩn thận) đọc công khai cho một static site. Trên cả hai là Block Public Access, một lớp ghi đè ở mức tài khoản/bucket ngăn các cấp quyền công khai bất kể policy — hãy để nó bật ở mọi nơi trừ một số ít bucket công khai có chủ đích.

Access PointsS3 Access Grants là các công cụ mới hơn để quản lý truy cập ở quy mô lớn (các endpoint mạng/policy được đặt tên cho mỗi ứng dụng, và ánh xạ các prefix S3tới các principal IAM/identity-center) — hãy dùng chúng khi một bucket policy duy nhất trở nên cồng kềnh trên nhiều app.

S3 static website hosting

S3 có thể phục vụ một static site (HTML/CSS/JS) trực tiếp: bật static-website hosting, đặt một tài liệu index và error, và các object trở thành trang web. Trong thực tế, hãy đặt CloudFront (CDN) ở phía trước để có HTTPS (endpoint website của S3 chỉ hỗ trợ HTTP), tên miền tùy chỉnh, caching, và chi phí egress thấp hơn, dùng Origin Access Control (OAC) để bucket vẫn riêng tư và chỉ CloudFront đọc được. Mẫu hiện đại là: bucket S3 riêng tư + CloudFront + OAC + chứng chỉ ACM — không bao giờ dùng một bucket công khai phục vụ trực tiếp.

S3 event notifications

S3 có thể phát ra event khi có hành động trên object (s3:ObjectCreated:*, s3:ObjectRemoved:*, và nhiều hơn) để kích hoạt xử lý phía sau:

Điều này biến S3 thành cửa trước của các pipeline hướng sự kiện — “ai đó vừa upload một file” trở thành một trigger, không phải một lần poll.

Khái niệm chính

EBS — Elastic Block Store

EBS cung cấp các block volume gắn qua mạng cho EC2 — các đĩa mà OS của bạn boot từ đó và database của bạn ghi vào. Một volume sống trong một AZ, gắn vào một instance tại một thời điểm (trừ io2 Multi-Attach), và tồn tại độc lập với vòng đời của instance.

LoạiClassTốt nhất choGhi chú
gp3SSD tổng quátMặc định cho hầu hết workloadCơ sở 3.000 IOPS / 125 MB/s, provision thêm độc lập; rẻ hơn gp2
gp2SSD tổng quát (cũ hơn)LegacyIOPS scale theo kích thước; ưu tiên gp3
io2 / io2 Block ExpressSSD IOPS provisionedDatabase nhạy độ trễTới 256.000 IOPS, độ bền 99,999%, Multi-Attach
st1HDD thông lượngTuần tự lớn (log, big data)Rẻ, hướng thông lượng, không dành cho I/O ngẫu nhiên
sc1HDD lạnhHiếm truy cập, ưu tiên chi phíBlock rẻ nhất; thông lượng thấp

gp3 là mặc định — nó tách IOPS/thông lượng khỏi dung lượng (với gp2 bạn phải over-provision kích thước để mua IOPS) và rẻ hơn. Chuyển gp2 → gp3 để có một lợi ích chi phí dễ dàng.

Snapshot là bản backup gia tăng, tại-một-thời-điểm của một volume được lưu trong S3 (đằng sau hậu trường). Chỉ các block thay đổi được lưu sau snapshot đầu tiên, nên chúng tiết kiệm không gian. Dùng chúng để backup, để nhân bản volume, để di chuyển dữ liệu qua AZ/Region (sao chép một snapshot), và để build AMI. Tự động hóa chúng với Data Lifecycle Manager hoặc AWS Backup. Snapshot (và volume) nên được mã hóa — bật mã hóa EBS mặc định trên toàn tài khoản.

EFS và FSx — file storage dùng chung

Khi nhiều instance (hoặc Lambda/container) cần cùng các file cùng lúc, block storage không đáp ứng được — bạn cần một hệ thống file dùng chung.

EFS (Elastic File System) là một hệ thống file NFS co giãn, được quản lý hoàn toàn cho Linux: mount nó từ hàng nghìn instance trên nhiều AZ đồng thời, và nó tự phình ra và co lại mà không cần provisioning. Nó có storage class (Standard, Infrequent Access, Archive) với quản lý lifecycle. Dùng nó cho nội dung dùng chung, thư mục home, persistent volume của container, và các app lift-and-shift mong đợi một hệ thống file POSIX. Đánh đổi: chi phí theo GB và độ trễ cao hơn EBS; không dành cho database IOPS cao chạy một instance.

FSx là một họ các hệ thống file được quản lý cho các hệ sinh thái cụ thể:

Chọn EFS cho NFS Linux tổng quát; chọn hương vị FSx phù hợp khi bạn cần Windows/SMB, thông lượng lớp Lustre, hoặc các tính năng của một hệ thống file nhà cung cấp cụ thể.

AWS Backup

AWS Backup tập trung hóa việc backup trên nhiều service — EBS, EFS, RDS, DynamoDB, S3, FSx, EC2, và nhiều hơn — dưới một service dựa trên policy duy nhất. Bạn định nghĩa backup plan (lịch, thời gian giữ, lifecycle sang cold storage), gán tài nguyên theo tag, và có một nơi duy nhất để giám sát, khôi phục và chứng minh tuân thủ. Nó hỗ trợ sao chép cross-Region và cross-account (quan trọng cho DR và để sống sót qua một vụ xâm nhập tài khoản) và Backup Vault Lock (backup bất biến, WORM mà ngay cả admin cũng không thể xóa sớm). Hãy ưu tiên nó hơn các script snapshot tự viết theo từng service — một policy, một audit trail, một console khôi phục.

Storage Gateway

Storage Gateway bắc cầu các ứng dụng on-premises tới storage AWS thông qua một virtual appliance cục bộ:

Chi phí truyền dữ liệu và egress

Cái bẫy bắt gặp mọi người: dữ liệu vào AWS thì miễn phí; dữ liệu ra internet thì mất tiền, và nó cộng dồn lên. Các quy tắc chính:

Hệ quả thực tế: giữ các service nói chuyện nhiều trong một AZ nơi khả năng chịu lỗi cho phép, dùng VPC Gateway Endpoint cho S3/DynamoDB để né chi phí NAT, phục vụ tải xuống qua CloudFront, và theo dõi phí xử lý dữ liệu của NAT Gateway (chúng đo mỗi GB đi qua). Egress là mục chi thầm lặng — hãy mô hình hóa nó trước khi xây các pipeline nặng dữ liệu.

Ví dụ thực hành: S3 bucket policy + lifecycle

Bucket policy — ép buộc TLS và cấp cho một tài khoản khác quyền đọc, trong khi các principal của chính bạn dùng IAM:

{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "DenyInsecureTransport",
      "Effect": "Deny",
      "Principal": "*",
      "Action": "s3:*",
      "Resource": [
        "arn:aws:s3:::my-data-bucket",
        "arn:aws:s3:::my-data-bucket/*"
      ],
      "Condition": { "Bool": { "aws:SecureTransport": "false" } }
    },
    {
      "Sid": "AllowPartnerAccountRead",
      "Effect": "Allow",
      "Principal": { "AWS": "arn:aws:iam::222233334444:root" },
      "Action": ["s3:GetObject", "s3:ListBucket"],
      "Resource": [
        "arn:aws:s3:::my-data-bucket",
        "arn:aws:s3:::my-data-bucket/shared/*"
      ]
    }
  ]
}

Statement đầu tiên là một Deny toàn diện trên mọi request không TLS — một deny tường minh luôn thắng, nên nó làm cứng cả bucket bất kể identity policy. Statement thứ hai cấp cho tài khoản 2222... quyền đọc chỉ trên prefix shared/; lưu ý các cấp quyền cross-account như thế này yêu cầu một resource (bucket) policy — một IAM policy trong tài khoản của bạn không thể cấp quyền cho principal của người khác.

Cấu hình lifecycle — chuyển log xuống bậc thấp và cho hết hạn, đồng thời dọn dẹp các phiên bản cũ:

{
  "Rules": [
    {
      "ID": "log-tiering-and-expiry",
      "Filter": { "Prefix": "logs/" },
      "Status": "Enabled",
      "Transitions": [
        { "Days": 30,  "StorageClass": "STANDARD_IA" },
        { "Days": 90,  "StorageClass": "GLACIER" }
      ],
      "Expiration": { "Days": 365 }
    },
    {
      "ID": "expire-old-versions",
      "Filter": {},
      "Status": "Enabled",
      "NoncurrentVersionTransitions": [
        { "NoncurrentDays": 30, "StorageClass": "GLACIER" }
      ],
      "NoncurrentVersionExpiration": { "NoncurrentDays": 90 },
      "AbortIncompleteMultipartUpload": { "DaysAfterInitiation": 7 }
    }
  ]
}

Rule đầu tiên chuyển các object logs/ sang Standard-IA ở mốc 30 ngày, Glacier ở mốc 90, và xóa chúng ở mốc một năm — đường cong chi phí mẫu mực cho dữ liệu observability. Rule thứ hai dọn dẹp sau versioning (các phiên bản cũ sang Glacier ở mốc 30 ngày, biến mất ở mốc 90) và hủy các multipart upload bị đình trệ sau 7 ngày, một nguồn phổ biến của storage vô hình bị tính phí. Áp dụng bằng aws s3api put-bucket-lifecycle-configuration --bucket my-data-bucket --lifecycle-configuration file://lifecycle.json.

Best Practices

  1. Bật Block Public Access trên toàn tài khoản, và giữ nó bật — bật nó ở mức tài khoản để không bucket nào vô tình công khai được. Chỉ cấp quyền công khai cho một số ít bucket có chủ đích (và ngay cả khi đó cũng ưu tiên CloudFront + OAC). Bucket S3 công khai là loại vi phạm dữ liệu AWS phổ biến nhất.
  2. Mã hóa mọi thứ, và ép buộc TLS khi truyền — dựa vào mã hóa mặc định của S3 (SSE-KMS cho dữ liệu nhạy cảm với audit trail), bật mã hóa EBS mặc định trên toàn tài khoản, và từ chối các request không TLS qua bucket policy. Mã hóa lúc nghỉ gần như miễn phí và đóng lại cả một lớp phát hiện.
  3. Áp lifecycle policy cho mọi bucket — thiếu chúng, dữ liệu lạnh nằm mãi trong Standard đắt đỏ và các phiên bản cũ tích lũy âm thầm. Chuyển xuống IA/Glacier và cho hết hạn theo lịch; đồng thời hủy các multipart upload chưa hoàn tất, một chi phí ẩn phổ biến.
  4. Dùng Intelligent-Tiering khi mẫu truy cập chưa biết — nó tự tối ưu chi phí lưu trữ không có phí truy xuất cho các tier tức thì và không phải đoán. Với dữ liệu rõ ràng nóng hoặc rõ ràng lạnh, chọn class cụ thể; với mọi thứ còn lại, dùng Intelligent-Tiering.
  5. Bật versioning trên các bucket quan trọng, ghép với hết hạn lifecycle — versioning là nút undo của bạn chống lại xóa nhầm và ransomware. Thêm MFA Delete hoặc Object Lock cho dữ liệu tuân thủ, và cho hết hạn các phiên bản noncurrent để nó không nhân bội hóa đơn.
  6. Ưu tiên IAM policy cho các principal của chính bạn, bucket policy cho các rule ở mức tài nguyên — giữ “role nào của tôi chạm được bucket nào” trong identity policy; để dành bucket policy cho chia sẻ cross-account và ép buộc toàn diện (TLS, mã hóa). Đừng rải logic truy cập ra cả hai mà không có lý do.
  7. Đặt CloudFront ở phía trước nội dung tĩnh và tải xuống — không bao giờ phục vụ một website S3 công khai trực tiếp. CloudFront cho HTTPS, tên miền tùy chỉnh, caching, egress rẻ hơn, và (qua OAC) để bucket vẫn riêng tư. Nó vừa an toàn hơn vừa rẻ hơn ở quy mô lớn.
  8. Mặc định EBS volume là gp3 và chuyển khỏi gp2 — gp3 tách IOPS/thông lượng khỏi kích thước và rẻ hơn cho cùng hiệu năng. Việc chuyển đổi diễn ra online và là một trong những lợi ích chi phí dễ nhất hiện có.
  9. Tự động hóa và mã hóa snapshot; sao chép các snapshot quan trọng qua Region — dùng Data Lifecycle Manager hoặc AWS Backup cho các snapshot theo lịch, được giữ, được mã hóa, và sao chép các backup quan trọng cho DR sang một Region thứ hai và lý tưởng là một tài khoản riêng.
  10. Tập trung hóa backup với AWS Backup và khóa vault — một service dựa trên policy trên EBS/EFS/RDS/DynamoDB/S3 thắng các script theo từng service. Dùng Backup Vault Lock (WORM) và sao chép cross-account để một tài khoản bị xâm nhập không thể phá hủy bản khôi phục cuối cùng của bạn.
  11. Right-size hình dạng storage theo mẫu truy cập — đĩa của một instance → EBS; nhiều máy chia sẻ file → EFS/FSx; blob theo key qua HTTP → S3. Đừng ép một hệ thống file dùng chung lên thứ lẽ ra phải là object storage, hoặc ngược lại.
  12. Dùng VPC Gateway Endpoint cho S3 và DynamoDB — chúng định tuyến lưu lượng riêng tư và miễn phí, tránh cả internet lẫn phí xử lý dữ liệu của NAT Gateway. Với việc truy cập S3 khối lượng lớn từ các private subnet, đây là một khoản tiết kiệm lớn, dễ dàng.
  13. Mô hình hóa egress trước khi xây các pipeline nặng dữ liệu — egress internet và truyền cross-AZ/Region là các mục chi thầm lặng, đáng kể. Giữ lưu lượng nói chuyện nhiều trong-AZ nơi khả năng chịu lỗi cho phép, phục vụ qua CloudFront, và theo dõi phí xử lý theo GB của NAT Gateway.
  14. Bật S3 access logging hoặc CloudTrail data event trên các bucket nhạy cảm — bạn không thể điều tra một truy cập bạn chưa từng ghi lại. Ghi log các lần đọc/ghi trên dữ liệu được quy định, và giám sát bằng S3 Storage Lens để phát hiện bất thường về chi phí và sử dụng.
  15. Dùng S3 event notification để xây các pipeline hướng sự kiện — để các lần upload kích hoạt Lambda/SQS/EventBridge thay vì polling. Nó rẻ hơn, nhanh hơn, và là cách chuẩn mực để xử lý dữ liệu ngay khi nó đáp xuống.
  16. Gắn tag cho storage để phân bổ chi phí và tự động hóa lifecycle — tag trên bucket, volume và snapshot điều khiển các báo cáo chi phí, gán backup, và rule dọn dẹp. Storage không gắn tag là không thể quy trách nhiệm và có xu hướng trở thành rác bị bỏ quên, bị tính phí.

Tài liệu tham khảo

Part of the Cloud knowledge base — AWS deep dive. Complements DevOps Roadmap.

Overview

AWS storage divides into three physical shapes, and choosing the wrong one is a mistake you feel every day. Object storage (S3) holds unstructured blobs addressed by key over HTTP — infinitely scalable, cheap, and the default home for almost everything that isn’t a live disk or a database. Block storage (EBS) presents raw disks to a single EC2 instance — the volume your OS and databases sit on. File storage (EFS, FSx) offers a shared POSIX or SMB filesystem that many machines mount at once. Around these sit backup (AWS Backup), hybrid bridges (Storage Gateway), and the ever-present question of data-transfer cost, which quietly dominates many bills.

This page goes deep on S3 — the crown jewel — covering buckets, storage classes, lifecycle, versioning, encryption, the bucket-policy-vs-IAM distinction, static hosting, and event notifications. It then covers EBS volume types and snapshots, EFS and FSx for shared filesystems, AWS Backup for centralized data protection, Storage Gateway for hybrid, and the egress cost traps that catch everyone. It closes with a worked S3 bucket policy and lifecycle configuration.

A useful framing: pick storage by access pattern, not by size. One machine needs a disk? EBS. Many machines need the same files? EFS/FSx. Something reads and writes blobs by name over the network and never needs a filesystem? S3 — and it is almost always S3. Getting this right up front avoids the painful re-architectures that come from bolting a filesystem onto what should have been object storage.

Why storage decisions compound

Storage is where data gravity accumulates. A petabyte in the wrong storage class costs 10× what it should; a bucket without lifecycle rules grows forever; unencrypted-by-default data becomes an audit finding; and egress you never metered turns a data pipeline into a five-figure surprise. Unlike compute, you cannot simply terminate storage — the data is the point. Decisions about class, lifecycle, encryption, and transfer path made early save real money and real incidents later, because moving data at scale is slow and expensive.

Fundamentals

S3 — Simple Storage Service

S3 is object storage: you put objects (data + metadata, up to 5 TB each) into buckets (globally-unique-named containers that live in one Region), addressed by a key (the full path-like name). There are no folders — the / in a key is just a character, and the “folder” view is a console convenience. S3 offers eleven nines (99.999999999%) of durability by redundantly storing across multiple AZs, and it scales to effectively unlimited capacity with no provisioning.

Storage classes trade retrieval cost/latency against storage price. Choosing the right class per object is the single biggest S3 cost lever.

ClassUse forRetrievalRelative cost
S3 StandardHot, frequently accessed dataInstantBaseline (highest storage $)
S3 Intelligent-TieringUnknown/changing access patternsInstantAuto-moves between tiers; small monitoring fee
S3 Standard-IA (Infrequent Access)Warm data accessed monthly-ishInstant~40% cheaper storage; per-GB retrieval fee
S3 One Zone-IARe-creatable warm dataInstantCheaper still; single AZ (no AZ resilience)
S3 Glacier Instant RetrievalArchive needing instant accessInstant (ms)Cheap storage; higher retrieval fee
S3 Glacier Flexible RetrievalArchives, minutes-to-hours OKMinutes–12hVery cheap
S3 Glacier Deep ArchiveCompliance/cold, rarely read12–48hCheapest storage on AWS

Intelligent-Tiering is the safe default when you don’t know the access pattern — it automatically moves objects between frequent, infrequent, and archive tiers based on actual usage, with no retrieval fees for the instant tiers. Use it and stop guessing.

Lifecycle policies automate class transitions and expiration by rule (prefix, tag, age). A typical policy: keep logs in Standard for 30 days, transition to Standard-IA at 30, to Glacier at 90, and delete at 365. Lifecycle rules are how you keep storage costs from growing unboundedly — without them, cold data sits in expensive Standard forever.

Versioning keeps every version of an object (including delete markers) so overwrites and deletes are recoverable. It is the foundation for protecting against accidental deletion and ransomware. Combine with MFA Delete (require MFA to permanently remove versions) and Object Lock (WORM — write-once-read-many, for compliance retention). Note: versioning multiplies storage, so pair it with lifecycle rules that expire old noncurrent versions.

Encryption — S3 encrypts all new objects at rest by default (SSE-S3) as of 2023. Options:

Bucket policies vs IAM policies

This distinction confuses everyone, so pin it down: both are IAM policy documents, but they attach to different things and answer different questions.

IAM policyBucket policy
Attached toA user, group, or role (the identity)The bucket (the resource)
Answers”What can this identity do?""Who may access this bucket, and how?”
Cross-accountGrants only within your accountCan grant access to other accounts / the public
Best forStandard “which of my principals can touch which buckets”Cross-account sharing, public/anonymous access, enforcing TLS/encryption org-agnostically

Access is granted if either an applicable identity policy or the resource policy allows it (and no explicit Deny or SCP blocks it). Use IAM policies for the normal case (your own principals), and bucket policies when the rule is about the bucket itself — cross-account access, blanket TLS enforcement, or (carefully) public read for a static site. On top of both sits Block Public Access, an account/bucket-level override that prevents public grants regardless of policy — leave it on everywhere except the deliberate handful of public buckets.

Access Points and S3 Access Grants are newer tools for managing access at scale (named network/policy endpoints per application, and mapping S3 prefixes to IAM/identity-center principals) — reach for them when a single bucket policy grows unwieldy across many apps.

S3 static website hosting

S3 can serve a static site (HTML/CSS/JS) directly: enable static-website hosting, set an index and error document, and objects become web pages. In practice, front it with CloudFront (CDN) for HTTPS (S3 website endpoints are HTTP-only), custom domains, caching, and lower egress cost, using Origin Access Control (OAC) so the bucket stays private and only CloudFront can read it. The modern pattern is: private S3 bucket + CloudFront + OAC + ACM certificate — never a public bucket serving directly.

S3 event notifications

S3 can emit events on object actions (s3:ObjectCreated:*, s3:ObjectRemoved:*, and more) to trigger downstream processing:

This turns S3 into the front door of event-driven pipelines — “someone uploaded a file” becomes a trigger, not a poll.

Key Concepts

EBS — Elastic Block Store

EBS provides network-attached block volumes for EC2 — the disks your OS boots from and your databases write to. A volume lives in one AZ, attaches to one instance at a time (except io2 Multi-Attach), and persists independently of the instance’s life.

TypeClassBest forNotes
gp3General SSDDefault for most workloadsBaseline 3,000 IOPS / 125 MB/s, provision more independently; cheaper than gp2
gp2General SSD (older)LegacyIOPS scale with size; prefer gp3
io2 / io2 Block ExpressProvisioned IOPS SSDLatency-critical databasesUp to 256,000 IOPS, 99.999% durability, Multi-Attach
st1Throughput HDDBig sequential (logs, big data)Cheap, throughput-oriented, not for random I/O
sc1Cold HDDRarely accessed, cost-firstCheapest block; low throughput

gp3 is the default — it decouples IOPS/throughput from capacity (with gp2 you had to over-provision size to buy IOPS) and is cheaper. Migrate gp2 → gp3 for an easy cost win.

Snapshots are incremental, point-in-time backups of a volume stored in S3 (behind the scenes). Only changed blocks are stored after the first snapshot, so they’re space-efficient. Use them for backup, to clone volumes, to move data across AZs/Regions (copy a snapshot), and to build AMIs. Automate them with Data Lifecycle Manager or AWS Backup. Snapshots (and volumes) should be encrypted — enable default EBS encryption account-wide.

EFS and FSx — shared file storage

When many instances (or Lambda/containers) need the same files at once, block storage won’t do — you need a shared filesystem.

EFS (Elastic File System) is a fully managed, elastic NFS filesystem for Linux: mount it from thousands of instances across AZs simultaneously, and it grows and shrinks automatically with no provisioning. It has storage classes (Standard, Infrequent Access, Archive) with lifecycle management. Use it for shared content, home directories, container persistent volumes, and lift-and-shift apps expecting a POSIX filesystem. Trade-off: higher per-GB cost and latency than EBS; not for single-instance high-IOPS databases.

FSx is a family of managed filesystems for specific ecosystems:

Choose EFS for general Linux NFS; choose the matching FSx flavor when you need Windows/SMB, Lustre-class throughput, or a specific vendor filesystem’s features.

AWS Backup

AWS Backup centralizes backup across services — EBS, EFS, RDS, DynamoDB, S3, FSx, EC2, and more — under one policy-driven service. You define backup plans (schedule, retention, lifecycle to cold storage), assign resources by tag, and get a single place to monitor, restore, and prove compliance. It supports cross-Region and cross-account copy (critical for DR and for surviving an account compromise) and Backup Vault Lock (immutable, WORM backups that even admins cannot delete early). Prefer it over hand-rolled per-service snapshot scripts — one policy, one audit trail, one restore console.

Storage Gateway

Storage Gateway bridges on-premises applications to AWS storage via a local virtual appliance:

Data transfer and egress costs

The trap that catches everyone: data into AWS is free; data out to the internet costs money, and it adds up. Key rules:

Practical consequences: keep chatty services in one AZ where resilience allows, use VPC Gateway Endpoints for S3/DynamoDB to dodge NAT costs, serve downloads through CloudFront, and watch NAT Gateway data-processing charges (they meter every GB through them). Egress is the silent line item — model it before building data-heavy pipelines.

Worked example: S3 bucket policy + lifecycle

Bucket policy — enforce TLS and grant one other account read access, while your own principals use IAM:

{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "DenyInsecureTransport",
      "Effect": "Deny",
      "Principal": "*",
      "Action": "s3:*",
      "Resource": [
        "arn:aws:s3:::my-data-bucket",
        "arn:aws:s3:::my-data-bucket/*"
      ],
      "Condition": { "Bool": { "aws:SecureTransport": "false" } }
    },
    {
      "Sid": "AllowPartnerAccountRead",
      "Effect": "Allow",
      "Principal": { "AWS": "arn:aws:iam::222233334444:root" },
      "Action": ["s3:GetObject", "s3:ListBucket"],
      "Resource": [
        "arn:aws:s3:::my-data-bucket",
        "arn:aws:s3:::my-data-bucket/shared/*"
      ]
    }
  ]
}

The first statement is a blanket Deny on any non-TLS request — an explicit deny always wins, so this hardens the whole bucket regardless of identity policies. The second grants account 2222... read on the shared/ prefix only; note cross-account grants like this require a resource (bucket) policy — an IAM policy in your account can’t grant to someone else’s principal.

Lifecycle configuration — tier logs down and expire them, and clean up old versions:

{
  "Rules": [
    {
      "ID": "log-tiering-and-expiry",
      "Filter": { "Prefix": "logs/" },
      "Status": "Enabled",
      "Transitions": [
        { "Days": 30,  "StorageClass": "STANDARD_IA" },
        { "Days": 90,  "StorageClass": "GLACIER" }
      ],
      "Expiration": { "Days": 365 }
    },
    {
      "ID": "expire-old-versions",
      "Filter": {},
      "Status": "Enabled",
      "NoncurrentVersionTransitions": [
        { "NoncurrentDays": 30, "StorageClass": "GLACIER" }
      ],
      "NoncurrentVersionExpiration": { "NoncurrentDays": 90 },
      "AbortIncompleteMultipartUpload": { "DaysAfterInitiation": 7 }
    }
  ]
}

The first rule moves logs/ objects to Standard-IA at 30 days, Glacier at 90, and deletes them at a year — the canonical cost curve for observability data. The second cleans up after versioning (old versions to Glacier at 30 days, gone at 90) and aborts stalled multipart uploads after 7 days, a common source of invisible, billed-for storage. Apply with aws s3api put-bucket-lifecycle-configuration --bucket my-data-bucket --lifecycle-configuration file://lifecycle.json.

Best Practices

  1. Turn on Block Public Access account-wide, and keep it on — enable it at the account level so no bucket can accidentally go public. Grant public access only to the deliberate few buckets (and prefer CloudFront + OAC even then). Public S3 buckets are the single most common AWS data breach.
  2. Encrypt everything, and enforce TLS in transit — rely on S3 default encryption (SSE-KMS for sensitive data with an audit trail), enable default EBS encryption account-wide, and deny non-TLS requests via bucket policy. Encryption at rest is nearly free and closes an entire finding class.
  3. Apply lifecycle policies to every bucket — without them, cold data sits in expensive Standard forever and old versions accumulate silently. Tier down to IA/Glacier and expire on a schedule; also abort incomplete multipart uploads, a common hidden cost.
  4. Use Intelligent-Tiering when the access pattern is unknown — it auto-optimizes storage cost with no retrieval fees for instant tiers and no guessing. For clearly hot or clearly cold data, pick the specific class; for everything else, Intelligent-Tiering.
  5. Enable versioning on important buckets, paired with lifecycle expiry — versioning is your undo button against accidental deletes and ransomware. Add MFA Delete or Object Lock for compliance data, and expire noncurrent versions so it doesn’t multiply your bill.
  6. Prefer IAM policies for your own principals, bucket policies for the resource-level rules — keep “which of my roles can touch which buckets” in identity policies; reserve bucket policies for cross-account sharing and blanket enforcement (TLS, encryption). Don’t scatter access logic across both without reason.
  7. Front static content and downloads with CloudFront — never serve a public S3 website directly. CloudFront gives HTTPS, custom domains, caching, cheaper egress, and (via OAC) lets the bucket stay private. It is both more secure and cheaper at scale.
  8. Default EBS volumes to gp3 and migrate off gp2 — gp3 decouples IOPS/throughput from size and is cheaper for the same performance. The migration is online and one of the easiest cost wins available.
  9. Automate and encrypt snapshots; copy critical ones cross-Region — use Data Lifecycle Manager or AWS Backup for scheduled, retained, encrypted snapshots, and copy DR-critical backups to a second Region and ideally a separate account.
  10. Centralize backups with AWS Backup and lock the vault — one policy-driven service across EBS/EFS/RDS/DynamoDB/S3 beats per-service scripts. Use Backup Vault Lock (WORM) and cross-account copy so a compromised account can’t destroy your last-resort restore.
  11. Right-size the storage shape to the access pattern — one instance’s disk → EBS; many machines sharing files → EFS/FSx; blobs by key over HTTP → S3. Don’t force a shared filesystem onto what should be object storage, or vice versa.
  12. Use VPC Gateway Endpoints for S3 and DynamoDB — they route traffic privately and free, avoiding both the internet and NAT Gateway data-processing charges. For high-volume S3 access from private subnets this is a large, easy saving.
  13. Model egress before building data-heavy pipelines — internet egress and cross-AZ/Region transfer are silent, sizeable line items. Keep chatty traffic in-AZ where resilience allows, serve via CloudFront, and watch NAT Gateway per-GB processing.
  14. Enable S3 access logging or CloudTrail data events on sensitive buckets — you cannot investigate access you never recorded. Log reads/writes on regulated data, and monitor with S3 Storage Lens for cost and usage anomalies.
  15. Use S3 event notifications to build event-driven pipelines — let uploads trigger Lambda/SQS/EventBridge instead of polling. It’s cheaper, faster, and the idiomatic way to process data as it lands.
  16. Tag storage for cost allocation and lifecycle automation — tags on buckets, volumes, and snapshots drive cost reports, backup assignment, and cleanup rules. Untagged storage is unaccountable and tends to become orphaned, billed-for cruft.

References