Incident Management & PostmortemIncident Management & Postmortems
Thuộc bộ kiến thức Engineering Manager Roadmap.
Tổng quan
Khi production sập lúc 2 giờ sáng, engineering manager (EM) hiếm khi là người trực tiếp gõ lệnh vào terminal. Đó là chủ đích. Phần việc “tay trên bàn phím” — chẩn đoán, khoanh vùng (containment), xử lý sự cố, cơ chế kỹ thuật của triage, containment, eradication và xử lý bằng chứng forensic — đã được trình bày chi tiết trong DevSecOps — Incident Response & Digital Forensics. Note này chủ đích không lặp lại nội dung đó. Nó tập trung vào công việc thực sự của EM trong và sau một incident: giữ cho phản ứng (response) có tổ chức mà không cản trở nó, quản lý con người và các stakeholder xung quanh phần việc kỹ thuật, và — sau khi lửa đã tắt — đảm bảo tổ chức thực sự tốt lên đo lường được, chứ không chỉ đơn giản là chuyển sang việc khác.
Hai kiểu thất bại lặp đi lặp lại ở những manager mới làm quen với incident. Kiểu thứ nhất là manager không cưỡng lại được việc trở thành người có chuyên môn kỹ thuật cao nhất trong phòng, việc này kéo họ vào việc gõ phím thay vì làm công việc điều phối và che chắn (shielding) mà chỉ họ mới làm được. Kiểu thứ hai là manager coi postmortem là một thủ tục hành chính — một tài liệu để lưu trữ rồi quên đi — thay vì thời điểm có đòn bẩy học hỏi lớn nhất mà một tổ chức kỹ thuật có được. Cả hai kiểu thất bại đều có thể tránh, và tránh được chúng chính là phần lớn nội dung của note này.
Sợi dây kết nối xuyên suốt incident response, sức khỏe on-call, và postmortem chính là trust (niềm tin). Người phản ứng (responder) làm việc tốt hơn khi họ tin rằng manager đứng sau lưng họ trong lúc xử lý sự cố. Đội nhóm chỉ viết postmortem trung thực khi họ tin rằng sự trung thực đó sẽ không bị dùng để chống lại họ. Và tổ chức chỉ nhanh hơn theo thời gian khi họ tin rằng action item từ incident trước đó đáng để làm, bởi vì action item từ incident trước nữa đã thực sự được hoàn thành. Xói mòn niềm tin ở bất kỳ điểm nào trong ba điểm này, cả hệ thống sẽ suy thoái: responder ngừng escalate sớm, postmortem bị “làm sạch” (sanitize), và cùng một loại incident lại tái diễn.
Kiến thức nền tảng
Vai trò của manager trong một incident: không nhất thiết là Incident Commander
Một bản năng dễ hiểu nhưng thường gặp ở manager mới là cho rằng “làm chủ” đội nhóm nghĩa là làm chủ luôn incident. Trong một quy trình incident response trưởng thành, đó thường là mô hình sai. Vai trò Incident Commander (IC) — được trình bày chi tiết trong note incident response của DevSecOps — là một vai trò xoay vòng (rotating), thường mang tính kỹ thuật, với nhiệm vụ điều phối chính bản thân response: xác định severity, chỉ đạo technical lead, quyết định khi nào escalate, và quyết định khi nào incident được coi là đã giải quyết xong. Vai trò này chủ đích không gắn với cấp bậc trong sơ đồ tổ chức, vì IC phù hợp nhất cho một incident hỏng database có thể là một database engineer senior thấp hơn EM ba cấp, và việc ép manager vào ghế đó hoặc làm chậm response, hoặc biến manager thành nút thắt cổ chai (bottleneck) cho những quyết định mà họ không phải người phù hợp nhất để đưa ra.
Công việc thực sự của EM trong một incident đang diễn ra chạy song song với IC, chứ không nằm dưới quyền IC:
- Che chắn responder khỏi những phân tâm. Mỗi luồng Slack thêm, mỗi DM “chỉ hỏi nhanh update thôi?”, hay một exec ghé qua hỏi trạng thái đều là sự tập trung bị đánh cắp từ người đang cố gắng ngăn chảy máu hệ thống. Đóng góp thời gian thực có giá trị nhất của manager thường là tự mình hấp thụ toàn bộ những gián đoạn đó, để IC và technical lead chỉ cần nói chuyện với một người duy nhất.
- Chịu trách nhiệm giao tiếp với stakeholder song song. Cần có ai đó cập nhật đều đặn cho leadership, support, sales, và (với incident ảnh hưởng khách hàng) status page, bằng ngôn ngữ dễ hiểu, mà không cần IC phải dừng lại dịch chi tiết kỹ thuật thành tác động kinh doanh mỗi mười phút. Đây thường là công việc của manager, phối hợp chặt chẽ với — nhưng tách biệt khỏi — vai trò Communications Lead được mô tả trong note DevSecOps dành riêng cho security incident.
- Đưa ra các quyết định về nguồn lực (resourcing). Có nên gọi thêm người từ team khác không? Có nên tạm dừng một release không quan trọng để kỹ sư không phải context-switch không? Có nên bảo ai đó dừng lại và đi ngủ vì họ đã on call sáu tiếng và bắt đầu mắc lỗi không? Đây là các quyết định quản lý, không phải kỹ thuật, và chúng thường vô hình với tất cả mọi người trừ người không phải tự đưa ra quyết định đó vì manager đã làm giúp.
- Giữ tổ chức bình tĩnh mà không giảm nhẹ vấn đề. Một manager biểu hiện hoảng loạn rõ rệt sẽ khuếch đại sự hoảng loạn xuyên suốt chuỗi báo cáo; một manager giả vờ mọi thứ ổn sẽ mất uy tín ngay khi sự thật lộ ra. Tông giọng đúng là bình tĩnh, trung thực, cụ thể: chúng ta biết gì, chưa biết gì, đang làm gì, và khi nào có update tiếp theo.
Không điều nào ở trên thay thế phán đoán kỹ thuật — một manager có nền tảng kỹ thuật mạnh hoàn toàn có thể nhảy vào làm individual contributor nếu đội nhóm thiếu người, đặc biệt ở công ty nhỏ. Nhưng đó là một đánh đổi có ý thức so với công việc điều phối và che chắn nêu trên, chứ không phải kỳ vọng mặc định của vai trò.
Phân loại severity: góc nhìn quản lý
Tiêu chí kỹ thuật cho severity (blast radius, độ nhạy cảm của dữ liệu, compromise đã xác nhận hay mới nghi ngờ) đã được trình bày trong bảng severity của note DevSecOps. Từ góc độ quản lý, điều quan trọng về phân loại severity là nó kích hoạt (trigger) điều gì — vì đó chính là đòn bẩy mà manager thực sự nắm được:
| Severity | Ai bị page | Ai được thông báo | Hành động điển hình của manager |
|---|---|---|---|
| Critical (SEV-1) — outage lớn hoặc breach đang hoạt động ảnh hưởng nghiêm trọng đến kinh doanh | Kỹ sư on-call, IC, EM, thường thêm một responder phụ | Leadership cấp cao ngay lập tức, khách hàng/status page nếu ảnh hưởng khách hàng, legal/PR nếu là security incident | EM tham gia kênh incident theo thời gian thực, đảm nhận comms, có thể gọi thêm người |
| High (SEV-2) — suy giảm đáng kể, outage một phần, security event đã được khoanh vùng | Kỹ sư on-call, IC | EM, các team lead liên quan, đội support | EM theo dõi, check-in định kỳ, đảm bảo on-call không đơn độc |
| Medium (SEV-3) — bug hoặc bất thường ảnh hưởng hạn chế | Kỹ sư on-call | Kênh của team, theo dõi dưới dạng ticket | EM biết qua tóm tắt async, thường không bị page |
| Low (SEV-4) — vấn đề thẩm mỹ hoặc không khẩn cấp | Không page | Backlog | Manager không cần can thiệp ngoài lập kế hoạch bình thường |
Ánh xạ (mapping) này bị sai theo bất kỳ hướng nào cũng gây tổn thất thực sự. Page quá mạnh tay cho các vấn đề severity thấp huấn luyện mọi người bỏ qua page (“alert fatigue”), chính là kiểu thất bại biến một SEV-1 thực sự thành một phản ứng chậm chạp vì ai cũng nghĩ đó chỉ là nhiễu. Escalate quá ít — giữ leadership trong bóng tối về điều thực sự nghiêm trọng — khiến manager mất uy tín ngay khi sự việc tự lộ ra, và tước đi khả năng của leadership trong việc đưa ra quyết định thông tin đầy đủ (hoãn launch, gọi thêm nguồn lực, chủ động giao tiếp với khách hàng) trong khi vẫn còn thời gian để làm điều đó với chi phí thấp.
Sự công bằng và sức khỏe của rotation on-call
Một quy trình incident response chỉ hoạt động nếu rotation on-call phía sau nó bền vững. Một rotation làm kiệt sức con người sẽ tạo ra các phản ứng chậm hơn, dễ sai sót hơn theo thời gian dù quy trình trên giấy có tốt đến đâu — responder mệt mỏi bỏ sót thứ này thứ kia, escalate trễ vì không muốn “làm phiền” ai đó, và cuối cùng rời team. Khía cạnh well-being được trình bày trong Team Culture & Well-Being; những phần đặc thù cho incident management mà manager nên chủ động theo dõi là:
- Số lượng page trên mỗi người, theo thời gian. Một rotation page ai đó năm lần mỗi đêm, mỗi đêm họ trực, là không bền vững dù morale trông có vẻ tốt trong buổi 1:1 — hãy sửa các alert gây nhiễu ở gốc, đừng chỉ yêu cầu mọi người chịu đựng chúng.
- Sự công bằng trong thành phần rotation. Có phải cùng một hoặc hai người phải phụ trách disproportionately đêm khuya, cuối tuần, ngày lễ vì họ là “người hiểu hệ thống”, trong khi những người khác thoải mái không cần biết hệ thống đó? Đó là một single point of failure được ngụy trang thành sự tiện lợi, và cũng là con đường nhanh nhất để đốt cháy kỹ sư am hiểu nhất của bạn.
- Thời gian phục hồi sau một ca on-call tồi tệ. Một đêm xử lý incident vất vả nên có một hệ quả rõ ràng, không mơ hồ: nghỉ bù, ngày hôm sau nhẹ nhàng hơn, không kỳ vọng người đó tươi tỉnh xuất hiện trong buổi họp lập kế hoạch 9 giờ sáng sau khi vừa xử lý xong lúc 3 giờ sáng. Nếu không có cơ chế phục hồi hiển nhiên, mọi người sẽ học được rằng giỏi xử lý incident bị “phạt” bằng nhiều việc hơn, chứ không phải được thưởng.
- Onboarding người mới vào rotation một cách chủ đích, shadowing trước khi trở thành primary, thay vì ném ai đó vào mà chưa được thực hành khi sự cố đầu tiên xảy ra. Điều này giữ quy mô rotation lành mạnh theo thời gian thay vì co lại thành “ai ở đây lâu nhất.”
Một manager coi danh sách rotation chỉ là bài toán xếp lịch, thay vì bài toán well-being và thiết kế hệ thống, cuối cùng sẽ hết người sẵn lòng tham gia.
Khái niệm chính
Vì sao blameless postmortem quan trọng — cam kết văn hóa, không phải định dạng tài liệu
Điều quan trọng nhất cần hiểu về blameless postmortem là chữ “blameless” mô tả một cam kết văn hóa, không phải một tiêu đề mục. Đội nhóm đôi khi áp dụng template — Summary, Timeline, Action Items — trong khi vẫn tổ chức các buổi họp mang tính đổ lỗi, và template không làm gì để sửa điều đó. Cơ chế mà tính blameless bảo vệ khỏi rất đơn giản và đã được ghi nhận rõ trong tài liệu về incident response: những người sợ bị trừng phạt vì kể lại trung thực điều đã xảy ra sẽ bỏ sót thông tin, giảm nhẹ vai trò của mình trong sự việc, hoặc né tránh việc nêu lên mối lo ngại ngay từ đầu — và mỗi mẩu thông tin bị thiếu là một phần của hệ thống mà tổ chức không thể sửa, vì tổ chức không bao giờ biết đến nó.
Cụ thể, trong một văn hóa mang tính đổ lỗi:
- Kỹ sư đã gõ lệnh kích hoạt outage sẽ nói càng ít càng tốt về chuỗi sự kiện, vì sợ bị chỉ mặt điểm tên — nên chuỗi quyết định thực sự dẫn đến sự cố không bao giờ lộ ra.
- Kỹ sư nhận thấy điều gì đó “không ổn” một giờ trước incident, nhưng không nói gì vì không chắc chắn và không muốn tỏ ra hoang mang, cũng sẽ im lặng về điều đó trong postmortem — nên tín hiệu cảnh báo sớm có thể ngăn tái diễn bị mất.
- Các incident tương lai giải quyết chậm hơn, vì bản năng của responder dưới áp lực trở thành “đừng làm mọi thứ tệ hơn cho bản thân mình” thay vì “nêu ra mọi thứ liên quan nhanh nhất có thể,” ngay cả giữa lúc đang xử lý incident.
Blameless là giải pháp thực tế: nó coi incident là triệu chứng của một hệ thống đã để cho một người hợp lý, có thiện ý gây ra tổn hại, và hướng toàn bộ năng lượng phân tích vào hệ thống thay vì vào cá nhân. Điều này không có nghĩa là trách nhiệm cá nhân biến mất — sự cẩu thả lặp đi lặp lại hoặc cố ý vi phạm chính sách là một cuộc trò chuyện về hiệu suất (performance), được tổ chức riêng biệt, không lồng ghép vào buổi review blameless. Nó có nghĩa là giả định mặc định, với đại đa số incident, là người liên quan đã làm việc một cách hợp lý với thông tin và công cụ họ có, và nhiệm vụ của postmortem là tìm ra vì sao hệ thống lại để hành động hợp lý đó gây ra outage.
Sống theo tinh thần này với tư cách manager đòi hỏi nỗ lực chủ động, lặp đi lặp lại, không phải một lần thông báo:
- Tự mình làm gương. Lần đầu tiên manager nói “ừ, nếu bạn kiểm tra kỹ hơn trước khi deploy” trong một postmortem — dù thông cảm, dù chỉ một lần — cả phòng sẽ tự điều chỉnh lại và postmortem tiếp theo sẽ mỏng đi.
- Hỏi “vì sao hệ thống cho phép điều này xảy ra” trước khi hỏi “ai đã làm việc này.” Chuyển mọi bản năng gọi tên một người thành một câu hỏi về guardrail đã thiếu.
- Bảo vệ buổi họp khỏi trở thành bằng chứng đánh giá hiệu suất. Postmortem không nên được viện dẫn trong performance review như bằng chứng chống lại những người liên quan. Nếu mọi người tin rằng điều họ nói trong postmortem có thể quay lại trong review tiếp theo, họ sẽ ngừng nói thật.
- Cảm ơn mọi người vì đã nêu lên sự thật khó nghe, công khai nếu phù hợp, đặc biệt khi sự thật đó bất tiện hoặc chỉ ra một quyết định mà chính manager đã đưa ra.
Root cause analysis: kỹ thuật “5 Whys” và cái bẫy “human error”
Kỹ thuật 5 Whys là một điểm khởi đầu đơn giản, hiệu quả cho root cause analysis: nêu vấn đề, hỏi “vì sao điều đó xảy ra,” và với mỗi câu trả lời, lại hỏi “vì sao” tiếp, thường là năm lần, cho đến khi đạt được điều thực sự có thể hành động để sửa. Nó hiệu quả chính vì buộc đội nhóm vượt qua câu trả lời đầu tiên, hời hợt nhất.
Cách phổ biến nhất khiến 5 Whys đi sai hướng là dừng lại ở “human error.” Xét một chuỗi điển hình:
- Vì sao site sập? — Một cấu hình (config) sai đã được deploy lên production.
- Vì sao config sai lại được deploy? — Kỹ sư đã không test nó trong staging trước.
- Vì sao họ không test trong staging trước? — Họ tự tin rằng đó là một thay đổi tầm thường và staging cảm giác như là thủ tục không cần thiết dưới áp lực thời gian.
Một đội dừng lại ở đây sẽ kết luận “human error: kỹ sư lẽ ra nên test trong staging” — và action item trở thành “nhắc kỹ sư test trong staging,” điều này không sửa được gì, vì kỹ sư mệt mỏi hoặc vội vàng tiếp theo sẽ đưa ra cùng một lựa chọn hợp lý-trong-khoảnh-khắc đó. Bước đi đúng đắn là tiếp tục hỏi vì sao hệ thống lại để quyết định của con người đó gây ra tác động:
- Vì sao việc bỏ qua staging lại dẫn đến outage toàn diện thay vì một lỗi bị bắt kịp thời? — Không có cổng validate tự động (automated validation gate) nào giữa “config được commit” và “config chạy thật trên production.”
- Vì sao không có validation gate? — Config deploy chưa bao giờ được đưa vào cùng các rào chắn an toàn CI/CD như code deploy, vì ban đầu chúng được coi là “thay đổi ad hoc, rủi ro thấp.”
Bây giờ các action item mới thực sự và bền vững: thêm automated config validation, yêu cầu thay đổi config phải đi qua cùng pipeline như code, thêm canary hoặc staged rollout cho thay đổi config. “Human error” gần như không bao giờ là root cause — nó là triệu chứng cuối cùng có thể nhìn thấy của một hệ thống không có guardrail để chặn một lỗi bình thường của con người trước khi nó gây hại. Công việc của manager khi review bất kỳ postmortem nào là tiếp tục đẩy qua câu trả lời đầu tiên gọi tên hành động của một người, hướng tới thuộc tính hệ thống đã để hành động đó gây ảnh hưởng.
Hai cái bẫy liên quan cần lưu ý khi manager điều phối phân tích này:
- Nhiều root cause, không phải một. Các incident thực tế gần như không bao giờ chỉ có một nguyên nhân; chúng thường là hai hoặc ba điểm yếu độc lập tình cờ trùng nhau (một bug, cộng thêm một test còn thiếu, cộng thêm một alert đã bị tắt tiếng vì lý do “hay báo giả” đã biết, cộng thêm một kỹ sư on-call chưa quen với subsystem đó). Khăng khăng kể một câu chuyện “root cause là X” duy nhất sẽ loại bỏ thông tin hữu ích — template postmortem bên dưới chủ đích dùng “Root cause(s)” ở dạng số nhiều.
- Contributing factor không giống root cause. Một điều gì đó có thể làm incident tệ hơn đáng kể (ví dụ: một runbook lỗi thời làm chậm response 20 phút) mà không phải là lý do incident bắt đầu. Cả hai đều đáng ghi lại, nhưng nhập nhằng chúng làm mờ action item — sửa runbook ngăn một response chậm lần sau, không ngăn lần tái diễn tiếp theo của bug gốc.
Incident postmortem vs. project postmortem
Việc nhìn lại một cách blameless không chỉ dành cho outage. Cùng những nguyên tắc — kể lại trung thực, tập trung vào hệ thống thay vì đổ lỗi cá nhân, action item thực sự được theo dõi — cũng áp dụng tốt cho một dự án hoặc sáng kiến đã được triển khai (hoặc triển khai tệ). Một dự án trễ ba tháng, ship với scope bị giảm đáng kể, hoặc cần một cuộc “giải cứu” tốn kém vào phút chót xứng đáng nhận cùng mức độ nghiêm túc như một outage, vì câu hỏi cốt lõi là giống hệt nhau: tổ chức đã học được gì, và sẽ thay đổi gì nhờ đó?
| Khía cạnh | Incident postmortem | Project postmortem |
|---|---|---|
| Kích hoạt | Một outage, suy giảm dịch vụ, hoặc security event — thường có khung thời gian ngắn, cấp tính | Một cột mốc dự án/sáng kiến: launch, hủy bỏ, hoặc trễ lịch/hụt scope đáng kể |
| Khung thời gian phân tích điển hình | Vài phút đến vài ngày (timeline của incident) | Vài tuần đến vài tháng (timeline của dự án) |
| Trọng tâm chính | Điều gì hỏng về mặt kỹ thuật, và điều gì để nó hỏng | Những quyết định lập kế hoạch, ước lượng, giao tiếp, hoặc phân định scope nào dẫn đến kết quả |
| Root cause thường gặp | Thiếu guardrail, test chưa đủ, khoảng trống trong alerting, quyền sở hữu (ownership) không rõ trong lúc xử lý | Yêu cầu không rõ ràng, ước lượng phi thực tế, dependency bất ngờ, scope creep, thiếu nhân sự, khoảng trống phối hợp liên team |
| Action item điển hình | Thêm monitoring/alerting, thêm automated validation, cập nhật runbook, sửa loại bug cụ thể | Thay đổi quy trình ước lượng, điều chỉnh cách theo dõi dependency, sửa lại cách phê duyệt thay đổi scope, thay đổi mô hình nhân sự cho dự án tương lai |
| Ai nên tham dự | Responder, IC, chủ sở hữu dịch vụ liên quan, EM | Project lead, EM, các contributor chính, đôi khi cả stakeholder yêu cầu dự án (PM, sales, exec sponsor) |
| Kiểu thất bại thường gặp | Dừng lại ở “human error”; không theo dõi action item | Biến thành buổi đổ lỗi “ai làm trễ deadline”; coi đó là một buổi retro một lần rồi không theo dõi tiếp |
Bản năng bỏ qua project postmortem rất mạnh chính vì không có outage nào ép buộc cuộc trò chuyện — không có gì đang cháy, nên rất dễ để đội nhóm chuyển thẳng sang dự án tiếp theo. Cưỡng lại bản năng đó là một trong những việc có đòn bẩy cao nhất mà một EM có thể làm, vì sự rối loạn chức năng ở cấp dự án (ước lượng thiếu thời gian mãn tính, một vấn đề dependency liên team lặp lại, một quy trình lập kế hoạch không bao giờ tính đến thời gian onboarding) tích lũy âm thầm qua mọi dự án tiếp theo, theo cách mà một outage đơn lẻ hiếm khi làm được.
Template postmortem blameless đầy đủ
Template dưới đây áp dụng trực tiếp cho incident postmortem; với project postmortem, hãy điều chỉnh “Impact” thành schedule/scope/cost và “Timeline” thành các điểm quyết định chính của dự án thay vì timestamp incident theo từng phút.
# Postmortem: [Tên Incident/Dự án]
Trạng thái: [Draft / In Review / Final]
Ngày xảy ra incident: [YYYY-MM-DD]
Ngày viết postmortem: [YYYY-MM-DD]
Tác giả: [Tên]
Severity: [SEV-1 / SEV-2 / ...]
## Summary
Tóm tắt 2–4 câu, ngôn ngữ dễ hiểu, về điều gì đã xảy ra, ai bị ảnh hưởng,
và cách nó được giải quyết. Người ngoài team không có context thêm cũng
phải hiểu được.
## Impact
- Thời lượng: [thời điểm bắt đầu] đến [thời điểm kết thúc] ([tổng thời lượng])
- Người dùng/khách hàng bị ảnh hưởng: [số lượng hoặc phần trăm, hoặc "chỉ nội bộ"]
- Dịch vụ bị ảnh hưởng: [danh sách]
- Tác động doanh thu/kinh doanh: [nếu áp dụng và đã biết]
- Tác động dữ liệu: [mất, hỏng, hoặc lộ dữ liệu nếu có]
- Tác động SLA/SLO: [error budget đã dùng, vi phạm SLA nếu có]
## Timeline
Tất cả thời gian theo [múi giờ]. Bao gồm phát hiện, escalation, và mọi
hành động/quyết định quan trọng, không chỉ phần khắc phục.
| Thời gian | Sự kiện |
|-------|-------|
| 14:02 | Alert kích hoạt: tỷ lệ lỗi tăng cao ở checkout-service |
| 14:05 | Kỹ sư on-call xác nhận page |
| 14:12 | Incident được khai báo SEV-2, IC được chỉ định |
| 14:20 | Giả thuyết root cause: config deploy sai lúc 13:55 |
| 14:25 | Bắt đầu rollback config |
| 14:31 | Tỷ lệ lỗi trở về mức bình thường |
| 14:45 | Incident được khai báo đã giải quyết; tiếp tục theo dõi |
## Root Cause(s)
Liệt kê mọi nguyên nhân đóng góp độc lập được xác định qua root cause
analysis (ví dụ: 5 Whys). Đa số incident có nhiều hơn một.
1. [Root cause 1 — khoảng trống kỹ thuật/quy trình, không phải tên một người]
2. [Root cause 2, nếu có]
## Contributing Factors
Những yếu tố làm tác động của incident tệ hơn hoặc làm chậm phản ứng,
mà không phải là lý do nó bắt đầu.
- [ví dụ: "Runbook cho dịch vụ này đã lỗi thời 8 tháng"]
- [ví dụ: "Alert đã bị tắt tiếng vì lý do 'hay báo giả' đã biết và vẫn bị tắt"]
## What Went Well
Những điều cụ thể, rõ ràng đã hiệu quả, đáng được củng cố và lặp lại.
- [ví dụ: "Rollback tự động rút ngắn thời gian phục hồi từ ~30 phút xuống ~6 phút"]
- [ví dụ: "Bàn giao IC lúc đổi ca diễn ra suôn sẻ, không mất context"]
## What Went Poorly
Những điều cụ thể, rõ ràng chưa hiệu quả — được đóng khung ở cấp hệ thống.
- [ví dụ: "Không có alert nào kích hoạt trên leading indicator thực sự;
chúng ta chỉ bị page khi tỷ lệ lỗi ảnh hưởng khách hàng đã cao"]
- [ví dụ: "Mất 12 phút để xác định ai sở hữu dịch vụ bị ảnh hưởng"]
## Action Items
Mỗi mục đều có owner và ngày hạn. Không có owner, không có ngày hạn —
không phải là action item.
| Action item | Owner | Ngày hạn | Ưu tiên | Trạng thái |
|---|---|---|---|---|
| Thêm cổng validate tự động cho config deploy | @alice | 2026-08-01 | P1 | Chưa bắt đầu |
| Cập nhật runbook on-call cho checkout-service | @bob | 2026-07-25 | P2 | Đang thực hiện |
| Thêm alert leading-indicator cho queue depth | @carol | 2026-08-08 | P1 | Chưa bắt đầu |
## Lessons Learned
Một tổng hợp tường thuật ngắn — team khác, không chỉ team này, nên rút ra
điều gì từ incident này? Nên cảnh giác với mẫu hình (pattern) gì ở nơi khác?
[ví dụ: "Bất kỳ dịch vụ nào chấp nhận thay đổi config ngoài pipeline
deploy code thông thường đều là điểm mù. Chúng ta nên rà soát các dịch
vụ khác có cùng khoảng trống này."]
Best Practices
- Tách bạch rõ ràng vai trò IC khỏi vai trò manager, ngay cả ở công ty nhỏ nơi manager đủ năng lực kỹ thuật để làm IC. Nói rõ ràng, trong kênh incident, ai đang đóng vai trò nào, để responder biết escalate quyết định kỹ thuật cho ai và escalate nhu cầu nguồn lực/giao tiếp cho ai.
- Viết bản cập nhật cho khách hàng/exec trước khi bị hỏi. Một nhịp độ chủ động, đều đặn (ví dụ: mỗi 30 phút cho SEV-1) ngăn những tin nhắn “có update gì chưa?” tự chúng trở thành một phân tâm cho quá trình xử lý.
- Tổ chức postmortem khi ký ức còn tươi mới nhưng cảm xúc đã lắng xuống — thường trong vài ngày làm việc sau khi giải quyết. Quá sớm thì mọi người vẫn còn căng thẳng và phòng thủ; quá muộn thì chi tiết bị tái dựng lại thay vì được nhớ chính xác.
- Điều phối, đừng thẩm vấn. Với tư cách manager trong phòng, câu hỏi của bạn nên nghe như “thông tin gì lẽ ra sẽ giúp bạn đưa ra quyết định khác vào lúc đó?” chứ không phải “sao lúc đó bạn không kiểm tra X trước?” Cách đặt câu hỏi thứ hai, dù không cố ý, vẫn là đổ lỗi kèm dấu hỏi chấm.
- Theo dõi action item trong cùng backlog với công việc feature, với cùng mức độ hiển thị và grooming, không phải trong một tài liệu “action item postmortem” riêng biệt bị lãng quên. Một action item không được ưu tiên hóa so với công việc feature sẽ phải cạnh tranh sự chú ý và thường xuyên thua cuộc. Nếu leadership không chịu ưu tiên các action item về độ tin cậy so với feature trên roadmap, đó là một cuộc trò chuyện về nguồn lực mà manager cần chủ động nêu ra, chứ không phải che đậy bằng cách viết action item mà không ai bao giờ xem lại.
- Gán owner thực sự và ngày hạn thực sự cho mọi action item, và theo dõi những mục quá hạn giống như cách bạn theo dõi một deadline feature bị trễ — công khai, trong cùng các diễn đàn, với cùng mức độ nghiêm túc.
- Chia sẻ postmortem rộng rãi, không chỉ trong team bị ảnh hưởng. Một bản tin nội bộ ngắn, một diễn đàn “review postmortem” định kỳ xuyên team, hoặc một kho lưu trữ postmortem có thể tìm kiếm ngăn cùng một loại incident tái diễn ở một dịch vụ khác sáu tháng sau, chỉ vì không ai ngoài team gốc từng nghe về nó.
- Theo dõi các root cause lặp lại qua nhiều postmortem trong một hoặc hai quý — ba incident không liên quan mà đều truy về “không có staging environment tương đồng với production” là một tín hiệu chiến lược, không phải ba ngày xui xẻo riêng lẻ, và nó xứng đáng có một sáng kiến riêng thay vì ba action item rời rạc.
- Chạy project postmortem theo một nhịp độ cố định (ví dụ: sau mỗi lần launch lớn, bất kể kết quả tốt hay xấu), không chỉ khi có điều gì đó rõ ràng đi sai — chỉ nhìn lại khi thất bại dạy tổ chức rằng những việc “ổn” không có gì để học, điều này hiếm khi đúng.
- Bảo vệ tính bền vững của on-call như một đầu vào liên tục cho chất lượng incident, không phải một vấn đề HR riêng biệt — xem Team Culture & Well-Being để có bức tranh burnout rộng hơn, và theo dõi các chỉ số sức khỏe rotation đều đặn như các chỉ số incident.
- Phân biệt một incident đơn lẻ với một khủng hoảng (crisis). Khi một incident leo thang thành thứ đe dọa doanh nghiệp một cách rộng hơn — data breach lớn có rủi ro pháp lý, outage kéo dài nhiều ngày, sự kiện gây tổn hại danh tiếng đáng kể — phản ứng chuyển từ incident management tiêu chuẩn sang phạm vi business continuity; xem Crisis Leadership & Business Continuity để có góc nhìn rộng hơn đó.
Tài liệu tham khảo
- Google SRE Book — Postmortem Culture: Learning from Failure
- Google SRE Workbook — Postmortem Culture: Learning from Failure
- PagerDuty — Incident Response Documentation
- PagerDuty — Postmortems
- Atlassian — Incident Postmortems
- NIST SP 800-61 — Computer Security Incident Handling Guide
- DevSecOps — Incident Response & Digital Forensics
Part of the Engineering Manager Roadmap knowledge base.
Overview
When production breaks at 2 a.m., the engineering manager is rarely the person typing commands into a terminal. That is by design. The hands-on-keyboard work of diagnosing, containing, and fixing an incident — the technical mechanics of triage, containment, eradication, and forensic evidence handling — is covered in depth in DevSecOps — Incident Response & Digital Forensics. This note deliberately does not repeat that material. It covers the job the EM actually does during and after an incident: keeping the response organized without getting in its way, managing the humans and the stakeholders around the technical work, and — once the fire is out — making sure the organization gets measurably better instead of just moving on.
Two failure modes show up again and again in managers who are new to incidents. The first is the manager who cannot resist becoming the most senior person in the room technically, which pulls them onto the keyboard and away from the coordination and shielding work that only they can do. The second is the manager who treats the postmortem as a compliance exercise — a document to file and forget — rather than the single highest-leverage moment an engineering organization gets to actually learn something. Both failure modes are avoidable, and avoiding them is most of what this note is about.
The connecting thread across incident response, on-call health, and postmortems is trust. Responders perform better when they trust that the manager has their back during the incident. Teams write honest postmortems only when they trust that honesty will not be used against them. And organizations only get faster over time when they trust that action items from the last incident were worth doing, because the ones from the incident before that actually got done. Erode trust at any of these three points and the whole system degrades: responders stop escalating early, postmortems get sanitized, and the same incident class recurs.
Fundamentals
The manager’s role during an incident: not the Incident Commander
A common and understandable instinct for a new manager is to assume that being “in charge” of the team means being in charge of the incident. In a mature incident response setup, that is usually the wrong model. The Incident Commander (IC) role — defined in detail in the DevSecOps incident response note — is a rotating, often technical role whose job is coordination of the response itself: declaring severity, directing the technical leads, deciding when to escalate, and deciding when the incident is resolved. It is deliberately not tied to org-chart seniority, because the best IC for a database corruption incident might be a senior database engineer three levels below the EM, and forcing the manager into that seat either slows the response down or turns the manager into a bottleneck for decisions they are not best positioned to make.
The EM’s actual job during a live incident runs in parallel to the IC’s, not underneath it:
- Shield the responders from distraction. Every additional Slack thread, “just a quick update?” DM, or drive-by exec asking for a status check is attention stolen from someone trying to stop the bleeding. The manager’s most valuable real-time contribution is often absorbing every one of those interruptions personally so the IC and technical leads only have to talk to one person.
- Own stakeholder communication in parallel. Someone needs to keep leadership, support, sales, and (for customer-facing incidents) the status page updated on a steady cadence, in plain language, without needing the IC to stop and translate technical detail into business impact every ten minutes. This is frequently the manager’s job, done in close coordination with — but separate from — the Communications Lead role described in the DevSecOps note for security incidents specifically.
- Make resourcing calls. Should someone be paged in from another team? Should a non-critical release be halted so engineers aren’t context-switching? Should someone be told to stop and go to sleep because they’ve been on for six hours and are making mistakes? These are management calls, not technical ones, and they are usually invisible to everyone except the person who didn’t have to make them because the manager already did.
- Keep the org calm without minimizing the problem. A manager who is visibly panicked amplifies panic through the whole reporting chain; a manager who pretends everything is fine loses credibility the moment the truth comes out. The right tone is steady, honest, specific: what we know, what we don’t yet know, what we’re doing about it, when the next update comes.
None of this replaces technical judgment — a manager with strong technical background may well jump in as an individual contributor if the team is short-staffed, especially at a smaller company. But that is a conscious trade-off against the coordination and shielding work described above, not the default expectation of the role.
Severity classification: a management lens
The technical criteria for severity (blast radius, data sensitivity, confirmed vs. suspected compromise) are covered in the DevSecOps note’s severity table. From a management standpoint, what matters about severity classification is what it triggers — because that is the lever a manager actually pulls:
| Severity | Who gets paged | Who gets informed | Manager’s typical action |
|---|---|---|---|
| Critical (SEV-1) — major outage or active breach with real business impact | On-call engineer, IC, EM, often a secondary responder | Executive leadership immediately, customers/status page if customer-facing, legal/PR for security incidents | EM joins the incident channel in real time, takes over comms, may pull in extra hands |
| High (SEV-2) — significant degradation, partial outage, contained security event | On-call engineer, IC | EM, affected team leads, support team | EM monitors, checks in periodically, ensures on-call isn’t isolated |
| Medium (SEV-3) — limited-impact bug or anomaly | On-call engineer | Team channel, tracked as a ticket | EM aware via async summary, not usually paged |
| Low (SEV-4) — cosmetic or non-urgent issue | No page | Backlog | No manager involvement beyond normal planning |
Getting this mapping wrong in either direction has a real cost. Paging too aggressively for low-severity issues trains people to ignore pages (“alert fatigue”), which is exactly the failure mode that turns a real SEV-1 into a slow response because everyone assumes it’s noise. Under-escalating — keeping leadership in the dark on something that’s actually serious — costs the manager credibility the moment it surfaces on its own, and removes leadership’s ability to make informed calls (delaying a launch, pulling in more resources, getting ahead of customer communication) while there was still time to do so cheaply.
On-call fairness and rotation health
An incident response process only works if the on-call rotation behind it is sustainable. A rotation that burns people out produces slower, more error-prone responses over time even if the process on paper is excellent — tired responders miss things, escalate late out of a desire not to “bother” anyone, and eventually leave the team. This is covered from the well-being angle in Team Culture & Well-Being; the incident-management-specific pieces a manager should actively track are:
- Page volume per person, over time. A rotation that pages someone five times a night, every night they’re on, is not sustainable no matter how good morale looks in a 1:1 — fix the underlying noisy alerts, don’t just ask people to tolerate them.
- Fairness of rotation composition. Is the same one or two people disproportionately covering nights, weekends, and holidays because they’re “the ones who know the system,” while everyone else stays comfortably unfamiliar with it? That is a single point of failure disguised as a convenience, and it is also the fastest route to burning out your most knowledgeable engineer.
- Recovery time after a bad on-call shift. A rough overnight incident should have a visible, unambiguous consequence: time off in lieu, a lighter next day, no expectation of showing up bright-eyed to a 9 a.m. planning meeting after a 3 a.m. resolution. If there’s no visible recovery mechanism, people learn that being good at incidents is punished with more work, not rewarded.
- Onboarding new people into the rotation deliberately, shadowing before going primary, rather than throwing someone in without practice the first time something breaks. This keeps rotation size healthy over time instead of shrinking to “whoever’s been here longest.”
A manager who treats the rotation roster as a scheduling problem, rather than a well-being and system-design problem, will eventually run out of people willing to be on it.
Key Concepts
Why blameless postmortems matter — a culture commitment, not a document format
The single most important thing to understand about blameless postmortems is that “blameless” describes a culture commitment, not a section heading. Teams sometimes adopt the template — Summary, Timeline, Action Items — while still running blame-driven meetings, and the template does nothing to fix that. The mechanism blamelessness protects against is simple and well documented in the incident-response literature: people who fear punishment for an honest account of what happened will omit information, soften their role in events, or avoid raising concerns in the first place — and every piece of missing information is a piece of the system the organization cannot fix, because it never finds out about it.
Concretely, in a blame-driven culture:
- The engineer who typed the command that triggered the outage says as little as possible about the sequence of events, for fear of being singled out — so the actual sequence of contributing decisions never surfaces.
- The engineer who noticed something felt “off” an hour before the incident, but didn’t say anything because they weren’t sure and didn’t want to look alarmist, stays quiet about it in the postmortem too — so the early warning signal that could prevent recurrence is lost.
- Future incidents get slower to resolve, because the responders’ instinct under pressure becomes “don’t make it worse for myself” rather than “surface everything relevant as fast as possible,” even mid-incident.
Blamelessness is the practical fix: it treats the incident as a symptom of a system that allowed a reasonable, well-intentioned person to cause harm, and directs all analytical energy at the system rather than the person. This does not mean individual accountability disappears — repeated negligence or willful policy violation is a performance conversation, held separately, not smuggled into the blameless review. It means the default assumption, for the overwhelming majority of incidents, is that the person involved was doing a reasonable job with the information and tools they had, and the postmortem’s job is to find out why the system let their reasonable actions cause an outage.
Living this out as a manager takes active, repeated effort, not a one-time announcement:
- Model it yourself. The first time a manager says “well, if you’d just double-checked before deploying” in a postmortem — even sympathetically, even once — the room recalibrates and the next postmortem gets thinner.
- Ask “why did the system allow this” before “who did this.” Reframe every instinct to name a person into a question about the guardrail that was missing.
- Protect the meeting from become performance evidence. Postmortems should not be referenced in performance reviews as evidence against the people involved. If people believe what they say in a postmortem might resurface in their next review, they will stop saying it honestly.
- Thank people for surfacing hard truths, publicly if appropriate, especially when the hard truth is inconvenient or points at a decision the manager themselves made.
Root cause analysis: the 5 Whys, and the trap of “human error”
The 5 Whys technique is a simple, effective starting point for root cause analysis: state the problem, ask “why did that happen,” and for each answer, ask “why” again, typically five times, until you reach something that is actually actionable to fix. It works well precisely because it forces a team past the first, most superficial answer.
The most common way 5 Whys goes wrong is stopping at “human error.” Consider a typical chain:
- Why did the site go down? — A bad configuration was deployed to production.
- Why was a bad configuration deployed? — The engineer didn’t test it in staging first.
- Why didn’t they test it in staging first? — They were confident it was a trivial change and staging felt like unnecessary overhead under time pressure.
A team that stops here concludes “human error: the engineer should have tested it” — and the action item becomes “remind engineers to test in staging,” which fixes nothing, because the next tired or rushed engineer will make the same reasonable-in-the-moment call. The productive move is to keep asking why the system allowed that human decision to cause impact:
- Why did skipping staging lead to a full outage rather than a caught mistake? — There was no automated validation gate between “config committed” and “config live in production.”
- Why was there no validation gate? — Config deploys were never brought under the same CI/CD safety rails as code deploys, because they were originally considered “low risk, ad hoc changes.”
Now the action items are real and durable: add automated config validation, require config changes to go through the same pipeline as code, add a canary or staged rollout for config changes. “Human error” is almost never the root cause — it is the last visible symptom of a system that had no guardrail to catch a normal human mistake before it caused harm. A manager’s job in reviewing any postmortem is to keep pushing past the first answer that names a person’s action and toward the system property that let that action matter.
Two related traps to watch for as a manager facilitating this analysis:
- Multiple root causes, not one. Real incidents are almost never one cause; they are usually two or three independent weaknesses that happened to line up (a bug, plus a missing test, plus an alert that had been silenced for a known-flaky reason, plus an on-call engineer unfamiliar with that subsystem). Insisting on a single “the root cause was X” narrative discards useful information — the postmortem template below explicitly uses a plural “Root cause(s).”
- Contributing factors are not the same as root causes. Something can meaningfully worsen an incident (e.g., a stale runbook slowed the response by 20 minutes) without being why the incident started. Both are worth recording, but conflating them muddies the action items — fixing the runbook prevents a slow response next time, not the next occurrence of the underlying bug.
Incident postmortems vs. project postmortems
Blameless retrospection is not only for outages. The same principles — honest accounting, system focus over individual blame, action items that actually get tracked — apply just as well to a shipped (or badly shipped) project or initiative. A project that ran three months late, shipped with a materially reduced scope, or required a costly last-minute rescue deserves the same rigor an outage does, because the underlying question is identical: what did the organization learn, and what will it change as a result?
| Dimension | Incident postmortem | Project postmortem |
|---|---|---|
| Trigger | An outage, degradation, or security event — usually time-boxed and acute | A project/initiative milestone: launch, cancellation, or a significant schedule/scope miss |
| Typical timeframe of analysis | Minutes to days (the incident timeline) | Weeks to months (the project timeline) |
| Primary focus | What broke technically, and what let it break | What planning, estimation, communication, or scoping decisions led to the outcome |
| Common root causes | Missing guardrails, insufficient testing, alerting gaps, unclear ownership during response | Unclear requirements, unrealistic estimates, dependency surprises, scope creep, understaffing, cross-team coordination gaps |
| Typical action items | Add monitoring/alerting, add automated validation, update runbooks, fix the specific bug class | Change the estimation process, adjust how dependencies are tracked, revise how scope changes get approved, change staffing model for future projects |
| Who should attend | Responders, IC, affected service owners, EM | Project lead, EM, key contributors, sometimes the requesting stakeholder (PM, sales, exec sponsor) |
| Common failure mode | Stopping at “human error”; not tracking action items | Turning it into a blame session about “who slipped the date”; treating it as a one-time retro with no follow-through |
The instinct to skip project postmortems is strong precisely because there’s no outage forcing the conversation — nothing is on fire, so it’s easy to let the team move straight to the next project. Resisting that instinct is one of the higher-leverage things an EM can do, because project-level dysfunction (chronic underestimation, a recurring cross-team dependency problem, a planning process that never accounts for onboarding time) compounds silently across every project that follows, in a way a single outage rarely does.
Full blameless postmortem template
The template below applies to incident postmortems directly; for project postmortems, adapt “Impact” to schedule/scope/cost and “Timeline” to the project’s key decision points rather than minute-by-minute incident timestamps.
# Postmortem: [Incident/Project Name]
Status: [Draft / In Review / Final]
Date of incident: [YYYY-MM-DD]
Date of postmortem: [YYYY-MM-DD]
Author(s): [Names]
Severity: [SEV-1 / SEV-2 / ...]
## Summary
A 2–4 sentence, plain-language summary of what happened, who was affected,
and how it was resolved. Should be understandable by someone outside the
team with no additional context.
## Impact
- Duration: [start time] to [end time] ([total duration])
- Users/customers affected: [number or percentage, or "internal only"]
- Services affected: [list]
- Revenue/business impact: [if applicable and known]
- Data impact: [any data loss, corruption, or exposure]
- SLA/SLO impact: [error budget consumed, SLA breach if any]
## Timeline
All times in [timezone]. Include detection, escalation, and every
significant action/decision, not just the fix.
| Time | Event |
|-------|-------|
| 14:02 | Alert fired: elevated error rate on checkout-service |
| 14:05 | On-call engineer acknowledges page |
| 14:12 | Incident declared SEV-2, IC assigned |
| 14:20 | Root cause hypothesis: bad config deploy at 13:55 |
| 14:25 | Config rollback initiated |
| 14:31 | Error rate returns to baseline |
| 14:45 | Incident declared resolved; monitoring continues |
## Root Cause(s)
List every independent contributing cause identified through root cause
analysis (e.g., 5 Whys). Most incidents have more than one.
1. [Root cause 1 — the technical/process gap, not a person's name]
2. [Root cause 2, if applicable]
## Contributing Factors
Things that worsened the incident's impact or slowed the response, without
being why it started.
- [e.g., "Runbook for this service was 8 months out of date"]
- [e.g., "Alert had been muted for a known-flaky reason and stayed muted"]
## What Went Well
Concrete, specific things that worked, worth reinforcing and repeating.
- [e.g., "Rollback automation cut recovery time from ~30 min to ~6 min"]
- [e.g., "IC handoff at shift change was smooth, no context lost"]
## What Went Poorly
Concrete, specific things that didn't work — framed at the system level.
- [e.g., "No alert fired on the actual leading indicator; we only got
paged once customer-facing errors were already high"]
- [e.g., "Took 12 minutes to identify who owned the affected service"]
## Action Items
Every item has an owner and a due date. No owner, no date — no action item.
| Action item | Owner | Due date | Priority | Status |
|---|---|---|---|---|
| Add automated validation gate for config deploys | @alice | 2026-08-01 | P1 | Not started |
| Update on-call runbook for checkout-service | @bob | 2026-07-25 | P2 | In progress |
| Add leading-indicator alert for queue depth | @carol | 2026-08-08 | P1 | Not started |
## Lessons Learned
A short narrative synthesis — what should other teams, not just this one,
take away from this incident? What pattern should we watch for elsewhere?
[e.g., "Any service accepting config changes outside the normal code
deploy pipeline is a blind spot. We should audit for other services with
this same gap."]
Best Practices
- Separate the IC role from the manager role explicitly, even at a small company where the manager is technically capable of being IC. Say out loud, in the incident channel, who is playing which role, so responders know who to escalate technical decisions to and who to escalate resourcing/communication needs to.
- Write the customer/exec-facing update before you’re asked for it. A steady, proactive cadence (e.g., every 30 minutes for a SEV-1) prevents the “any update?” pings that themselves become a distraction to the response.
- Hold the postmortem while memory is fresh but emotions have settled — typically within a few business days of resolution. Too soon and people are still stressed and defensive; too late and details are reconstructed rather than remembered.
- Facilitate, don’t interrogate. As the manager in the room, your questions should sound like “what information would have helped you make a different call in the moment?” not “why didn’t you just check X first?” The second framing, however unintentional, is blame with a question mark.
- Track action items in the same backlog as feature work, with the same visibility and grooming, not in a separate “postmortem action items” graveyard doc. An action item that isn’t prioritized against feature work competes for attention and reliably loses. If leadership won’t prioritize reliability action items against roadmap features, that’s a resourcing conversation the manager needs to have explicitly, not something to paper over by writing action items nobody ever revisits.
- Assign a real owner and a real date to every action item, and follow up on overdue ones the same way you’d follow up on a slipping feature deadline — visibly, in the same forums, with the same seriousness.
- Share postmortems broadly, not just within the affected team. A short internal newsletter, a recurring “postmortem review” forum across teams, or a searchable postmortem archive prevents the same incident class from recurring in a different service six months later because nobody outside the original team ever heard about it.
- Watch for recurring root causes across postmortems over a quarter or two — three unrelated incidents that all trace back to “no staging environment parity” is a strategic signal, not three separate bad days, and it deserves a dedicated initiative rather than three isolated action items.
- Run project postmortems on a fixed cadence (e.g., after every major launch, regardless of how it went), not only when something visibly goes wrong — retrospecting only on failures teaches the org that things which “went fine” have nothing to teach, which is rarely true.
- Protect on-call sustainability as an ongoing input to incident quality, not a separate HR concern — see Team Culture & Well-Being for the broader burnout picture, and treat rotation health metrics with the same regularity as incident metrics.
- Distinguish a single incident from a crisis. When an incident escalates into something threatening the business more broadly — major data breach with regulatory exposure, extended multi-day outage, significant reputational event — the response shifts from standard incident management into business continuity territory; see Crisis Leadership & Business Continuity for that broader lens.
References
- Google SRE Book — Postmortem Culture: Learning from Failure
- Google SRE Workbook — Postmortem Culture: Learning from Failure
- PagerDuty — Incident Response Documentation
- PagerDuty — Postmortems
- Atlassian — Incident Postmortems
- NIST SP 800-61 — Computer Security Incident Handling Guide
- DevSecOps — Incident Response & Digital Forensics