← Quản lý kỹ thuật← Engineering Manager
Quản lý kỹ thuậtEngineering Manager19 Th7, 2026Jul 19, 202634 phút đọc25 min read

Agile Process & DeliveryAgile Process & Delivery

Thuộc bộ kiến thức Engineering Manager Roadmap.

Tổng quan

“Agile” có lẽ là từ bị hiểu sai nhiều nhất trong quản lý phần mềm. Phần lớn các team nói “chúng tôi làm agile” thực ra chỉ đang nói “chúng tôi chạy các ceremony của Scrum” — daily standup, sprint hai tuần, một backlog trên Jira — trong khi tinh thần thực sự của Agile Manifesto đã bị lãng quên đâu đó trên đường đi. Với vai trò engineering manager (EM), việc của bạn không phải là chạy ceremony cho đúng quy trình. Việc của bạn là đảm bảo team biến ý tưởng thành phần mềm chạy được một cách đáng tin cậy, thích ứng khi thực tế mâu thuẫn với kế hoạch, và không bị chìm trong process không còn xứng đáng với chi phí duy trì nó.

Note này coi “agile process” là phương tiện, không phải mục đích. Nó sẽ đi qua Manifesto thực sự nói gì (khác với những gì “làm agile” đã trở thành trong thực tế), cách chọn và tiến hóa một framework process (Scrum, Kanban, Scrumban), cách chạy planning và estimation mà không biến chúng thành màn kịch, cách quản lý release và chất lượng, và — quan trọng nhất — cách nói về timeline và velocity một cách trung thực thay vì phóng chiếu một sự chính xác giả tạo cho các stakeholder, những người sau này sẽ giữ bạn đúng theo con số bạn đưa ra dưới áp lực.

Sợi chỉ xuyên suốt: process tồn tại để phục vụ việc delivery và những con người đang thực hiện delivery đó. Khoảnh khắc một ceremony, một metric, hay một document ngừng làm được điều đó, nó là ứng viên để loại bỏ — bất kể nó đã tồn tại lâu đến đâu hay ai là người áp đặt nó.

Kiến thức nền tảng

Agile Manifesto, đọc thật kỹ

Agile Manifesto chỉ gồm bốn tuyên bố giá trị ngắn gọn và mười hai nguyên tắc, được viết năm 2001 bởi mười bảy practitioner đã chán ngán các process nặng nề, dựa trên documentation (Waterfall, RUP). Đáng để đọc lại các giá trị này đầy đủ, vì phần lớn team chưa từng đọc quá cái buzzword:

Chúng tôi đang khám phá những cách tốt hơn để phát triển phần mềm bằng cách tự làm và giúp người khác làm. Qua công việc này, chúng tôi đã đi đến việc coi trọng:

Con người và sự tương tác hơn là quy trình và công cụ Phần mềm chạy được hơn là tài liệu đầy đủ Hợp tác với khách hàng hơn là đàm phán hợp đồng Phản ứng với thay đổi hơn là bám theo kế hoạch

Dòng quan trọng nhất, và cũng thường bị bỏ qua nhất, nằm ngay sau đó: “Nghĩa là, dù những điều bên phải vẫn có giá trị, chúng tôi coi trọng những điều bên trái hơn.” Manifesto không nói rằng công cụ, tài liệu, hợp đồng, kế hoạch là vô giá trị — nó nói rằng khi những thứ đó xung đột với con người, phần mềm chạy được, hợp tác, và khả năng thích ứng, thì vế bên trái thắng. Một team có process được document tuyệt đẹp nhưng chưa ship được phần mềm nào không phải là agile. Một team ship phần mềm lỗi thật nhanh mà không có hợp tác hay feedback loop cũng vậy.

Mười hai nguyên tắc đằng sau Manifesto làm rõ thêm điều này — “phần mềm chạy được là thước đo tiến độ chính”, “chào đón thay đổi requirement, kể cả khi đã muộn trong quá trình phát triển”, “kiến trúc, requirement và thiết kế tốt nhất xuất hiện từ những team tự tổ chức”, “định kỳ, team nhìn lại cách làm việc hiệu quả hơn, rồi điều chỉnh hành vi của mình cho phù hợp”. Hãy để ý những gì không xuất hiện: không có một dòng nào nhắc đến Scrum, sprint, story point, hay standup trong Manifesto hay các nguyên tắc của nó. Đó là những cách triển khai mà các cộng đồng sau này xây dựng để hiện thực hóa các giá trị — hữu ích, nhưng không phải bản thân giá trị. Sự phân biệt này quan trọng vì nó chính là “giấy phép” để bạn, với tư cách EM, được quyền thay đổi hoặc bỏ bất kỳ practice cụ thể nào không còn phục vụ giá trị nền tảng, mà không bị coi là “từ bỏ agile”.

Scrum, Kanban và Scrumban

Ba framework thống trị agile delivery trong thực tế. Chúng không thể thay thế lẫn nhau, và chọn sai framework cho dạng công việc của team sẽ gây ra friction kinh niên mà mọi người thường đổ lỗi cho “thực thi kém” thay vì nhìn ra vấn đề là do mismatch giữa framework và bản chất công việc.

Scrum chia công việc thành các iteration có độ dài cố định (sprint, thường 1–4 tuần) với các role được định nghĩa rõ (Product Owner, Scrum Master, Developers) và một nhịp ceremony (sprint planning, daily scrum, sprint review, sprint retrospective) xoay quanh một backlog. Nó được mô tả chính thức trong Scrum Guide. Scrum hoạt động tốt nhất khi công việc có thể được gom thành các batch vừa với kích thước sprint, với scope đủ ổn định để có thể commit, và khi team hưởng lợi từ một nhịp kiểm tra bên ngoài đều đặn (demo cho stakeholder) và phản tư nội bộ (retro).

Kanban là một phương pháp dựa trên flow, không có iteration cố định: các work item liên tục di chuyển qua một board với giới hạn work-in-progress (WIP) rõ ràng cho từng cột, và team tối ưu cho luồng liên tục (continuous flow) thay vì commitment theo sprint. Nó bắt nguồn từ Lean manufacturing và được ghi chép kỹ trong Kanban Guide for Scrum Teams cũng như tài liệu Kanban nói chung (ví dụ tài nguyên từ Kanban University). Kanban phù hợp với team có công việc biến động cao, bị ngắt quãng (interrupt-driven), hoặc thiên về support/vận hành — nơi việc commit “bộ item này trong hai tuần tới” gần như là chuyện viễn tưởng vì priority thay đổi hàng ngày (on-call, production support, yêu cầu đột xuất).

Scrumban là mô hình lai: giữ nhịp và ceremony của Scrum (planning, retro) nhưng đặt trên nền continuous flow và WIP limit của Kanban thay vì commitment sprint cố định. Nó thường thấy ở các team đang chuyển từ Scrum sang Kanban, hoặc team muốn giữ nhịp phản tư của Scrum mà không giả vờ rằng mọi công việc đến đều có thể lên kế hoạch trước hai tuần.

Khía cạnhScrumKanbanScrumban
CadenceSprint độ dài cố định (1–4 tuần)Liên tục, không iteration cố địnhCadence cố định cho planning/retro, flow liên tục cho công việc
RoleProduct Owner, Scrum Master, Developers (định nghĩa trong Scrum Guide)Không quy định role; team tự tổ chức quanh boardThường không chính thức, mượn lỏng lẻo role của Scrum
Artifact chínhProduct backlog, sprint backlog, incrementKanban board, WIP limit, cumulative flow diagramBacklog + Kanban board, WIP limit
Ràng buộc chínhSprint commitment (scope khóa trong sprint)WIP limit theo từng cộtWIP limit, không khóa scope
Đơn vị planningSprint (batch story được commit trước)Từng item được pull khi có capacity trốngPlanning nhẹ, just-in-time
Phù hợp nhất vớiFeature team có công việc lên kế hoạch/gom batch được và hưởng lợi từ nhịp demo/review đều đặnSupport, ops, platform, hoặc team có tỷ lệ interrupt cao và công việc đến khó đoánTeam đang chuyển framework, hoặc muốn nhịp phản tư của Scrum mà không cần mô hình commitment của Scrum
Failure mode nếu áp dụng saiSprint trở thành hình thức — team chuyển dời hơn nửa backlog qua mỗi sprintTeam mất đi cơ chế ép buộc cho phản tư/cải tiếnCó thể thành “tệ nhất của cả hai” nếu áp dụng không chủ đích (không commitment thật, không kỷ luật flow thật)

Chọn giữa ba mô hình này là một bài tập chẩn đoán, không phải sở thích cá nhân. Hãy hỏi: công việc đến có biến động nhiều không? Stakeholder có cần một nhịp demo dự đoán được không? Team có làm một phần đáng kể công việc ngoài kế hoạch/phản ứng (on-call, incident, yêu cầu gấp) không? Một platform team liên tục nhận yêu cầu interrupt-driven mà bị ép vào commitment sprint hai tuần sẽ “thất bại” sprint một cách kinh niên — không phải vì team làm việc kém, mà vì framework không khớp với hình dạng công việc. Ngược lại, một feature team dùng Kanban mà không có cadence nào cả có thể trôi dạt vì thiếu cơ chế ép buộc cho phản tư và căn chỉnh với stakeholder.

Sprint planning: capacity, estimation, và ai sở hữu commitment

Sprint planning có ba việc riêng biệt mà các cách triển khai tệ hay gộp lẫn vào nhau: tính xem team thực sự có bao nhiêu capacity, estimate kích thước tương đối của các công việc ứng viên, và quyết định commit vào những gì. Mỗi việc cần được xử lý riêng.

Capacity planning bắt đầu từ số person-day khả dụng, không phải headcount. Một team năm người trong sprint hai tuần không phải là “50 person-day capacity” — phải trừ PTO, ngày nghỉ lễ, lịch on-call, meeting, phỏng vấn, và “thuế” của việc support production đang diễn ra. Các team lên kế hoạch dựa trên headcount danh nghĩa thay vì khả dụng thực tế đang tự tạo ra một nguồn gây thất bại sprint kinh niên, chẳng liên quan gì đến độ chính xác của estimation. Một phép tính capacity đơn giản và trung thực (số ngày khả dụng × focus factor, với focus factor thường 60–80% sau khi trừ meeting, interrupt, và chi phí context-switching) luôn tốt hơn phép nhân headcount lạc quan.

Estimation tồn tại để hỗ trợ quyết định planning, không phải để tạo ra một dự báo chính xác về tương lai — không phương pháp estimation nào làm được điều đó một cách đáng tin cậy, và việc coi estimate là commitment chính là nguyên nhân phổ biến nhất gây rối loạn chức năng estimation. Các cách tiếp cận phổ biến:

Cách tiếp cậnCách hoạt độngĐiểm mạnhGiới hạn
Story point (sizing tương đối, thường theo dãy giống Fibonacci: 1, 2, 3, 5, 8, 13)Team gán một kích thước không đơn vị so với một story tham chiếuTách rời size khỏi thời gian; ép buộc cuộc trò chuyện “cái này so với X thì thế nào” giúp lộ ra độ phức tạp ẩnDễ bị game/lạm phát theo thời gian (“point inflation”); vô nghĩa khi so giữa các team; có thể thành nghi thức quan liêu nếu cuộc trò chuyện thực chất biến mất
Planning pokerMỗi người chọn giá trị point cùng lúc (giấu kín), lật ra cùng nhau, thảo luận các giá trị khác biệtNgăn hiện tượng anchoring vào tiếng nói ồn ào/senior nhất; làm bất đồng lộ ra nhanhChậm nếu lạm dụng trên item tầm thường; thành màn kịch nếu mọi người ngừng suy nghĩ thật và chỉ copy theo số đông
T-shirt sizing (XS/S/M/L/XL)Sizing tương đối thô, thường cho estimate sớm/thôNhanh, không giả vờ chính xác, tốt cho triage backlogQuá thô cho commitment sprint; cuối cùng vẫn cần dịch sang point/giờ
#NoEstimates / đếm itemTheo dõi throughput số item bất kể kích thước (giả định slicing tương đối đồng đều)Loại bỏ hoàn toàn overhead estimation; hiệu quả khi có kỷ luật slicing story nhỏĐòi hỏi kỷ luật thật trong việc chia nhỏ và đồng đều công việc; không giúp được với công việc lớn, cồng kềnh
Estimation theo giờ/ngàyEstimate thời gian trực tiếp cho từng taskDễ hiểu với người không phải kỹ sư/stakeholderHệ thống bị đánh giá thấp (xem planning fallacy bên dưới); mời gọi micromanage theo giờ

Dù dùng cách nào, estimate là một phát biểu xác suất về một hoạt động phức tạp, bất định — không phải một lời hứa. Rối loạn chức năng số một cần đề phòng với vai trò EM là estimate bị diễn giải lại ở tầng trên (bởi một PM, một director, một đội sales) thành cam kết deadline. Nếu “5 point” âm thầm biến thành “xong trước thứ Năm” trong slide của ai đó, thì quy trình estimation đã bị vũ khí hóa chống lại team, và bạn cần can thiệp trước khi nó lặp lại.

Commitment phải là quyết định của team, không phải mệnh lệnh của manager. Đây là điểm mà phần lớn EM mới mắc lỗi khi có áp lực delivery. Nếu một manager (hoặc một stakeholder bên ngoài nói qua manager) đặt scope sprint đơn phương — “sprint này chúng ta làm mười lăm item này” — thì hai chuyện xảy ra: tín hiệu thực sự của team về capacity của chính họ bị bỏ qua, và trách nhiệm giải trình âm thầm chuyển từ “team quyết định điều này khả thi” sang “manager bảo chúng tôi làm vậy”, điều này ăn mòn tinh thần sở hữu (ownership) và cho mọi người một cái cớ sẵn có khi sprint thất bại. Việc của EM trong planning là mang bối cảnh (priority, deadline, ràng buộc, nhu cầu stakeholder) vào phòng họp, không phải viết sẵn câu trả lời lên bảng trước khi thảo luận bắt đầu. Khi một deadline bên ngoài thực sự mâu thuẫn với những gì team tin là khả thi, đó là một cuộc đàm phán về scope hoặc timeline — được nêu ra rõ ràng, chứ không phải giải quyết bằng cách ghi đè lên đánh giá capacity của chính team.

Khái niệm chính

Tiến hóa process mà không thành “process theater”

“Process theater” là bất kỳ ceremony, document, công cụ, hay metric nào sống sót qua thời điểm nó từng giải quyết một vấn đề thật — được giữ lại nhờ quán tính, “xưa nay vẫn làm vậy”, hoặc vì bỏ nó đi cảm giác như thừa nhận thất bại. Đây là một trong những cách đáng tin cậy nhất để velocity và tinh thần của team âm thầm suy giảm trong khi mọi dashboard vẫn trông ổn.

Kỷ luật để tránh điều này: bắt đầu tối giản, chỉ thêm process khi có một điểm đau thực sự, lặp lại được chứng minh, và định kỳ rà soát process hiện có xem nó còn xứng đáng với chi phí hay không.

Phép thử cho bất kỳ hạng mục process nào, cũ hay mới: bạn có thể gọi tên failure mode cụ thể mà nó ngăn chặn không, và failure mode đó còn khả dĩ với team này, ngay lúc này, không? Nếu câu trả lời là không, đó là ứng viên để xóa bỏ, không phải một “con bò thiêng”.

Tài liệu process và onboarding

Một cách “chúng ta làm việc như thế nào” không được document là một process chỉ tồn tại trong đầu các thành viên senior, nghĩa là mỗi người mới hoặc phải làm phiền ai đó cả tiếng đồng hồ, hoặc phải reverse-engineer process bằng cách quan sát chuyện gì xảy ra. Một document “team working agreement” ngắn gọn và còn sống — độ dài sprint và lịch ceremony, Definition of Ready/Done, quy ước branching và review, kỳ vọng on-call, đường escalation, tooling — nên là thứ đầu tiên một engineer mới đọc, và nó cần được đối xử như một đoạn code: nó sẽ cũ đi trừ khi có ai đó chịu trách nhiệm giữ nó cập nhật.

Cơ chế thực tế để giữ nó sống thay vì trở thành hóa thạch: review và cập nhật nó như một mục rõ ràng ngay sau khi mỗi thay đổi process được áp dụng (chứ không phải “để sau”), phân công quyền sở hữu luân phiên thay vì để mặc cho người viết ban đầu, và dùng nó chủ động trong onboarding — cho người mới đọc nó ngay ngày đầu tiên và gắn cờ bất cứ điều gì khó hiểu hoặc lỗi thời, việc này đồng thời cũng là một hình thức audit nhẹ nhàng độ chính xác của nó. Một working agreement mà không ai chỉnh sửa suốt một năm hoặc là một team cực kỳ ổn định (hiếm), hoặc là một document không ai còn tin tưởng nữa (phổ biến hơn) — đáng để kiểm tra xem là trường hợp nào.

Development và release workflow

Chiến lược branching đặt nhịp điệu cho cách công việc được merge và ship; xem ../devops/en/04-version-control.md để biết đầy đủ cơ chế của các mô hình branching Git (trunk-based, Git Flow, GitHub Flow). Với vai trò EM, quyết định quan trọng nhất là trunk-based development so với feature branch tồn tại lâu dài: trunk-based (merge nhỏ, thường xuyên vào main, dùng feature flag cho bất cứ thứ gì chưa sẵn sàng) giữ friction tích hợp thấp và cho phép continuous delivery, trong khi branch tồn tại lâu dài (kiểu Git Flow) giảm cảm giác rủi ro của mỗi lần merge nhưng lại dồn friction tích hợp vào các lần merge lớn, hiếm, và có blast radius cao. Phần lớn team hiệu suất cao hội tụ về trunk-based development với branch tồn tại ngắn, vì nó khớp với tinh thần “batch nhỏ, feedback nhanh” của cả agile lẫn các phát hiện của DORA về nhóm hiệu suất tinh hoa.

Definition of Ready (DoR)Definition of Done (DoD) là hai rào chắn giữ scope của sprint trung thực. DoR là ngưỡng một backlog item phải vượt qua trước khi nó vào sprint planning như một ứng viên — acceptance criteria đã viết, dependency đã xác định, design/mockup đã đính kèm nếu cần, kích thước ước lượng tương đối. Công việc không đạt DoR mà vẫn bị kéo vào sprint là nguyên nhân phổ biến nhất của việc phát hiện scope giữa sprint (“à, cái này thực ra cần thay đổi backend mà chúng ta chưa scope”). DoD là ngưỡng một item phải vượt qua để được coi là hoàn thành — thường gồm: code đã merge, test pass, đã deploy lên môi trường liên quan, acceptance criteria đã verify, tài liệu đã cập nhật, monitoring/alerting đã có nếu liên quan. Một team không có DoD rõ ràng sẽ có bất đồng kinh niên về việc liệu thứ gì đó có “thực sự xong” chưa, thường lộ ra trong lúc demo khi stakeholder tìm thấy một edge case chưa từng được test.

Cả DoR và DoD nên ngắn gọn (một checklist, không phải một tài liệu chính sách), do team sở hữu (không bị áp đặt), và được xem lại khi chúng gây friction — xem ./08-engineering-practices-oversight.md để biết DoD tương tác thế nào với chuẩn code review, ngưỡng testing, và việc giám sát thực hành kỹ thuật nói chung.

Quản lý release

Cadence release là một sự đánh đổi thật sự, không phải một bài toán đã có một đáp án đúng duy nhất:

Cách tiếp cậnMô tảƯu điểmChi phí
Continuous deploymentMọi thay đổi vượt qua pipeline tự động ship lên productionBatch size nhỏ nhất có thể, feedback nhanh nhất, rủi ro mỗi lần release thấp nhất (nghiên cứu của DORA liên tục gắn điều này với hiệu suất tinh hoa)Đòi hỏi test coverage tự động mạnh, feature flag, và kỷ luật monitoring/rollback — tốn kém nếu áp đặt ngược lên một team chưa có những nền tảng đó
Release theo lịch (ví dụ hàng tuần, hai tuần một lần)Các thay đổi tích lũy và ship theo batch theo lịch cố địnhDự đoán được với stakeholder, support, và giao tiếp release note; dễ điều phối dependency chéo team hơnBatch lớn hơn nghĩa là diện tích ảnh hưởng lớn hơn mỗi lần release, khó truy nguyên nguyên nhân hơn khi có sự cố, và mô hình “chuyến tàu” nơi thay đổi trễ hoặc lỡ chuyến hoặc làm trễ nó
Release train (lịch cố định, code freeze trước đó)Mô hình lai: cadence cố định, nhưng có một điểm cắt rõ ràng và cửa sổ ổn định hóaCân bằng giữa dự đoán được và một phần kiểm soát batch sizeBản thân cửa sổ freeze thường là nguồn gây friction và công việc bị chặn

Feature flag là cơ chế tách rời việc deploy code khỏi việc release một tính năng cho người dùng — code có thể merge vào main và deploy lên production ở trạng thái ẩn (flag tắt), rồi được bật dần (người dùng nội bộ, rollout theo phần trăm, khách hàng cụ thể) độc lập với pipeline deploy. Đây chính là điều làm cho continuous deployment tương thích với việc rollout tính năng có kiểm soát, dần dần, thay vì “deploy = mọi người thấy ngay bây giờ”. Xem ../devops/en/11-ci-cd.md để biết cơ chế pipeline giúp điều này an toàn (progressive delivery, canary release, các trigger rollback tự động). Với vai trò EM, feature flag còn là một công cụ lập lịch: một tính năng hoàn thành 90% có thể merge dần đằng sau một flag thay vì sống trên một branch dài ngày tích lũy merge conflict, đây là một trong những cách ít được đánh giá đúng mức nhất để giữ trunk-based development khả thi cho các tính năng lớn, kéo dài nhiều sprint.

Quản lý milestone và ước lượng timeline

Thói quen gây hại nhất trong giao tiếp timeline phần mềm là sự chính xác giả tạo (false precision) — tuyên bố “cái này ship ngày 14 tháng 3” khi câu trả lời trung thực là “nhiều khả năng nhất trong khoảng cuối tháng 2 đến giữa tháng 4, với một xác suất nhỏ trễ hơn nữa”. Planning fallacy (Kahneman & Tversky) mô tả xu hướng gần như phổ quát, được ghi chép kỹ, là đánh giá thấp thời gian hoàn thành công việc ngay cả khi đã tính đến các rủi ro đã biết, vì con người lên kế hoạch theo con đường thực thi tốt nhất và hệ thống hóa việc hạ thấp xác suất của những điều bất ngờ — một dependency bị trễ, một thiết kế hóa ra sai, một sự cố production ngoài kế hoạch ngốn mất một tuần. Đây không phải là lỗi kỷ luật riêng của những kỹ sư tệ; đó là một thiên kiến nhận thức xuất hiện ngay cả ở những team có quy trình estimation xuất sắc, đó là lý do vì sao chỉ riêng process (ticket chi tiết hơn, planning poker nhiều hơn) không sửa được nó — chỉ có việc mô hình hóa sự bất định một cách rõ ràng mới làm được.

Các kỹ thuật thực tế để giao tiếp trung thực thay vì phóng chiếu sự tự tin giả tạo:

Xem ./15-strategy-business-case-and-prioritization.md để biết những giao tiếp timeline này nuôi dưỡng các cuộc trò chuyện prioritization và business-case rộng hơn với stakeholder như thế nào.

Velocity: nó dùng để làm gì và không dùng để làm gì

Velocity — thường là số story point (hoặc item) hoàn thành mỗi sprint — là một công cụ capacity planning và không hơn gì. Công dụng chính đáng duy nhất của nó: giúp một team cụ thể dự báo họ có thể nhận bao nhiêu công việc trong một sprint tương lai, dựa trên lịch sử gần đây, nhất quán của chính họ, và giúp chính team đó có một cuộc trò chuyện sớm về việc liệu một milestone có khả thi hay không dựa trên throughput hiện tại.

Những gì velocity không đáng tin cậy để dùng, và tại sao nó phản tác dụng khi bị lạm dụng:

Rào chắn thực tế: velocity ở lại bên trong team, được team dùng, cho việc planning của chính team. Ngay khoảnh khắc nó vượt ra ngoài ranh giới đó — báo cáo cho lãnh đạo như một KPI, so sánh giữa các team, gắn với chu kỳ review — nó ngừng đo cái nó từng đo và bắt đầu đo việc mọi người đã học được cách game nó tốt đến đâu.

Quality metrics so với vanity metrics

Một quy trình delivery lành mạnh cần theo dõi chất lượng, không chỉ throughput, và các metric đáng theo dõi là những metric tương quan với outcome thực, chứ không phải những metric dễ tính toán hay dễ báo cáo cho đẹp mặt.

Chương trình DORA (DevOps Research and Assessment) đã xác định bốn metric giúp phân biệt đáng tin cậy các tổ chức kỹ thuật hiệu suất tinh hoa, dựa trên nhiều năm nghiên cứu với mẫu lớn:

MetricĐo cái gìVì sao quan trọng
Deployment frequency (tần suất deploy)Code deploy thành công lên production thường xuyên đến đâuProxy cho batch size — deploy thường xuyên hơn nghĩa là thay đổi nhỏ hơn, rủi ro thấp hơn
Lead time for changes (thời gian từ commit đến chạy production)Thời gian từ khi code được commit đến khi chạy trên productionProxy cho lượng friction/lãng phí trong chính pipeline delivery
Change failure rate (tỷ lệ thay đổi gây lỗi)Phần trăm deployment gây ra lỗi production cần khắc phụcTín hiệu chất lượng trực tiếp về quy trình release, không chỉ về code
Time to restore service (thời gian khôi phục dịch vụ)Mất bao lâu để phục hồi sau một lỗi productionĐo khả năng phục hồi và phản ứng sự cố, không chỉ việc tránh lỗi

Ngoài DORA, các tín hiệu chất lượng thực sự hữu ích khác gồm escaped defect rate (bug phát hiện ở production so với bắt được trước khi release — thước đo trực tiếp việc testing và review có bắt đúng thứ cần bắt hay không) và cycle time (thời gian từ khi công việc bắt đầu đến khi ship, một “họ hàng” chi tiết hơn của lead time, làm nổi bật nơi công việc bị kẹt trong chính quy trình của một team, ví dụ nằm ở code review ba ngày).

Đối lập với điều này là các vanity metric phổ biến — những con số dễ báo cáo và trông đẹp trên slide nhưng không tương quan đáng tin cậy với sức khỏe delivery thực: số dòng code viết ra, số commit, số ticket đóng mỗi người, số story point hoàn thành (như đã nói ở trên), hay số giờ log. Những metric này dễ dàng bị game và, tệ hơn, chủ động khuyến khích hành vi sai khi ai đó tối ưu trực tiếp cho chúng (nhiều commit hơn bằng cách chia nhỏ thay đổi tầm thường, nhiều ticket hơn bằng cách chọn việc dễ, nhiều dòng code hơn bằng cách không refactor theo hướng đơn giản hóa). Phép thử cho bất kỳ metric nào trước khi áp dụng: nó có thể bị game bởi ai đó hành động thuần túy vì lợi ích ngắn hạn của bản thân không, và nếu tối ưu trực tiếp cho nó, nó có còn chỉ về phía outcome bạn thực sự muốn không? Bốn metric của DORA vượt qua phép thử này khá tốt vì game một trong số chúng thường làm hại rõ rệt đến một metric khác (ví dụ deploy thường xuyên hơn mà không sửa chất lượng sẽ hiện ngay ra trong change failure rate).

Best Practices

Chọn và chạy process

Tình huốngKhuyến nghị
Thành lập team mới hoặc chuyển đổi frameworkBắt đầu với bộ ceremony tối thiểu khả thi; chỉ thêm process khi có điểm đau có tên, lặp lại
Công việc mang tính interrupt-driven cao (ops, support, platform)Ưu tiên Kanban hoặc Scrumban hơn Scrum; ép sprint commitment lên công việc không thể lên kế hoạch sẽ tạo ra “sprint thất bại” kinh niên
Sprint commitment đang bị áp đặt từ trên xuốngDừng lại — mang bối cảnh và ràng buộc vào planning, để team sở hữu commitment; commitment bị áp đặt âm thầm chuyển trách nhiệm giải trình ra khỏi team
Estimate đang bị báo cáo lên trên như deadlineCan thiệp ngay lập tức; định khung lại estimate là input cho capacity planning, không phải lời hứa, mỗi lần điều này xảy ra
Một ceremony hoặc document ngừng tạo giá trịGọi tên failure mode mà nó từng ngăn chặn; nếu failure mode đó không còn khả dĩ, cắt bỏ nó — rà soát theo lịch, không chỉ khi có phàn nàn
Timeline đang được stakeholder yêu cầuĐưa ra khoảng với độ tin cậy và yếu tố biến động rõ ràng, không phải một ngày duy nhất; một khoảng nói trung thực tốn ít niềm tin hơn một điểm estimate bị trễ
Velocity đang được ai đó ngoài team yêu cầuChuyển hướng sang metric dựa trên outcome và chất lượng (DORA, escaped defect, cycle time); đừng để velocity trở thành metric hiển thị ra ngoài hoặc để so sánh
Một engineer mới gia nhập teamChỉ họ đến team working agreement còn sống (ceremony, DoR/DoD, branching, escalation) ngay ngày đầu; dùng câu hỏi của họ để audit độ chính xác của document
Quyết định cadence releaseKhớp cadence với độ trưởng thành test/rollback của team — continuous deployment đòi hỏi feature flag, test tự động mạnh, và rollback nhanh; nếu chưa có, release theo lịch hoặc release train là lựa chọn trung thực, không phải thất bại của tham vọng

Checklist ngắn cho EM

Tài liệu tham khảo

Part of the Engineering Manager Roadmap knowledge base.

Overview

Agile is the most misunderstood word in software management. Most teams that say “we do agile” mean “we do Scrum ceremonies” — daily standups, two-week sprints, a backlog in Jira — while the actual point of the Agile Manifesto gets lost somewhere along the way. As an engineering manager, your job is not to run ceremonies correctly. Your job is to make sure the team reliably turns ideas into working software, adapts when reality contradicts the plan, and doesn’t drown in process that no longer earns its keep.

This note treats “agile process” as a means, not an end. It covers what the Manifesto actually says (as opposed to what “doing agile” has come to mean in practice), how to choose and evolve a process framework (Scrum, Kanban, Scrumban), how to run planning and estimation without turning them into theater, how to manage releases and quality, and — critically — how to talk about timelines and velocity honestly instead of projecting false precision to stakeholders who will hold you to a number you invented under pressure.

The throughline is this: process exists to serve delivery and the humans doing the delivering. The moment a ceremony, a metric, or a document stops doing that, it’s a candidate for removal — regardless of how established it is or who mandated it.

Fundamentals

The Agile Manifesto, actually

The Agile Manifesto is four short value statements and twelve principles, written in 2001 by seventeen practitioners tired of heavyweight, documentation-driven processes (Waterfall, RUP). It is worth reading the values in full because most teams have never actually read past the buzzword:

We are uncovering better ways of developing software by doing it and helping others do it. Through this work we have come to value:

Individuals and interactions over processes and tools Working software over comprehensive documentation Customer collaboration over contract negotiation Responding to change over following a plan

The crucial, most-often-missed line follows immediately after: “That is, while there is value in the items on the right, we value the items on the left more.” The Manifesto does not say tools, documentation, contracts, and plans are worthless — it says that when they conflict with people, working software, collaboration, and adaptability, the left side wins. A team with a beautifully documented process and no shipped software is not agile. Neither is a team that ships broken software fast with no collaboration or feedback loop.

The twelve principles behind the Manifesto flesh this out — “working software is the primary measure of progress,” “welcome changing requirements, even late in development,” “the best architectures, requirements, and designs emerge from self-organizing teams,” “at regular intervals, the team reflects on how to become more effective, then tunes and adjusts its behavior accordingly.” Notice what’s absent: there is no mention of Scrum, sprints, story points, or standups anywhere in the Manifesto or its principles. Those are implementations that later communities built to embody the values — useful, but not the values themselves. This distinction matters because it’s your license, as an EM, to change or drop any specific practice that stops serving the underlying value, without that being “abandoning agile.”

Scrum, Kanban, and Scrumban

Three frameworks dominate real-world agile delivery. They are not interchangeable, and picking the wrong one for your team’s work pattern causes chronic friction that people usually blame on “bad execution” rather than framework mismatch.

Scrum structures work into fixed-length iterations (sprints, typically 1–4 weeks) with defined roles (Product Owner, Scrum Master, Developers) and a cadence of ceremonies (sprint planning, daily scrum, sprint review, sprint retrospective) built around a backlog. It’s described authoritatively in the Scrum Guide. Scrum works best when work can be meaningfully batched into sprint-sized chunks with a stable-enough scope to commit to, and when the team benefits from a regular cadence of external inspection (stakeholder demos) and internal reflection (retros).

Kanban is a flow-based method with no fixed iterations: work items move continuously through a board with explicit work-in-progress (WIP) limits per column, and the team optimizes for continuous flow rather than sprint commitments. It’s rooted in Lean manufacturing and documented well by the Kanban Guide for Scrum Teams and general Kanban literature (e.g., the Kanban University resources). Kanban suits teams with highly variable, interrupt-driven, or support/operations-heavy work — where committing to “this set of items by two weeks from now” is fiction because priorities shift daily (on-call, production support, ad hoc requests).

Scrumban is a hybrid: Scrum’s cadence and ceremonies (planning, retro) layered on top of Kanban’s continuous flow and WIP limits instead of fixed sprint commitments. It’s common in teams transitioning from Scrum to Kanban, or in teams that want the reflective cadence of Scrum without pretending that all incoming work is plannable two weeks out.

DimensionScrumKanbanScrumban
CadenceFixed-length sprints (1–4 weeks)Continuous, no fixed iterationFixed cadence for planning/retro, continuous flow for work
RolesProduct Owner, Scrum Master, Developers (defined in the Scrum Guide)No prescribed roles; team self-organizes around the boardUsually informal, borrows Scrum roles loosely
Core artifactsProduct backlog, sprint backlog, incrementKanban board, WIP limits, cumulative flow diagramBacklog + Kanban board, WIP limits
Key constraintSprint commitment (scope locked for the sprint)WIP limits per columnWIP limits, no scope lock
Planning unitSprint (batch of stories committed up front)Single item pulled when capacity frees upLightweight, just-in-time planning
Best fitFeature teams with plannable, batchable work and value in a regular demo/review cadenceSupport, ops, platform, or teams with high interrupt rates and unpredictable inbound workTeams transitioning frameworks, or wanting Scrum’s reflection cadence without Scrum’s commitment model
Failure mode if misappliedSprints become a formality — team carries over half the backlog every sprintTeam loses any forcing function for reflection/improvementCan become “worst of both” if adopted without intent (no real commitment, no real flow discipline)

Choosing between these is a diagnostic exercise, not a preference. Ask: how volatile is inbound work? Do stakeholders need a predictable demo cadence? Does the team do a meaningful share of unplanned/reactive work (on-call, incidents, urgent requests)? A platform team fielding constant interrupt-driven requests forced into two-week sprint commitments will chronically “fail” sprints — not because the team is underperforming, but because the framework doesn’t match the work shape. Conversely, a product feature team on Kanban with no cadence at all can drift without a forcing function for retrospection and stakeholder alignment.

Sprint planning: capacity, estimation, and commitment ownership

Sprint planning has three distinct jobs that get conflated in bad implementations: figuring out how much capacity the team actually has, estimating the relative size of candidate work, and deciding what to commit to. Each deserves separate treatment.

Capacity planning starts from available person-days, not headcount. A five-person team in a two-week sprint is not “50 person-days of capacity” — subtract PTO, holidays, on-call rotations, meetings, interviews, and the tax of ongoing production support. Teams that plan against nominal headcount rather than real availability build in a source of chronic sprint failure that has nothing to do with estimation accuracy. A simple, honest capacity calculation (days available × focus factor, where focus factor is typically 60–80% once you subtract meetings, interrupts, and context-switching overhead) beats an optimistic headcount multiplication every time.

Estimation exists to support planning decisions, not to produce an accurate forecast of the future — no estimation approach reliably does that, and treating estimates as commitments is the single most common cause of estimation dysfunction. Common approaches:

ApproachHow it worksStrengthLimitation
Story points (relative sizing, often Fibonacci-like: 1, 2, 3, 5, 8, 13)Team assigns a unitless size relative to a reference storyDecouples size from time; forces a “how does this compare to X” conversation that surfaces hidden complexityEasily gamed/inflated over time (“point inflation”); meaningless across teams; can become bureaucratic ritual if the conversation dies out
Planning pokerEach person picks a point value simultaneously (hidden), reveals together, discusses outliersPrevents anchoring on the loudest/most senior voice; surfaces disagreement fastSlow if overused on trivial items; theater if people stop actually thinking and just copy the room
T-shirt sizing (XS/S/M/L/XL)Coarse relative sizing, often for early/rough estimatesFast, low-precision-signaling (doesn’t pretend false accuracy), good for backlog triageToo coarse for sprint commitment; needs translation to points/hours eventually
#NoEstimates / counting itemsTrack throughput of items regardless of size (assumes reasonably uniform slicing)Removes estimation overhead entirely; works well with disciplined small-story slicingRequires real discipline in breaking work small and uniformly; doesn’t help with large, lumpy work
Hours/days estimationDirect time estimate per taskIntuitive to non-engineers/stakeholdersSystematically underestimates (see planning fallacy below); invites micromanagement of hours

Whichever approach is used, the estimate is a probabilistic statement about a complex, uncertain activity — not a promise. The number one dysfunction to guard against as an EM is estimates being reinterpreted upstream (by a PM, a director, a sales team) as a deadline commitment. If “5 points” quietly becomes “done by Thursday” in someone else’s slide deck, the estimation process has been weaponized against the team, and you need to intervene before it happens again.

Commitment must be a team decision, not a manager mandate. This is the point most new EMs get wrong under delivery pressure. If a manager (or an outside stakeholder speaking through the manager) sets the sprint scope unilaterally — “we’re doing these fifteen items this sprint” — two things happen: the team’s actual signal about their own capacity is discarded, and accountability quietly shifts from “the team decided this was achievable” to “the manager told us to do this,” which corrodes ownership and gives everyone a built-in excuse when the sprint fails. The EM’s job in planning is to bring context (priorities, deadlines, constraints, stakeholder needs) into the room, not to write the answer on the whiteboard before the discussion starts. When a real external deadline conflicts with what the team believes is achievable, that’s a negotiation about scope or timeline — surfaced explicitly, not resolved by overriding the team’s own capacity judgment.

Key Concepts

Evolving process without process theater

Process theater is any ceremony, document, tool, or metric that survives past the point it was solving a real problem — kept alive by inertia, “that’s how we’ve always done it,” or because removing it feels like admitting failure. It is one of the most reliable ways for a team’s velocity and morale to quietly decay while every dashboard still looks fine.

The discipline that avoids this: start minimal, add process only in response to a demonstrated, recurring pain point, and periodically audit existing process for whether it still earns its keep.

The test for any piece of process, old or new: can you name the specific failure mode it prevents, and is that failure mode still plausible for this team, right now? If the answer is no, it’s a deletion candidate, not a sacred cow.

Process documentation and onboarding

An undocumented “how we work” is a process that only exists in senior team members’ heads, which means every new hire either interrupts someone for an hour or reverse-engineers the process by watching what happens. A living, short “team working agreement” document — sprint length and ceremony schedule, Definition of Ready/Done, branching and review conventions, on-call expectations, escalation paths, tooling — should be the first thing a new engineer reads, and it should be treated as a piece of code: it goes stale unless someone owns keeping it current.

Practical mechanics that keep it alive rather than becoming a fossil: review and update it as an explicit item after any process change lands (not “someday”), assign rotating ownership rather than leaving it to whoever wrote it originally, and use it actively in onboarding — have new hires read it in their first day and flag anything confusing or out of date, which doubles as a lightweight audit of its accuracy. A working agreement nobody has edited in a year is either a perfectly stable team (rare) or a document nobody trusts anymore (common) — worth checking which.

Development and release workflow

Branching strategy sets the rhythm of how work merges and ships; see ../devops/en/04-version-control.md for the full mechanics of Git branching models (trunk-based, Git Flow, GitHub Flow). As an EM the decision that matters most is trunk-based development versus long-lived feature branches: trunk-based (small, frequent merges to main, feature flags for anything not ready) keeps integration pain low and enables continuous delivery, while long-lived branches (Git Flow style) reduce the feeling of risk per merge but concentrate integration pain into large, infrequent, high-blast-radius merges. Most high-performing teams converge on trunk-based development with short-lived branches, because it matches the “small batches, fast feedback” spirit of both agile and DORA’s findings on elite performers.

Definition of Ready (DoR) and Definition of Done (DoD) are the two guardrails that keep a sprint’s scope honest. DoR is the bar a backlog item must clear before it enters sprint planning as a candidate — acceptance criteria written, dependencies identified, design/mockups attached if needed, roughly sized. Work that doesn’t meet DoR and gets pulled into a sprint anyway is the single most common cause of mid-sprint scope discovery (“oh, this actually needs backend changes we didn’t scope”). DoD is the bar an item must clear to be considered complete — typically: code merged, tests passing, deployed to the relevant environment, acceptance criteria verified, documentation updated, monitoring/alerting in place if relevant. A team with no explicit DoD will have chronic disagreement about whether something is “actually done,” usually surfacing during a demo when a stakeholder finds an edge case that was never tested.

Both DoR and DoD should be short (a checklist, not a policy document), owned by the team (not imposed), and revisited when they cause friction — see ./08-engineering-practices-oversight.md for how DoD interacts with code review standards, testing bars, and technical practice oversight more broadly.

Release management

Release cadence is a genuine trade-off, not a solved problem with one right answer:

ApproachDescriptionAdvantageCost
Continuous deploymentEvery change that passes the pipeline ships to production automaticallySmallest possible batch size, fastest feedback, lowest per-release risk (DORA’s research consistently links this to elite performance)Requires strong automated test coverage, feature flags, and monitoring/rollback discipline — expensive to retrofit onto a team without those foundations
Scheduled releases (e.g., weekly, biweekly)Changes accumulate and ship in a batch on a fixed schedulePredictable for stakeholders, support, and release-note communication; easier to coordinate cross-team dependenciesLarger batches mean more surface area per release, harder root-causing when something breaks, and a “train” model where late changes either miss the train or delay it
Release trains (fixed schedule, code freeze before)Hybrid: fixed cadence, but with an explicit cutoff and stabilization windowBalances predictability with some batch-size controlThe freeze window itself is often a source of friction and blocked work

Feature flags are the mechanism that decouples deploying code from releasing a feature to users — code can merge to main and deploy to production dark (flag off), then be enabled progressively (internal users, percentage rollout, specific customers) independent of the deploy pipeline. This is what makes continuous deployment compatible with controlled, gradual feature rollout instead of “deploy = everyone sees it now.” See ../devops/en/11-ci-cd.md for the pipeline mechanics that make this safe (progressive delivery, canary releases, automated rollback triggers). As an EM, feature flags are also a scheduling tool: a feature that’s 90% done can merge incrementally behind a flag rather than living on a long-lived branch accumulating merge conflicts, which is one of the more underrated ways to keep trunk-based development viable for larger, multi-sprint features.

Milestone management and timeline estimation

The single most damaging habit in software timeline communication is false precision — stating “this ships March 14th” when the honest answer is “most likely between late February and mid-April, with a small chance of slipping further.” The planning fallacy (Kahneman & Tversky) describes the well-documented, near-universal tendency to underestimate task duration even when accounting for known risks, because people plan against a best-case execution path and systematically discount the probability of the unexpected — a dependency that slips, a design that turns out to be wrong, an unplanned production incident that eats a week. This isn’t a discipline failure specific to bad engineers; it’s a cognitive bias that shows up even in teams with excellent estimation processes, which is why process alone (more granular tickets, more planning poker) doesn’t fix it — only explicitly modeling uncertainty does.

Practical techniques for communicating honestly instead of projecting false confidence:

See ./15-strategy-business-case-and-prioritization.md for how these timeline communications feed into broader prioritization and business-case conversations with stakeholders.

Velocity: what it’s for and what it isn’t

Velocity — typically, story points (or items) completed per sprint — is a capacity-planning tool and nothing more. Its one legitimate use: helping a specific team forecast how much work it can plausibly take on in a future sprint, based on its own recent, consistent history, and helping that same team have an early conversation about whether a milestone looks achievable given current throughput.

What velocity is reliably not good for, and why it backfires when misused:

The practical guardrail: velocity stays inside the team, used by the team, for the team’s own planning. The instant it leaves that boundary — reported to leadership as a KPI, compared across teams, tied to review cycles — it stops measuring what it used to measure and starts measuring how well people have learned to game it.

Quality metrics vs. vanity metrics

A healthy delivery process needs to track quality, not just throughput, and the metrics worth tracking are ones that correlate with real outcomes rather than ones that are easy to compute or flattering to report.

The DORA (DevOps Research and Assessment) program identified four metrics that reliably distinguish elite-performing engineering organizations, based on years of large-sample research:

MetricWhat it measuresWhy it matters
Deployment frequencyHow often code successfully deploys to productionProxy for batch size — more frequent deploys mean smaller, lower-risk changes
Lead time for changesTime from code committed to code running in productionProxy for how much friction/waste is in the delivery pipeline itself
Change failure ratePercentage of deployments causing a production failure requiring remediationDirect quality signal on the release process, not just the code
Time to restore serviceHow long it takes to recover from a production failureMeasures resilience and incident response, not just failure avoidance

Beyond DORA, other genuinely useful quality signals include the escaped defect rate (bugs found in production versus caught pre-release — a direct measure of whether testing and review are catching what they should) and cycle time (time from work starting to work shipping, a finer-grained sibling of lead time that highlights where work gets stuck within a single team’s process, e.g., sitting in code review for three days).

Contrast this with common vanity metrics — numbers that are easy to report and look good on a slide but don’t reliably correlate with real delivery health: lines of code written, number of commits, tickets closed per person, story points completed (as covered above), or hours logged. These metrics are trivially gameable and, worse, actively reward the wrong behavior when anyone optimizes for them directly (more commits by splitting trivial changes, more tickets by picking easy work, more lines of code by not refactoring toward simplicity). The test for any metric before adopting it: can it be gamed by someone acting purely in their own short-term interest, and if optimized directly, does it still point toward the outcome you actually want? DORA’s four metrics pass this test reasonably well because gaming one of them tends to visibly hurt another (e.g., deploying more often without fixing quality shows up immediately in change failure rate).

Best Practices

Choosing and running the process

SituationRecommendation
Standing up a new team or migrating frameworksStart with the minimal viable ceremony set; add process only against a named, recurring pain point
Work is highly interrupt-driven (ops, support, platform)Prefer Kanban or Scrumban over Scrum; forcing sprint commitments onto unplannable work manufactures chronic “failed sprints”
Sprint commitment is being set top-downStop — bring context and constraints into planning, let the team own the commitment; a mandated commitment quietly shifts accountability away from the team
Estimates are being reported upstream as deadlinesIntervene immediately; re-frame estimates as capacity-planning inputs, not promises, every time this happens
A ceremony or document has stopped adding valueName the failure mode it was meant to prevent; if that mode is no longer plausible, cut it — review this on a schedule, not just on complaint
Timelines are being requested by stakeholdersGive ranges with explicit confidence and drivers of variance, not single dates; a range stated honestly costs less trust than a point estimate that slips
Velocity is being requested by someone outside the teamRedirect to outcome- and quality-based metrics (DORA, escaped defects, cycle time); do not let velocity become an externally visible or comparison metric
A new engineer joins the teamPoint them to a living team working agreement (ceremonies, DoR/DoD, branching, escalation) on day one; use their questions to audit the document’s accuracy
Deciding release cadenceMatch cadence to the team’s test/rollback maturity — continuous deployment requires feature flags, strong automated tests, and fast rollback; without those, scheduled releases or release trains are the honest choice, not a failure of ambition

A short checklist for the EM

References