← Quản lý kỹ thuật← Engineering Manager
Quản lý kỹ thuậtEngineering Manager19 Th7, 2026Jul 19, 202623 phút đọc17 min read

Giám sát Engineering PracticesEngineering Practices Oversight

Thuộc bộ kiến thức Engineering Manager Roadmap.

Tổng quan

Một engineering manager (EM) thường không viết code production hàng ngày, và ngay cả khi có viết, họ cũng không thể tự mình review mọi pull request, chạy mọi test suite, hay theo dõi mọi lần deploy. Điều giúp một manager tin tưởng rằng một team tám, hai mươi, hay tám mươi engineer đang ship an toàn mà không cần loại giám sát cá nhân liên tục đó chính là sự tồn tại của các engineering practices tốt: CI/CD, code review, testing, và security work đủ mạnh để đóng vai trò guardrail. Các practice này thực hiện việc thực thi hàng ngày; công việc của manager là quyết định guardrail trông như thế nào, nhận ra khi nào chúng bắt đầu xuống cấp, và đưa ra các quyết định đầu tư giữ cho chúng khỏe mạnh.

Đây là một sự chuyển đổi tư thế tinh tế nhưng quan trọng. Một manager mới cố gắng bù đắp cho quy mô tổ chức bằng cách tự mình review nhiều code hơn, hoặc trở thành chốt chặn cuối cùng trước mỗi lần deploy, không phải đang scale — họ đang trở thành chính cái bottleneck mà lẽ ra họ phải loại bỏ. Giải pháp thay thế là coi practices như hạ tầng: thứ bạn thiết kế, cấp ngân sách, đo lường, và giao việc vận hành hàng ngày cho tech lead cùng senior engineer, giống như bạn giao việc vận hành một service cho một on-call rotation thay vì chính bạn ngồi nhìn dashboard cả ngày.

Note này cố tình tập trung vào tầng giám sát (oversight layer) — những gì EM quyết định, thực thi, và đo lường — chứ không phải cơ chế kỹ thuật của việc pipeline được xây dựng ra sao, một công cụ diff hoạt động thế nào, hay một test framework được cấu hình ra sao. Những cơ chế đó nằm ở nơi khác trong bộ kiến thức này và được liên kết xuyên suốt: CI/CD, Version Control & Collaboration, Testing, và DevSecOps knowledge base. Đọc note này cùng với Agile Process & Delivery, nói về cách công việc luân chuyển qua team, và Technical Decision-Making & Architecture, nói về các lựa chọn thiết kế ở tầng cao hơn quyết định việc duy trì các practice này khó hay dễ ngay từ đầu.

Kiến thức nền tảng

Mô hình giám sát: guardrail, không phải gate

Mô hình tư duy cốt lõi là: một manager đóng vai trò gate (chốt chặn) tự mình đứng giữa mọi công việc và việc release nó, và tốc độ của team bị giới hạn bởi băng thông của manager. Một manager xây dựng guardrail (hàng rào bảo vệ) — CI phải xanh, code review phải xảy ra, security scanning phải pass — cho phép team di chuyển theo nhịp riêng của mình bên trong các ranh giới manager đã thiết kế, mà không cần có mặt trong phòng. Guardrail scale được; gate thì không.

Điều này có nghĩa là sự tham gia thực tế hàng ngày của EM vào engineering practices trông không giống “review code” mà giống hơn với:

Câu hỏi EM đặt raNó tiết lộ điều gì
Pipeline CI của chúng ta nhanh và đáng tin, hay là thứ mọi người tìm cách né tránh?Guardrail có thực sự chịu lực hay không
Code review có diễn ra trong một SLA hợp lý, hay đang xếp hàng nhiều ngày?Review có phải là một chốt chất lượng thật sự hay chỉ là hình thức gây nghẽn
Team có biết, với tư cách tập thể, test coverage của mình mang lại gì — và ở đâu thì không?Đầu tư testing có khớp với rủi ro thực tế hay không
Security là việc một người làm ở cuối, hay việc mọi người làm xuyên suốt?”Shift-left” là thật hay chỉ là khẩu hiệu
Các chỉ số DORA của chúng ta nói gì về hệ thống, không phải về một engineer cụ thể?Các practice có đang tạo ra kết quả mà chúng được thiết kế để tạo ra hay không

Không câu hỏi nào trong số này đòi hỏi EM phải mở một diff hay đọc một stack trace. Chúng đòi hỏi EM nhìn vào xu hướng, trò chuyện với tech lead, và quyết định nên dành thời gian engineering khan hiếm vào đâu.

Ai sở hữu cái gì: cơ chế vs. giám sát

Một bảng hữu ích cần ghi nhớ khi giao việc là ai sở hữu cơ chế hàng ngày của một practice so với điều manager chịu trách nhiệm cụ thể:

PracticeSở hữu cơ chế hàng ngàyEM sở hữu
CI/CD pipelineTech lead / platform hoặc DevOps engineer cấu hình các stage, tooling, environmentĐặt tiêu chuẩn độ tin cậy/tốc độ, cấp ngân sách cho pipeline work so với feature, escalate khi hỏng dai dẳng
Code reviewSenior engineer và reviewer thực hiện review thực tếĐặt chuẩn mực SLA, ngăn bottleneck, dùng review để phát triển con người
Chiến lược testingEngineer viết test; QA/test engineer có thể sở hữu frameworkQuyết định mức đầu tư dựa trên rủi ro của team, bảo vệ thời gian cho nó
Security practicesSecurity champion / AppSec làm threat modeling, scanning, remediationĐưa security vào định nghĩa “done”, bảo vệ ngân sách cho security work vô hình

Mô hình lặp lại: EM không vắng mặt khỏi các practice này, nhưng sự tham gia của họ nằm ở tầng chính sách, ngân sách, và escalation — không phải ở tầng thực thi.

CI/CD dưới góc nhìn quản lý

Cơ chế kỹ thuật sâu của pipeline — các stage, build-once-deploy-many, test pyramid bên trong CI, chiến lược deployment — được trình bày trong CI/CD. Từ ghế của EM, pipeline không phải là một artifact kỹ thuật để cấu hình; nó là thứ quyết định việc ship tốn kém đến mức nào cho cả team, mỗi ngày.

“Tốt” dưới góc nhìn quản lý có nghĩa là ba điều cùng lúc:

Cái giá của việc làm sai điều này không hề trừu tượng. Một pipeline flaky hoặc chậm là một loại thuế lặp lại đánh vào tinh thần và tốc độ: engineer chuyển ngữ cảnh trong lúc chờ, mất trạng thái tập trung (flow), phát triển sự bất lực học được đối với CI (“nó luôn hỏng, cứ chạy lại”), và cuối cùng ngừng tin vào build xanh — lúc đó pipeline đã ngừng làm việc của nó dù về mặt kỹ thuật vẫn đang chạy. Đây là kiểu xuống cấp vô hình trong một sprint đơn lẻ nhưng tàn phá trong một quý, và đó chính xác là lý do nó cần sự chú ý quản lý có chủ đích thay vì giả định nó sẽ tự sửa.

Quyết định cụ thể mà EM sở hữu là khi nào rút engineer khỏi feature work để đầu tư vào sức khỏe pipeline. Không có con số phổ quát, nhưng các điểm kích hoạt hữu ích bao gồm: thời gian feedback build/test dần vượt quá mức team coi là “nhanh” (thường là 10–15 phút cho vòng feedback nhanh), tỷ lệ flaky test cao đến mức engineer thường xuyên chạy lại mà không điều tra, hoặc một pattern các sự cố được truy ra là do những lỗ hổng lẽ ra pipeline phải bắt được. Khi bất kỳ dấu hiệu nào trong số này xuất hiện lặp lại, việc coi sức khỏe pipeline là một hạng mục backlog hạng nhất — với story point thật và ưu tiên thật so với feature — là một quyết định quản lý, không phải một điều “hay thì làm” của engineering.

Code review dưới góc nhìn quản lý

Chi tiết kỹ thuật của pull request, branching, và review tooling nằm trong Version Control & Collaboration. Việc của EM là thiết lập môi trường chính sách mà review diễn ra bên trong.

Ba đòn bẩy quan trọng nhất:

Đòn bẩy ít được dùng nhất, và đáng nêu ra rõ ràng, là dùng code review như một cơ chế lan truyền kiến thức chứ không chỉ là gác cổng. Ghép PR của một junior engineer với một senior reviewer, hoặc chủ động phân công review chéo team, mang lại nhiều giá trị cho kỹ năng tập thể và bus-factor của team hơn bất kỳ khối lượng đào tạo chính thức nào — và đó là một đòn bẩy chỉ có manager mới ở vị trí để chủ động kéo, vì từng engineer riêng lẻ sẽ mặc định chọn ai nhanh nhất hoặc rảnh nhất.

Chiến lược testing dưới góc nhìn quản lý

Phân loại kỹ thuật — unit, integration, end-to-end, contract test, cuộc tranh luận pyramid vs. trophy — được trình bày trong Testing. Điều EM thực sự quyết định là bao nhiêu thời gian của team đáng để dành cho tầng nào, xét theo hồ sơ rủi ro của thứ họ đang xây dựng.

Đây về cơ bản là một quyết định rủi ro-và-đầu tư, không phải một quyết định kỹ thuật. Một team xây dựng công cụ admin nội bộ dùng bởi năm người có chi phí thất bại chấp nhận được rất khác với một team xây dựng luồng checkout cho sản phẩm thanh toán. Vai trò của manager là làm cho việc phân loại rủi ro đó trở nên rõ ràng thay vì để mỗi engineer tự đoán mức độ cẩn thận cần có:

Bối cảnhTư thế chấp nhận đượcVì sao
Công cụ nội bộ, feature ít traffic, thay đổi đảo ngược được”Move fast, fix forward” — ship sau feature flag, theo dõi, rollback nếu saiChi phí một bug thấp và khôi phục được
Bề mặt sản phẩm hướng người dùng cốt lõiTest coverage tự động vững chắc + rollout theo giai đoạnChi phí một regression hiển thị và ảnh hưởng danh tiếng
Thanh toán, billing, đối soát tài chínhTest coverage dày, review bắt buộc, thường có cửa sổ đóng băng thay đổi hoặc sign-off bổ sungChi phí một bug là tiền bạc, đôi khi không đảo ngược được
Lĩnh vực bị quản lý (y tế, tuân thủ tài chính, hệ thống an toàn tính mạng)Kế hoạch test chính thức, audit trail, có thể cần review/chứng nhận bên ngoàiChi phí một bug có thể là pháp lý hoặc an toàn tính mạng, không chỉ tài chính

“Move fast and fix forward” là một văn hóa engineering chính đáng — nhưng chỉ khi manager đã chủ động quyết định phạm vi ảnh hưởng (blast radius) của thất bại là chấp nhận được. Áp dụng nó đồng loạt, không có sự phân loại rủi ro đó, chính là cách các team gặp phải sự cố production đúng ở những hệ thống mà họ ít có khả năng chịu đựng nhất. Đây là một trong những trường hợp rõ ràng nhất mà phán đoán của EM — chứ không phải sở thích engineering — đặt ra tiêu chuẩn.

Security là việc của tất cả mọi người

Cơ chế kỹ thuật sâu của security practices — SAST/DAST, quét dependency, threat modeling, secure coding, incident response — được trình bày xuyên suốt DevSecOps knowledge base. Trách nhiệm ở tầng quản lý hẹp hơn nhưng khó duy trì hơn: biến security thành thứ cả team sở hữu liên tục, không phải một checklist một người chạy trước mỗi lần release.

Hai điều thuộc hẳn về manager:

Khái niệm chính

Đặt ra tiêu chuẩn mà không trở thành bottleneck

Căng thẳng mà mọi EM cuối cùng đều gặp phải là: ai đó cần thực thi các practice, nhưng không thể luôn là manager, nếu không manager sẽ trở thành trần giới hạn cho throughput của team. Giải pháp là một mô hình hai tầng:

Giao quyền sở hữu chất lượng theo cách này có nghĩa là manager xem xét kết quả của các practice định kỳ — xu hướng sức khỏe pipeline, độ trễ review, postmortem sự cố, tuổi của security backlog — thay vì tự mình xét lại từng quyết định. Anti-pattern cần chú ý ở chính mình là trở thành reviewer hoặc approver cuối cùng trên thực tế cho mọi thứ, điều này không chỉ tạo ra bottleneck mà còn âm thầm nói với senior engineer rằng phán đoán của họ không được tin tưởng — điều ăn mòn sự phát triển của họ và toàn bộ lý do tuyển senior ngay từ đầu.

Đo lường sức khỏe practice: DORA four keys

Chương trình nghiên cứu DORA (DevOps Research and Assessment), được tóm tắt trong Accelerate và các báo cáo State of DevOps thường niên, đã chắt lọc nhiều năm dữ liệu thành bốn chỉ số tương quan với các tổ chức engineering hiệu suất cao. Với một EM, đây là tín hiệu chẩn đoán ở tầng hệ thống, không phải bảng xếp hạng để đánh giá từng engineer cá nhân.

Chỉ sốĐo lường điều gìSự suy giảm báo hiệu gì cho manager
Deployment frequency (tần suất deploy)Team ship lên production bao lâu một lầnKích thước batch đang lớn dần, ma sát trong quy trình release tăng, hoặc nỗi sợ deploy
Lead time for changes (thời gian từ commit đến production)Thời gian từ commit đến khi chạy trên productionBottleneck review, CI chậm/flaky, hoặc ma sát trong chuỗi approval
Change failure rate (tỷ lệ thay đổi gây lỗi)Phần trăm deployment gây lỗi productionLỗ hổng testing, chiều sâu review không đủ, hoặc phân loại rủi ro không khớp thực tế
Time to restore service (thời gian khôi phục dịch vụ)Mất bao lâu để hồi phục sau một sự cố productionObservability yếu, quy trình rollback không rõ ràng, hoặc quyền sở hữu sự cố không rõ ràng

Giá trị của bốn chỉ số này là chúng biến “engineering có khỏe mạnh không?” từ một cảm giác mơ hồ thành một đường xu hướng mà manager có thể hành động. Ví dụ, change failure rate tăng lên là một gợi ý để đi xem xét liệu đầu tư testing hay chiều sâu review có đang âm thầm xói mòn hay không — không phải một lý do để công khai chỉ trích ai đó có deploy gây ra sự cố gần nhất. Dùng tốt, các chỉ số DORA hướng manager đến practice nào cần chú ý; dùng sai, chúng bị gian lận (deployment frequency bị thổi phồng bằng commit vô nghĩa, lead time bị lách bằng cách gộp review) hoặc bị biến thành chỉ số hiệu suất cá nhân, cả hai đều phá hủy mục đích của chúng và làm tổn hại niềm tin. Kết hợp chúng với tín hiệu định tính từ 1:1 và retro — các con số cho bạn biết nên nhìn vào đâu, không phải toàn bộ câu chuyện.

Chi phí dồn tích của nợ practice

Practice yếu không thất bại ầm ĩ ngay ngày đầu; chúng tích tụ âm thầm rồi biểu hiện thành một cuộc khủng hoảng. Đáng để lần theo chuỗi này một cách rõ ràng, vì đó là lý lẽ cho việc tại sao một EM nên đầu tư vào practice trước khi chúng hiển nhiên bị hỏng:

Trong mọi trường hợp, pattern giống nhau: chi phí bỏ bê một practice bị hoãn lại, không bị tránh khỏi, và nó dồn tích. Đây là lý lẽ mạnh nhất để một EM coi sức khỏe practice là một hạng mục lặp lại trong lập kế hoạch, không phải một dự án dọn dẹp làm một lần mỗi năm.

Tín hiệu cần can thiệp trực tiếp

Hầu hết thời gian, việc giao cho tech lead là đủ. Một EM nên can thiệp trực tiếp hơn khi họ thấy: cùng một stage pipeline hỏng lặp đi lặp lại mà không ai sở hữu việc sửa nó, vi phạm SLA review trở thành chuẩn mực thay vì ngoại lệ, test coverage âm thầm giảm trên một hệ thống vừa được phân loại lại là rủi ro cao hơn, hoặc một security backlog bị hạ ưu tiên xuống dưới feature work trong nhiều chu kỳ lập kế hoạch liên tiếp. Đây là những khoảnh khắc mà “tin tưởng tech lead xử lý” đã ngừng hoạt động và manager cần tự mình đặt lại ưu tiên.

Best Practices

Khu vực practiceLink deep-diveTrách nhiệm cụ thể của EM
CI/CDCI/CDĐịnh nghĩa tiêu chuẩn tốc độ/độ tin cậy; quyết định khi nào cấp ngân sách cho pipeline health work thay vì feature; escalate khi flaky dai dẳng
Code reviewVersion Control & CollaborationĐặt chuẩn mực SLA review; ngăn bottleneck từ single-approver; dùng phân công review để lan truyền kiến thức
TestingTestingĐặt mức đầu tư testing dựa trên rủi ro của team; quyết định nơi nào “move fast, fix forward” chấp nhận được và nơi nào không
SecurityDevSecOps knowledge baseĐưa security vào Definition of Done; bảo vệ ngân sách cho security work vô hình
Luồng deliveryAgile Process & DeliveryGiữ WIP và kích thước batch hợp lý để chất lượng review/testing không xói mòn dưới áp lực flow
Kiến trúcTechnical Decision-Making & ArchitectureĐảm bảo các quyết định kiến trúc không âm thầm khiến testing hay deploy an toàn khó hơn

Ngoài bảng trên, một vài thói quen giúp mô hình giám sát này hoạt động tốt trong thực tế:

Tài liệu tham khảo

Part of the Engineering Manager Roadmap knowledge base.

Overview

An engineering manager (EM) usually does not write production code every day, and even when they do, they cannot personally review every pull request, run every test suite, or watch every deployment. What lets a manager trust that a team of eight, twenty, or eighty engineers is shipping safely without that kind of constant personal oversight is the presence of good engineering practices: CI/CD, code review, testing, and security work that are strong enough to act as guardrails. The practices do the day-to-day enforcement; the manager’s job is to decide what the guardrails look like, notice when they decay, and make the investment calls that keep them healthy.

This is a subtle but important shift in posture. A new manager who tries to compensate for organizational scale by reviewing more code personally, or by being the last gate before every deploy, is not scaling — they are becoming the bottleneck they were supposed to remove. The alternative is to treat practices as infrastructure: something you design, fund, measure, and delegate the daily operation of to tech leads and senior engineers, the same way you’d delegate the operation of a service to an on-call rotation rather than staring at a dashboard all day yourself.

This note is deliberately about the oversight layer — what an EM decides, enforces, and measures — not the technical mechanics of how a pipeline is built, how a diff tool works, or how a test framework is configured. Those mechanics live elsewhere in this knowledge base and are linked throughout: CI/CD, Version Control & Collaboration, Testing, and the DevSecOps knowledge base. Read this note alongside Agile Process & Delivery, which covers how work flows through the team, and Technical Decision-Making & Architecture, which covers the higher-level design choices that determine how hard these practices are to sustain in the first place.

Fundamentals

The oversight model: guardrails, not gates

The core mental model is this: a manager who is a gate personally sits between every piece of work and its release, and the team’s speed is capped by the manager’s bandwidth. A manager who builds guardrails — CI that must be green, code review that must happen, security scanning that must pass — lets the team move at its own pace inside boundaries the manager designed, without needing to be in the room. Guardrails scale; gates don’t.

This means the EM’s actual day-to-day involvement with engineering practices looks less like “reviewing code” and more like:

Question the EM asksWhat it reveals
Is our CI pipeline fast and trusted, or is it something people route around?Whether the guardrail is actually load-bearing
Are code reviews happening within a reasonable SLA, or queuing for days?Whether review is a genuine quality gate or a bottleneck theater
Do we know, as a team, what our test coverage buys us — and where it doesn’t?Whether testing investment matches actual risk
Is security something one person does at the end, or something everyone does throughout?Whether “shift-left” is real or aspirational
What do our DORA metrics say about the system, not about any one engineer?Whether the practices are producing the outcomes they’re meant to

None of these questions require the EM to open a diff or read a stack trace. They require the EM to look at trends, talk to tech leads, and decide where to spend scarce engineering time.

Who owns what: mechanics vs. oversight

A useful table to keep in mind when delegating is who owns the mechanics of a practice versus what the manager is specifically accountable for:

PracticeOwns the day-to-day mechanicsEM owns
CI/CD pipelineTech lead / platform or DevOps engineer configures stages, tooling, environmentsSetting reliability/speed bar, funding pipeline work vs. features, escalating chronic breakage
Code reviewSenior engineers and reviewers do the actual reviewsSetting SLA norms, preventing bottlenecks, using review to develop people
Testing strategyEngineers write tests; QA/test engineers may own frameworksDeciding the team’s risk-based investment level, protecting time for it
Security practicesSecurity champions / AppSec do threat modeling, scanning, remediationMaking security part of “done,” protecting budget for invisible security work

The pattern repeats: the EM is not absent from these practices, but their involvement is at the level of policy, budget, and escalation — not execution.

CI/CD from a management lens

The deep mechanics of pipelines — stages, build-once-deploy-many, test pyramids inside CI, deployment strategies — are covered in CI/CD. From an EM’s chair, the pipeline is not a technical artifact to configure; it is the thing that determines how expensive shipping is for the whole team, every day.

“Good” from a management perspective means three things at once:

The cost of getting this wrong is not abstract. A flaky or slow pipeline is a recurring tax on morale and velocity: engineers context-switch while waiting, lose flow state, develop a learned helplessness toward CI (“it’s always broken, just rerun it”), and eventually stop trusting green builds — at which point the pipeline has stopped doing its job even though it is technically still running. This is the kind of decay that is invisible in a single sprint and devastating over a quarter, which is exactly why it needs deliberate management attention rather than assuming it will fix itself.

The specific decision an EM owns is when to pull engineers off feature work to invest in pipeline health. There is no universal number, but useful trigger points include: build/test feedback time creeping past what the team considers “fast” (often 10–15 minutes for the fast feedback loop), a flaky-test rate high enough that engineers routinely re-run without investigating, or a pattern of incidents traced back to gaps the pipeline should have caught. When any of these show up repeatedly, treating pipeline health as a first-class backlog item — with real story points and real prioritization against features — is a management decision, not an engineering nicety.

Code review from a management lens

The nuts and bolts of pull requests, branching, and review tooling live in Version Control & Collaboration. The EM’s job is to set the policy environment that review happens inside.

Three levers matter most:

The most underused lever, and one worth calling out explicitly, is using code review as a knowledge-spreading mechanism rather than only a gatekeeping one. Pairing a junior engineer’s PR with a senior reviewer, or deliberately assigning cross-team reviews, does more for the team’s collective skill and bus-factor than any amount of formal training — and it’s a lever only the manager is positioned to pull deliberately, because individual engineers will default to whoever is fastest or most available.

Testing strategy from a management lens

The technical taxonomy — unit, integration, end-to-end, contract tests, the pyramid vs. trophy debate — is covered in Testing. What the EM actually decides is how much of the team’s time is worth spending on which layer, given the risk profile of what they’re building.

This is fundamentally a risk-and-investment decision, not a technical one. A team building an internal admin tool used by five people has a very different acceptable failure cost than a team building the checkout flow for a payments product. The manager’s role is to make that risk classification explicit rather than leaving each engineer to guess at how careful to be:

ContextAcceptable postureWhy
Internal tools, low-traffic features, reversible changes”Move fast, fix forward” — ship behind a flag, monitor, roll back if wrongCost of a bug is low and recoverable
Core user-facing product surfacesSolid automated test coverage + staged rolloutCost of a regression is visible and reputational
Payments, billing, financial reconciliationHeavy test coverage, mandatory review, often a change freeze window or additional sign-offCost of a bug is monetary, sometimes irreversible
Regulated domains (healthcare, finance compliance, safety-critical systems)Formal test plans, audit trails, possibly external review/certificationCost of a bug can be legal or safety-critical, not just financial

“Move fast and fix forward” is a legitimate engineering culture — but only where the manager has consciously decided the blast radius of failure is acceptable. Applying it uniformly, without that risk classification, is how teams end up with production incidents in exactly the systems that could least afford them. This is one of the clearest cases where the EM’s judgment call — not engineering preference — sets the bar.

Security as everyone’s job

Deep technical mechanics of security practices — SAST/DAST, dependency scanning, threat modeling, secure coding, incident response — are covered across the DevSecOps knowledge base. The management-level responsibility is narrower but harder to sustain: making security something the whole team owns continuously, not a checklist one person runs before a release.

Two things fall squarely on the manager:

Key Concepts

Setting a bar without becoming a bottleneck

The tension every EM eventually hits is: someone needs to enforce the practices, but it can’t always be the manager, or the manager becomes the ceiling on the team’s throughput. The resolution is a two-layer model:

Delegating quality ownership this way means the manager reviews the outcomes of the practices periodically — pipeline health trends, review latency, incident postmortems, security backlog age — rather than personally re-litigating each decision. The anti-pattern to watch for in yourself is becoming the de facto final reviewer or approver on everything, which not only creates a bottleneck but quietly tells senior engineers their judgment isn’t trusted, which is corrosive to their growth and to the whole point of hiring seniors in the first place.

Measuring practice health: the DORA four keys

The DORA (DevOps Research and Assessment) research program, summarized in Accelerate and the annual State of DevOps reports, distilled years of data into four metrics that correlate with high-performing engineering organizations. For an EM, these are system-level diagnostic signals, not a scoreboard for ranking individual engineers.

MetricWhat it measuresWhat a decline signals to a manager
Deployment frequencyHow often the team ships to productionBatch sizes growing, release process friction increasing, or fear of deploying
Lead time for changesTime from commit to running in productionReview bottlenecks, slow/flaky CI, or approval-chain friction
Change failure ratePercentage of deployments causing a production failureTesting gaps, insufficient review depth, or risk classification not matching reality
Time to restore serviceHow long it takes to recover from a production incidentWeak observability, unclear rollback procedures, or unclear incident ownership

The value of these four keys is that they turn “is engineering healthy?” from a vague feeling into a trend line the manager can act on. A rising change failure rate, for instance, is a prompt to go look at whether test investment or review depth has quietly eroded — not a reason to publicly call out whoever’s deploy caused the last incident. Used well, DORA metrics point the manager toward which practice needs attention; used badly, they get gamed (deploy frequency inflated with meaningless commits, lead time gamed by batching review) or turned into an individual performance metric, which both defeats their purpose and damages trust. Pair them with qualitative signals from 1:1s and retros — the numbers tell you where to look, not the whole story.

The compounding cost of practice debt

Weak practices don’t fail loudly on day one; they accumulate quietly and then show up as a crisis. It’s worth tracing the chain explicitly, because it’s the argument for why an EM should invest in practices before they’re visibly broken:

In every case, the pattern is the same: the cost of neglecting a practice is deferred, not avoided, and it compounds. This is the single strongest argument for an EM to treat practice health as a recurring line item in planning, not a cleanup project done once a year.

Signals that warrant direct intervention

Most of the time, delegation to tech leads is enough. An EM should step in more directly when they see: the same pipeline stage breaking repeatedly without anyone owning the fix, review SLA breaches becoming the norm rather than the exception, test coverage silently dropping on a system just reclassified as higher-risk, or a security backlog that has been reprioritized below feature work for multiple planning cycles in a row. These are the moments where “trust the tech lead to handle it” has stopped working and the manager needs to personally re-set the priority.

Best Practices

Practice areaDeep-dive linkEM’s specific responsibility
CI/CDCI/CDDefine the speed/reliability bar; decide when to fund pipeline health work over features; escalate chronic flakiness
Code reviewVersion Control & CollaborationSet review SLA norms; prevent single-approver bottlenecks; use review assignments to spread knowledge
TestingTestingSet the team’s risk-based testing investment; decide where “move fast, fix forward” is and isn’t acceptable
SecurityDevSecOps knowledge baseMake security part of Definition of Done; protect budget for invisible security work
Delivery flowAgile Process & DeliveryKeep WIP and batch size sane so review/testing quality doesn’t erode under flow pressure
ArchitectureTechnical Decision-Making & ArchitectureEnsure architectural decisions don’t quietly make testing or safe deployment harder

Beyond the table, a few habits keep the oversight model working in practice:

References