← Quản lý kỹ thuật← Engineering Manager
Quản lý kỹ thuậtEngineering Manager19 Th7, 2026Jul 19, 202627 phút đọc20 min read

Ra quyết định kỹ thuật & Quản trị kiến trúcTechnical Decision-Making & Architecture Governance

Thuộc bộ kiến thức Engineering Manager Roadmap.

Tổng quan

Đâu đó giữa vị trí senior engineer và engineering manager, có một sự chuyển dịch âm thầm xảy ra: bạn không còn là người hiểu sâu nhất về bất kỳ subsystem cụ thể nào trong phòng họp nữa, và bạn trở thành người chịu trách nhiệm đảm bảo cả phòng vẫn đi đến được một quyết định tốt dù không phải bạn là chuyên gia. Đây là một chuyển đổi khó chịu với nhiều manager đi lên từ công việc kỹ thuật hands-on, vì nó có thể cảm giác như mất đi quyền lực. Thực ra không phải vậy — đó là sự thay đổi về loại quyền lực bạn nắm giữ. Công việc của bạn không còn là có câu trả lời tốt nhất; nó là sở hữu quy trình để câu trả lời tốt nhất có thể được tìm ra, được quyết định, và được ghi nhớ lại.

Note này nói về quy trình đó: các quyết định kỹ thuật nên được đưa ra như thế nào trong một team hay tổ chức, cách phân biệt giữa các quyết định đáng để cân nhắc kỹ và các quyết định nên được chốt trong năm phút, cách ghi lại quyết định để lý do đằng sau nó tồn tại lâu hơn những người đã đưa ra nó, và cách thiết lập vừa đủ tiêu chuẩn hóa (standardization) để engineer có thể di chuyển linh hoạt giữa các hệ thống mà không quá cứng nhắc đến mức không ai chọn được công cụ phù hợp cho một bài toán thực tế. Note này không thay thế kiến thức architecture kỹ thuật chuyên sâu — phần đó xem tại ../backend/en/14-software-architecture.md cho các architectural style, pattern và tradeoff, và ../backend/en/06-apis.md cho chi tiết API design. Note này nằm ở một lớp cao hơn: nó nói về cách manager đảm bảo công việc architecture diễn ra, diễn ra với đúng người, và luôn trung thực với những gì business thực sự cần.

Kiến thức nền tảng

Vai trò của manager là sở hữu quy trình, không phải “biết tất cả về kỹ thuật”

Một failure mode phổ biến của manager mới là cố gắng vẫn là người có thẩm quyền kỹ thuật cao nhất cho mọi quyết định trong team mình, hoặc vì họ đã quen là người đó, hoặc vì họ không tin ai khác có thể đảm nhận vai trò này. Cách này không scale được khi team lớn hơn vài người, và nó thực sự gây hại cho sự phát triển của team — nếu manager luôn có tiếng nói kỹ thuật cuối cùng, các senior engineer sẽ ngừng phát triển judgment và ngừng cố gắng xây dựng luận điểm mạnh mẽ, vì họ biết quyết định thực ra không phải của họ.

Mô hình lành mạnh hơn: công việc của manager là đảm bảo người gần nhất với vấn đề có đủ context, đủ mandate (thẩm quyền), và đủ áp lực để đưa ra một quyết định có lý lẽ vững — và quyết định đó thực sự được chốt trong một khoảng thời gian giới hạn. Cụ thể, điều đó có nghĩa là:

Bạn là người nhận ra khi một cuộc tranh luận đã ngừng đưa ra thông tin mới mà chỉ lặp lại các lập trường cũ — và là người có đủ vị thế để nói “chúng ta có đủ thông tin để quyết rồi, quyết đi.”

Decision paralysis và unilateral decision là hai mặt của cùng một vấn đề

Các team có xu hướng thất bại theo một trong hai hướng cực đoan:

Tê liệt vì đồng thuận (paralysis by consensus): mọi lựa chọn architecture đều bị đem ra tranh luận lại trong một cuộc họp với quá nhiều stakeholder, không ai sở hữu kết quả, và “quyết định” thực ra chỉ là một cuộc trò chuyện không đi đến đâu rồi lại được bàn tiếp ở sprint sau. Việc này đốt hàng tuần cho những thứ không xứng đáng, và nó dạy cho engineer rằng đưa ra một ý kiến mạnh là cách tốt để trì hoãn quyết định vô thời hạn.

Quyết định đơn phương bỏ qua chuyên môn (unilateral decisions): một manager, tech lead, hoặc senior IC tự quyết một mình, không tham khảo những người sẽ phải gánh hậu quả (on-call, team duy trì downstream service, security) — và quyết định hóa ra sai theo cách mà một cuộc trò chuyện năm phút lẽ ra đã phát hiện được. Việc này thường được biện minh là “để đi nhanh hơn,” nhưng thực ra nó chỉ đẩy chi phí xuống downstream và tạo thêm một khoản “thuế niềm tin” cho lần tiếp theo người đó ra quyết định.

Cả hai failure mode đều xuất phát từ cùng một mảnh ghép còn thiếu: không ai định nghĩa rõ ràng ai là người quyết, ai bắt buộc phải được tham khảo, và ai chỉ cần được thông báo. Giải quyết việc này bằng một framework nhẹ nhàng xử lý được cả hai vấn đề cùng lúc — nó giới hạn cuộc thảo luận (để không trở thành tê liệt) trong khi vẫn yêu cầu đúng chuyên môn được tham khảo (để không trở thành quyết định đơn phương bỏ qua expertise).

Khái niệm chính

DACI và RACI: gọi tên ai là người quyết

Thói quen có đòn bẩy cao nhất một manager có thể cài đặt là: gọi tên, trước khi cuộc tranh luận bắt đầu, ai thực sự là người cầm bút quyết định. Hai framework phổ biến:

DACI (Driver, Approver, Contributors, Informed) hướng đến một quyết định cụ thể:

Vai tròÝ nghĩa
DriverSở hữu việc thúc đẩy quyết định tiến lên: thu thập input, định hình các phương án, đẩy nó đến khi chốt xong. Thường là tech lead hoặc senior engineer, không nhất thiết là manager.
ApproverNgười (đôi khi hơn một người, tốt nhất không quá hai) có quyền phê duyệt cuối cùng. Mọi người cần biết trước ai là người này.
ContributorsNhững người có chuyên môn hoặc có liên quan đến kết quả nên bắt buộc phải được tham khảo trước khi quyết định — engineer on-call, người review security, team sở hữu downstream dependency.
InformedNhững người cần biết kết quả nhưng không có quyền góp ý vào nó — họ được thông báo sau khi quyết định đã chốt, không phải trong lúc bàn.

RACI (Responsible, Accountable, Consulted, Informed) là phiên bản tổng quát hơn, thường dùng cho việc sở hữu công việc liên tục thay vì một quyết định đơn lẻ, và phân biệt “ai làm việc” (Responsible) với “ai chịu trách nhiệm cuối cùng” (Accountable) — một sự phân biệt mà DACI gộp chung vào Driver/Approver.

Với một quyết định architecture đơn lẻ, DACI thường phù hợp hơn vì nó có hình dạng của một quyết định chứ không phải của quyền sở hữu công việc. Việc của manager hiếm khi là làm Driver hay Approver — mà là đảm bảo các vai trò đó được gán trước khi cuộc tranh luận bắt đầu, công khai, để không ai bối rối giữa chừng về việc ý kiến phản đối của mình là một veto hay chỉ là một input.

Quyết định có thể đảo ngược (“cửa hai chiều”) vs. không thể đảo ngược (“cửa một chiều”)

Không phải quyết định kỹ thuật nào cũng xứng đáng với cùng một mức độ quy trình, và áp dụng quy trình nặng nề cho một quyết định nhẹ chính là một thất bại quản lý — nó dạy engineer coi quy trình là ma sát thay vì là sự bảo vệ. Lăng kính rõ ràng nhất để hiệu chỉnh mức độ quy trình là cách Jeff Bezos phân biệt giữa cửa một chiều (one-way door)cửa hai chiều (two-way door):

Công việc thực tế của manager là giúp team phân loại đúng họ đang đứng trước loại cửa nào trước khi quyết định áp dụng bao nhiêu quy trình — vì phân loại sai một cửa một chiều thành có thể đảo ngược là cách các team mắc kẹt trong một architecture mà không ai từng cân nhắc kỹ, còn phân loại sai một cửa hai chiều thành không thể đảo ngược là cách các team giậm chân tại chỗ vì những quyết định không đáng để tốn công như vậy.

Architecture Decision Records (ADR)

ADR là một tài liệu ngắn ghi lại một quyết định architecture có ý nghĩa quan trọng: đã quyết định gì, tại sao, những phương án thay thế nào đã được xem xét, và hậu quả kỳ vọng (bao gồm cả mặt trái) là gì. Điểm mấu chốt của ADR không phải là “cái gì” — thường cái đó đã hiện diện trong code hoặc trong hệ thống — mà là tại sao. Sáu tháng hay hai năm sau, khi có ai đó (có thể là chính engineer đó, nhưng thường là người mới) nhìn vào một lựa chọn thiết kế kỳ lạ và tự hỏi “tại sao mình lại làm thế này,” ADR chính là artifact trả lời câu hỏi đó thay vì bắt người ta phải khảo cổ qua các Slack thread cũ và PR đã đóng.

ADR bắt nguồn từ bài viết năm 2011 của Michael Nygard và từ đó đã trở thành một chuẩn de facto trong ngành, với một template được sử dụng rộng rãi duy trì tại adr GitHub organization.

Ví dụ template ADR

# ADR-0012: Sử dụng PostgreSQL làm datastore chính cho Orders service

## Status
Accepted (2026-03-14)
<!-- Các giá trị có thể: Proposed, Accepted, Rejected, Deprecated, Superseded by ADR-00XX -->

## Context
Orders service hiện đang lưu order state trong một in-memory store được backup
bằng snapshot định kỳ lên S3, kế thừa từ bản prototype ban đầu. Khi order volume
đã vượt 50k đơn/ngày, chúng ta bắt đầu mất dữ liệu khi pod restart giữa các
snapshot, và không có khả năng chạy query ad-hoc cho support và analytics mà
không phải replay snapshot vào một môi trường riêng.

Chúng ta cần một persistence layer:
- Đảm bảo durability theo từng write, không phải theo chu kỳ snapshot
- Hỗ trợ transactional consistency giữa một order và các line item của nó
- Cho phép team support chạy read query mà không cần engineering can thiệp
- Có thể vận hành được bởi team hiện tại mà không cần tuyển DBA chuyên biệt

Các phương án đã xem xét: PostgreSQL (RDS), DynamoDB, MongoDB Atlas, giữ nguyên
cách snapshot hiện tại với chu kỳ ngắn hơn.

## Decision
Chúng ta sẽ dùng PostgreSQL, host trên Amazon RDS với Multi-AZ bật, làm datastore
chính cho Orders service.

Lý do:
- Transactional guarantee ánh xạ trực tiếp vào yêu cầu consistency giữa
  order/line-item mà không cần thêm logic compensation ở tầng application.
- Team đã có kinh nghiệm vận hành PostgreSQL sâu từ Billing service; không cần
  bộ kỹ năng vận hành mới.
- DynamoDB bị loại: các access pattern cho query support/analytics quá ad-hoc và
  đang thay đổi để cam kết vào yêu cầu thiết kế key ngay từ đầu của DynamoDB.
- MongoDB Atlas bị loại: không có lợi thế đáng kể nào so với PostgreSQL cho
  workload này, và sẽ tạo thêm gánh nặng vận hành document-store thứ hai bên
  cạnh các hệ thống relational hiện có.
- Rút ngắn chu kỳ snapshot bị loại vì đây chỉ là giải pháp tạm không giải quyết
  được vấn đề durability hay khả năng query, chỉ thu hẹp cửa sổ lỗi.

## Consequences
Tích cực:
- Durability theo từng write; không còn mất dữ liệu theo cửa sổ snapshot.
- Query ad-hoc trên read replica, không cần engineering can thiệp.
- Tái sử dụng runbook vận hành và kinh nghiệm on-call sẵn có từ Billing.

Tiêu cực / rủi ro:
- Thêm một stateful dependency mới cần vận hành và backup cho service này.
- Schema migration giờ cần kỷ luật tương tự (backward-compatible, được review)
  như Billing đã áp dụng — team này chưa từng làm việc này và sẽ cần thời gian
  ramp-up.
- Khả năng scale write theo chiều ngang bị giới hạn bởi một primary duy nhất;
  nếu order volume tăng gấp 10 lần, chúng ta có thể cần xem xét lại (sharding,
  hoặc một write path dựa trên queue).

## Related
- Supersedes: none
- Related to: ADR-0009 (thiết lập PostgreSQL cho Billing service, được dùng lại
  làm template vận hành)

Khi nào nên viết ADR — và khi nào là thừa thãi

Hãy viết ADR khi quyết định thực sự là một cửa một chiều, hoặc gần như vậy: lựa chọn core data store, thay đổi một public API contract, áp dụng một architectural pattern mới (chuyển sang event-driven, đưa vào service mesh), một lựa chọn thiết kế liên quan đến security hoặc compliance, hoặc bất cứ điều gì mà một engineer mới sau một năm nữa sẽ hợp lý khi hỏi “tại sao lại xây theo cách này?”

Đừng viết ADR cho các lựa chọn có thể đảo ngược, ít rủi ro — dùng linter config nào trong một service riêng lẻ, chọn thư viện utility nào trong hai thư viện tương đương, quy ước đặt tên cho một module đơn lẻ. Ép buộc viết ADR cho mọi quyết định làm mất giá trị của những ADR thực sự quan trọng; nếu engineer phải viết ADR cho những lựa chọn nhỏ nhặt, họ sẽ hoặc ngừng viết hoàn toàn hoặc biến nó thành một nghi thức vô nghĩa. Một heuristic hợp lý: nếu việc đảo ngược quyết định sau này sẽ tốn hơn một ngày công hoặc ảnh hưởng đến team khác, nó có lẽ xứng đáng có một ADR.

Tài liệu architecture luôn phản ánh đúng thực tế

Sơ đồ architecture xuống cấp nhanh hơn hầu như bất kỳ artifact nào khác trong codebase, vì không giống code, không có gì bắt buộc nó phải đồng bộ với thực tế — một sơ đồ không làm build fail khi hệ thống nó mô tả thay đổi bên dưới. Một sơ đồ từ design doc viết hai năm trước, thể hiện một architecture “kỳ vọng” mà sau đó bị đơn giản hóa dưới áp lực deadline, còn tệ hơn cả việc không có sơ đồ nào: nó chủ động đánh lừa bất kỳ ai tin tưởng nó, và làm giảm niềm tin vào mọi sơ đồ khác trên cùng wiki.

Một mô hình tư duy hữu ích để giữ tài liệu architecture vừa đầy đủ vừa trung thực là mô hình C4 (Context, Containers, Components, Code), cấu trúc tài liệu thành bốn mức zoom:

MứcThể hiệnĐối tượng
ContextHệ thống như một khối duy nhất, người dùng của nó, và các hệ thống khác nó giao tiếpBất kỳ ai, kể cả stakeholder không kỹ thuật
ContainersCác đơn vị deploy chính bên trong hệ thống (service, database, web app) và cách chúng giao tiếpEngineer ở các team khác nhau, người mới
ComponentsCác khối cấu trúc chính bên trong một containerEngineer làm việc trong container đó
CodeChi tiết mức class/module (thường để cho chính code, hoặc sơ đồ tự sinh)Engineer đang trực tiếp sửa code đó

Điều manager thực sự quan tâm ở đây không phải là vẽ sơ đồ — mà là biến mức Context và Container thành một artifact sống thay vì một deliverable làm một lần: được review với nhịp độ nhẹ nhàng (ví dụ mỗi quý, hoặc bất cứ khi nào một ADR thay đổi topology), được sở hữu bởi một người hoặc team cụ thể thay vì “ai viết ban đầu,” và được coi là không đáng tin và gắn cờ cần cập nhật ngay khi ai đó nhận ra nó không khớp thực tế. Một sơ đồ có ngày “last verified” hiển thị rõ ràng hữu ích hơn nhiều so với một sơ đồ được vẽ đẹp mà không ai nhìn đến trong một năm.

Vai trò giám sát của manager trong architecture review

Bạn không được kỳ vọng là chuyên gia architecture sâu nhất trong phòng — phần đó được nói ở ../backend/en/14-software-architecture.md, và đó là việc của senior engineer, staff engineer, và architect để mang đến chiều sâu đó. Điều bạn sở hữu là đảm bảo review thực sự diễn ra vào đúng thời điểm (trước khi đầu tư đáng kể, không phải như một con dấu cao su sau khi hệ thống đã xây dựng nửa chừng), đảm bảo nó bao gồm đúng Contributors dưới một phân công kiểu DACI, và — đây là phần thực sự thuộc về bạn — đảm bảo cuộc thảo luận luôn bám vào mục tiêu business thay vì trôi dạt sang sự tinh tế kỹ thuật vì chính nó.

Hoàn toàn có thể xảy ra việc một phòng đầy các engineer giỏi hội tụ vào một giải pháp architecture đẹp về mặt kỹ thuật nhưng sai đối với business: over-engineer so với quy mô thực tế sản phẩm sẽ đạt trong hai năm tới, tối ưu cho một sự linh hoạt mà sản phẩm sẽ không bao giờ cần, hoặc tốn thời gian xây dựng không tương xứng với doanh thu hoặc rủi ro mà nó giải quyết. Một phần công việc của bạn trong các review này là liên tục đặt các câu hỏi neo giữ: giải pháp này giải quyết vấn đề business nào, chi phí của việc làm cái đơn giản hơn thay vào đó là bao nhiêu, chuyện gì xảy ra nếu chúng ta sai về giả định tăng trưởng đang thúc đẩy thiết kế này? Bạn không cần biết công nghệ sâu như engineer của mình để đặt những câu hỏi tốt như vậy — bạn cần biết business và các ràng buộc của team rõ hơn bất kỳ ai khác trong phòng, và đó chính xác là sự bất đối xứng khiến sự hiện diện của bạn trong review có giá trị thay vì thừa thãi.

Technical standards: tradeoff giữa tiêu chuẩn hóa và linh hoạt

Mọi tổ chức cuối cùng đều phải quyết định mức độ tiêu chuẩn hóa: một ngôn ngữ được chấp thuận duy nhất hay để team tự chọn, một linter và style guide bắt buộc ở mọi nơi hay style riêng theo team, một danh sách infrastructure/service được duyệt hay tự do lựa chọn, một API framework duy nhất hay bất cứ thứ gì team muốn dùng.

Standardization mang lại lợi ích thực sự, tích lũy theo thời gian: engineer có thể chuyển giữa các team mà không tốn chi phí ramp-up, code review dễ dàng hơn đáng kể vì reviewer đã quen với idiom, incident response nhanh hơn vì engineer on-call nhận ra được hình dạng của bất kỳ service nào mình bị page vào, và tổ chức có thể đầu tư tập trung vào tooling (CI template dùng chung, thư viện dùng chung, dashboard dùng chung) thay vì mỗi team tự phát minh lại. Chi phí cũng có thật: một tiêu chuẩn đúng với nhu cầu tổ chức ba năm trước có thể trở thành một chiếc áo bó cho một team đang giải quyết một bài toán thực sự khác biệt, và over-standardization gửi tín hiệu (đúng đắn) đến senior engineer rằng judgment của họ không được tin tưởng.

Giải pháp thực tế mà hầu hết tổ chức lành mạnh áp dụng là một mô hình phân tầng:

TầngVí dụCách tiếp cận tiêu chuẩn hóa
Nền tảng / blast-radius caoVersion control, nền tảng CI/CD, core auth/identity, cloud provider chínhBắt buộc toàn tổ chức; sai lệch cần escalation và một lý do thuyết phục
Dùng chung nhưng có thể thương lượngNgôn ngữ/runtime, họ datastore chính, API frameworkMột danh sách được duyệt ngắn (2-3 lựa chọn), không phải một lựa chọn bắt buộc duy nhất; lựa chọn mới đi qua một review nặng như ADR
Cấp teamThư viện nội bộ, cấu trúc project, các lựa chọn tooling không quan trọng, tinh chỉnh rule linter trong một baseline chungTeam tự quyết, miễn không tạo gánh nặng vận hành lên team khác

Công việc của manager là giữ cho phân tầng này rõ ràng và được xem xét lại — các tiêu chuẩn chưa từng được viết ra sẽ bị thực thi không nhất quán và gây khó chịu, còn các tiêu chuẩn đã viết ra nhưng không bao giờ được xem xét lại sẽ trở thành vấn đề sơ đồ architecture hai năm tuổi y hệt, chỉ là cho tooling thay vì cho hệ thống.

Chiến lược API ở cấp tổ chức

Internal API xứng đáng nhận cùng một tư duy quản trị như tech stack, và các pattern kỹ thuật sâu cho chính API design — REST vs RPC vs GraphQL, phương án versioning, pagination, định dạng error — được nói ở ../backend/en/06-apis.md. Ở cấp tổ chức, các câu hỏi liên quan đến manager lại khác: ai được phép dựng lên một internal API surface mới, chính sách versioning và deprecation nào là bắt buộc với bất kỳ ai expose một API mà team khác phụ thuộc vào, và — đòn bẩy lớn nhất — khi nào tổ chức bắt buộc dùng một shared API platform/gateway thay vì để mỗi team tự xây và expose API riêng.

Một shared platform (một API gateway chung, một pipeline sinh SDK nội bộ dùng chung, một schema registry bắt buộc) có lời một khi bạn có đủ số internal consumer để các quy ước không nhất quán giữa các team bắt đầu tạo ra chi phí tích hợp thực sự — mỗi team tiêu thụ phải học một auth pattern khác, một pagination convention khác, một error shape khác. Nó thường không đáng để bắt buộc với chỉ vài service ít internal consumer; overhead phối hợp của một shared platform có thể vượt quá chi phí không nhất quán mà nó nhằm ngăn chặn. Một dấu hiệu hợp lý để xem xét lại điều này với vai trò manager: khi bạn nhận thấy hơn một team đang xây các client wrapper riêng chỉ để chuẩn hóa lại sự không nhất quán giữa hai internal API của chính bạn, đó thường là dấu hiệu cho thấy khoản đầu tư vào platform đã trở nên đáng giá.

Dù tổ chức chọn hướng nào, có một thứ gần như luôn đáng để bắt buộc bất kể quy mô: một chính sách deprecation được nêu rõ cho bất kỳ API nào mà team khác phụ thuộc vào. “Bạn có thể thay đổi cái này, nhưng phải version nó và cho consumer N tuần thông báo trước khi gỡ bỏ phiên bản cũ” rẻ để viết ra và đắt để bỏ qua một khi breaking change đầu tiên được ship mà không cảnh báo và làm sập production của một team downstream.

Best Practices

Hiệu chỉnh mức độ quy trình theo loại cửa, không theo chức danh của người quyết

Failure mode phổ biến nhất trong quản trị là áp dụng quy trình cấp ủy ban cho một quyết định thực ra là cửa hai chiều, hoặc để một cửa một chiều bị quyết định trong một cuộc trò chuyện hành lang vì người quyết đủ senior nên không ai nghĩ đến việc kiểm tra lại. Trước khi cuộc tranh luận về quyết định bắt đầu, hãy hỏi rõ ràng thành lời: “nếu điều này hóa ra sai, việc đảo ngược nó tốn kém đến mức nào?” Chỉ một câu hỏi đó, được hỏi nhất quán, làm được nhiều hơn bất kỳ framework nào trong việc định cỡ đúng quy trình quyết định của tổ chức bạn.

Gán vai trò DACI trước cuộc tranh luận, không phải sau

Nếu một cuộc họp quyết định bắt đầu mà không có Driver và Approver được thống nhất, cuộc thảo luận sẽ hoặc kéo dài quá điểm có thông tin mới hữu ích, hoặc bị quyết định không chính thức bởi người senior nhất trong phòng — đó là một quyết định đơn phương khoác áo một cuộc họp nhóm. Gọi tên vai trò trong lời mời họp hoặc phần đầu tài liệu, trước khi mọi người bắt đầu tranh luận lập trường, tránh được cả hai failure mode.

Viết ADR khi context còn tươi mới, không phải nhiều tháng sau

ADR viết ngay tại thời điểm nắm bắt được các phương án thực sự đã được xem xét và loại bỏ, kể cả những phương án có vẻ hợp lý vào lúc đó. ADR viết hồi cứu, sau khi có ai đó hỏi “tại sao chúng ta làm thế này,” có xu hướng hợp lý hóa lựa chọn hiện tại thay vì ghi lại trung thực các tradeoff đã được cân nhắc — giá trị của artifact phần lớn bị mất đi.

Biến architecture review thành một cổng chắn, không phải một thủ tục hình thức

Nếu review chỉ diễn ra sau khi hệ thống đã được xây xong, nó không còn là một điểm quyết định nữa mà trở thành một ô cần tick trước khi merge. Lên lịch review vào thời điểm mà việc đổi hướng vẫn còn rẻ — sau design doc, trước khi đầu tư triển khai đáng kể — và trao cho Approver quyền thực sự để trả lại thiết kế, không chỉ để đóng dấu chấp thuận.

Xem xét lại tiêu chuẩn theo nhịp độ định kỳ, đừng chỉ thiết lập rồi bỏ quên

Một danh sách tech stack được duyệt hoặc style guide chưa được xem xét lại trong hai năm sẽ tích lũy các ngoại lệ âm thầm mà mọi người lặng lẽ lách qua, điều này còn tệ hơn cả không có tiêu chuẩn nào — nó tạo ảo tưởng về sự nhất quán trong khi sự nhất quán đã xói mòn từ lâu. Đặt một lịch review định kỳ, nhẹ nhàng (dù chỉ một hoặc hai lần một năm) thay vì chỉ xem xét lại tiêu chuẩn một cách phản ứng khi có ai đó phàn nàn.

Giữ sơ đồ Context và Container của C4 luôn cập nhật, coi sơ đồ Component/Code là có thể bỏ đi

Cố gắng giữ mọi mức sơ đồ architecture đồng bộ hoàn hảo với code là một trận chiến không thể thắng và thường không đáng công sức ở mức Component và Code, vốn thay đổi quá thường xuyên. Tập trung công sức bảo trì vào mức Context và Container, vốn ít thay đổi và là những mức mà người mới và người cộng tác liên team thực sự cần.

Phân biệt “tôi không được tham khảo” với “tôi không có quyền veto”

Khi input của một Contributor không được làm theo, họ vẫn nên cảm thấy input của mình đã được lắng nghe và cân nhắc thực sự — kể cả khi Approver cuối cùng chọn hướng khác. Đánh mất sự phân biệt này (cảm thấy bị bỏ qua so với cảm thấy bị gạt bỏ) là điều biến một quy trình DACI lành mạnh thành nguồn gốc của sự bất mãn âm thầm, và cuối cùng, sự thờ ơ với các cuộc thảo luận architecture trong tương lai.

Để biết các quyết định quản trị architecture này liên hệ thế nào với việc quản lý liên tục technical debt, deprecation, và tradeoff rủi ro theo thời gian, xem ./10-technical-roadmap-debt-and-risk.md.

Tài liệu tham khảo

Part of the Engineering Manager Roadmap knowledge base.

Overview

Somewhere between senior engineer and engineering manager, a quiet shift happens: you stop being the person in the room who knows the most about any given subsystem, and you become the person responsible for making sure the room reaches a good decision anyway. That is an uncomfortable transition for a lot of managers who came up through hands-on technical work, because it can feel like a loss of authority. It isn’t — it’s a change in the kind of authority you hold. Your job is no longer to have the best answer; it’s to own the process by which the best available answer gets found, gets decided, and gets remembered.

This note is about that process: how technical decisions should get made in a team or org, how to tell the difference between decisions worth agonizing over and decisions that should be made in five minutes, how to record decisions so the reasoning survives past the people who made it, and how to set just enough standardization that engineers can move fluidly between systems without so much rigidity that nobody can choose the right tool for a real problem. None of this replaces deep technical architecture knowledge — for that, see ../backend/en/14-software-architecture.md for architectural styles, patterns, and tradeoffs, and ../backend/en/06-apis.md for API design specifics. This note sits one layer above: it’s about how a manager makes sure architecture work happens, happens with the right people, and stays honest to what the business actually needs.

Fundamentals

The manager’s role is process ownership, not technical omniscience

A common failure mode for new engineering managers is trying to remain the top technical authority on every decision in their team, either because they’re used to being that person or because they distrust that anyone else can fill the role. This doesn’t scale past a handful of engineers, and it actively damages a team’s growth — if the manager always has the final technical word, senior engineers stop developing judgment and stop bothering to build a strong case, because they know the decision isn’t really theirs to influence.

The healthier model: the manager’s job is to make sure that whoever is closest to the problem has the context, the mandate, and the pressure to make a well-reasoned decision — and that the decision actually gets made in a bounded amount of time. Concretely, that means:

You are the person who notices when a debate has stopped surfacing new information and started just repeating positions — and who has the standing to say “we have enough to decide, let’s decide.”

Decision paralysis and unilateral decisions are two failure modes of the same problem

Teams tend to fail toward one of two extremes:

Paralysis by consensus: every architectural choice gets relitigated in a meeting with too many stakeholders, no one owns the outcome, and the “decision” is really just an inconclusive conversation that gets picked up again next sprint. This burns weeks on things that don’t warrant it, and it teaches engineers that raising a strong opinion is a good way to stall a decision indefinitely.

Unilateral decisions that bypass expertise: a manager, tech lead, or senior IC makes a call alone, without consulting the people who will operate the consequences (on-call, the team maintaining the downstream service, security) — and the decision turns out wrong in a way that a five-minute conversation would have caught. This is often justified as “moving fast,” but it usually just moves the cost downstream and adds a trust tax the next time that person makes a call.

Both failure modes come from the same missing piece: nobody has explicitly defined who decides, who must be consulted, and who is just kept informed. Fixing that with a lightweight framework solves both problems at once — it bounds the discussion (so it doesn’t become paralysis) while still requiring the right expertise to be consulted (so it doesn’t become a unilateral bypass).

Key Concepts

DACI and RACI: naming who decides

The single highest-leverage habit a manager can install is naming, before a debate starts, who actually holds the pen. Two common frameworks:

DACI (Driver, Approver, Contributors, Informed) is oriented around a specific decision:

RoleMeaning
DriverOwns moving the decision forward: gathers input, frames the options, drives it to closure. Usually a tech lead or senior engineer, not necessarily the manager.
ApproverThe person (sometimes more than one, ideally not more than two) who has final sign-off. Everyone should know in advance who this is.
ContributorsPeople whose expertise or stake in the outcome means they must be consulted before the decision is made — the on-call engineer, the security reviewer, the team that owns the downstream dependency.
InformedPeople who need to know the outcome but have no input into it — they hear about the decision after it’s made, not during.

RACI (Responsible, Accountable, Consulted, Informed) is the more general-purpose version, often used for ongoing ownership rather than a single decision, and distinguishes “who does the work” (Responsible) from “who is ultimately answerable for it” (Accountable) — a distinction DACI collapses into Driver/Approver.

For a single architectural decision, DACI is usually the better fit because it’s decision-shaped rather than ownership-shaped. The manager’s job is rarely to be the Driver or the Approver — it’s to make sure those roles are assigned before the debate starts, publicly, so nobody is confused mid-discussion about whether their objection is a veto or an input.

Reversible (“two-way door”) vs. irreversible (“one-way door”) decisions

Not all technical decisions deserve the same process weight, and applying heavyweight process to a lightweight decision is itself a management failure — it trains engineers to see process as friction rather than as protection. The clearest lens for calibrating process weight is Jeff Bezos’s framing of one-way doors versus two-way doors:

The manager’s practical job is to help the team correctly classify which door they’re standing in front of before deciding how much process to apply — because misclassifying a one-way door as reversible is how teams end up trapped in an architecture nobody chose carefully, and misclassifying a two-way door as irreversible is how teams grind to a halt over decisions that don’t matter enough to deserve it.

Architecture Decision Records (ADRs)

An ADR is a short document that captures a single architecturally significant decision: what was decided, why, what alternatives were considered, and what the consequences (including downsides) are expected to be. The point of an ADR is not the “what” — that’s usually visible in the code or the system itself — it’s the why. Six months or two years later, when someone (possibly the same engineer, more likely someone new) is staring at an odd design choice and wondering “why on earth did we do it this way,” the ADR is the artifact that answers that question instead of forcing an archaeology dig through old Slack threads and closed PRs.

ADRs originated from Michael Nygard’s 2011 writeup and have since become a de facto industry standard, with a widely used template maintained in the adr GitHub organization.

Example ADR template

# ADR-0012: Use PostgreSQL as the primary datastore for the Orders service

## Status
Accepted (2026-03-14)
<!-- Possible values: Proposed, Accepted, Rejected, Deprecated, Superseded by ADR-00XX -->

## Context
The Orders service currently persists order state in an in-memory store backed by
scheduled snapshots to S3, inherited from the original prototype. As order volume
has grown past 50k orders/day, we're seeing data loss on pod restarts between
snapshots, and we have no ability to run ad-hoc queries for support and analytics
without replaying snapshots into a separate environment.

We need a persistence layer that:
- Guarantees durability per-write, not per-snapshot-interval
- Supports transactional consistency across an order and its line items
- Allows the support team to run read queries without engineering involvement
- Can be operated by our current team without hiring specialized DBAs

Candidates considered: PostgreSQL (RDS), DynamoDB, MongoDB Atlas, keeping the
current snapshot approach with a shorter interval.

## Decision
We will use PostgreSQL, hosted on Amazon RDS with Multi-AZ enabled, as the primary
datastore for the Orders service.

Rationale:
- Transactional guarantees map directly onto the order/line-item consistency
  requirement without extra application-level compensation logic.
- The team already has deep operational PostgreSQL experience from the Billing
  service; no new operational skill set is required.
- DynamoDB was rejected: the access patterns for support/analytics queries are
  too ad hoc and evolving to commit to DynamoDB's upfront key-design requirement.
- MongoDB Atlas was rejected: no meaningful advantage over PostgreSQL for this
  workload, and it would introduce a second document-store operational burden
  alongside our existing relational systems.
- Shortening the snapshot interval was rejected as a stopgap that doesn't solve
  the durability or queryability problems, only shrinks the failure window.

## Consequences
Positive:
- Per-write durability; no more snapshot-window data loss.
- Ad hoc read queries against a read replica, without engineering involvement.
- Reuses existing operational runbooks and on-call expertise from Billing.

Negative / risks:
- Introduces a new stateful dependency to operate and back up for this service.
- Schema migrations now require the same discipline (backward-compatible,
  reviewed) as Billing already applies — this team hasn't done that before and
  will need ramp-up.
- Horizontal write scaling is bounded by a single primary; if order volume grows
  10x we will likely need to revisit (sharding, or a queue-backed write path).

## Related
- Supersedes: none
- Related to: ADR-0009 (Billing service PostgreSQL setup, reused as an
  operational template)

When to write an ADR — and when it’s overkill

Write one when the decision is a genuine one-way door, or close to it: choice of a core data store, a change to a public API contract, adoption of a new architectural pattern (moving to event-driven, introducing a service mesh), a security or compliance-relevant design choice, or anything that a new engineer would reasonably ask “why is it built this way?” about a year from now.

Don’t write one for reversible, low-stakes choices — which linter config to use inside a single service, which of two equivalent utility libraries to import, a naming convention for a single module. Forcing an ADR onto every decision devalues the ones that matter; if engineers have to write ADRs for trivial choices, they’ll either stop writing them at all or reduce them to meaningless ceremony. A reasonable heuristic: if reversing the decision later would take more than a day of work or affect another team, it probably deserves an ADR.

Architecture documentation that stays true

Architecture diagrams rot faster than almost any other artifact in a codebase, because unlike code, nothing forces them to stay in sync with reality — a diagram doesn’t fail a build when the system it describes changes underneath it. A diagram from a design doc written two years ago, showing an “aspirational” architecture that was later simplified under deadline pressure, is worse than no diagram at all: it actively misleads whoever trusts it, and it undermines confidence in every other diagram in the same wiki.

A useful mental model for keeping architecture documentation both comprehensive and honest is the C4 model (Context, Containers, Components, Code), which structures documentation as four levels of zoom:

LevelShowsAudience
ContextThe system as one box, its users, and the other systems it talks toAnyone, including non-technical stakeholders
ContainersThe major deployable units inside the system (services, databases, web apps) and how they communicateEngineers across teams, new hires
ComponentsThe major structural building blocks inside a single containerEngineers working within that container
CodeClass/module-level detail (usually left to the code itself, or generated diagrams)Engineers actively editing that code

The manager’s practical stake in this isn’t drawing the diagrams — it’s making the Context and Container levels a living artifact instead of a one-time deliverable: reviewed at a light cadence (e.g., each quarter, or whenever an ADR changes the topology), owned by a named person or team rather than “whoever wrote it originally,” and treated as untrustworthy and flagged for update the moment someone notices it doesn’t match reality. A diagram with a visible “last verified” date is more useful than a beautifully rendered one nobody has looked at in a year.

The manager’s oversight role in architecture reviews

You are not expected to be the deepest architecture expert in the room — that’s covered in ../backend/en/14-software-architecture.md, and it’s the job of your senior engineers, staff engineers, and architects to bring that depth. What you own is making sure the review happens at the right time (before significant investment, not as a rubber stamp after the system is half-built), making sure it includes the right Contributors under a DACI-style assignment, and — this is the part that’s genuinely yours — making sure the discussion stays anchored to business goals rather than drifting into technical elegance for its own sake.

It is entirely possible for a room of strong engineers to converge on an architecturally beautiful solution that is wrong for the business: over-engineered for the actual scale the product will see in the next two years, optimized for a flexibility the product will never need, or expensive in build time relative to the revenue or risk it addresses. Part of your job in these reviews is to keep asking the grounding questions: what business problem does this solve, what’s the cost of doing the simpler thing instead, what happens if we’re wrong about the growth assumption driving this design? You don’t need to know the technology as deeply as your engineers to ask good questions like these — you need to know the business and the team’s constraints better than anyone else in the room, and that’s exactly the asymmetry that makes your presence in the review valuable rather than redundant.

Technical standards: the standardization/flexibility tradeoff

Every org eventually has to decide how much to standardize: one approved language versus letting teams pick their own, one linter and style guide enforced everywhere versus per-team style, an approved list of infrastructure/services versus free choice, one API framework versus whatever a team wants to reach for.

Standardization has real, compounding benefits: engineers can move between teams without a ramp-up cost, code review is meaningfully easier because reviewers already know the idioms, incident response is faster because on-call engineers recognize the shape of any service they’re paged into, and the org can invest centrally in tooling (shared CI templates, shared libraries, shared dashboards) instead of every team reinventing it. The cost is real too: a standard that was right for the org’s needs three years ago can become a straitjacket for a team solving a genuinely different problem, and over-standardization signals (correctly) to senior engineers that their judgment isn’t trusted.

The practical resolution most healthy orgs land on is a tiered model:

TierExampleStandardization approach
Foundational / high blast-radiusVersion control, CI/CD platform, core auth/identity, primary cloud providerMandated org-wide; deviation requires an escalation and a strong justification
Shared but negotiableLanguage/runtime, primary datastore family, API frameworkA short approved list (2-3 options), not one mandatory choice; new choices go through an ADR-weight review
Team-levelInternal libraries, project structure, non-critical tooling choices, linter rule tuning within a shared baselineTeam’s call, as long as it doesn’t create an operational burden on other teams

The manager’s job is to keep this tiering explicit and revisited — standards that were never written down get enforced inconsistently and resented, and standards that are written down but never revisited become the two-year-old architecture diagram problem all over again, just for tooling instead of systems.

API strategy at the org level

Internal APIs deserve the same governance thinking as the technology stack, and the deep technical patterns for API design itself — REST vs RPC vs GraphQL, versioning schemes, pagination, error formats — are covered in ../backend/en/06-apis.md. At the org level, the manager-relevant questions are different: who is allowed to stand up a new internal API surface, what versioning and deprecation policy is mandatory for anyone exposing an API another team depends on, and — the biggest lever of all — when does the org mandate a shared API platform/gateway versus letting each team build and expose its own.

A shared platform (a common API gateway, a shared internal SDK generation pipeline, a mandatory schema registry) pays off once you have enough internal consumers that inconsistent conventions between teams start generating real integration cost — every consuming team having to learn a different auth pattern, a different pagination convention, a different error shape. It’s usually not worth mandating for a handful of services with few internal consumers; the coordination overhead of a shared platform can exceed the inconsistency cost it’s meant to prevent. A reasonable trigger for revisiting this as a manager: when you notice more than one team building bespoke client wrappers just to normalize away inconsistency between two of your own internal APIs, that’s usually a sign the platform investment has become worth its cost.

Whatever the org lands on, one thing is close to universally worth mandating regardless of scale: a stated deprecation policy for any API another team depends on. “You can change this, but you must version it and give consumers N weeks of notice before removing the old version” is cheap to write down and expensive to have skipped once the first breaking change ships without warning and takes down a downstream team’s production system.

Best Practices

Calibrate process weight to the door, not to the title of the person deciding

The most common governance failure is applying committee-level process to a decision that’s actually a two-way door, or letting a one-way door get decided in a hallway conversation because the person deciding is senior enough that nobody thought to check. Before a decision debate starts, explicitly ask out loud: “if this turns out wrong, how expensive is it to undo?” That one question, asked consistently, does more to right-size your org’s decision process than any framework.

Assign DACI roles before the debate, not after

If a decision meeting starts without an agreed Driver and Approver, the discussion will either drag on past the point of useful new information, or get informally decided by whoever’s most senior in the room — which is a unilateral decision wearing a group-meeting costume. Naming roles in the meeting invite or the doc header, before people start arguing positions, avoids both failure modes.

Write the ADR while the context is still fresh, not months later

ADRs written in the moment capture the alternatives that were actually considered and rejected, including the ones that seemed reasonable at the time. ADRs written retroactively, after someone asks “why did we do this,” tend to rationalize the existing choice rather than honestly record the tradeoffs that were weighed — the value of the artifact is largely lost.

Make architecture review a gate, not a formality

If review happens only after a system is already built, it stops being a decision point and becomes a box to check before merging. Schedule the review at the point where changing course is still cheap — after a design doc, before significant implementation investment — and give the Approver real authority to send the design back, not just to bless it.

Revisit standards on a cadence, don’t just set and forget

An approved tech stack list or style guide that hasn’t been revisited in two years accumulates silent exceptions that everyone quietly works around, which is worse than having no standard — it creates the illusion of consistency while consistency has already eroded. Put a recurring, lightweight review on the calendar (even just once or twice a year) rather than only revisiting standards reactively when someone complains.

Keep the C4 Context and Container diagrams current, treat Component/Code diagrams as disposable

Trying to keep every level of architecture diagram perfectly in sync with the code is a losing battle and usually not worth the effort at the Component and Code levels, which change too often. Concentrate the maintenance effort on the Context and Container levels, which change rarely and are the levels new hires and cross-team collaborators actually need.

Distinguish “I wasn’t consulted” from “I don’t get a veto”

When a Contributor’s input wasn’t followed, they should still feel that their input was heard and genuinely weighed — even if the Approver ultimately went a different direction. Losing this distinction (feeling ignored vs. feeling overruled) is what turns a healthy DACI process into a source of quiet resentment and, eventually, disengagement from future architecture discussions.

For how these architectural governance decisions connect to the ongoing management of technical debt, deprecations, and risk tradeoffs over time, see ./10-technical-roadmap-debt-and-risk.md.

References