← Quản lý kỹ thuật← Engineering Manager
Quản lý kỹ thuậtEngineering Manager19 Th7, 2026Jul 19, 202627 phút đọc21 min read

Ngân sách, Chi phí & Capacity PlanningBudgeting, Cost & Capacity Planning

Thuộc bộ kiến thức Engineering Manager Roadmap.

Tổng quan

Nằm đâu đó giữa bản strategy deck và sprint board là một lớp công việc kém hào nhoáng nhưng không thể tránh khỏi của engineering management: budget. Mọi cam kết roadmap đã chốt trong Chiến lược, Business Case & Ưu tiên hóa cuối cùng đều phải được dịch thành một con số mà đối tác finance có thể duyệt — headcount, chi phí cloud, tooling, contractor — và mỗi con số đó đều phải được bảo vệ, theo dõi, và đối chiếu với thực tế đã xảy ra. Những manager coi budgeting là bài tập giấy tờ hàng năm do finance áp xuống thường xuyên đánh mất đòn bẩy: họ chấp nhận bất kỳ con số nào được gán cho mình thay vì chủ động định hình nó, và giữa năm họ bị bất ngờ bởi cost overrun mà không hề có visibility trước đó vì không ai theo dõi đồng hồ.

Budgeting, cost management, và capacity planning thực chất là ba mặt của cùng một bài toán phân bổ nguồn lực, nhìn ở các khung thời gian khác nhau. Budgeting là bài tập hàng năm (hoặc hàng quý) quyết định, bằng tiền, tổ chức sẵn sàng chi bao nhiêu và chi vào đâu. Cost optimization là kỷ luật liên tục đảm bảo số tiền đó được chi hiệu quả sau khi đã phân bổ. Capacity planning là bài tập dịch ngược tiền và headcount thành số lượng công việc mà team thực sự có thể giao — và nó áp dụng cho con người (một team với quy mô nhất định có thể hấp thụ bao nhiêu roadmap work một cách thực tế) không kém gì cho infrastructure (một kiến trúc nhất định có thể phục vụ bao nhiêu tải). Một manager thành thạo cả ba mảng này có thể bước vào cuộc họp planning với một đề xuất có căn cứ vững chắc, bắt được một cloud bill đang tăng phi mã trước khi nó trở thành nỗi xấu hổ ở cấp board, và phản biện một cách đáng tin khi roadmap được lên kích cỡ dựa trên một team không có đủ capacity để giao.

Bài viết này giả định người đọc đã có các con số liên quan đến budget từ hai chủ đề trước: hiring plan từ Tuyển dụng & Thiết kế tổ chức, và cam kết roadmap từ Chiến lược, Business Case & Ưu tiên hóa. Bài viết bao quát cơ chế biến các con số đó thành một budget, bảo vệ budget đó, kiểm soát chi phí sau khi đã chi, và lập capacity planning — cho cả infrastructure lẫn con người — để cam kết và thực tế không âm thầm trôi xa nhau. Về cơ chế kỹ thuật của cloud cost optimization và infrastructure scaling, bài viết này trỏ tới Cloud, DevOps, và Network Engineer thay vì lặp lại — việc của manager ở đây là làm cho trade-off hiện rõ với business, không phải tự tay thực hiện bài tập rightsizing.

Kiến thức nền tảng

Một engineering budget thực sự bao gồm những gì

Manager mới thường ngạc nhiên trước hình dạng của một engineering budget: nó chủ yếu là con người, không phải infrastructure. Với hầu hết các tổ chức product engineering, chi phí compensation và headcount chiếm khoảng 60% đến 80% tổng budget, phần còn lại chia cho cloud infrastructure, tooling, và mọi thứ khác. Tỷ lệ này quan trọng vì nó định hướng lại nơi nên đầu tư nỗ lực cost-optimization — ám ảnh với một cloud bill chỉ chiếm 10% budget trong khi bỏ qua một cơ cấu tổ chức cồng kềnh đang đẩy dòng headcount chiếm 70% lên cao là đang tối ưu sai đòn bẩy.

Danh mụcTỷ trọng điển hìnhBao gồm
Headcount & compensation60–80%Lương, thưởng, equity, benefit, thuế người sử dụng lao động, cho cả nhân sự hiện tại lẫn kế hoạch tuyển mới
Cloud infrastructure8–20%Compute, storage, network egress, managed services, chi phí data platform
Tooling & SaaS licenses3–10%CI/CD, observability, project management, design, security scanning, công cụ giao tiếp
Contractor & vendor spendBiến động, 0–15%Staff augmentation, consultant chuyên biệt, outsource QA hoặc support
Training & development1–3%Conference, khóa học, chứng chỉ, chương trình học tập nội bộ

Tỷ lệ chính xác thay đổi theo giai đoạn công ty và mô hình kinh doanh — một business nặng về infrastructure (video streaming, ML training) có thể thấy chi phí cloud leo lên 30–40% engineering budget, trong khi một shop nặng về consulting có thể mang dòng contractor lớn hơn nhiều. Bảng này là điểm khởi đầu để sanity-check một budget, không phải quy tắc ép buộc.

Chu kỳ planning hàng năm

Hầu hết công ty vận hành engineering budgeting theo chu kỳ hàng năm với một điểm check-in hàng quý nhẹ hơn, và chu kỳ này có một nhịp điệu đáng để nội tâm hóa thay vì phản ứng với nó như thể mỗi năm là mới.

Giai đoạnThời điểm điển hìnhĐiều gì xảy ra
Strategic inputQ3 (cho budget theo năm dương lịch)Finance và leadership đặt mục tiêu tăng trưởng hoặc cắt giảm chi phí ở tầm cao; engineering leadership nhận một envelope sơ bộ để lên kế hoạch
Bottom-up buildQ3–Q4Manager xây dựng headcount và spend request gắn với cam kết roadmap
Negotiation & reconciliationQ4Bottom-up ask được đối chiếu với top-down envelope; trade-off được làm rõ
Approval & allocationCuối Q4 / đầu Q1Con số cuối cùng được chốt; budget được phân bổ theo cost center
Execution & monitoringSuốt cả nămActuals được theo dõi so với budget; variance được flag
Quarterly reforecastMỗi quýBudget được xem lại dựa trên tốc độ hiring thực tế, attrition, và chi tiêu; điều chỉnh được thực hiện mà không cần chờ chu kỳ hàng năm tiếp theo

Một manager chỉ xuất hiện ở giai đoạn “bottom-up build” đã đánh mất ảnh hưởng — top-line envelope được đặt ở giai đoạn strategic input quyết định còn bao nhiêu dư địa để đàm phán sau này. Chủ động trao đổi với finance và skip-level leadership trước khi bản ask chính thức đến hạn, để hiểu những giả định về tăng trưởng hoặc cắt giảm chi phí nào đã được đưa vào envelope, là điều phân biệt một đề xuất có căn cứ với một wish list.

Khái niệm chính

Xây dựng một budget request có căn cứ vững chắc

Hình thức yếu nhất của một budget ask là “chúng ta cần thêm người” hoặc “chúng ta cần budget cloud lớn hơn,” đưa ra mà không có chuỗi lập luận kết nối ask đó với một outcome mà business quan tâm. Một request có căn cứ vững chắc thay vào đó bắt đầu từ roadmap work đã cam kết hoặc ưu tiên, dịch công việc đó thành ước tính về capacity engineering cần thiết, so sánh với capacity hiện tại, và biểu đạt khoảng cách đó thành một request cụ thể, có định giá.

Một cấu trúc hữu ích cho bản request:

  1. Neo vào outcome đã cam kết. Nêu rõ cam kết roadmap hoặc canh bạc chiến lược mà budget này nhằm tài trợ — lý tưởng là những cam kết đã thống nhất trong business case từ Chiến lược, Business Case & Ưu tiên hóa. Một request đứng riêng, không gắn với roadmap đã đồng thuận, sẽ mời gọi câu hỏi “tại sao cần cái này nếu nó không gắn với thứ chúng ta đã đồng ý xây?”
  2. Trình bày phép toán capacity. Dịch công việc đã cam kết thành person-week hoặc person-quarter công sức, và so sánh với capacity hiện tại của team (xem phần capacity planning bên dưới để làm điều này một cách trung thực). Khoảng cách giữa capacity cần thiết và capacity hiện có chính là headcount ask.
  3. Định giá. Biến khoảng cách headcount và infrastructure thành tiền: chi phí fully-loaded mỗi lần hire (lương, benefit, overhead — thường 1.25–1.4 lần lương cơ bản), chi phí infra tăng thêm ở mức traffic hoặc data volume mà roadmap ngụ ý, chi phí tooling mỗi seat mới.
  4. Chỉ ra phương án thay thế. Điều gì xảy ra nếu budget không được duyệt — cam kết nào bị trễ, trễ bao lâu, và chi phí business của việc trễ đó là gì? Điều này định khung lại ask từ “engineering muốn nhiều hơn” thành “business đã đồng ý X; đây là chi phí để giao X, và đây là điều xảy ra nếu không được tài trợ.”
  5. Chỉ ra độ nhạy. Một khoảng (ví dụ “với N engineer chúng ta đạt mốc Q3; với N-2 chúng ta đạt Q4 và bỏ tính năng phụ”) đáng tin và hữu ích hơn cho đối tác finance so với một con số cố định duy nhất, vì nó cho họ một đòn bẩy để kéo nếu top-line envelope không đủ giãn ra cho toàn bộ ask.

Dự đoán trước sự phản đối

Đối tác finance và executive leadership thường hỏi một bộ câu hỏi có thể đoán trước khi review một engineering budget request, và chuẩn bị câu trả lời trước sẽ biến một cuộc vật lộn phòng thủ thành một cuộc trao đổi tự tin.

Câu hỏi có thể gặpCâu trả lời tốt trông như thế nào
”Tại sao team hiện tại không thể tự hấp thụ việc này?”Một bảng phân tích capacity (xem bên dưới) cho thấy team hiện tại đã ở mức tải bền vững hoặc cao hơn, có tính đến on-call, meeting, và onboarding
”Chúng ta đã nhận được gì từ việc tăng headcount năm ngoái?”Một bản retrospective ngắn gắn hiring năm ngoái với outcome đã giao, không chỉ “chúng ta đã tuyển 5 người"
"Có thể làm việc này bằng contractor thay vì full-time hire không?”Một trade-off trung thực (xem bên dưới) thay vì một câu “không” theo phản xạ
”Xu hướng chi phí cloud là gì, và có đang kiểm soát được không?”Một đường xu hướng, không phải một snapshot đơn lẻ, cộng với bằng chứng về một cost-optimization practice đang hoạt động
”Điều gì xảy ra nếu chúng ta cắt cái này 20%?”Một câu trả lời cụ thể, đã suy nghĩ trước (cam kết nào bị trễ) thay vì ứng biến ngay tại chỗ
”Đây là chi phí một lần hay lặp lại?”Phân biệt rõ ràng giữa chi phí một lần (ví dụ dự án migration, contractor đợt) và chi phí lặp lại (ví dụ một dòng headcount lâu dài)

Điểm chung xuyên suốt các câu hỏi này là gánh nặng chứng minh nằm ở người đề xuất, và những câu trả lời mơ hồ hoặc ứng biến ngay tại chỗ sẽ xói mòn niềm tin cho chu kỳ budget tiếp theo dù chu kỳ này có được duyệt hay không.

Cost optimization lớn hơn cloud bill

Manager lần đầu dễ bị cám dỗ đánh đồng “tối ưu chi phí engineering” với “tối ưu chi phí cloud,” vì cloud bill dễ thấy, có hạng mục rõ ràng, và tấn công vào nó cho cảm giác thỏa mãn. Nhưng vì headcount chiếm 60–80% budget, đòn bẩy cost-optimization lớn nhất gần như luôn nằm ở tổ chức, không phải infrastructure.

Headcount efficiency nói về việc cơ cấu tổ chức và quy trình có cho phép những người đã có trên bảng lương tạo ra giá trị tối đa họ có khả năng hay không — một team ngập trong overhead phối hợp, ownership không rõ ràng, hoặc công việc trùng lặp là một vấn đề chi phí dù không có dòng nào ghi rõ điều đó. Đây là phần mở rộng trực tiếp của công việc thiết kế tổ chức trong Tuyển dụng & Thiết kế tổ chức: một team được định hình kém là đắt đỏ ngay cả với một con số headcount cố định, vì nó lãng phí capacity đã được trả tiền.

Tool và SaaS consolidation là một dòng thường bị đánh giá thấp. Khi tổ chức lớn lên, ba team độc lập mua ba nền tảng observability khác nhau, hoặc công ty trả tiền cho năm công cụ project-management vì mỗi team tự chọn một công cụ trong giai đoạn thiếu giám sát, là chuyện thường thấy. Một cuộc audit SaaS hàng năm — liệt kê mọi công cụ, chi phí, chủ sở hữu, và mức trùng lặp với công cụ khác — thường xuyên tìm ra khoản tiết kiệm 10–30% dòng tooling mà không ảnh hưởng đến delivery, chỉ đơn giản bằng cách hợp nhất các subscription trùng lặp và đàm phán giá enterprise cho những công cụ còn lại thay vì trả giá bán lẻ theo từng team.

Chi phí cloud infrastructure là có thật và thường là dòng dễ thấy nhất, nhưng đây cũng là dòng đã có playbook kỹ thuật trưởng thành nhất ở nơi khác trong knowledge base này — rightsizing instance, reserved instance và savings plan, autoscaling, spot/preemptible capacity, phân tầng storage, v.v. Thay vì lặp lại nội dung đó, xem CloudDevOps để nắm cơ chế; việc của manager là đảm bảo practice đó tồn tại, có người sở hữu, và được review theo nhịp độ — không phải tự tay thực hiện rightsizing.

Trade-off giữa cắt giảm chi phí và velocity/morale

Cost optimization có một ngưỡng trần mà qua đó việc cắt giảm thêm không còn miễn phí mà bắt đầu đánh đổi trực tiếp với velocity hoặc morale, và việc của manager là biết ngưỡng đó ở đâu trước khi vượt qua nó dưới áp lực. Cắt một công cụ SaaS thực sự dư thừa là miễn phí. Cắt tooling observability giúp engineer debug production incident nhanh thì không — nó trông như một khoản tiết kiệm trên spreadsheet và sau đó xuất hiện dưới dạng thời gian xử lý incident lâu hơn, on-call đau đầu hơn, và attrition. Tương tự, cắt giảm headcount được đóng khung thuần túy là “efficiency” mà không có mức giảm tương ứng về phạm vi roadmap chỉ đơn giản là đẩy cùng lượng công việc lên ít người hơn, đọc như một khoản tiết kiệm cho một quý và trở thành vấn đề burnout và attrition cho ba quý tiếp theo. Kỷ luật ở đây là làm cho trade-off hiện rõ thay vì giả vờ nó không tồn tại: mọi khoản cắt chi phí nên đi kèm câu trả lời trung thực cho “điều này khiến chúng ta mất gì về velocity, reliability, hoặc retention,” không chỉ “nó tiết kiệm được bao nhiêu.”

Resource allocation và forecasting, nhìn lại từ góc độ budget

Capacity và resource forecasting đã được bàn ở góc độ delivery planning ở nơi khác trong knowledge base này; góc nhìn budget thêm vào lớp dịch sang tiền trên nền hiring plan.

Lập kế hoạch attrition. Một headcount budget giả định attrition bằng 0 là sai ngay từ đầu — hầu hết tổ chức engineering thấy voluntary attrition hàng năm ở mức 8–15% ngay cả với team khỏe mạnh, và budget nên lập kế hoạch backfill hiring cho những người dự kiến rời đi tách biệt với hiring tăng trưởng net-new. Gộp chung hai loại này là nguồn gốc phổ biến của bất ngờ giữa năm: một team “có budget cho 3 hire mới” nhưng mất 2 người do attrition thực chất chỉ tăng trưởng 1 người, trong khi vẫn phải gánh chi phí recruiting và onboarding của 5 quy trình tuyển dụng.

Contractor vs. full-time hire. Trade-off này lặp lại mỗi chu kỳ budget và xứng đáng có một câu trả lời chủ động thay vì mặc định. Contractor thường tốn hơn mỗi giờ làm việc nhưng không mang nghĩa vụ dài hạn, ramp nhanh hơn cho công việc đã scope rõ, và có thể giảm quy mô mà không cần layoff; full-time hire tốn ít hơn mỗi giờ trong khung thời gian nhiều năm, xây dựng institutional knowledge bền vững, và là lựa chọn đúng cho bất cứ điều gì cốt lõi với sự khác biệt hóa dài hạn của sản phẩm. Một heuristic hợp lý: dùng contractor cho công việc có giới hạn thời gian, được đặc tả rõ, có ngày kết thúc rõ ràng (một dự án migration, một đợt tuân thủ compliance, một đợt tăng capacity QA) và full-time hire cho bất cứ điều gì sẽ cần ownership liên tục sau khi dự án hiện tại kết thúc.

Khía cạnhContractorFull-time hire
Chi phí mỗi giờCao hơnThấp hơn (khấu hao theo thời gian gắn bó)
Thời gian ramp-upNhanh (cho công việc đã scope rõ)Chậm hơn, nhưng xây context lâu dài
Linh hoạt giảm quy môCao, chi phí severance/pháp lý thấpThấp, layoff mang chi phí và ảnh hưởng morale
Institutional knowledgeRa đi cùng hợp đồngTích lũy và compound
Phù hợp nhấtCông việc có giới hạn thời gian, scope rõOwnership sản phẩm cốt lõi, liên tục

Dịch hiring plan thành một dòng budget nghĩa là đi xa hơn “5 engineer mới” để mô hình hóa khi nào 5 người đó thực sự lên bảng lương — thời gian recruiting và onboarding (thường 2–4 tháng từ lúc duyệt req đến ngày bắt đầu) có nghĩa là một headcount được duyệt ở Q1 hiếm khi tạo ra đầy đủ một quý chi phí hoặc output ngay trong Q1. Budget giả định hiring tức thời sẽ thổi phồng cả chi phí gần hạn lẫn capacity giao hàng gần hạn.

Scaling infrastructure: một trade-off capacity-vs-cost, không phải một lựa chọn công nghệ

Khi một hệ thống tiệm cận giới hạn capacity, cám dỗ là đóng khung cuộc trò chuyện như một vấn đề kỹ thuật — database nào, caching layer nào, autoscaling policy nào. Đó là những quyết định thật, nhưng chúng thuộc về engineer và architect thực hiện công việc, với chiều sâu kỹ thuật được bàn trong Cloud, DevOps, và Network Engineer (load balancing, network capacity, CDN strategy, v.v.). Việc riêng biệt của manager nằm ở một tầng cao hơn: làm cho hàm ý chi phí của một quyết định capacity hiện rõ với business trước khi nó được đưa ra, không phải sau đó.

Cụ thể, điều này nghĩa là dịch “chúng ta cần scale tầng database” thành một tuyên bố mà business hiểu được: “xử lý traffic dự phóng Q4 ở tốc độ tăng trưởng hiện tại tốn thêm ước tính $X/tháng infrastructure, hoặc thay vào đó chúng ta có thể trì hoãn $X bằng cách chấp nhận hiệu năng suy giảm vào giờ cao điểm — business ưu tiên trade-off nào?” Không có khung này, quyết định capacity bị đưa ra đơn phương bởi engineering (rủi ro một bill mà không ai duyệt) hoặc bị đưa ra đơn phương bởi finance (rủi ro một outage mà không ai được cảnh báo trước). Đòn bẩy của manager nằm ở việc phơi bày trade-off đủ sớm để nó trở thành một quyết định thay vì một bất ngờ.

Capacity planning cho con người, không chỉ infrastructure

Capacity planning cho infrastructure có một bộ từ vựng trưởng thành — headroom, peak load, autoscaling threshold — và đáng để mượn cùng kỷ luật đó cho việc lập capacity planning cho team, vì sai lầm kinh điển là coi một sprint hay một quý như thể 100% thời gian đó dành cho “pure feature time.” Không bao giờ như vậy.

Một mô hình capacity trung thực hơn trừ đi tải đã biết, lặp lại khỏi thời gian lịch trước khi đến với feature work:

Nguồn tiêu tốn thời gianTỷ trọng capacity điển hìnhGhi chú
On-call & interrupt work5–20%Biến động mạnh theo độ trưởng thành của service và quy mô on-call rotation; tăng vọt khi có incident
Meeting & ceremony10–20%Standup, planning, review, 1:1, cross-team sync
Onboarding drag (cho team có hire mới)20–50% capacity của một hire mới trong 1–2 quý đầu, cộng với thuế lên người mentor/reviewer hỗ trợ họThời gian ramp cũng không miễn phí với phần còn lại của team
PTO, ngày lễ, nghỉ ốm8–12% tính theo nămBiến động theo khu vực và chính sách công ty
Tech debt, tooling, và bảo trì ngoài roadmap10–20%Thường vô hình trong planning vì không phải một roadmap item được đặt tên
Còn lại cho roadmap work đã cam kếtThường chỉ 40–60% capacity danh nghĩaĐây là con số nên dùng để định kích cỡ cam kết, không phải con số đếm headcount

Về mặt thực tiễn, điều này nghĩa là một team 6 engineer không có “6 engineer capacity” cho roadmap planning — sau khi trừ on-call, meeting, PTO, thời gian ramp cho bất kỳ ai được hire trong hai quý gần nhất, và một khoản dự phòng tech-debt hợp lý, con số trung thực có thể gần với 3.5–4 engineer capacity cam kết hơn. Roadmap và budget request được định kích cỡ dựa trên con số headcount danh nghĩa thay vì con số đã chiết khấu này là nguồn gốc phổ biến nhất của cuộc trò chuyện “tại sao cái này bị trễ” về sau — team không thiếu người so với kế hoạch, kế hoạch đã bị quá cỡ so với capacity thực của team.

Điều này cũng có hệ quả budget trực tiếp: nếu một roadmap thực sự cần 6 engineer capacity cam kết và mô hình capacity trung thực nói rằng team giao được 4, khoảng cách đó cần một headcount request (được tài trợ, với lead time từ phần trên) hoặc một khoản cắt scope — không phải hy vọng rằng mọi người sẽ đơn giản làm việc chăm chỉ hơn.

Best Practices

Coi budget là một dự báo sống, không phải một tài liệu hàng năm

Con số hàng năm là một ước tính khởi đầu, không phải một hợp đồng với thực tế. Review actuals so với budget hàng tháng hoặc hàng quý, và flag variance sớm, biến một khủng hoảng cuối năm tiềm ẩn thành một điều chỉnh giữa chặng đường có thể quản lý được. Chờ đến chu kỳ hàng năm mới nhận ra một cloud bill đã trôi lên 40% so với kế hoạch nghĩa là tám hoặc chín tháng chi tiêu vượt mức không được quản lý trước khi ai đó nhìn lại con số.

Sở hữu một cost dashboard, đừng chờ finance phát hiện bất ngờ

Một manager có thể mở ra chi phí cloud hiện tại, chi phí tooling, và tốc độ đốt headcount bất cứ lúc nào — mà không cần chờ báo cáo từ finance — có vị thế đàm phán và lập kế hoạch mạnh hơn về căn bản so với người biết về việc chi vượt qua kênh gián tiếp. Điều này không cần phức tạp; một spreadsheet cập nhật hàng tháng đơn giản hoặc một dashboard nhẹ được nạp từ billing export là đủ để bắt hầu hết bất ngờ sớm. FinOps practice (xem tài liệu tham khảo) chính thức hóa điều này thành một kỷ luật chia sẻ, liên tục giữa engineering, finance, và product thay vì một bài tập một-năm-một-lần do finance sở hữu.

Tách biệt rõ ràng chi phí một lần và chi phí lặp lại

Một dự án migration, một đợt contractor, hoặc một khoản mua tooling một lần không bao giờ nên bị âm thầm gộp vào “run rate” — làm vậy sẽ thổi phồng baseline năm sau và khiến so sánh tương lai vô nghĩa. Gắn nhãn chi tiêu là một-lần vs. lặp-lại ngay tại thời điểm phát sinh, không phải hồi tố khi ai đó hỏi tại sao budget trông khác kỳ vọng.

Định nghĩa KPI hiệu quả chi phí một cách chủ động, và theo dõi những gì nó tối ưu bỏ đi

Theo dõi các chỉ số hiệu quả chi phí là hữu ích, nhưng mỗi chỉ số mang một rủi ro bóp méo cần được nêu rõ cùng với chỉ số, không phải phát hiện ra sau khi sự đã rồi.

KPIĐo lường gìRủi ro bóp méo nếu bị tối ưu quá mức
Cost per feature / cost per story pointHiệu quả delivery so với chi tiêuThưởng cho việc ship tính năng nhỏ, dễ thay vì tính năng khó hơn, giá trị cao hơn; phạt đầu tư vào platform work không gắn tính năng ngay lập tức
Infra cost dưới dạng % doanh thuLiệu chi phí infrastructure có scale dưới tuyến tính so với business hay khôngCó thể che giấu việc thiếu đầu tư vào reliability hoặc headroom scalability nếu doanh thu tình cờ tăng nhanh hơn xu hướng chi phí
Cost per customer / cost per transactionKinh tế đơn vị của sản phẩmCó thể bị game bằng cách trì hoãn đầu tư infra cần thiết (redundancy, DR capacity) không sinh lời cho đến khi có incident xảy ra
Chi phí engineer trên mỗi đơn vị outputHiệu quả headcountGây áp lực khiến team báo cáo thiếu hoặc giấu công việc “vô hình” (on-call, mentoring, tech debt) không tính vào metric output hiển thị

Nguyên tắc chung: bất kỳ chỉ số hiệu quả chi phí nào, nếu đẩy đủ xa, đều có thể được cải thiện bằng cách âm thầm cắt đầu tư reliability, trì hoãn bảo trì, hoặc quá tải con người — không cái nào trong số đó xuất hiện trong bản thân chỉ số cho đến khi chi phí bị trì hoãn ập đến, thường dưới dạng một incident, một outage, hoặc attrition. Ghép mỗi cost KPI với một chỉ số chất lượng hoặc sức khỏe tương ứng (tỷ lệ incident, tải on-call, retention) là guardrail chuẩn: báo cáo chúng cùng nhau, và coi một cải thiện chi phí đi kèm suy giảm đồng thời ở chỉ số ghép cặp là một cờ đỏ, không phải một chiến thắng.

Xây budget từ dưới lên dựa trên mô hình capacity, không phải từ trên xuống dựa trên năm ngoái cộng thêm tăng trưởng

“Lấy budget năm ngoái cộng thêm 15%” là một lối tắt phổ biến, và nó thường thất bại một cách hệ thống trong việc bắt được những tình huống mà khoảng cách capacity thực của team lớn hơn hoặc nhỏ hơn một phần trăm cố định gợi ý. Xây dựng từ mô hình capacity trung thực trong phần Khái niệm chính — công việc roadmap nào đã cam kết, cần capacity gì, team thực sự có capacity gì sau khi trừ on-call, meeting, onboarding, và PTO — cho ra một con số có căn cứ trong cuộc họp được mô tả ở bảng phản đối bên trên, vì nó có thể truy nguyên tới các cam kết cụ thể thay vì một phép ngoại suy.

Một ví dụ minh họa

Bảng dưới đây phác thảo một budget hàng năm sơ bộ cho một team product engineering quy mô vừa khoảng 25 người (chia hợp lý cho hai hoặc ba squad cộng thêm một chức năng platform), chỉ để minh họa tỷ trọng — con số thực tế biến động rất lớn theo giai đoạn công ty, khu vực địa lý, và ngành.

Danh mụcChi tiêu hàng năm sơ bộ% tổngGhi chú
Headcount (25 FTE, fully loaded)$4.5M~72%Lương, benefit, thuế, equity theo hệ số fully-loaded trung bình
Cloud infrastructure$900K~14%Compute, storage, data platform, network egress
Tooling & SaaS licenses$500K~8%CI/CD, observability, project management, security scanning, giao tiếp
Contractor & vendor spend$250K~4%Một dự án migration có giới hạn thời gian cộng với tăng cường QA theo mùa
Training & development$100K~2%Conference, khóa học, học bổng học tập nội bộ
Tổng~$6.25M100%

Đọc bảng này theo cách một đối tác finance sẽ đọc: headcount áp đảo mọi thứ khác, nghĩa là cuộc trò chuyện về chi phí có đòn bẩy cao nhất gần như luôn xoay quanh cơ cấu tổ chức và tốc độ hiring, không phải dòng cloud; dòng cloud và tooling cộng lại vẫn xứng đáng một cost-optimization practice chủ động (cắt 15% ở đó là tiền thật — khoảng $200K trong ví dụ này — dù nó chỉ là một phần nhỏ của dòng headcount); và dòng contractor nên gắn với một dự án cụ thể, được đặt tên, có giới hạn thời gian thay vì trở thành một cố định vĩnh viễn, nếu không nó thực chất là một full-time hire được trả theo giá premium của contractor.

Tài liệu tham khảo

Part of the Engineering Manager Roadmap knowledge base.

Overview

Somewhere between the strategy deck and the sprint board sits a less glamorous but unavoidable layer of engineering management: the budget. Every roadmap commitment made in Strategy, Business Case & Prioritization eventually has to be translated into a number a finance partner can approve — headcount, cloud spend, tooling, contractors — and every one of those numbers has to be defended, tracked, and reconciled against what actually happened. Managers who treat budgeting as an annual paperwork exercise handed down from finance consistently lose leverage: they end up accepting whatever number is assigned instead of shaping it, and they get surprised mid-year by cost overruns they had no visibility into because nobody was watching the meter.

Budgeting, cost management, and capacity planning are really three faces of the same underlying resource-allocation problem, viewed at different time horizons. Budgeting is the annual (or quarterly) exercise of deciding, in dollars, how much the organization is willing to spend and on what. Cost optimization is the ongoing discipline of making sure that money is spent efficiently once it’s allocated. Capacity planning is the exercise of translating dollars and headcount back into a number of things the team can actually deliver — and it applies as much to people (how much roadmap work a team of a given size can realistically absorb) as it does to infrastructure (how much load a given architecture can serve). A manager who is fluent in all three can walk into a planning conversation with a defensible ask, catch a runaway cloud bill before it becomes a board-level embarrassment, and push back credibly when a roadmap is sized against a team that doesn’t have the capacity to deliver it.

This note assumes the reader already has budget-adjacent numbers from two upstream topics: headcount plans from Hiring & Organization Design, and roadmap commitments from Strategy, Business Case & Prioritization. It covers the mechanics of turning those into a budget, defending that budget, keeping costs under control once spent, and planning capacity — for both infrastructure and people — so that commitments and reality don’t quietly diverge. For the technical mechanics of cloud cost optimization and infrastructure scaling, this note points to Cloud, DevOps, and Network Engineer rather than duplicating them — the manager’s job here is to make trade-offs visible to the business, not to perform the rightsizing exercise personally.

Fundamentals

What’s actually in an engineering budget

New managers are often surprised by the shape of an engineering budget: it is overwhelmingly people, not infrastructure. For most product engineering organizations, compensation and headcount-related costs make up somewhere between 60% and 80% of the total budget, with cloud infrastructure, tooling, and everything else splitting the remainder. This ratio matters because it reorients where cost-optimization effort should go — obsessing over a cloud bill that is 10% of the budget while ignoring a bloated org structure that drives the 70% headcount line is optimizing the wrong lever.

CategoryTypical share of budgetWhat it includes
Headcount & compensation60–80%Salaries, bonuses, equity, benefits, employer taxes, for both current staff and planned new hires
Cloud infrastructure8–20%Compute, storage, network egress, managed services, data platform costs
Tooling & SaaS licenses3–10%CI/CD, observability, project management, design, security scanning, communication tools
Contractor & vendor spendVariable, 0–15%Staff augmentation, specialized consultants, outsourced QA or support
Training & development1–3%Conferences, courses, certifications, internal learning programs

The exact split varies by company stage and business model — an infrastructure-heavy business (video streaming, ML training) can see cloud costs climb toward 30–40% of the engineering budget, while a consulting-heavy shop might carry a much larger contractor line. The table is a starting point for sanity-checking a budget, not a rule to force-fit.

The annual planning cycle

Most companies run engineering budgeting on an annual cycle with a lighter quarterly check-in, and the cycle has a rhythm worth internalizing rather than reacting to each year as if it were new.

PhaseTypical timingWhat happens
Strategic inputQ3 (for a calendar-year budget)Finance and leadership set top-line growth or cost targets; engineering leadership gets a rough envelope to plan against
Bottom-up buildQ3–Q4Managers build headcount and spend requests tied to roadmap commitments
Negotiation & reconciliationQ4Bottom-up asks get reconciled against the top-down envelope; trade-offs get made explicit
Approval & allocationLate Q4 / early Q1Final numbers are locked; budgets are allocated to cost centers
Execution & monitoringThroughout the yearActuals are tracked against budget; variances are flagged
Quarterly reforecastEach quarterBudget is revisited against actual hiring pace, attrition, and spend; adjustments are made without waiting for the next annual cycle

A manager who only shows up at the “bottom-up build” phase has already ceded influence — the top-line envelope set in the strategic input phase determines how much room exists to negotiate later. Engaging with finance and skip-level leadership before the formal ask is due, to understand what growth or cost assumptions are already baked into the envelope, is what separates a defensible request from a wish list.

Key Concepts

Building a defensible budget request

The weakest form of a budget ask is “we need more people” or “we need a bigger cloud budget,” offered without a chain of reasoning connecting the ask to an outcome the business cares about. A defensible request instead starts from committed or prioritized roadmap work, translates that work into an estimate of engineering capacity required, compares that to current capacity, and expresses the gap as a specific, costed request.

A useful structure for the request itself:

  1. Anchor to committed outcomes. State the roadmap commitments or strategic bets this budget is meant to fund — ideally the same commitments already agreed in the business case from Strategy, Business Case & Prioritization. A request that stands alone, disconnected from an agreed roadmap, invites the question “why do you need this if it’s not tied to anything we’ve agreed to build?”
  2. Show the capacity math. Translate committed work into person-weeks or person-quarters of effort, and show current team capacity against it (see the capacity-planning section below for how to do this honestly). The gap between required and available capacity is the headcount ask.
  3. Cost it out. Turn the headcount and infrastructure gap into dollars: fully loaded cost per hire (salary, benefits, overhead — typically 1.25–1.4x base salary), incremental infra spend at the traffic or data volume the roadmap implies, tooling costs per new seat.
  4. Show the alternative. What happens if the budget is not approved — which commitments slip, by how much, and what’s the business cost of that slip? This reframes the ask from “engineering wants more” to “the business already agreed to X; here’s what it costs to deliver it, and here’s what happens if it isn’t funded.”
  5. Show sensitivity. A range (“with N engineers we hit the Q3 date; with N-2 we hit Q4 and drop the secondary feature”) is more credible and more useful to a finance partner than a single fixed number, because it gives them a lever to pull if the top-line envelope doesn’t stretch to the full ask.

Anticipating pushback

Finance partners and executive leadership ask a predictable set of questions when reviewing an engineering budget request, and preparing answers in advance turns a defensive scramble into a confident conversation.

Likely questionWhat a strong answer looks like
”Why can’t the existing team absorb this?”A capacity breakdown (see below) showing current team is already at or above sustainable load, with on-call, meetings, and onboarding accounted for
”What did we get for last year’s headcount increase?”A short retrospective connecting last year’s hires to delivered outcomes, not just “we hired 5 people"
"Could this be done with contractors instead of full-time hires?”An honest trade-off (see below) rather than a reflexive “no"
"What’s the cloud cost trend, and is it under control?”A trend line, not a single snapshot, plus evidence of an active cost-optimization practice
”What happens if we cut this by 20%?”A specific, pre-thought-through answer (which commitment slips) rather than an improvised one in the room
”Is this a one-time cost or does it recur?”Clear separation of one-time (e.g., migration project, contractor burst) vs. ongoing (e.g., a new permanent headcount line) spend

The common thread across all of these is that the burden of proof sits with the requester, and vague or improvised answers in the room erode trust for the next budget cycle even if this one gets approved.

Cost optimization is bigger than the cloud bill

It’s tempting for a first-time manager to equate “engineering cost optimization” with “cloud cost optimization,” because cloud bills are visible, itemized, and satisfying to attack. But given that headcount is 60–80% of the budget, the largest cost-optimization lever is almost always organizational, not infrastructural.

Headcount efficiency is about whether the org structure and process let the people already on payroll produce as much value as they’re capable of — a team drowning in coordination overhead, unclear ownership, or duplicated effort is a cost problem even though no line item says so. This is a direct extension of the organization design work in Hiring & Organization Design: a poorly shaped team is expensive even at a fixed headcount number, because it wastes the capacity that’s already been paid for.

Tool and SaaS consolidation is a frequently underweighted line. As organizations grow, it’s common for three teams to independently procure three different observability platforms, or for a company to be paying for five project-management tools because each team adopted its own during a period of low oversight. An annual SaaS audit — listing every tool, its cost, its owner, and its overlap with other tools — routinely finds savings of 10–30% of the tooling line with no impact on delivery, purely by consolidating redundant subscriptions and negotiating enterprise pricing on the survivors instead of paying retail per-team rates.

Cloud infrastructure cost is real and often the most visible line, but it is also the one with the most mature technical playbook already available elsewhere in this knowledge base — rightsizing instances, reserved instances and savings plans, autoscaling, spot/preemptible capacity, storage tiering, and so on. Rather than duplicating that material, see Cloud and DevOps for the mechanics; the manager’s job is to make sure the practice exists, is owned by someone, and is reviewed on a cadence — not to perform the rightsizing personally.

The cost-cutting vs. velocity/morale trade-off

Cost optimization has a ceiling past which further cuts stop being free and start trading directly against velocity or morale, and a manager’s job is to know where that ceiling is before crossing it under pressure. Cutting a genuinely redundant SaaS tool is free. Cutting the observability tooling that lets engineers debug production incidents quickly is not — it looks like a cost saving on the spreadsheet and shows up later as longer incident resolution times, more on-call pain, and attrition. Similarly, headcount reductions framed purely as “efficiency” without a corresponding reduction in roadmap scope simply push the same amount of work onto fewer people, which reads as a cost saving for one quarter and a burnout and attrition problem for the next three. The discipline here is to make the trade-off explicit rather than pretending it doesn’t exist: every cost cut should come with an honest answer to “what does this cost us in velocity, reliability, or retention,” not just “how much does it save.”

Resource allocation and forecasting, revisited from the budget lens

Capacity and resource forecasting is covered from a delivery-planning angle elsewhere in this knowledge base; the budget lens adds the dollar translation on top of the headcount plan.

Attrition planning. A headcount budget that assumes zero attrition is wrong on arrival — most engineering organizations see annual voluntary attrition somewhere in the 8–15% range even in healthy teams, and budgets should plan backfill hiring for expected departures separately from net-new growth hiring. Conflating the two is a common source of mid-year surprise: a team that “has budget for 3 new hires” but loses 2 people to attrition has effectively only grown by 1, while still carrying the recruiting and onboarding cost of 5 hiring processes.

Contractor vs. full-time hire. This trade-off recurs every budget cycle and deserves a deliberate answer rather than a default. Contractors typically cost more per hour worked but carry no long-term obligation, ramp faster for well-scoped work, and can be scaled down without a layoff; full-time hires cost less per hour over a multi-year horizon, build durable institutional knowledge, and are the right choice for anything core to the product’s long-term differentiation. A reasonable heuristic: use contractors for time-boxed, well-specified work with a clear end date (a migration, a compliance push, a burst of QA capacity) and full-time hires for anything that will need ongoing ownership past the current project.

DimensionContractorFull-time hire
Cost per hourHigherLower (amortized over tenure)
Ramp-up timeFast (for well-scoped work)Slower, but builds lasting context
Flexibility to scale downHigh, low severance/legal costLow, layoffs carry cost and morale impact
Institutional knowledgeLeaves with the contractAccumulates and compounds
Best fitTime-boxed, well-specified workCore, ongoing product ownership

Translating a hiring plan into a budget line means going beyond “5 new engineers” to modeling when those 5 people actually land on payroll — recruiting and onboarding lead time (often 2–4 months from req approval to start date) means a headcount approved in Q1 rarely produces a full quarter of cost or output in Q1. Budgets that assume instant hiring overstate both the near-term cost and the near-term delivered capacity.

Scaling infrastructure: a capacity-vs-cost trade-off, not a technology choice

When a system approaches its capacity limits, the temptation is to frame the conversation as a technical one — which database, which caching layer, which autoscaling policy. Those are real decisions, but they belong to the engineers and architects doing the work, with the technical depth covered in Cloud, DevOps, and Network Engineer (load balancing, network capacity, CDN strategy, and so on). The manager’s distinct job is one level up: making the cost implication of a capacity decision visible to the business before it’s made, not after.

Concretely, this means translating “we need to scale the database tier” into a business-legible statement: “handling projected Q4 traffic at current growth rates costs an estimated $X/month more in infrastructure, or alternatively we can defer $X by accepting degraded performance at peak — which trade-off does the business prefer?” Without that framing, capacity decisions get made unilaterally by engineering (risking a bill nobody signed off on) or get made unilaterally by finance (risking an outage nobody warned them about). The manager’s leverage is in surfacing the trade-off early enough that it’s a decision rather than a surprise.

Capacity planning for people, not just infrastructure

Infrastructure capacity planning has a mature vocabulary — headroom, peak load, autoscaling thresholds — and it’s worth borrowing the same discipline for planning team capacity, because the classic mistake is treating a sprint or a quarter as if 100% of it is available for “pure feature time.” It never is.

A more honest capacity model subtracts known, recurring load from raw calendar time before ever getting to feature work:

Time sinkTypical share of capacityNotes
On-call & interrupt work5–20%Varies heavily by service maturity and on-call rotation size; spikes during incidents
Meetings & ceremonies10–20%Standups, planning, reviews, 1:1s, cross-team syncs
Onboarding drag (for teams with recent hires)20–50% of a new hire’s first 1–2 quarters, plus a tax on the mentors/reviewers supporting themRamp time is not zero-cost to the rest of the team either
PTO, holidays, sick time8–12% annualizedVaries by region and company policy
Tech debt, tooling, and non-roadmap maintenance10–20%Often invisible in planning because it’s not a named roadmap item
Remaining for committed roadmap workOften only 40–60% of nominal capacityThis is the number that should be used to size commitments, not the headcount count itself

Practically, this means a team of 6 engineers does not have “6 engineers of capacity” for roadmap planning — after subtracting on-call, meetings, PTO, ramp time for anyone hired in the last two quarters, and a reasonable tech-debt allowance, the honest number might be closer to 3.5–4 engineers of committed-work capacity. Roadmaps and budget requests sized against the nominal headcount number instead of this discounted number are the single most common source of the “why did this slip” conversation later — the team wasn’t understaffed relative to the plan, the plan was oversized relative to the team’s real capacity.

This also has a direct budget consequence: if a roadmap genuinely requires 6 engineers of committed-work capacity and the honest capacity model says the team delivers 4, the gap either needs a headcount request (funded, with the lead time from the section above) or a scope cut — not a hope that people will simply work harder.

Best Practices

Treat the budget as a living forecast, not an annual artifact

The annual number is a starting estimate, not a contract with reality. Reviewing actuals against budget monthly or quarterly, and flagging variance early, turns a potential year-end crisis into a manageable mid-course correction. Waiting for the annual cycle to notice a cloud bill has crept 40% over plan means eight or nine months of unmanaged overspend before anyone looks at the number again.

Own a cost dashboard, don’t wait for finance to surface a surprise

A manager who can pull up current cloud spend, tooling cost, and headcount burn rate at any time — without waiting for a finance report — is in a fundamentally stronger negotiating and planning position than one who finds out about overspend secondhand. This doesn’t need to be sophisticated; a simple monthly-updated spreadsheet or a lightweight dashboard fed from billing exports is enough to catch most surprises early. FinOps practice (see references) formalizes this as a shared, continuous discipline between engineering, finance, and product rather than a once-a-year finance-owned exercise.

Separate one-time from recurring costs explicitly

A migration project, a contractor burst, or a one-time tooling purchase should never be quietly folded into the “run rate” — doing so inflates next year’s baseline and makes future comparisons meaningless. Tag spend as one-time vs. recurring at the point it’s incurred, not retroactively when someone asks why the budget looks different than expected.

Define cost-efficiency KPIs deliberately, and watch what they optimize away

It’s useful to track cost-efficiency metrics, but each one carries a distortion risk that should be named alongside the metric, not discovered after the fact.

KPIWhat it measuresDistortion risk if over-optimized
Cost per feature / cost per story pointEfficiency of delivery relative to spendRewards shipping small, easy features over harder, higher-value ones; penalizes investment in platform work with no immediate feature attached
Infra cost as % of revenueWhether infrastructure spend scales sub-linearly with the businessCan mask under-investment in reliability or scalability headroom if revenue happens to grow faster than the cost trend
Cost per customer / cost per transactionUnit economics of the productCan be gamed by deferring necessary infra investment (redundancy, DR capacity) that doesn’t pay off until an incident occurs
Engineer cost per unit of outputHeadcount efficiencyPressures teams to under-report or hide the “invisible” work (on-call, mentoring, tech debt) that doesn’t count toward the visible output metric

The general principle: any cost-efficiency metric, taken far enough, can be improved by quietly cutting reliability investment, deferring maintenance, or overloading people — none of which show up in the metric itself until the deferred cost arrives, usually as an incident, an outage, or attrition. Pairing every cost KPI with a corresponding quality or health metric (incident rate, on-call load, retention) is the standard guardrail: report them together, and treat a cost improvement that comes with a simultaneous decline in the paired metric as a red flag, not a win.

Build the budget bottom-up from the capacity model, not top-down from last year plus growth

“Take last year’s budget and add 15%” is a common shortcut, and it systematically fails to catch situations where the team’s real capacity gap is larger or smaller than a flat percentage would suggest. Building from the honest capacity model in the Key Concepts section — what roadmap work is committed, what capacity it requires, what capacity the team actually has after subtracting on-call, meetings, onboarding, and PTO — produces a number that’s defensible in the room described in the pushback table above, because it’s traceable to specific commitments rather than an extrapolation.

A worked mini-example

The table below sketches a rough annual budget for a mid-size product engineering team of around 25 people (a plausible split across two or three squads plus a platform function), purely to illustrate proportions — real numbers vary enormously by company stage, geography, and industry.

CategoryRough annual spend% of totalNotes
Headcount (25 FTE, fully loaded)$4.5M~72%Salary, benefits, taxes, equity at a blended fully-loaded multiplier
Cloud infrastructure$900K~14%Compute, storage, data platform, network egress
Tooling & SaaS licenses$500K~8%CI/CD, observability, project management, security scanning, communication
Contractor & vendor spend$250K~4%A time-boxed migration project plus seasonal QA augmentation
Training & development$100K~2%Conferences, courses, internal learning stipend
Total~$6.25M100%

Reading this table the way a finance partner will: headcount dwarfs everything else, which means the highest-leverage cost conversation is almost always about org shape and hiring pace, not the cloud line; the cloud and tooling lines together are still worth an active optimization practice (a 15% cut there is real money — roughly $200K in this example — even though it’s a fraction of the headcount line); and the contractor line should map to a specific, named, time-boxed project rather than being a permanent fixture, or it’s really a full-time hire being paid at a contractor premium.

References