← Kỹ sư dữ liệu← Data Engineer
Kỹ sư dữ liệuData Engineer19 Th7, 2026Jul 19, 202622 phút đọc17 min read

Monitoring & ObservabilityMonitoring & Observability

Thuộc bộ kiến thức Data Engineer Roadmap.

Tổng quan

Một pipeline có thể chạy xong với dấu tick xanh, không có lỗi nào trong log, và thời gian chạy nằm hoàn toàn trong khoảng bình thường — nhưng vẫn đưa cho team business một dashboard được xây trên dữ liệu cũ (stale), bị thiếu (truncated), hoặc âm thầm hỏng. Đây chính là vấn đề trọng tâm mà note này giải quyết: monitoring truyền thống chỉ cho bạn biết job có chạy hay không; nó không nói gì về việc dữ liệu mà job tạo ra có đúng hay không. Một câu SELECT COUNT(*) trên bảng đáng lẽ phải nhận 2 triệu dòng qua đêm nhưng chỉ nhận 4.000 dòng sẽ không làm bất kỳ exception nào bật lên ở đâu cả. Một thay đổi schema ở upstream đổi tên user_id thành userId sẽ không làm task Airflow nào fail; nó chỉ khiến mọi query downstream âm thầm trả về NULL cho cột đó, và người đầu tiên phát hiện ra chuyện này là một analyst bối rối hỏi tại sao số liệu cohort tuần trước lại tụt về 0.

Khoảng cách giữa “pipeline chạy thành công” và “pipeline tạo ra dữ liệu đúng” chính là lý do data engineering phát triển riêng một mảng gọi là data observability, khác biệt — nhưng được xây trên nền — các thực hành observability hạ tầng và ứng dụng tổng quát đã nói ở DevOps — Monitoring & Observability. Note đó bàn kỹ về ba trụ cột (metrics, logs, traces), Prometheus, SLO, và thiết kế alerting; nên đọc trước nếu các thuật ngữ đó còn xa lạ. Note này giả định bạn đã có nền tảng đó và đặt một câu hỏi hẹp hơn, sắc hơn: một khi đã biết job có chạy, làm sao biết được dữ liệu nó tạo ra có đáng tin không?

Cách đóng khung đã trở thành chuẩn ngành cho câu hỏi này là năm trụ cột của data observability do Monte Carlo Data phổ biến: freshness (độ mới), volume (khối lượng), schema (cấu trúc), distribution (phân phối giá trị), và lineage (dòng chảy dữ liệu). Mỗi trụ cột nhắm vào một kiểu lỗi âm thầm cụ thể mà pipeline có thể mắc phải — đến muộn, đến sai số lượng, đổi hình dạng mà không báo trước, trôi lệch thống kê ra ngoài phạm vi kỳ vọng, hoặc làm hỏng thứ gì đó downstream mà không ai ngờ tới. Note này sẽ đi qua từng trụ cột, đối chiếu trực tiếp operational monitoring với data monitoring, tóm tắt ngắn gọn cách bộ công cụ observability tổng quát (logs, metrics, Prometheus, Datadog) áp dụng cho code pipeline, rồi bàn cách thiết kế alerting sao cho bắt được sự cố dữ liệu mà không nhấn chìm cả team trong nhiễu. Về các quy tắc chất lượng dữ liệu và hệ thống metadata mà công cụ data observability thường xây dựng dựa trên, xem ./14-data-quality-governance-and-metadata.md; về scheduler và DAG mà các check này được gắn vào, xem ./09-workflow-orchestration.md.

Kiến thức nền tảng

Operational monitoring vs. data monitoring

Sự phân biệt quan trọng nhất trong toàn bộ chủ đề này là “monitoring” thực ra mang hai nghĩa hoàn toàn khác nhau tùy vào việc ai đang hỏi.

Operational monitoringData monitoring (data observability)
Câu hỏi nó trả lờiJob có chạy không? Chạy mất bao lâu? Có lỗi không?Dữ liệu job này tạo ra có đúng, đầy đủ, và mới không?
Tín hiệu điển hìnhExit code, thời gian chạy task, CPU/memory, số lần retry, trạng thái DAG runSố dòng dữ liệu, tỷ lệ null, diff schema, dịch chuyển phân phối, độ trễ freshness
Nơi lưu trữUI orchestrator (Airflow/Dagster), Prometheus, DatadogNền tảng data observability, dbt tests, SQL check tự viết, bảng audit trong warehouse
Kiểu lỗi nó bắt đượcTask crash, timeout, OOM, lỗi mạngAPI upstream âm thầm đổi field; nguồn dữ liệu đột nhiên chỉ gửi 90% khối lượng thường ngày; bảng không update suốt 8 tiếng nhưng không task nào fail
”Xanh” thực sự chứng minh điều gìCode đã chạy mà không throw lỗiKhông chứng minh điều gì về tính đúng đắn — một job có thể “thành công” trong khi ghi 0 dòng, toàn null, hoặc trùng lặp dữ liệu

Lý do sự phân biệt này quan trọng trong thực tế: một dashboard giám sát pipeline — loại được xây từ ba trụ cột trong note DevOps — có thể hoàn toàn xanh trong khi một dashboard business xây trên output của pipeline đó lại âm thầm sai. Một câu INSERT chạy trên một bảng nguồn rỗng “thành công” xét về mọi khía cạnh operational: exit code trả về 0, không throw lỗi, chạy trong thời gian bình thường. Nó cũng ghi 0 dòng, và trừ khi có gì đó đang kiểm tra rõ ràng dữ liệu, không ai biết chuyện này cho tới khi một stakeholder để ý thấy biểu đồ nằm ngang vài ngày sau. Data observability tồn tại chính là để lấp khoảng trống này — giám sát nội dung (payload), chứ không chỉ tiến trình mang nó.

Logs, metrics, traces cho code pipeline — tóm tắt ngắn

Ba trụ cột từ note DevOps vẫn áp dụng trực tiếp cho việc thực thi pipeline, và là nền tảng cần thiết trước khi data observability có ý nghĩa — bạn không thể lý luận về “dữ liệu có bị muộn không” nếu chưa có timestamp, và không thể debug tại sao job fail nếu không có structured log. Trong bối cảnh data engineering cụ thể:

CAP theorem, nhìn lại từ góc độ observability

Khi một alert bật lên báo “dữ liệu đang trông không nhất quán”, sẽ rất hữu ích nếu bạn đã biết trước liệu hệ thống đó có từng được cam kết consistency chặt chẽ hay không. CAP theorem và hệ quả thực tiễn của nó — sự đánh đổi giữa consistency và availability khi có network partition, cùng phổ các mô hình consistency (strong, eventual, causal) mà các data store phân tán lựa chọn — đã được bàn kỹ ở Backend — NoSQL Databases. Điểm cần rút ra ở góc độ observability: nếu hệ thống nguồn được xây trên một store eventually-consistent (nhiều NoSQL database, một số cache replicated, hầu hết setup multi-region), thì việc một replica trễ vài giây và tạm thời trả về dữ liệu cũ là hành vi kỳ vọng, không phải sự cố. Phân biệt được “store này eventually consistent và đây là độ trễ replication bình thường” với “store này lẽ ra phải strongly consistent và có gì đó thực sự hỏng” giúp tránh cả báo động giả (page on-call vì độ trễ vốn dự kiến) lẫn sự chủ quan giả (bỏ qua một vấn đề staleness thật sự chỉ vì “eventual consistency mà”). Freshness monitoring, bàn ở phần sau, thực chất là một công cụ ở tầng business phản ánh chính sự căng thẳng giữa consistency và tính kịp thời này.

Horizontal vs. vertical scaling, sync vs. async — lý luận về nghẽn cổ chai

Hai khái niệm hệ thống phân tán tổng quát xuất hiện liên tục khi diễn giải tại sao một alert data monitoring bật lên. Vertical scaling (máy lớn hơn — thêm CPU/RAM cho một node warehouse hoặc Spark driver duy nhất) có trần và một điểm lỗi duy nhất (single point of failure); horizontal scaling (thêm worker/node/shard) là cách hầu hết nền tảng dữ liệu hiện đại (Spark, Snowflake, BigQuery, Kafka) hấp thụ khối lượng tăng, nhưng nó dịch chuyển nghẽn cổ chai sang phối hợp (coordination), shuffle dữ liệu, và lệch phân vùng (partition skew) — một anomaly về volume đôi khi thực ra là vấn đề “một partition nhận gấp 100 lần dữ liệu so với các partition khác” chứ không phải nguồn thực sự giảm khối lượng. Giao tiếp synchronous (một bước pipeline block chờ phản hồi API) khiến độ trễ cộng dồn và một dependency chậm làm nghẽn cả chuỗi, đe dọa trực tiếp SLA về freshness; giao tiếp asynchronous (queue, event stream, pub/sub) tách rời tốc độ producer và consumer nhưng lại đưa câu hỏi freshness trở lại dưới một hình thức khác — “consumer đang chậm bao xa so với đầu stream” (consumer lag) tự nó chính là một metric freshness. Các khái niệm này được bàn tổng quát hơn từ góc độ thiết kế hệ thống ở Backend; điểm cần lưu ý ở đây hẹp hơn: trước khi coi một alert freshness hoặc volume là “dữ liệu sai”, hãy kiểm tra xem đó có phải thực ra là “hạ tầng bên dưới đã scale hoặc giao tiếp theo cách làm thay đổi khái niệm ‘bình thường’” hay không.

Khái niệm chính

Năm trụ cột của data observability

Monte Carlo Data đã phổ biến một khung năm trụ cột từ đó trở thành từ vựng mặc định của mảng này (được các công cụ cạnh tranh như Bigeye, Soda, Metaplane áp dụng, với vài biến thể nhỏ). Mỗi trụ cột tương ứng với một kiểu lỗi âm thầm khác nhau — một cách pipeline trông ổn về mặt operational trong khi dữ liệu bên dưới lại sai.

Freshness (độ mới)

Freshness đặt câu hỏi: dữ liệu có đến đúng thời điểm cần thiết không? Một bảng mà dashboard BI kỳ vọng cập nhật mỗi giờ, nhưng lần update gần nhất cách đây sáu tiếng, là một lỗi freshness — và đó là lỗi âm thầm, vì job upstream vốn thường điền dữ liệu vào bảng đó có thể đơn giản là ngừng được schedule, bị kẹt trong queue, hoặc bị bỏ qua âm thầm, mà không throw lỗi ở bất kỳ đâu. Giám sát freshness thường hoạt động bằng cách theo dõi, cho từng bảng, timestamp của lần update gần nhất (hoặc giá trị max của cột updated_at/loaded_at) và cảnh báo khi độ trễ giữa “hiện tại” và “lần update cuối” vượt ngưỡng kỳ vọng. Bản thân ngưỡng cần được học hoặc cấu hình riêng cho từng bảng — một bảng update mỗi 15 phút và một bảng update mỗi ngày đều cần giám sát freshness, nhưng với khung cảnh báo khác nhau hoàn toàn. Đây là trụ cột liên quan trực tiếp nhất tới kịch bản “họp BI 8 giờ sáng”: nếu dashboard cấp điều hành cần fct_orders phải mới tính đến 8 giờ, freshness monitoring chính là thứ page ai đó lúc 7 giờ thay vì để VP là người phát hiện đầu tiên.

Volume (khối lượng)

Volume đặt câu hỏi: lượng dữ liệu có nằm trong khoảng kỳ vọng không? Một nguồn thường gửi 500.000 event mỗi ngày mà đột nhiên chỉ gửi 50.000 — giảm 90% — gần như chắc chắn báo hiệu một vấn đề upstream đáng để điều tra: một tích hợp API bị hỏng, một đối tác ngừng gửi feed, một filter upstream bị áp nhầm. Tương tự, một cú tăng đột biến 10 lần về volume có thể chỉ ra lỗi xử lý trùng lặp (cùng một batch bị ingest lại) chứ không phải tăng trưởng thật. Giám sát volume thường được triển khai bằng cách theo dõi số dòng theo từng bảng hoặc từng lần load, so sánh với baseline lịch sử (trung bình trượt, baseline có nhận biết ngày trong tuần, hoặc ngưỡng phần trăm thay đổi đơn giản), và đây là một trong những check rẻ nhất để thêm vào vì số dòng gần như luôn đã sẵn có từ chính job load — điểm đã nói ở trên về việc log rows_read/rows_written trong structured log chính là nguồn cấp trực tiếp cho trụ cột này.

Schema (cấu trúc)

Schema đặt câu hỏi: hình dạng của dữ liệu có thay đổi mà không ai được báo không? Một cột bị đổi tên, một kiểu dữ liệu chuyển từ INT sang STRING, một field lồng nhau được thêm hoặc bỏ, một API nguồn thêm một field bắt buộc mới — bất kỳ điều nào trong số này đều có thể âm thầm phá vỡ các transformation downstream, hoặc bằng cách tạo ra NULL ở nơi một JOIN hay tham chiếu cột không còn khớp, hoặc bằng cách khiến một phép cast fail sâu bên trong một transformation mà không ai theo dõi đủ sát để nhận ra kịp thời. Giám sát schema hoạt động bằng cách chụp snapshot schema (tên cột, kiểu, cấu trúc lồng nhau) ở mỗi lần load và diff với snapshot trước đó, cảnh báo khi có thay đổi bất ngờ — trong khi vẫn cho phép một luồng rõ ràng, đã được review, cho việc tiến hóa schema có chủ đích (một migration, một cột mới có chủ ý) đi qua mà không báo động. Trụ cột này chồng lấn trực tiếp với các thực hành schema validation và contract testing ở ./14-data-quality-governance-and-metadata.md; note đó bàn về khía cạnh governance (ai sở hữu một schema, thay đổi được review và truyền đạt như thế nào), còn trụ cột này bàn về khía cạnh phát hiện (bắt được một thay đổi xảy ra mà hoàn toàn không đi qua quy trình đó).

Distribution (phân phối giá trị)

Distribution đặt câu hỏi: các giá trị bên trong dữ liệu có bình thường về mặt thống kê không? Số dòng có thể chính xác tuyệt đối và schema có thể không hề bị đụng tới trong khi giá trị của từng cột trôi lệch theo những cách khiến dữ liệu sai về bản chất: một cột discount_pct lẽ ra nằm trong khoảng 0–100 đột nhiên chứa giá trị âm; tỷ lệ null trên email vốn thường dưới 1% nhảy vọt lên 40% vì một form upstream ngừng bắt buộc field đó; một cột phân loại (categorical) lẽ ra chỉ chứa năm trạng thái đã biết bắt đầu chứa giá trị thứ sáu chưa được ánh xạ. Giám sát distribution so sánh các đặc tính thống kê — tỷ lệ null, giới hạn min/max, mean/độ lệch chuẩn, cardinality của các trường phân loại, tính duy nhất của key — với baseline lịch sử, và là trụ cột chồng lấn nhiều nhất với kiểm thử chất lượng dữ liệu kiểu cổ điển (dbt test, Great Expectations, Soda checks): khác biệt trong thực tế chủ yếu nằm ở việc các check được viết tay theo từng cột (data quality testing) hay được tự động suy ra và giám sát liên tục bởi một nền tảng học “bình thường” từ lịch sử (công cụ data observability).

Lineage (dòng chảy dữ liệu)

Lineage đặt câu hỏi: nếu cái này hỏng, cái gì khác sẽ hỏng theo? Freshness, volume, schema, và distribution đều phát hiện rằng có gì đó sai; lineage là thứ cho bạn biết phạm vi ảnh hưởng (blast radius) — những bảng downstream, dashboard, feature ML, và báo cáo nào phụ thuộc vào bảng vừa gặp sự cố, và ai sở hữu chúng. Không có lineage, một bảng nguồn hỏng duy nhất chỉ tạo ra một alert và một số nạn nhân downstream âm thầm không xác định: các dashboard cách ba bước join lặng lẽ trở nên cũ hoặc sai, và những người muốn biết chuyện này không bao giờ được báo, vì chẳng ai từng vẽ ra dependency đó từ đầu. Lineage thường được thu thập bằng cách parse các lệnh SQL/dbt ref()/source() để tự động xây dựng đồ thị phụ thuộc, hoặc qua metadata do orchestrator và warehouse đẩy ra (column-level lineage trong các công cụ như OpenLineage, DAG tích hợp sẵn của dbt, hoặc lineage từ lịch sử query gốc của warehouse). Trụ cột này vừa mang tính metadata/governance vừa mang tính observability, và được bàn thêm từ góc độ cataloging ở ./14-data-quality-governance-and-metadata.md.

Tóm tắt các trụ cột — điều gì hỏng nếu bỏ qua, cách giám sát điển hình

Trụ cộtĐiều gì hỏng nếu bỏ quaCách giám sát điển hình
FreshnessDashboard và báo cáo âm thầm trở nên cũ; quyết định được đưa ra dựa trên số liệu của hôm qua (hoặc tuần trước) mà không ai nhận raTheo dõi max(updated_at)/loaded_at theo từng bảng; cảnh báo khi hiện tại trừ lần update cuối vượt ngưỡng kỳ vọng riêng cho từng bảng
VolumeSự cố upstream, tích hợp hỏng, hoặc load trùng lặp không bị phát hiện vì job vẫn “thành công”Theo dõi số dòng theo từng lần load, so với baseline lịch sử/có nhận biết ngày trong tuần; cảnh báo khi lệch phần trăm
SchemaJoin/cast downstream âm thầm fail hoặc trả về NULL; transformation hỏng theo cách trông giống vấn đề dữ liệu chứ không phải vấn đề schemaChụp snapshot schema theo từng lần load, diff với snapshot trước, cảnh báo khi có thay đổi chưa được review; kiểm tra contract/schema registry ở khâu ingestion
DistributionGiá trị sai về mặt thống kê (null, số ngoài khoảng, category chưa được ánh xạ) lan vào các aggregate và model mà không bị phát hiệnTheo dõi tỷ lệ null, min/max, mean/stddev, cardinality theo từng cột so với baseline lịch sử; dbt tests / Great Expectations / Soda checks
LineageMột sự cố upstream duy nhất gây thiệt hại downstream không xác định, không được ánh xạ — không ai biết cần kiểm tra gì hoặc báo cho aiĐồ thị phụ thuộc tự động từ việc parse SQL/dbt hoặc metadata kiểu OpenLineage; column-level lineage trong catalog

Best Practices

Thiết kế alert quanh SLA, không phải ngưỡng thô

Cách bền vững nhất để tránh cả bỏ sót sự cố lẫn alert fatigue là neo alert vào một cam kết business thực sự thay vì một ngưỡng thống kê tùy tiện. Nếu buổi họp sáng của team finance phụ thuộc vào việc fct_revenue phải mới tính đến 8:00 sáng, đó chính là SLA: “bảng này phải được update trước 08:00” là một điều kiện cụ thể, có thể kiểm chứng, ánh xạ trực tiếp sang một check freshness, và nó cho bạn biết chính xác mức độ khẩn cấp khi bị vi phạm (rất khẩn cấp — ai đó sắp trình bày số liệu cũ) so với một tín hiệu chung chung kiểu “bảng này hôm nay chậm hơn 20% so với bình thường”, vốn có thể hoàn toàn ổn. Hãy viết rõ các SLA này cho từng bảng quan trọng — hạn chót freshness, khoảng volume chấp nhận được, chủ sở hữu, và các bên tiêu thụ downstream — theo cách một team SRE viết SLO cho service, và xử lý vi phạm SLA dữ liệu với mức độ nghiêm túc của incident response tương tự như những gì note DevOps mô tả cho error-budget burn.

Tinh chỉnh ngưỡng theo anomaly thật, không phải nhiễu

Một check volume bật lên mỗi khi số dòng hàng ngày dao động hơn 2% sẽ thỉnh thoảng đúng và thường xuyên sai, và những báo động giả liên tục sẽ huấn luyện đội on-call bỏ qua kênh cảnh báo hoàn toàn — tương đương alert fatigue trong data observability mà phần alerting của note DevOps đã mô tả. Ưu tiên các baseline có tính đến biến động đã biết, kỳ vọng trước: tính mùa vụ theo ngày trong tuần (volume thứ Hai khác volume thứ Bảy với hầu hết doanh nghiệp consumer), đợt tăng vọt cuối tháng/cuối quý, và xu hướng tăng trưởng dần dần, thay vì một dải phần trăm phẳng tính so với riêng hôm qua. Khi có thể, để vài tuần lịch sử tự thiết lập khoảng “bình thường” thay vì chọn tay ngưỡng ngay từ ngày đầu — đây chính là điều các nền tảng data observability chuyên dụng (Monte Carlo, Bigeye, Metaplane) tự động hóa, so với các check dbt/SQL viết tay, vốn cần tinh chỉnh ngưỡng thủ công nhiều hơn theo từng bảng.

Kết hợp data observability với pipeline observability thay vì thay thế nó

Data observability không làm cho operational monitoring trở nên thừa thãi — nó nằm chồng lên trên đó. Một pipeline được instrument tốt vẫn cần các metrics ở cấp task, structured log, và alerting DAG-run đã nói ở note DevOps; data observability thêm vào một lớp check thứ hai, độc lập, trên output của các task đó. Trong thực tế, hầu hết các team gắn cả hai vào cùng một orchestrator: hook on_failure_callback/SLA-miss của chính DAG Airflow xử lý các lỗi operational, trong khi một task data-quality ở cuối DAG (một lần chạy dbt test, một checkpoint Great Expectations, hoặc một lời gọi API tới nền tảng data observability) xử lý lớp tính đúng đắn của dữ liệu, và cả hai đều đẩy alert vào cùng một kênh on-call để không ai phải kiểm tra hai dashboard riêng biệt mới biết được liệu lần chạy đêm qua có thực sự đáng tin hay không.

Làm cho lineage và ownership dễ tra cứu trước sự cố, không phải trong lúc sự cố

Thời điểm tệ nhất để tìm hiểu “ai sở hữu bảng này và cái gì phụ thuộc vào nó” là giữa lúc sự cố đang diễn ra. Hãy đầu tư vào một đồ thị lineage được tạo tự động (qua DAG của dbt, OpenLineage, hoặc một công cụ catalog) và đảm bảo mọi bảng quan trọng đều có chủ sở hữu được gán từ trước, trước khi có bất kỳ sự cố thực sự nào xảy ra — điều này biến một cuộc điều tra root-cause dựa trên lineage từ một cuộc săn lùng trên Slack thành một lần tra cứu đồ thị mất hai phút, và là thứ thực sự khiến trụ cột “phạm vi ảnh hưởng” hữu ích về mặt vận hành thay vì chỉ mang tính lý thuyết.

Tài liệu tham khảo

Part of the Data Engineer Roadmap knowledge base.

Overview

A pipeline can finish with a green checkmark, zero errors in the logs, and a runtime well within its usual window — and still hand the business team a dashboard built on stale, truncated, or silently corrupted data. This is the central problem this note addresses: traditional monitoring tells you whether the job ran; it says nothing about whether the data the job produced is actually correct. A SELECT COUNT(*) on a table that should have received 2 million rows overnight but received 4,000 will not raise an exception anywhere. An upstream schema change that renames user_id to userId will not fail an Airflow task; it will just make every downstream query silently return NULL for that column, and the first anyone hears about it is a confused analyst asking why last week’s cohort numbers dropped to zero.

This gap between “the pipeline ran successfully” and “the pipeline produced correct data” is why data engineering has grown its own discipline of data observability, distinct from — but built on top of — the general infrastructure and application observability practices covered in DevOps — Monitoring & Observability. That note covers the three pillars (metrics, logs, traces), Prometheus, SLOs, and alerting design in depth; read it first if those terms are unfamiliar. This note assumes that foundation and asks a narrower, sharper question: once you know a job ran, how do you know the data it produced is trustworthy?

The framing that has become the industry standard here is Monte Carlo Data’s five pillars of data observability: freshness, volume, schema, distribution, and lineage. Each pillar targets a specific way pipelines fail silently — arriving late, arriving in the wrong quantity, changing shape without warning, drifting statistically out of expected ranges, or breaking something downstream that nobody realized depended on it. This note works through each pillar, contrasts operational monitoring with data monitoring directly, briefly recaps how the general observability toolkit (logs, metrics, Prometheus, Datadog) applies to pipeline code, and then covers how to design alerting so that data incidents get caught without burying the team in noise. For the data-quality rules and metadata systems that data observability tooling often builds on, see ./14-data-quality-governance-and-metadata.md; for the schedulers and DAGs that these checks get wired into, see ./09-workflow-orchestration.md.

Fundamentals

Operational monitoring vs. data monitoring

The most important distinction in this entire topic is that “monitoring” means two genuinely different things depending on who’s asking.

Operational monitoringData monitoring (data observability)
Question it answersDid the job run? How long did it take? Did it throw an error?Is the data this job produced correct, complete, and fresh?
Typical signalsExit codes, task duration, CPU/memory, retry counts, DAG run statusRow counts, null rates, schema diffs, distribution shifts, freshness lag
Where it livesOrchestrator UI (Airflow/Dagster), Prometheus, DatadogData observability platform, dbt tests, custom SQL checks, warehouse audit tables
Failure mode it catchesCrashed task, timeout, OOM, network errorUpstream API silently changed a field; a source stopped sending 90% of its usual volume; a table hasn’t updated in 8 hours but no task failed
What a “green” status actually provesThe code executed without throwingNothing about correctness — a job can succeed while writing zero rows, all nulls, or duplicated data

The reason this distinction matters in practice: a pipeline monitoring dashboard — the kind built from the three pillars in the DevOps note — can be entirely green while a business dashboard built on top of that pipeline’s output is quietly wrong. An INSERT that runs against an empty source table “succeeds” in every operational sense: it returns an exit code of 0, it doesn’t throw, it completes in normal time. It also inserts zero rows, and unless something is explicitly checking the data, nobody finds out until a stakeholder notices a chart flatlining days later. Data observability exists to close exactly this gap — to monitor the payload, not just the process that carries it.

Logs, metrics, and traces for pipeline code — a brief recap

The three pillars from the DevOps note still apply directly to pipeline execution, and are necessary groundwork before data-specific observability makes sense — you can’t reason about “is the data late” without first having timestamps, and you can’t debug why a job failed without structured logs. In a data engineering context specifically:

CAP theorem, revisited from an observability angle

When an alert fires saying “the data looks inconsistent right now,” it helps to already know whether the system in question was ever promised strict consistency in the first place. The CAP theorem and its practical consequences — the trade-off between consistency and availability under a network partition, and the spectrum of consistency models (strong, eventual, causal) that distributed data stores choose from — are covered in depth in Backend — NoSQL Databases. The observability-relevant takeaway: if a source system is built on an eventually-consistent store (many NoSQL databases, some replicated caches, most multi-region setups), then a replica lagging by a few seconds and briefly serving stale reads is expected behavior, not an incident. Distinguishing “this store is eventually consistent and this is normal replication lag” from “this store should be strongly consistent and something is actually broken” prevents both false alarms (paging on-call for expected lag) and false calm (dismissing a real staleness problem as “just eventual consistency”). Freshness monitoring, discussed below, is in effect a business-level instrument of exactly this same underlying tension between consistency and timeliness.

Horizontal vs. vertical scaling, sync vs. async — reasoning about bottlenecks

Two general distributed-systems concepts recur constantly when interpreting why a data monitoring alert fired. Vertical scaling (a bigger machine — more CPU/RAM on a single warehouse node or Spark driver) has a ceiling and a single point of failure; horizontal scaling (more workers/nodes/shards) is how most modern data platforms (Spark, Snowflake, BigQuery, Kafka) absorb growing volume, but it shifts bottlenecks toward coordination, shuffling, and partition skew — a volume anomaly is sometimes really a “one partition got 100x the data of its siblings” problem rather than a true source-volume drop. Synchronous communication (a pipeline step blocks waiting on an API response) makes latency additive and a single slow dependency stalls the whole chain, directly threatening freshness SLAs; asynchronous communication (queues, event streams, pub/sub) decouples producer and consumer speed but reintroduces the freshness question in a different form — “how far behind is the consumer from the tip of the stream?” (consumer lag) is itself a freshness metric. These concepts are covered more generally from a systems-design angle in Backend; the point here is narrower: before treating a data-freshness or volume alert as a “the data is wrong” problem, check whether it’s actually a “the underlying infrastructure scaled or communicated in a way that changes what normal looks like” problem.

Key Concepts

The five pillars of data observability

Monte Carlo Data popularized a five-pillar framework that has since become the de facto vocabulary for this space (adopted, with minor variations, by Bigeye, Soda, Metaplane, and most competing tools). Each pillar corresponds to a different silent-failure mode — a way a pipeline can look fine operationally while the data underneath it is wrong.

Freshness

Freshness asks: is the data arriving when it’s supposed to? A table that a BI dashboard expects to refresh hourly, but which last updated six hours ago, is a freshness failure — and it is a silent one, because the upstream job that normally populates it may have simply stopped being scheduled, gotten stuck in a queue, or been silently skipped, without throwing any error anywhere. Freshness monitoring typically works by tracking, per table, the timestamp of the most recent update (or the max value of a updated_at/loaded_at column) and alerting when the lag between “now” and “last update” exceeds an expected threshold. The threshold itself needs to be learned or configured per table — a table that updates every 15 minutes and one that updates once a day both need freshness monitoring, but with wildly different alert windows. This is the pillar most directly tied to the “6am BI meeting” scenario: if the executive dashboard needs fct_orders to be current as of 8am, freshness monitoring is what pages someone at 7am instead of letting the VP notice first.

Volume

Volume asks: is the amount of data within the range you’d expect? A source that normally sends 500,000 events a day suddenly sending 50,000 — a 90% drop — signals an upstream problem almost certainly worth investigating: an API integration broke, a partner stopped sending a feed, an upstream filter got applied by mistake. Equally, a sudden 10x increase in volume can indicate a duplicate-processing bug (the same batch got re-ingested) rather than genuine growth. Volume monitoring is usually implemented as row-count tracking per table or per load, compared against a historical baseline (a trailing average, a day-of-week-aware baseline, or a simple percentage-change threshold), and it is one of the cheapest checks to add because the row count is almost always already available from the load job itself — the earlier point about logging rows_read/rows_written in structured logs feeds directly into this pillar.

Schema

Schema asks: did the shape of the data change without anyone being told? A column gets renamed, a type changes from INT to STRING, a nested field gets added or dropped, a source API adds a new required field — any of these can silently break downstream transformations, either by producing NULLs where a JOIN or column reference no longer matches, or by causing a cast to fail deep inside a transformation that nobody is watching closely enough to notice quickly. Schema monitoring works by snapshotting the schema (column names, types, nested structure) at each load and diffing it against the previous snapshot, alerting on any unexpected change — while allowing an explicit, reviewed path for intentional schema evolution (a migration, a deliberate new column) to pass through without alerting. This pillar overlaps directly with the schema-validation and contract-testing practices in ./14-data-quality-governance-and-metadata.md; that note covers the governance side (who owns a schema, how changes get reviewed and communicated), while this pillar covers the detection side (catching a change that happened without going through that process at all).

Distribution

Distribution asks: are the values inside the data statistically normal? Row counts can be exactly right and the schema can be untouched while individual column values drift in ways that make the data wrong in substance: a discount_pct column that should range 0–100 suddenly contains negative values; a null rate on email that is normally under 1% jumps to 40% because an upstream form stopped requiring the field; a categorical column that should only ever contain five known statuses starts containing a sixth, unmapped value. Distribution monitoring compares statistical properties — null rates, min/max bounds, mean/standard deviation, cardinality of categorical fields, uniqueness of keys — against historical baselines, and is the pillar with the most overlap with classic data-quality testing (dbt test, Great Expectations, Soda checks): the difference in practice is mostly about whether the checks are hand-written per column (data quality testing) or automatically inferred and monitored continuously by a platform that learns “normal” from history (data observability tooling).

Lineage

Lineage asks: if this breaks, what else breaks with it? Freshness, volume, schema, and distribution all detect that something is wrong; lineage is what tells you the blast radius — which downstream tables, dashboards, ML features, and reports depend on the table that just failed, and who owns them. Without lineage, a single broken source table generates one alert and an unknown number of silent downstream casualties: dashboards built three joins away from the break quietly go stale or wrong, and the people who’d want to know are never told, because nobody mapped the dependency in the first place. Lineage is typically captured either by parsing SQL/dbt ref()/source() calls to build a dependency graph automatically, or via metadata pushed by the orchestrator and warehouse (column-level lineage in tools like OpenLineage, dbt’s built-in DAG, or a warehouse’s native query-history lineage). This pillar is as much a metadata/governance concern as an observability one, and is covered further from the cataloging angle in ./14-data-quality-governance-and-metadata.md.

Pillar summary — what breaks if you ignore it, how it’s monitored

PillarWhat breaks if ignoredHow it’s typically monitored
FreshnessDashboards and reports silently go stale; decisions get made on yesterday’s (or last week’s) numbers with nobody realizing itTrack max(updated_at)/loaded_at per table; alert when now − last-update exceeds an expected, table-specific threshold
VolumeUpstream outages, broken integrations, or duplicate loads go unnoticed because the job still “succeeds”Row-count tracking per load, compared to a historical/day-of-week-aware baseline; alert on percentage deviation
SchemaDownstream joins/casts silently fail or return NULLs; transformations break in ways that look like data problems rather than schema problemsSnapshot schema per load, diff against previous snapshot, alert on unreviewed changes; contract/schema-registry checks at ingestion
DistributionStatistically wrong values (nulls, out-of-range numbers, unmapped categories) propagate into aggregates and models undetectedTrack null rate, min/max, mean/stddev, cardinality per column against historical baselines; dbt tests / Great Expectations / Soda checks
LineageA single upstream break causes unknown, unmapped downstream damage — nobody knows what to check or who to notifyAutomated dependency graph from SQL/dbt parsing or OpenLineage-style metadata; column-level lineage in the catalog

Best Practices

Design alerts around SLAs, not raw thresholds

The most durable way to avoid both missed incidents and alert fatigue is to anchor alerts to an actual business commitment rather than an arbitrary statistical threshold. If the finance team’s morning stand-up depends on fct_revenue being fresh by 8:00 AM, that is the SLA: “this table must be updated by 08:00” is a concrete, testable condition that maps directly to a freshness check, and it tells you exactly how urgent a violation is (very — someone is about to present stale numbers) versus a generic “this table is 20% slower than usual today” signal, which might be entirely fine. Write these SLAs down explicitly per critical table — freshness deadline, acceptable volume range, owner, and downstream consumers — the same way an SRE team writes SLOs for services, and treat data SLA breaches with the same incident-response seriousness the DevOps note describes for error-budget burn.

Tune thresholds to genuine anomalies, not noise

A volume check that fires every time daily row counts vary by more than 2% will be right occasionally and wrong constantly, and constant false alarms train the on-call rotation to ignore the channel entirely — the data-observability equivalent of alert fatigue covered in the DevOps note’s alerting section. Prefer baselines that account for known, expected variance: day-of-week seasonality (Monday volumes differ from Saturday volumes for most consumer businesses), month-end/quarter-end spikes, and gradual trend growth, rather than a flat percentage band computed against yesterday alone. Where possible, let a few weeks of history establish the “normal” range automatically rather than hand-picking a threshold on day one — this is exactly what dedicated data observability platforms (Monte Carlo, Bigeye, Metaplane) automate, versus hand-rolled dbt/SQL checks, which tend to need more manual threshold-tuning per table.

Combine data observability with pipeline observability rather than replacing it

Data observability does not make operational monitoring unnecessary — it sits on top of it. A well-instrumented pipeline still needs the task-level metrics, structured logs, and DAG-run alerting covered in the DevOps note; data observability adds a second, independent layer of checks on the output of those tasks. In practice, most teams wire both into the same orchestrator: an Airflow DAG’s own on_failure_callback/SLA-miss hooks handle operational failures, while a data-quality task at the end of the DAG (a dbt test run, a Great Expectations checkpoint, or a call out to a data observability platform’s API) handles the data-correctness layer, and both feed alerts into the same on-call channel so nobody has to check two separate dashboards to know whether last night’s run was actually trustworthy.

Make lineage and ownership discoverable before an incident, not during one

The worst time to figure out “who owns this table and what depends on it” is in the middle of a live incident. Invest in an automatically generated lineage graph (via dbt’s DAG, OpenLineage, or a catalog tool) and make sure every critical table has an assigned owner, ahead of any actual break — this turns a lineage-driven root-cause investigation from a Slack scavenger hunt into a two-minute graph lookup, and is what actually makes the “blast radius” pillar useful operationally rather than theoretical.

References