Data Quality, Governance & MetadataData Quality, Governance & Metadata
Thuộc bộ kiến thức Data Engineer Roadmap.
Tổng quan
Mọi ghi chú khác trong bộ kiến thức này — ingestion (xem 03 — Data Sources & Ingestion), pipeline (xem 08 — ETL, ELT & Data Pipelines), warehouse và lake (xem 06 — Data Modeling & Warehousing và 07 — Data Lakes & Modern Architectures) — suy cho cùng đều phục vụ một mục tiêu duy nhất: một người hoặc một hệ thống đưa ra quyết định dựa trên một con số mà pipeline tạo ra. Nếu con số đó sai, mọi sự tinh tế kỹ thuật phía trên nó đều vô nghĩa. Một pipeline chạy đúng lịch, scale thoải mái, chi phí vận hành gần như bằng không, code đọc vào là thích — nhưng âm thầm trả về con số doanh thu sai — thì tệ hơn hẳn một pipeline xấu xí, chậm chạp, đáng xấu hổ nhưng cho ra con số đúng. Đây là câu ngạn ngữ lâu đời nhất trong ngành máy tính, “garbage in, garbage out”, nhưng đáng để phát biểu lại ở dạng sắc bén nhất dành riêng cho data engineering: tính đúng đắn (correctness) không phải là một thuộc tính chất lượng trong nhiều thuộc tính; nó chính là sản phẩm. Mọi thứ khác — độ trễ, chi phí, sự tinh gọn — chỉ là cải thiện trải nghiệm trên nền một sản phẩm vốn dĩ phải đúng thì mới có lý do tồn tại.
Điều khiến dữ liệu sai trở nên nguy hiểm một cách bất đối xứng, so với hầu hết mọi loại lỗi phần mềm khác, là cách nó phá hủy niềm tin. Một service bị crash thì dễ thấy, được page, được sửa, rồi bị quên đi. Một dashboard âm thầm hiển thị sai số trong ba tuần liền thường không dễ thấy — cho đến khi một stakeholder ra quyết định dựa trên nó, phát hiện lỗi bằng cách khác, và mất niềm tin không chỉ vào dashboard đó mà vào toàn bộ nền tảng dữ liệu phía sau. Một khi một VP đã từng “bị bỏng” bởi một con số sai, phản ứng hợp lý của họ là ngừng tin bất kỳ con số nào từ hệ thống đó, kể cả 99% dashboard vẫn luôn đúng, và bắt đầu yêu cầu một analyst “kiểm tra lại bằng Excel” trước mỗi cuộc họp. Hành vi đó — âm thầm tính lại số bằng tay vì không ai tin nền tảng nữa — là triệu chứng rõ ràng nhất cho thấy niềm tin dữ liệu của một tổ chức đã sụp đổ, và việc xây dựng lại nó cực kỳ tốn kém: niềm tin bị phá hủy chỉ bởi một sự cố tồi tệ, nhưng chỉ được xây lại bằng một quá trình dài, nhàm chán của việc luôn luôn đúng. Đây chính là lý do data quality, data governance và metadata management không phải là những mục “nên có” trong maturity model để làm sau này — chúng là cơ chế giúp một nền tảng dữ liệu giành được và giữ được thứ duy nhất khiến nó có giá trị: được tin tưởng.
Ghi chú này bao quát ba lĩnh vực, cùng nhau giữ cho dữ liệu đáng tin cậy ở quy mô lớn: data quality (dữ liệu có tuân thủ những gì nó phải là không — có accurate, complete, consistent, timely, unique, valid không?), data governance (ai sở hữu dữ liệu này, ai được phép truy cập, và những quy tắc nào áp dụng cho cách nó được tạo ra và sử dụng?), và metadata management (chúng ta biết gì về dữ liệu — schema, lineage, ý nghĩa, độ mới (freshness) của nó — và mọi người có thể tìm ra điều đó không?). Ba lĩnh vực này được trình bày riêng để rõ ràng nhưng thực tế gắn bó chặt chẽ với nhau: bạn không thể governance dữ liệu mà bạn không thể mô tả được, bạn không thể tự động hóa quality check nếu không có metadata về ý nghĩa của “đúng”, và một data catalog thực chất chỉ là metadata management được làm cho có thể tìm kiếm được.
Kiến thức nền tảng
Các chiều (dimension) của Data Quality
“Data quality” không phải là một thuộc tính đúng/sai duy nhất; nó tách thành nhiều chiều độc lập, và một dataset có thể đạt điểm cao ở chiều này trong khi thất bại thảm hại ở chiều khác (một bảng có thể hoàn toàn consistent — tự nhất quán bên trong — trong khi hoàn toàn inaccurate, tức là sai một cách nhất quán). Việc chỉ rõ chiều nào đang gãy là điều biến một lời phàn nàn mơ hồ (“dữ liệu trông có gì đó sai sai”) thành một khẳng định cụ thể, kiểm chứng được, và sửa được.
| Chiều (Dimension) | Định nghĩa | Ví dụ vi phạm |
|---|---|---|
| Accuracy | Dữ liệu phản ánh đúng giá trị hoặc sự kiện trong thế giới thực mà nó đại diện | Trường quốc gia của khách hàng ghi “Canada” nhưng địa chỉ thanh toán và IP geolocation đều chỉ “USA” |
| Completeness | Toàn bộ dữ liệu lẽ ra phải có mặt thì đều có mặt — không thiếu trường bắt buộc hoặc mất bản ghi | Bảng orders thiếu 3.000 dòng của ngày lễ vì một job ingestion upstream âm thầm lỗi vào ngày lượng traffic giảm mà không ai để ý |
| Consistency | Cùng một sự thật khớp nhau giữa các bảng, hệ thống, hoặc trường — không mâu thuẫn | Bảng orders cho thấy tổng chi tiêu trọn đời của một khách hàng là 4.200 USD, nhưng cộng từng dòng đơn hàng trong bảng order_items lại ra 3.850 USD |
| Timeliness | Dữ liệu sẵn sàng và cập nhật trong khoảng thời gian mà consumer cần | Một dashboard “real-time” chống gian lận được nạp dữ liệu bởi một batch job chỉ refresh mỗi 6 giờ, nên nó luôn stale đối với các analyst cần bắt gian lận trong ngày |
| Uniqueness | Không có bản ghi trùng lặp ngoài ý muốn cho cùng một thực thể thực tế | Một job import CRM chạy hai lần do bug retry-không-idempotent, tạo ra hai dòng cho cùng một khách hàng với hai customer ID khác nhau |
| Validity | Dữ liệu tuân thủ các quy tắc cú pháp và miền giá trị (format, kiểu dữ liệu, khoảng giá trị, giá trị cho phép) | Cột discount_percent chứa giá trị 150, về mặt cú pháp là một số nguyên hợp lệ nhưng về mặt ngữ nghĩa là bất khả thi cho một mức giảm giá phần trăm |
Hai trong số các chiều này đáng được lưu ý thêm vì chúng tương tác theo cách dễ gây nhầm lẫn. Accuracy và validity không phải là một: một giá trị có thể hoàn toàn valid (đúng kiểu, đúng format, trong khoảng cho phép) trong khi vẫn inaccurate (sai giá trị so với sự thật thực tế cụ thể đó) — ngày sinh 1990-01-01 là một ngày hoàn toàn hợp lệ, nhưng nó sai nếu người đó thực ra sinh năm 1985. Kiểm tra validity rẻ để tự động hóa vì chỉ cần biết quy tắc (một regex, một kiểu dữ liệu, một khoảng giá trị); kiểm tra accuracy đắt đỏ vì cần một nguồn sự thật độc lập để đối chiếu, đây chính là lý do phần lớn công cụ data quality tự động dựa nặng vào validity, completeness và consistency check, và coi việc xác minh accuracy thực sự là một bài toán khó hơn, thường mang tính thủ công hoặc thống kê.
Quality Check nằm ở đâu trong Pipeline
Quality không phải là một cổng kiểm tra (gate) duy nhất ở cuối pipeline; nó được kiểm tra tại nhiều điểm, và vị trí kiểm tra làm thay đổi hình dạng của thất bại:
- Tại ingestion (schema và validity check) — reject hoặc quarantine một bản ghi ngay khi nó vi phạm một quy tắc đã biết, trước khi nó làm ô nhiễm bất cứ thứ gì downstream. Đây là điểm rẻ nhất để bắt lỗi, vì chưa có gì tiêu thụ dữ liệu xấu cả.
- Giữa pipeline, giữa các bước transform (consistency và referential integrity check) — xác minh rằng một join không âm thầm làm mất dòng, một foreign key vẫn resolve được, số dòng của một phép aggregation khớp kỳ vọng, trước khi chuyển dữ liệu sang bước tiếp theo.
- Tại serving layer, ngay trước khi BI/ML tiêu thụ (completeness, freshness, distribution check) — tuyến phòng thủ cuối cùng: dù có gì đó drift ở phía trên, bắt được nó ở đây ít nhất cũng ngăn nó chạm tới dashboard hoặc model.
Cách tiếp cận nhiều lớp này giống với defense-in-depth trong bảo mật: không có kiểm tra đơn lẻ nào được giả định là đủ, và bắt được vấn đề càng sớm luôn rẻ hơn bắt được nó càng muộn, vì “bán kính ảnh hưởng” (blast radius) của dữ liệu xấu lớn dần theo từng job downstream đã tiêu thụ nó.
Khái niệm chính
Automated Data Quality Checks và Quality Gates
Cơ chế thực tế để thực thi các chiều nói trên là viết automated data quality checks như một phần tường minh, được version hóa của pipeline — không phải một buổi review spreadsheet thủ công mà ai đó làm khi nhớ ra. Các loại check phổ biến là:
- Row count check — lượng dữ liệu load hôm nay có xấp xỉ số dòng kỳ vọng, so với baseline lịch sử hoặc số dòng ở nguồn upstream không? Một cú sụt 90% (hoặc tăng vọt 300%) gần như luôn là bug, không phải sự kiện kinh doanh thật.
- Null check — các trường bắt buộc có thực sự được điền không? Một
customer_idbỗng NULL trên 5% số dòng thường nghĩa là một join đã đổi hoặc một trường upstream bị đổi tên. - Uniqueness constraint — một cột lẽ ra là primary key có thực sự không chứa giá trị trùng lặp không?
- Referential integrity check — mọi giá trị foreign key trong bảng con có thực sự tồn tại trong bảng cha không? (Một
order.customer_idtrỏ tới một khách hàng không tồn tại trong bảngcustomersnghĩa là hoặc một hard deletion đã xảy ra upstream, hoặc có bug thứ tự ingestion.) - Distribution và anomaly check — mean, độ lệch chuẩn, hoặc khoảng giá trị của một cột số có giống dữ liệu hôm nay không, hay trông khác biệt về mặt thống kê so với 30 ngày qua (một tín hiệu kinh điển của việc đổi đơn vị âm thầm ở upstream, ví dụ một trường tiền tệ âm thầm chuyển từ dollar sang cent)?
Hai công cụ thống trị cách các team triển khai những check này trong thực tế, và chúng giải quyết vấn đề từ hai tầng khác nhau của stack:
Great Expectations là một framework bằng Python được xây dựng chuyên cho data quality. Bạn định nghĩa các “Expectation” — các khẳng định khai báo, dễ đọc cho con người như “cột này không bao giờ được null” hoặc “giá trị cột này phải nằm giữa 0 và 100” — trên một dataset, và Great Expectations sẽ validate dữ liệu, cho ra kết quả pass/fail, và có thể tự sinh báo cáo HTML dễ đọc gọi là “Data Docs”. Nó không phụ thuộc engine (chạy được với pandas DataFrame, SQL database, Spark) và thường được chạy như một bước tường minh trong một pipeline được orchestrate (xem 09 — Workflow Orchestration), thường qua tích hợp Airflow của nó.
import great_expectations as gx
context = gx.get_context()
validator = context.sources.pandas_default.read_csv("orders.csv")
validator.expect_column_values_to_not_be_null("customer_id")
validator.expect_column_values_to_be_unique("order_id")
validator.expect_column_values_to_be_between("discount_percent", min_value=0, max_value=100)
results = validator.validate()
if not results.success:
raise ValueError("Data quality checks failed — halting pipeline")
dbt tests giải quyết cùng vấn đề nhưng sống ngay bên trong transformation layer, được định nghĩa khai báo bằng YAML ngay cạnh model mà chúng kiểm tra — nghĩa là quy tắc quality và transformation mà nó bảo vệ sống trong cùng một codebase, được version-control và review cùng nhau. dbt có sẵn bốn generic test dựng sẵn (not_null, unique, accepted_values, relationships), và package cộng đồng dbt-utils cùng dbt-expectations mở rộng thêm distribution và statistical check, về cơ bản mang phong cách khẳng định của Great Expectations vào cú pháp YAML của dbt.
# models/schema.yml
models:
- name: orders
columns:
- name: order_id
tests:
- unique
- not_null
- name: customer_id
tests:
- not_null
- relationships:
to: ref('customers')
field: customer_id
- name: discount_percent
tests:
- dbt_utils.accepted_range:
min_value: 0
max_value: 100
Chạy dbt test thực thi mọi test đã khai báo dưới dạng truy vấn SQL trên các model đã build và báo cáo cái nào thất bại — và quan trọng hơn, dbt build chạy test xen kẽ với việc build model, nên một test thất bại trên một model upstream có thể ngăn các model downstream phụ thuộc vào nó được build với dữ liệu xấu bên dưới ngay từ đầu.
Đây chính là bản chất của một quality gate: một điểm kiểm tra, được nối trực tiếp vào luồng điều khiển của pipeline, chặn dữ liệu tiến sang bước tiếp theo — và do đó không bao giờ chạm tới dashboard hay consumer downstream — nếu nó không đạt một check đã định nghĩa. Quality gate là một lựa chọn thiết kế có chủ đích để ưu tiên một thất bại rõ ràng, ồn ào (một pipeline run thất bại, một kỹ sư on-call bị page) hơn một thất bại thầm lặng, vô hình (một con số sai âm thầm được đưa tới dashboard của một VP). Sự đánh đổi này gần như luôn đúng: một pipeline hỏng thì phiền một ngày; một con số sai được tin tưởng suốt một tháng là một sự cố niềm tin.
Schema Drift: Kẻ giết người thầm lặng
Schema drift là hiện tượng xảy ra khi một hệ thống nguồn upstream thay đổi shape dữ liệu của nó — một cột bị đổi tên, một kiểu dữ liệu đổi từ integer sang string, một trường bị xóa, một trường bắt buộc mới xuất hiện — mà không có sự phối hợp nào với các team tiêu thụ dữ liệu đó ở downstream. Đây là một trong những failure mode nguy hiểm nhất trong data engineering chính vì nó có hai kiểu biểu hiện rất khác nhau, và kiểu thầm lặng còn tệ hơn nhiều so với kiểu ồn ào:
- Schema drift ồn ào: một cột bị xóa hoặc đổi tên, một truy vấn downstream tham chiếu tới nó, và pipeline ném exception rồi fail hẳn. Điều này tệ, nhưng đây là phiên bản tốt của vấn đề — thất bại rõ ràng, kỹ sư on-call bị page, và không ai tiêu thụ dữ liệu sai vì pipeline đơn giản là dừng lại.
- Schema drift thầm lặng: kiểu dữ liệu của một cột thay đổi theo cách không gây lỗi nhưng làm thay đổi ý nghĩa — một trường
amountupstream chuyển từ dollar sang cent, hoặc một trường boolean vốn dùngtrue/falsebắt đầu gửi1/0/nulltrong khinulltrước đây có nghĩa là “false” nhưng bây giờ thực sự có nghĩa là “không xác định”. Không có gì crash cả. Pipeline chạy xanh. Các con số chỉ đơn giản là sai, và chúng vẫn sai cho tới khi ai đó downstream nhận ra một báo cáo không khớp — điều này, theo phần Tổng quan ở trên, chính xác là cơ chế khiến niềm tin bị xói mòn.
Ba chiến lược cụ thể giảm thiểu schema drift:
- Schema validation tại ingestion — thực thi một schema kỳ vọng tường minh (qua JSON Schema, Avro schema, hoặc check của Great Expectations/dbt) ngay tại thời điểm dữ liệu đi vào pipeline, để một shape bất ngờ được bắt tại điểm sớm nhất, rẻ nhất thay vì lan truyền tiếp.
- Cảnh báo khi có thay đổi schema bất ngờ — các công cụ như Monte Carlo và các nền tảng catalog/observability hiện đại có thể diff schema của một bảng qua từng lần chạy và page một data engineer ngay khi một cột xuất hiện, biến mất, hoặc đổi kiểu, thậm chí trước khi bất kỳ test tường minh nào được viết để bắt riêng thay đổi đó.
- Contract testing với các team upstream — cách sửa bền vững nhất mang tính tổ chức, không phải kỹ thuật: thiết lập một data contract tường minh với team hoặc hệ thống tạo ra dữ liệu nguồn — một schema được version hóa, công bố, mà upstream cam kết không phá vỡ mà không thông báo trước — để một thay đổi schema trở thành một sự kiện được thương lượng, thông báo, thay vì một bất ngờ được phát hiện bởi một pipeline hỏng ba bước phía sau. Đây là cùng ý tưởng với một API contract, áp dụng cho dữ liệu thay vì endpoint.
Data Lineage
Data lineage là bản ghi cho biết một mẩu dữ liệu đến từ đâu và đã trải qua những phép biến đổi nào để đạt tới hình dạng hiện tại — về bản chất, một đồ thị có hướng nối mọi bảng nguồn, bước transformation, và bảng hoặc dashboard downstream. Lineage trả lời hai câu hỏi mà nếu không có nó thì tốn kém hoặc bất khả thi nếu chỉ dựa vào việc kiểm tra thủ công:
- Impact analysis, hỏi theo chiều tiến: “Nếu tôi thay đổi hoặc xóa cột nguồn này, cái gì sẽ hỏng ở downstream?” Không có lineage, trả lời câu này nghĩa là hoặc grep qua từng file SQL trong codebase hy vọng tìm ra mọi tham chiếu, hoặc đơn giản là thực hiện thay đổi rồi chờ xem ai phàn nàn — cả hai đều chậm và rủi ro ở bất kỳ quy mô thực tế nào.
- Root-cause debugging, hỏi theo chiều lùi: “Con số này trên dashboard điều hành bị sai — nó đến từ đâu, và tại bước biến đổi nào trong mười lăm bước giữa nguồn thô và dashboard này thì nó bị sai?” Lineage biến câu này từ một cuộc điều tra pháp y kéo dài nhiều ngày thành một phép duyệt đồ thị (graph traversal).
- Compliance và audit: các quy định như GDPR và các cuộc kiểm toán ngành thường xuyên yêu cầu tổ chức chứng minh chính xác một mẩu dữ liệu cá nhân đã chảy qua hệ thống của họ như thế nào, những phép biến đổi nào đã chạm vào nó, và hiện nó nằm ở đâu — lineage là chứng cứ trả lời trực tiếp yêu cầu đó (xem 15 — Data Security & Compliance).
Lineage có thể được ghi nhận theo hai cách, và sự khác biệt về độ tin cậy giữa chúng là rất rõ rệt:
- Automated lineage, được suy ra trực tiếp từ code chạy pipeline. dbt là ví dụ chủ đạo rõ ràng nhất: vì SQL của mỗi model khai báo tường minh các dependency upstream của nó qua
ref()vàsource(), dbt có thể dựng nên một DAG (lineage graph) chính xác, luôn cập nhật của mọi model mà không tốn thêm công sức tài liệu hóa — lineage chính là code, nên nó không thể lỗi thời so với những gì thực sự chạy. Các công cụ data catalog và observability chuyên dụng (Atlan, Collibra, DataHub, các nền tảng dựa trên OpenLineage) mở rộng điều này xa hơn bằng cách parse query log hoặc instrument các orchestrator để dựng lineage xuyên qua các hệ thống mà dbt không chạm tới — job ingestion, job Spark, truy vấn của công cụ BI — khâu nối thành một lineage graph liên hệ thống thay vì chỉ giới hạn trong một công cụ. - Tài liệu được duy trì thủ công — một trang wiki, một sơ đồ, một spreadsheet mà ai đó vẽ mô tả dữ liệu chảy như thế nào. Điều này mong manh theo cách dễ bị đánh giá thấp: nó đúng vào ngày được viết ra và bắt đầu xuống cấp ngay khi bất kỳ pipeline nào thay đổi, vì không có gì buộc sơ đồ phải được cập nhật đồng bộ với code. Tài liệu lineage được duy trì thủ công, trong thực tế, là một bức ảnh chụp cách pipeline hoạt động tại một thời điểm nào đó trong quá khứ, và khoảng cách giữa bức ảnh đó và thực tế chỉ có tăng lên.
Hàm ý thực tiễn rất rõ ràng: ưu tiên lineage được suy ra từ các artifact thực thi thực tế (code, query log, metadata của orchestrator) hơn là lineage được viết ra bởi con người, vì lineage suy ra không thể trôi lệch (drift) khỏi pipeline mà nó mô tả, trong khi lineage viết tay thì cuối cùng luôn trôi lệch.
Metadata Management và Data Catalog
Metadata là dữ liệu về dữ liệu, và nó tách rõ ràng thành ba loại trả lời ba câu hỏi khác nhau:
| Loại metadata | Nó nắm bắt điều gì | Ví dụ |
|---|---|---|
| Technical metadata | Schema, kiểu dữ liệu, tên bảng/cột, partitioning, storage format | orders.order_date là cột TIMESTAMP, được partition theo ngày, lưu dưới dạng Parquet |
| Business metadata | Ý nghĩa cho con người, định nghĩa, ownership, thuật ngữ trong glossary | ”Khách hàng đang hoạt động” (active customer) được định nghĩa là: có giao dịch mua trong 90 ngày gần nhất, theo mục glossary của team Marketing, thuộc quyền sở hữu của domain Growth |
| Operational metadata | Freshness, lịch sử chạy, SLA, khối lượng dữ liệu theo thời gian, lịch sử pass/fail của quality check | Bảng orders được refresh lần cuối 14 phút trước, pipeline của nó có tỷ lệ thành công 99,2% trong 90 ngày qua, và SLA của nó là “cập nhật trong vòng 1 giờ” |
Bất kỳ loại nào trong ba loại này đứng riêng lẻ đều hữu ích nhưng không đầy đủ: biết kiểu dữ liệu của một cột (technical) không cho bạn biết bạn có được phép dùng nó trong một báo cáo hướng khách hàng hay không (business/ownership); biết định nghĩa kinh doanh không cho bạn biết dữ liệu đằng sau nó có thực sự mới ngay bây giờ hay không (operational). Một data catalog — các công cụ như DataHub, Amundsen, Atlan, hoặc Collibra — tồn tại để làm giao diện tìm kiếm duy nhất gắn kết cả ba loại lại với nhau: một analyst tìm kiếm “customer churn” nên tìm ra đúng bảng, thấy schema của nó, thấy ai sở hữu nó, thấy định nghĩa kinh doanh của nó, thấy nó có mới hay không, và thấy nó có đang pass quality check hay không, tất cả trong một lần tra cứu thay vì năm tin nhắn Slack riêng biệt tới năm team khác nhau. Đây là điều khiến một catalog vừa là công cụ governance (nó là bề mặt thực thi cho ownership và access policy) vừa là công cụ discoverability (nó trả lời “chúng ta có dữ liệu gì, và tôi có thể tin nó không?”) cùng một lúc — hai mục đích chia sẻ cùng một nền tảng metadata bên dưới, chính là lý do chúng thường được xây dựng thành một sản phẩm thay vì hai.
Data Interoperability: Schema và Schema Registry
Hình thức quan trọng nhất về mặt vận hành của metadata chính là schema, và các hệ thống dữ liệu hiện đại ngày càng đưa schema đó ra bên ngoài thành một hợp đồng (contract) chung, máy đọc được, thay vì để nó ngầm hiểu trong bất kỳ code nào tình cờ tạo ra hoặc tiêu thụ dữ liệu. Apache Parquet và Apache Avro đều nhúng thông tin schema trực tiếp cùng với dữ liệu chúng lưu trữ — Parquet là một columnar format tối ưu cho đọc phân tích, Avro là một binary format gọn nhẹ với hỗ trợ mạnh cho schema evolution, thường dùng cho dữ liệu streaming theo hàng (row-oriented). JSON Schema cung cấp hợp đồng tương đương cho payload JSON, cho phép producer và consumer thống nhất về trường bắt buộc, kiểu dữ liệu, và ràng buộc mà không bên nào cần kiểm tra code của bên kia.
Điều này quan trọng nhất trong các hệ thống streaming, nơi một producer và hàng chục consumer độc lập có thể không bao giờ phối hợp deployment trực tiếp với nhau. Một schema registry — phổ biến nhất là Confluent Schema Registry dùng cùng Apache Kafka — nằm giữa producer và consumer và thực thi rằng mọi message publish lên một topic phải tuân theo một schema đã đăng ký (bằng Avro, Protobuf, hoặc JSON Schema) trước khi được phép vào topic đó. Quan trọng hơn, registry còn thực thi các quy tắc tương thích schema (compatibility rules) khi evolution: một producer muốn thay đổi schema của một topic (thêm trường, đổi kiểu) phải làm theo cách được khai báo là tương thích với các consumer hiện có (ví dụ, tương thích BACKWARD yêu cầu schema mới vẫn phải đọc được bởi consumer đang dùng schema cũ, thường bằng cách yêu cầu trường mới phải có giá trị mặc định). Điều này biến schema drift, trong bối cảnh streaming, từ một thất bại runtime thầm lặng thành một lần ghi bị từ chối ngay tại thời điểm publish — chính xác cùng nguyên tắc “bắt lỗi tại điểm sớm nhất, rẻ nhất” từ validation lúc ingestion đã bàn ở trên, áp dụng riêng cho trường hợp streaming.
Governance như một thực hành Socio-Technical
Rất dễ để coi governance là một vấn đề công cụ — mua một catalog, cấu hình vài access control list, xong. Trong thực tế, governance hiệu quả mang tính socio-technical: công cụ là cần thiết nhưng không đủ, vì những vấn đề khó thực sự mang tính tổ chức. Ai chịu trách nhiệm khi dữ liệu của một bảng trở nên stale — có một owner được chỉ định rõ ràng, hay trách nhiệm bị pha loãng cho không ai cả? Chính sách truy cập nào áp dụng cho một cột chứa PII, và ai quyết định điều đó? Một “thuật ngữ glossary” như “active customer” có thực sự được thống nhất giữa Marketing, Sales, và Finance, hay mỗi team âm thầm duy trì một định nghĩa hơi khác nhau khiến ra ba con số khác nhau cho cùng một metric trên ba dashboard khác nhau?
Đây chính xác là vấn đề mang tính tổ chức mà phần thảo luận về data mesh trong 07 — Data Lakes & Modern Architectures đề cập từ góc độ kiến trúc: governance thành công khi nó federated (phân tán, liên bang) — các domain team sở hữu quality, tài liệu, và quy tắc truy cập cho dữ liệu họ tạo ra, vì họ là những người duy nhất có đủ context để làm điều đó tốt — thay vì là một chức năng trung tâm thuần túy, nơi một team governance cố hiểu và gác cổng (gatekeep) mọi dataset trong công ty và không thể tránh khỏi việc trở thành cả điểm nghẽn lẫn một con dấu cao su (rubber stamp, phê duyệt những thứ mà nó không có đủ context để thực sự đánh giá). Nguyên tắc thứ tư của data mesh theo Dehghani, federated computational governance, chính xác là điều này: các quy tắc toàn cục (một trường PII phải luôn được mask, mọi dataset phải có một owner được chỉ định, mọi bảng được publish phải pass quality gate) được thống nhất ở cấp trung tâm rồi được thực thi tự động, mang tính tính toán (computational), tại tầng platform — thay vì thủ công, bởi một con người review từng dataset trước khi nó được ship. Metadata là thứ khiến việc thực thi này khả thi ngay từ đầu: bạn không thể thực thi mang tính tính toán quy tắc “PII phải được mask” trừ khi có gì đó trong tầng metadata của bạn thực sự đánh dấu cột nào chứa PII ngay từ đầu. Đây là ý nghĩa mà quality, governance, và metadata là một hệ thống liên kết duy nhất chứ không phải ba sáng kiến riêng biệt: metadata là nền tảng, quality check là các test tự động dựa trên metadata, và governance là việc thực thi chính sách dựa trên metadata, tất cả đều dựa trên cùng một nền tảng cơ bản “chúng ta biết gì về dữ liệu này”. Việc kiểm thử toàn bộ hệ thống này đầu-cuối, bao gồm cả các pipeline thực thi những check này, được bàn ở 16 — Testing for Data Pipelines.
Best Practices
Bảng dưới đây cô đọng các chiều từ phần Kiến thức nền tảng thành check tự động cụ thể thường bắt được từng loại vi phạm trong thực tế — cách ánh xạ mà kỹ sư thực sự dùng khi quyết định triển khai cái gì trước.
| Chiều data quality | Ví dụ vi phạm | Check tự động điển hình |
|---|---|---|
| Accuracy | Trường quốc gia của khách hàng mâu thuẫn với địa chỉ thanh toán / IP geolocation | Check đối chiếu chéo nguồn (cross-source reconciliation) với một tham chiếu độc lập (khó tự động hóa hơn; thường là một query audit định kỳ thay vì test theo từng dòng) |
| Completeness | Ingestion upstream âm thầm mất dòng vào ngày lễ, bảng orders thiếu 3.000 dòng | Row count check so với baseline lịch sử hoặc số dòng ở nguồn upstream |
| Consistency | Tổng chi tiêu trong orders mâu thuẫn với tổng của order_items | Check đối chiếu / aggregate liên bảng (dbt_utils.equality, hoặc một assertion SQL tùy chỉnh) |
| Timeliness | Dashboard “real-time” được nạp bởi một batch job refresh mỗi 6 giờ | Freshness check trên max(loaded_at) so với ngưỡng SLA (dbt source freshness của dbt, hoặc bộ giám sát freshness cấp catalog) |
| Uniqueness | Dòng khách hàng trùng lặp từ một import bị retry không idempotent | Test unique trên cột khóa tự nhiên hoặc surrogate key |
| Validity | Cột discount_percent chứa giá trị 150 | Check range/accepted_values, hoặc ràng buộc JSON Schema / Avro schema được thực thi tại ingestion |
Ngoài bảng ánh xạ đó, một số thực hành luôn phân biệt các team giữ được niềm tin của stakeholder với các team đánh mất nó:
- Đặt quality gate càng sớm càng tốt trong pipeline, và làm cho thất bại rõ ràng, ồn ào. Một pipeline run thất bại khiến ai đó bị page là một kết quả tốt so với một con số sai thầm lặng đến được dashboard; thiết kế theo hướng “fail nhanh và rõ ràng”, không phải “không bao giờ fail”.
- Viết test ngay cạnh phép transformation mà nó bảo vệ, không phải như một suy nghĩ muộn màng. Mô hình của dbt là đặt
tests:YAML cạnh định nghĩa model tồn tại chính xác để quy tắc quality và logic mà nó bảo vệ được review, version hóa, và cập nhật cùng nhau — một test được viết sáu tháng sau model, bởi người khác, có khả năng cao hơn nhiều sẽ trôi lệch khỏi những gì model thực sự làm. - Ưu tiên lineage suy ra (derived) hơn tài liệu duy trì thủ công, mọi lúc. Một sơ đồ lineage không được sinh ra từ code pipeline thực tế hoặc query log là một bức ảnh chụp bắt đầu xuống cấp ngay khi nó được vẽ ra.
- Chỉ định một owner rõ ràng, được nêu tên, cho mọi dataset mà các team khác phụ thuộc vào. “Thuộc sở hữu của team data” không phải là một owner; nó pha loãng trách nhiệm cho không ai cả. Một owner cụ thể là người bị page, người ký duyệt thay đổi schema, và người chịu trách nhiệm cho SLA chất lượng của dữ liệu.
- Thiết lập data contract với các team upstream cho bất cứ thứ gì bạn phụ thuộc vào mà bạn không kiểm soát. Một thỏa thuận bằng lời rằng “chúng tôi sẽ không đổi trường đó” không phải là một contract; một schema được version hóa, được test, mà upstream cam kết và CI thực thi, mới là.
- Phân tán (federate) governance cho các domain owner thay vì tập trung nó vào một team gác cổng duy nhất, và đầu tư nỗ lực trung tâm còn lại vào platform, catalog, và việc thực thi chính sách mang tính tính toán khiến việc phân tán trở nên an toàn — đây là mô hình metadata-first, phù hợp với data mesh đã mô tả ở trên, không phải một quyết định thuần túy về công cụ.
- Coi data catalog là một sản phẩm với vấn đề adoption của riêng nó, không phải một công việc thiết lập một lần. Một catalog mà không ai cập nhật hoặc tìm kiếm sẽ xuống cấp thành đúng loại tài liệu lỗi thời mà nó vốn được tạo ra để thay thế; tài liệu hóa ownership và freshness cần là một bước bắt buộc khi ship một dataset mới, không phải một việc tùy chọn làm sau.
- Dùng schema registry cho bất kỳ pipeline streaming nào có nhiều hơn một consumer. Chi phí thực thi tính tương thích tại thời điểm publish thấp hơn nhiều so với chi phí debug một consumer bị hỏng thầm lặng sau ba lần deploy kể từ khi producer đổi kiểu một trường.
Tài liệu tham khảo
- Great Expectations — Documentation
- dbt — Data Tests Documentation
- dbt — Lineage / DAG Documentation
- Confluent — Schema Registry Documentation
- JSON Schema — Official Specification
- Zhamak Dehghani, How to Move Beyond a Monolithic Data Lake to a Distributed Data Mesh (martinfowler.com, 2019)
- DataHub — Open Source Metadata Platform Documentation
- Monte Carlo — Data Observability
- roadmap.sh — Data Engineer Roadmap
Part of the Data Engineer Roadmap knowledge base.
Overview
Every other note in this knowledge base — ingestion (see 03 — Data Sources & Ingestion), pipelines (see 08 — ETL, ELT & Data Pipelines), warehouses and lakes (see 06 — Data Modeling & Warehousing and 07 — Data Lakes & Modern Architectures) — is, in the end, in service of one outcome: a person or a system makes a decision based on a number a pipeline produced. If that number is wrong, none of the engineering elegance upstream of it matters. A pipeline that runs on schedule, scales effortlessly, costs nothing to operate, and is a joy to read — but silently delivers the wrong revenue figure — is strictly worse than an ugly, slow, embarrassing pipeline that delivers the right one. This is the oldest aphorism in computing, “garbage in, garbage out,” but it is worth restating in its sharpest form for data engineering specifically: correctness is not one quality attribute among many; it is the product. Everything else — latency, cost, elegance — is a quality-of-life improvement on top of a product that has to be correct to exist at all.
What makes bad data uniquely dangerous, compared to almost any other class of software bug, is how asymmetrically it destroys trust. A crashed service is visible, gets paged, gets fixed, and is forgotten. A dashboard that quietly shows the wrong number for three weeks is often not visible — until a business stakeholder makes a decision on it, discovers the error some other way, and loses confidence not just in that one dashboard but in the entire data platform behind it. Once a VP has been burned by a wrong number, the rational response is to stop trusting any number from that system, including the 99% of dashboards that were always correct, and start asking an analyst to “double check it in Excel” before every meeting. That behavior — quietly re-deriving numbers by hand because nobody trusts the platform — is the single clearest symptom that an organization’s data trust has collapsed, and it is enormously expensive to rebuild once lost: trust is destroyed by one bad incident and rebuilt only by a long, boring track record of being right. This is why data quality, governance, and metadata management are not “nice to have” maturity-model line items to get to eventually — they are the mechanism by which a data platform earns and keeps the only thing that makes it useful at all: being believed.
This note covers the three disciplines that, together, keep data trustworthy at scale: data quality (does the data conform to what it should — is it accurate, complete, consistent, timely, unique, valid?), data governance (who owns this data, who may access it, and what rules apply to how it is produced and used?), and metadata management (what do we know about the data — its schema, its lineage, its meaning, its freshness — and can people find that out?). They are presented separately for clarity but are deeply intertwined in practice: you cannot govern data you cannot describe, you cannot automate quality checks without metadata about what “correct” means, and a data catalog is really just metadata management made searchable.
Fundamentals
The Dimensions of Data Quality
“Data quality” is not a single yes/no property; it decomposes into several independent dimensions, and a dataset can score well on one while failing badly on another (a table can be perfectly consistent — internally self-coherent — while being completely inaccurate, i.e., consistently wrong). Being precise about which dimension is broken is what turns a vague complaint (“the data looks off”) into a specific, testable, fixable claim.
| Dimension | Definition | Example violation |
|---|---|---|
| Accuracy | Data correctly reflects the real-world value or event it represents | A customer’s country field says “Canada” but their billing address and IP geolocation both say “USA” |
| Completeness | All data that should be present, is present — no missing required fields or dropped records | An orders table is missing 3,000 rows for July 4th because an upstream ingestion job silently failed on a holiday with reduced traffic and nobody noticed the row count drop |
| Consistency | The same fact agrees across tables, systems, or fields — no contradictions | The orders table shows a customer’s total lifetime spend as $4,200, but summing their individual order rows in the order_items table yields $3,850 |
| Timeliness | Data is available and up to date within the window consumers need | A “real-time” fraud dashboard is fed by a batch job that only refreshes every 6 hours, so it is reliably stale for the analysts relying on it to catch same-day fraud |
| Uniqueness | No unintended duplicate records for the same real-world entity | A CRM import runs twice due to a retry-without-idempotency bug, creating two rows for the same customer with two different customer IDs |
| Validity | Data conforms to the syntactic and domain rules it’s supposed to (format, type, range, allowed values) | A discount_percent column contains a value of 150, which is syntactically a valid integer but semantically impossible for a percentage discount |
Two of these dimensions deserve a further note because they interact in a way that trips people up. Accuracy and validity are not the same thing: a value can be perfectly valid (right type, right format, within an allowed range) while still being inaccurate (the wrong value for that specific real-world fact) — a birth date of 1990-01-01 is a perfectly valid date, but it’s wrong if the person was actually born in 1985. Validity checks are cheap to automate because they only require knowing the rule (a regex, a type, a range); accuracy checks are expensive because they require an independent source of truth to compare against, which is exactly why most automated data quality tooling leans heavily on validity, completeness, and consistency checks, and treats true accuracy verification as a harder, often manual or statistical, problem.
Where Quality Checks Live in the Pipeline
Quality is not a single gate at the end of a pipeline; it is checked at multiple points, and where you check it changes what failure looks like:
- At ingestion (schema and validity checks) — reject or quarantine a record the moment it violates a known rule, before it pollutes anything downstream. This is the cheapest place to catch a problem, because nothing has consumed the bad data yet.
- Mid-pipeline, between transformation stages (consistency and referential integrity checks) — verify that a join didn’t silently drop rows, that a foreign key still resolves, that an aggregation’s row count matches expectations, before passing data to the next stage.
- At the serving layer, just before BI/ML consumption (completeness, freshness, distribution checks) — the last line of defense: even if something upstream drifted, catching it here at least stops it from reaching a dashboard or a model.
This layered approach mirrors defense-in-depth in security: no single check is assumed sufficient, and catching an issue earlier is always cheaper than catching it later, because the “blast radius” of bad data grows with every downstream job that has already consumed it.
Key Concepts
Automated Data Quality Checks and Quality Gates
The practical mechanism for enforcing the dimensions above is to write automated data quality checks as an explicit, versioned part of the pipeline — not a manual spreadsheet review someone does when they remember to. The common check types are:
- Row count checks — did today’s load produce roughly the expected number of rows, compared to a historical baseline or an upstream source count? A sudden 90% drop (or 300% spike) is almost always a bug, not a real business event.
- Null checks — are required fields actually populated? A
customer_idthat’s suddenly NULL on 5% of rows usually means a join changed or an upstream field was renamed. - Uniqueness constraints — does a column that should be a primary key actually contain no duplicates?
- Referential integrity checks — does every foreign key value in a child table actually exist in the parent table? (An
order.customer_idthat references a customer that doesn’t exist in thecustomerstable means either a hard deletion happened upstream or an ingestion ordering bug.) - Distribution and anomaly checks — does a numeric column’s mean, standard deviation, or range look like today’s data, or does it look statistically different from the last 30 days (a classic signal of an upstream unit change, e.g., a currency field silently switching from dollars to cents)?
Two tools dominate how teams implement these checks in practice, and they solve the problem from different layers of the stack:
Great Expectations is a Python-based framework purpose-built for data quality. You define “Expectations” — declarative, human-readable assertions like “this column’s values must never be null” or “this column’s values must be between 0 and 100” — against a dataset, and Great Expectations validates the data, produces a pass/fail result, and can auto-generate human-readable HTML “Data Docs” reporting the outcome. It’s engine-agnostic (works against pandas DataFrames, SQL databases, Spark) and is typically run as an explicit step in an orchestrated pipeline (see 09 — Workflow Orchestration), often via its Airflow integration.
import great_expectations as gx
context = gx.get_context()
validator = context.sources.pandas_default.read_csv("orders.csv")
validator.expect_column_values_to_not_be_null("customer_id")
validator.expect_column_values_to_be_unique("order_id")
validator.expect_column_values_to_be_between("discount_percent", min_value=0, max_value=100)
results = validator.validate()
if not results.success:
raise ValueError("Data quality checks failed — halting pipeline")
dbt tests solve the same problem but live natively inside the transformation layer, defined declaratively in YAML right alongside the models they check — which means the quality rule and the transformation it protects live in the same codebase, version-controlled and reviewed together. dbt ships four built-in generic tests out of the box (not_null, unique, accepted_values, relationships), and the community package dbt-utils and dbt-expectations extend this with distribution and statistical checks, effectively bringing Great-Expectations-style assertions into dbt’s YAML syntax.
# models/schema.yml
models:
- name: orders
columns:
- name: order_id
tests:
- unique
- not_null
- name: customer_id
tests:
- not_null
- relationships:
to: ref('customers')
field: customer_id
- name: discount_percent
tests:
- dbt_utils.accepted_range:
min_value: 0
max_value: 100
Running dbt test executes every declared test as a SQL query against the built models and reports which ones failed — and critically, dbt build runs tests interleaved with model builds, so a failing test on an upstream model can stop its downstream dependents from being built with bad data underneath them at all.
This is the essence of a quality gate: a checkpoint, wired directly into the pipeline’s control flow, that halts data from progressing to the next stage — and therefore from ever reaching a dashboard or a downstream consumer — if it fails a defined check. A quality gate is a deliberate design choice to prefer a visible, loud failure (a failed pipeline run, a paged on-call engineer) over a silent, invisible one (a wrong number quietly shipped to a VP’s dashboard). This trade-off is almost always correct: a broken pipeline is annoying for a day; a wrong number believed for a month is a trust incident.
Schema Drift: The Silent Killer
Schema drift is what happens when an upstream source system changes its data shape — a column gets renamed, a type changes from integer to string, a field is dropped, a new required field appears — without any coordination with the teams consuming that data downstream. It is one of the most dangerous failure modes in data engineering precisely because it has two very different failure signatures, and the quiet one is far worse than the loud one:
- Loud schema drift: a column is dropped or renamed, a downstream query references it, and the pipeline throws an exception and fails outright. This is bad, but it’s the good version of the problem — the failure is visible, the on-call engineer gets paged, and nobody consumes wrong data because the pipeline simply stopped.
- Silent schema drift: a column’s type changes in a way that doesn’t error but changes meaning — an upstream
amountfield switches from dollars to cents, or a boolean field that used to betrue/falsestarts sending1/0/nullwherenullused to mean “false” but now genuinely means “unknown.” Nothing crashes. The pipeline runs green. The numbers are simply wrong, and they stay wrong until someone downstream notices a report doesn’t reconcile — which, per the Overview above, is exactly the mechanism by which trust erodes.
Three concrete strategies mitigate schema drift:
- Schema validation on ingestion — enforce an explicit expected schema (via a JSON Schema, an Avro schema, or a Great Expectations/dbt check) at the moment data enters the pipeline, so an unexpected shape is caught at the earliest, cheapest point rather than propagating.
- Alerting on unexpected schema changes — tools like Monte Carlo and modern catalog/observability platforms can diff a table’s schema run-over-run and page a data engineer the moment a column appears, disappears, or changes type, even before any explicit test was written to catch that specific change.
- Contract testing with upstream teams — the most durable fix is organizational, not technical: establish an explicit data contract with the team or system producing the source data — a versioned, published schema that upstream commits not to break without notice — so a schema change becomes a negotiated, communicated event instead of a surprise discovered by a broken pipeline three hops downstream. This is the same idea as an API contract, applied to data instead of endpoints.
Data Lineage
Data lineage is the record of where a piece of data came from and every transformation it passed through to reach its current form — effectively, a directed graph connecting every source table, transformation step, and downstream table or dashboard. Lineage answers two questions that are otherwise expensive or impossible to answer by inspection alone:
- Impact analysis, asked forward: “If I change or drop this source column, what breaks downstream?” Without lineage, answering this means either grepping through every SQL file in the codebase hoping you find every reference, or simply making the change and waiting to see who complains — both are slow and risky at any real scale.
- Root-cause debugging, asked backward: “This number on the executive dashboard is wrong — where did it come from, and at which of the fifteen transformation steps between the raw source and this dashboard did it go wrong?” Lineage turns this from a multi-day forensic investigation into a graph traversal.
- Compliance and audit: regulations like GDPR and industry audits routinely require organizations to demonstrate exactly how a piece of personal data flowed through their systems, what transformations touched it, and where it currently lives — lineage is the artifact that answers that demand directly (see 15 — Data Security & Compliance).
Lineage can be captured two ways, and the difference in reliability between them is stark:
- Automated lineage, derived directly from the code that runs the pipeline. dbt is the clearest mainstream example: because every model’s SQL explicitly declares its upstream dependencies via
ref()andsource(), dbt can construct an exact, always-current DAG (lineage graph) of every model with zero additional documentation effort — the lineage is the code, so it cannot go stale relative to what actually runs. Dedicated data catalog and observability tools (Atlan, Collibra, DataHub, OpenLineage-based platforms) extend this further by parsing SQL query logs or instrumenting orchestrators to build lineage across systems dbt doesn’t touch — ingestion jobs, Spark jobs, BI tool queries — stitching together a cross-system lineage graph rather than one scoped to a single tool. - Manually maintained documentation — a wiki page, a diagram, a spreadsheet someone draws describing how data flows. This is fragile in a way that’s easy to underestimate: it is correct on the day it’s written and begins decaying the moment any pipeline changes, because nothing forces the diagram to be updated in lockstep with the code. Manually maintained lineage documentation is, in practice, a snapshot of how the pipeline worked at some point in the past, and the gap between that snapshot and reality only ever grows.
The practical implication is unambiguous: prefer lineage that is derived from the actual execution artifacts (code, query logs, orchestrator metadata) over lineage that is authored by a human, because derived lineage cannot drift out of sync with the pipeline it describes, while authored lineage always eventually does.
Metadata Management and Data Catalogs
Metadata is data about data, and it splits cleanly into three categories that answer three different questions:
| Metadata type | What it captures | Example |
|---|---|---|
| Technical metadata | Schema, data types, table/column names, partitioning, storage format | orders.order_date is a TIMESTAMP column, partitioned by day, stored as Parquet |
| Business metadata | Human meaning, definitions, ownership, glossary terms | ”Active customer” is defined as: made a purchase in the trailing 90 days, per the Marketing team’s glossary entry, owned by the Growth domain |
| Operational metadata | Freshness, run history, SLAs, data volume over time, quality check pass/fail history | The orders table was last refreshed 14 minutes ago, its pipeline has a 99.2% success rate over the last 90 days, and its SLA is “fresh within 1 hour” |
Any one of these categories alone is useful but incomplete: knowing a column’s type (technical) doesn’t tell you whether you’re allowed to use it in a customer-facing report (business/ownership); knowing a business definition doesn’t tell you whether the data behind it is actually fresh right now (operational). A data catalog — tools like DataHub, Amundsen, Atlan, or Collibra — exists to be the single searchable interface that ties all three together: an analyst searching for “customer churn” should find the right table, see its schema, see who owns it, see its business definition, see whether it’s fresh, and see whether it’s passing its quality checks, all in one lookup instead of five separate Slack messages to five different teams. This is what makes a catalog both a governance tool (it’s the enforcement surface for ownership and access policy) and a discoverability tool (it answers “what data do we have, and can I trust it?”) at the same time — the two purposes share the same underlying metadata, which is precisely why they’re usually built as one product rather than two.
Data Interoperability: Schemas and Schema Registries
Metadata’s most operationally critical form is the schema itself, and modern data systems increasingly externalize that schema into a shared, machine-readable contract rather than leaving it implicit in whatever code happens to produce or consume the data. Apache Parquet and Apache Avro both embed schema information directly alongside the data they store — Parquet as a columnar format optimized for analytical reads, Avro as a compact binary format with strong support for schema evolution, commonly used for row-oriented streaming data. JSON Schema provides the equivalent contract for JSON payloads, letting a producer and consumer agree on required fields, types, and constraints without either side needing to inspect the other’s code.
This matters most in streaming systems, where a producer and dozens of independent consumers may never coordinate deployments directly. A schema registry — most commonly the Confluent Schema Registry used with Apache Kafka — sits between producers and consumers and enforces that every message published to a topic conforms to a registered schema (in Avro, Protobuf, or JSON Schema) before it’s allowed onto the topic at all. Crucially, the registry also enforces schema compatibility rules on evolution: a producer wanting to change a topic’s schema (add a field, change a type) must do so in a way that’s declared compatible with existing consumers (e.g., BACKWARD compatibility requires new schemas to still be readable by consumers using the old schema, typically by requiring new fields to have defaults). This turns schema drift, in a streaming context, from a silent runtime failure into a rejected write at publish time — the exact same “catch it at the earliest, cheapest point” principle from the ingestion-time validation discussed above, applied specifically to the streaming case.
Governance as a Socio-Technical Practice
It’s tempting to treat governance as a tooling problem — buy a catalog, configure some access control lists, done. In practice, governance that works is socio-technical: the tools are necessary but not sufficient, because the actual hard problems are organizational. Who is accountable when a table’s data goes stale — is there a named owner, or does responsibility diffuse across nobody? What access policy applies to a column containing PII, and who decided that? Is a “glossary term” like “active customer” actually agreed upon across Marketing, Sales, and Finance, or does each team quietly maintain a slightly different definition that produces three different numbers for the same metric in three different dashboards?
This is exactly the organizational problem that 07 — Data Lakes & Modern Architectures’s discussion of data mesh addresses from the architecture side: governance succeeds when it is federated — domain teams own the quality, documentation, and access rules for the data they produce, because they’re the only ones with the context to do it well — rather than being a purely central function where one governance team tries to understand and gatekeep every dataset in the company and inevitably becomes both a bottleneck and a rubber stamp (approving things it doesn’t have the context to actually evaluate). Dehghani’s fourth data mesh principle, federated computational governance, is precisely this: global rules (a PII field must always be masked, every dataset must have a named owner, every published table must pass its quality gate) are agreed on centrally and then enforced automatically, computationally, at the platform layer — rather than manually, by a human reviewing every dataset before it ships. Metadata is what makes this enforcement possible at all: you cannot computationally enforce “PII must be masked” unless something in your metadata layer actually flags which columns contain PII in the first place. This is the sense in which quality, governance, and metadata are one connected system rather than three separate initiatives: metadata is the substrate, quality checks are metadata-driven automated tests, and governance is metadata-driven policy enforcement, all resting on the same underlying “what do we know about this data” foundation. Testing this entire system end-to-end, including the pipelines that enforce these checks, is covered in 16 — Testing for Data Pipelines.
Best Practices
The table below distills the dimensions from the Fundamentals section into the specific automated check that typically catches each kind of violation in practice — the mapping engineers actually reach for when deciding what to implement first.
| Data quality dimension | Example violation | Typical automated check |
|---|---|---|
| Accuracy | Customer country field disagrees with billing address / IP geolocation | Cross-source reconciliation check against an independent reference (harder to automate; often a periodic audit query rather than a per-row test) |
| Completeness | Upstream ingestion silently drops rows on a holiday, orders table missing 3,000 rows | Row count check against a historical baseline or upstream source count |
| Consistency | Total spend in orders disagrees with sum of order_items | Reconciliation / cross-table aggregate check (dbt_utils.equality, or a custom SQL assertion) |
| Timeliness | ”Real-time” dashboard fed by a batch job refreshing every 6 hours | Freshness check on max(loaded_at) against an SLA threshold (dbt’s dbt source freshness, or catalog-level freshness monitors) |
| Uniqueness | Duplicate customer rows from a non-idempotent retried import | unique test on the natural or surrogate key column |
| Validity | discount_percent contains 150 | Range/accepted_values check, or a JSON Schema / Avro schema constraint enforced at ingestion |
Beyond that mapping, a handful of practices consistently separate teams that keep stakeholder trust from teams that lose it:
- Put quality gates as early in the pipeline as possible, and make failures loud. A failed pipeline run that pages someone is a good outcome compared to a silently wrong number reaching a dashboard; design for “fail fast and visibly,” not “never fail.”
- Write tests alongside the transformation they protect, not as an afterthought. dbt’s model of colocating
tests:YAML with the model definition exists specifically so the quality rule and the logic it protects are reviewed, versioned, and updated together — a test written six months after the model, by someone else, is far more likely to drift out of sync with what the model actually does. - Prefer derived lineage over hand-maintained documentation, every time. A lineage diagram that isn’t generated from the actual pipeline code or query logs is a snapshot that starts decaying the moment it’s drawn.
- Assign an explicit, named owner to every dataset that other teams depend on. “Owned by the data team” is not an owner; it diffuses accountability to nobody. A specific owner is who gets paged, who signs off on schema changes, and who is accountable for the data’s quality SLA.
- Establish data contracts with upstream teams for anything you depend on that you don’t control. A verbal agreement that “we won’t change that field” is not a contract; a versioned, tested schema that upstream commits to and CI enforces, is.
- Federate governance to domain owners rather than centralizing it in one gatekeeping team, and invest the resulting central effort into the platform, catalog, and computational policy enforcement that makes federation safe — this is the metadata-first, mesh-aligned model described above, not a purely tooling decision.
- Treat the data catalog as a product with its own adoption problem, not a one-time setup task. A catalog nobody updates or searches degrades into exactly the kind of stale documentation it was meant to replace; documenting ownership and freshness needs to be a required step in shipping a new dataset, not an optional follow-up.
- Use schema registries for any streaming pipeline with more than one consumer. The cost of enforcing compatibility at publish time is far lower than the cost of debugging a consumer that broke silently three deploys after a producer changed a field type.
References
- Great Expectations — Documentation
- dbt — Data Tests Documentation
- dbt — Lineage / DAG Documentation
- Confluent — Schema Registry Documentation
- JSON Schema — Official Specification
- Zhamak Dehghani, How to Move Beyond a Monolithic Data Lake to a Distributed Data Mesh (martinfowler.com, 2019)
- DataHub — Open Source Metadata Platform Documentation
- Monte Carlo — Data Observability
- roadmap.sh — Data Engineer Roadmap