← Kỹ sư dữ liệu← Data Engineer
Kỹ sư dữ liệuData Engineer19 Th7, 2026Jul 19, 202630 phút đọc23 min read

Data Quality, Governance & MetadataData Quality, Governance & Metadata

Thuộc bộ kiến thức Data Engineer Roadmap.

Tổng quan

Mọi ghi chú khác trong bộ kiến thức này — ingestion (xem 03 — Data Sources & Ingestion), pipeline (xem 08 — ETL, ELT & Data Pipelines), warehouse và lake (xem 06 — Data Modeling & Warehousing07 — Data Lakes & Modern Architectures) — suy cho cùng đều phục vụ một mục tiêu duy nhất: một người hoặc một hệ thống đưa ra quyết định dựa trên một con số mà pipeline tạo ra. Nếu con số đó sai, mọi sự tinh tế kỹ thuật phía trên nó đều vô nghĩa. Một pipeline chạy đúng lịch, scale thoải mái, chi phí vận hành gần như bằng không, code đọc vào là thích — nhưng âm thầm trả về con số doanh thu sai — thì tệ hơn hẳn một pipeline xấu xí, chậm chạp, đáng xấu hổ nhưng cho ra con số đúng. Đây là câu ngạn ngữ lâu đời nhất trong ngành máy tính, “garbage in, garbage out”, nhưng đáng để phát biểu lại ở dạng sắc bén nhất dành riêng cho data engineering: tính đúng đắn (correctness) không phải là một thuộc tính chất lượng trong nhiều thuộc tính; nó chính là sản phẩm. Mọi thứ khác — độ trễ, chi phí, sự tinh gọn — chỉ là cải thiện trải nghiệm trên nền một sản phẩm vốn dĩ phải đúng thì mới có lý do tồn tại.

Điều khiến dữ liệu sai trở nên nguy hiểm một cách bất đối xứng, so với hầu hết mọi loại lỗi phần mềm khác, là cách nó phá hủy niềm tin. Một service bị crash thì dễ thấy, được page, được sửa, rồi bị quên đi. Một dashboard âm thầm hiển thị sai số trong ba tuần liền thường không dễ thấy — cho đến khi một stakeholder ra quyết định dựa trên nó, phát hiện lỗi bằng cách khác, và mất niềm tin không chỉ vào dashboard đó mà vào toàn bộ nền tảng dữ liệu phía sau. Một khi một VP đã từng “bị bỏng” bởi một con số sai, phản ứng hợp lý của họ là ngừng tin bất kỳ con số nào từ hệ thống đó, kể cả 99% dashboard vẫn luôn đúng, và bắt đầu yêu cầu một analyst “kiểm tra lại bằng Excel” trước mỗi cuộc họp. Hành vi đó — âm thầm tính lại số bằng tay vì không ai tin nền tảng nữa — là triệu chứng rõ ràng nhất cho thấy niềm tin dữ liệu của một tổ chức đã sụp đổ, và việc xây dựng lại nó cực kỳ tốn kém: niềm tin bị phá hủy chỉ bởi một sự cố tồi tệ, nhưng chỉ được xây lại bằng một quá trình dài, nhàm chán của việc luôn luôn đúng. Đây chính là lý do data quality, data governance và metadata management không phải là những mục “nên có” trong maturity model để làm sau này — chúng là cơ chế giúp một nền tảng dữ liệu giành được và giữ được thứ duy nhất khiến nó có giá trị: được tin tưởng.

Ghi chú này bao quát ba lĩnh vực, cùng nhau giữ cho dữ liệu đáng tin cậy ở quy mô lớn: data quality (dữ liệu có tuân thủ những gì nó phải là không — có accurate, complete, consistent, timely, unique, valid không?), data governance (ai sở hữu dữ liệu này, ai được phép truy cập, và những quy tắc nào áp dụng cho cách nó được tạo ra và sử dụng?), và metadata management (chúng ta biết gì về dữ liệu — schema, lineage, ý nghĩa, độ mới (freshness) của nó — và mọi người có thể tìm ra điều đó không?). Ba lĩnh vực này được trình bày riêng để rõ ràng nhưng thực tế gắn bó chặt chẽ với nhau: bạn không thể governance dữ liệu mà bạn không thể mô tả được, bạn không thể tự động hóa quality check nếu không có metadata về ý nghĩa của “đúng”, và một data catalog thực chất chỉ là metadata management được làm cho có thể tìm kiếm được.

Kiến thức nền tảng

Các chiều (dimension) của Data Quality

“Data quality” không phải là một thuộc tính đúng/sai duy nhất; nó tách thành nhiều chiều độc lập, và một dataset có thể đạt điểm cao ở chiều này trong khi thất bại thảm hại ở chiều khác (một bảng có thể hoàn toàn consistent — tự nhất quán bên trong — trong khi hoàn toàn inaccurate, tức là sai một cách nhất quán). Việc chỉ rõ chiều nào đang gãy là điều biến một lời phàn nàn mơ hồ (“dữ liệu trông có gì đó sai sai”) thành một khẳng định cụ thể, kiểm chứng được, và sửa được.

Chiều (Dimension)Định nghĩaVí dụ vi phạm
AccuracyDữ liệu phản ánh đúng giá trị hoặc sự kiện trong thế giới thực mà nó đại diệnTrường quốc gia của khách hàng ghi “Canada” nhưng địa chỉ thanh toán và IP geolocation đều chỉ “USA”
CompletenessToàn bộ dữ liệu lẽ ra phải có mặt thì đều có mặt — không thiếu trường bắt buộc hoặc mất bản ghiBảng orders thiếu 3.000 dòng của ngày lễ vì một job ingestion upstream âm thầm lỗi vào ngày lượng traffic giảm mà không ai để ý
ConsistencyCùng một sự thật khớp nhau giữa các bảng, hệ thống, hoặc trường — không mâu thuẫnBảng orders cho thấy tổng chi tiêu trọn đời của một khách hàng là 4.200 USD, nhưng cộng từng dòng đơn hàng trong bảng order_items lại ra 3.850 USD
TimelinessDữ liệu sẵn sàng và cập nhật trong khoảng thời gian mà consumer cầnMột dashboard “real-time” chống gian lận được nạp dữ liệu bởi một batch job chỉ refresh mỗi 6 giờ, nên nó luôn stale đối với các analyst cần bắt gian lận trong ngày
UniquenessKhông có bản ghi trùng lặp ngoài ý muốn cho cùng một thực thể thực tếMột job import CRM chạy hai lần do bug retry-không-idempotent, tạo ra hai dòng cho cùng một khách hàng với hai customer ID khác nhau
ValidityDữ liệu tuân thủ các quy tắc cú pháp và miền giá trị (format, kiểu dữ liệu, khoảng giá trị, giá trị cho phép)Cột discount_percent chứa giá trị 150, về mặt cú pháp là một số nguyên hợp lệ nhưng về mặt ngữ nghĩa là bất khả thi cho một mức giảm giá phần trăm

Hai trong số các chiều này đáng được lưu ý thêm vì chúng tương tác theo cách dễ gây nhầm lẫn. Accuracy và validity không phải là một: một giá trị có thể hoàn toàn valid (đúng kiểu, đúng format, trong khoảng cho phép) trong khi vẫn inaccurate (sai giá trị so với sự thật thực tế cụ thể đó) — ngày sinh 1990-01-01 là một ngày hoàn toàn hợp lệ, nhưng nó sai nếu người đó thực ra sinh năm 1985. Kiểm tra validity rẻ để tự động hóa vì chỉ cần biết quy tắc (một regex, một kiểu dữ liệu, một khoảng giá trị); kiểm tra accuracy đắt đỏ vì cần một nguồn sự thật độc lập để đối chiếu, đây chính là lý do phần lớn công cụ data quality tự động dựa nặng vào validity, completeness và consistency check, và coi việc xác minh accuracy thực sự là một bài toán khó hơn, thường mang tính thủ công hoặc thống kê.

Quality Check nằm ở đâu trong Pipeline

Quality không phải là một cổng kiểm tra (gate) duy nhất ở cuối pipeline; nó được kiểm tra tại nhiều điểm, và vị trí kiểm tra làm thay đổi hình dạng của thất bại:

Cách tiếp cận nhiều lớp này giống với defense-in-depth trong bảo mật: không có kiểm tra đơn lẻ nào được giả định là đủ, và bắt được vấn đề càng sớm luôn rẻ hơn bắt được nó càng muộn, vì “bán kính ảnh hưởng” (blast radius) của dữ liệu xấu lớn dần theo từng job downstream đã tiêu thụ nó.

Khái niệm chính

Automated Data Quality Checks và Quality Gates

Cơ chế thực tế để thực thi các chiều nói trên là viết automated data quality checks như một phần tường minh, được version hóa của pipeline — không phải một buổi review spreadsheet thủ công mà ai đó làm khi nhớ ra. Các loại check phổ biến là:

Hai công cụ thống trị cách các team triển khai những check này trong thực tế, và chúng giải quyết vấn đề từ hai tầng khác nhau của stack:

Great Expectations là một framework bằng Python được xây dựng chuyên cho data quality. Bạn định nghĩa các “Expectation” — các khẳng định khai báo, dễ đọc cho con người như “cột này không bao giờ được null” hoặc “giá trị cột này phải nằm giữa 0 và 100” — trên một dataset, và Great Expectations sẽ validate dữ liệu, cho ra kết quả pass/fail, và có thể tự sinh báo cáo HTML dễ đọc gọi là “Data Docs”. Nó không phụ thuộc engine (chạy được với pandas DataFrame, SQL database, Spark) và thường được chạy như một bước tường minh trong một pipeline được orchestrate (xem 09 — Workflow Orchestration), thường qua tích hợp Airflow của nó.

import great_expectations as gx

context = gx.get_context()
validator = context.sources.pandas_default.read_csv("orders.csv")

validator.expect_column_values_to_not_be_null("customer_id")
validator.expect_column_values_to_be_unique("order_id")
validator.expect_column_values_to_be_between("discount_percent", min_value=0, max_value=100)

results = validator.validate()
if not results.success:
    raise ValueError("Data quality checks failed — halting pipeline")

dbt tests giải quyết cùng vấn đề nhưng sống ngay bên trong transformation layer, được định nghĩa khai báo bằng YAML ngay cạnh model mà chúng kiểm tra — nghĩa là quy tắc quality và transformation mà nó bảo vệ sống trong cùng một codebase, được version-control và review cùng nhau. dbt có sẵn bốn generic test dựng sẵn (not_null, unique, accepted_values, relationships), và package cộng đồng dbt-utils cùng dbt-expectations mở rộng thêm distribution và statistical check, về cơ bản mang phong cách khẳng định của Great Expectations vào cú pháp YAML của dbt.

# models/schema.yml
models:
  - name: orders
    columns:
      - name: order_id
        tests:
          - unique
          - not_null
      - name: customer_id
        tests:
          - not_null
          - relationships:
              to: ref('customers')
              field: customer_id
      - name: discount_percent
        tests:
          - dbt_utils.accepted_range:
              min_value: 0
              max_value: 100

Chạy dbt test thực thi mọi test đã khai báo dưới dạng truy vấn SQL trên các model đã build và báo cáo cái nào thất bại — và quan trọng hơn, dbt build chạy test xen kẽ với việc build model, nên một test thất bại trên một model upstream có thể ngăn các model downstream phụ thuộc vào nó được build với dữ liệu xấu bên dưới ngay từ đầu.

Đây chính là bản chất của một quality gate: một điểm kiểm tra, được nối trực tiếp vào luồng điều khiển của pipeline, chặn dữ liệu tiến sang bước tiếp theo — và do đó không bao giờ chạm tới dashboard hay consumer downstream — nếu nó không đạt một check đã định nghĩa. Quality gate là một lựa chọn thiết kế có chủ đích để ưu tiên một thất bại rõ ràng, ồn ào (một pipeline run thất bại, một kỹ sư on-call bị page) hơn một thất bại thầm lặng, vô hình (một con số sai âm thầm được đưa tới dashboard của một VP). Sự đánh đổi này gần như luôn đúng: một pipeline hỏng thì phiền một ngày; một con số sai được tin tưởng suốt một tháng là một sự cố niềm tin.

Schema Drift: Kẻ giết người thầm lặng

Schema drift là hiện tượng xảy ra khi một hệ thống nguồn upstream thay đổi shape dữ liệu của nó — một cột bị đổi tên, một kiểu dữ liệu đổi từ integer sang string, một trường bị xóa, một trường bắt buộc mới xuất hiện — mà không có sự phối hợp nào với các team tiêu thụ dữ liệu đó ở downstream. Đây là một trong những failure mode nguy hiểm nhất trong data engineering chính vì nó có hai kiểu biểu hiện rất khác nhau, và kiểu thầm lặng còn tệ hơn nhiều so với kiểu ồn ào:

Ba chiến lược cụ thể giảm thiểu schema drift:

Data Lineage

Data lineage là bản ghi cho biết một mẩu dữ liệu đến từ đâu và đã trải qua những phép biến đổi nào để đạt tới hình dạng hiện tại — về bản chất, một đồ thị có hướng nối mọi bảng nguồn, bước transformation, và bảng hoặc dashboard downstream. Lineage trả lời hai câu hỏi mà nếu không có nó thì tốn kém hoặc bất khả thi nếu chỉ dựa vào việc kiểm tra thủ công:

Lineage có thể được ghi nhận theo hai cách, và sự khác biệt về độ tin cậy giữa chúng là rất rõ rệt:

Hàm ý thực tiễn rất rõ ràng: ưu tiên lineage được suy ra từ các artifact thực thi thực tế (code, query log, metadata của orchestrator) hơn là lineage được viết ra bởi con người, vì lineage suy ra không thể trôi lệch (drift) khỏi pipeline mà nó mô tả, trong khi lineage viết tay thì cuối cùng luôn trôi lệch.

Metadata Management và Data Catalog

Metadata là dữ liệu về dữ liệu, và nó tách rõ ràng thành ba loại trả lời ba câu hỏi khác nhau:

Loại metadataNó nắm bắt điều gìVí dụ
Technical metadataSchema, kiểu dữ liệu, tên bảng/cột, partitioning, storage formatorders.order_date là cột TIMESTAMP, được partition theo ngày, lưu dưới dạng Parquet
Business metadataÝ nghĩa cho con người, định nghĩa, ownership, thuật ngữ trong glossary”Khách hàng đang hoạt động” (active customer) được định nghĩa là: có giao dịch mua trong 90 ngày gần nhất, theo mục glossary của team Marketing, thuộc quyền sở hữu của domain Growth
Operational metadataFreshness, lịch sử chạy, SLA, khối lượng dữ liệu theo thời gian, lịch sử pass/fail của quality checkBảng orders được refresh lần cuối 14 phút trước, pipeline của nó có tỷ lệ thành công 99,2% trong 90 ngày qua, và SLA của nó là “cập nhật trong vòng 1 giờ”

Bất kỳ loại nào trong ba loại này đứng riêng lẻ đều hữu ích nhưng không đầy đủ: biết kiểu dữ liệu của một cột (technical) không cho bạn biết bạn có được phép dùng nó trong một báo cáo hướng khách hàng hay không (business/ownership); biết định nghĩa kinh doanh không cho bạn biết dữ liệu đằng sau nó có thực sự mới ngay bây giờ hay không (operational). Một data catalog — các công cụ như DataHub, Amundsen, Atlan, hoặc Collibra — tồn tại để làm giao diện tìm kiếm duy nhất gắn kết cả ba loại lại với nhau: một analyst tìm kiếm “customer churn” nên tìm ra đúng bảng, thấy schema của nó, thấy ai sở hữu nó, thấy định nghĩa kinh doanh của nó, thấy nó có mới hay không, và thấy nó có đang pass quality check hay không, tất cả trong một lần tra cứu thay vì năm tin nhắn Slack riêng biệt tới năm team khác nhau. Đây là điều khiến một catalog vừa là công cụ governance (nó là bề mặt thực thi cho ownership và access policy) vừa là công cụ discoverability (nó trả lời “chúng ta có dữ liệu gì, và tôi có thể tin nó không?”) cùng một lúc — hai mục đích chia sẻ cùng một nền tảng metadata bên dưới, chính là lý do chúng thường được xây dựng thành một sản phẩm thay vì hai.

Data Interoperability: Schema và Schema Registry

Hình thức quan trọng nhất về mặt vận hành của metadata chính là schema, và các hệ thống dữ liệu hiện đại ngày càng đưa schema đó ra bên ngoài thành một hợp đồng (contract) chung, máy đọc được, thay vì để nó ngầm hiểu trong bất kỳ code nào tình cờ tạo ra hoặc tiêu thụ dữ liệu. Apache ParquetApache Avro đều nhúng thông tin schema trực tiếp cùng với dữ liệu chúng lưu trữ — Parquet là một columnar format tối ưu cho đọc phân tích, Avro là một binary format gọn nhẹ với hỗ trợ mạnh cho schema evolution, thường dùng cho dữ liệu streaming theo hàng (row-oriented). JSON Schema cung cấp hợp đồng tương đương cho payload JSON, cho phép producer và consumer thống nhất về trường bắt buộc, kiểu dữ liệu, và ràng buộc mà không bên nào cần kiểm tra code của bên kia.

Điều này quan trọng nhất trong các hệ thống streaming, nơi một producer và hàng chục consumer độc lập có thể không bao giờ phối hợp deployment trực tiếp với nhau. Một schema registry — phổ biến nhất là Confluent Schema Registry dùng cùng Apache Kafka — nằm giữa producer và consumer và thực thi rằng mọi message publish lên một topic phải tuân theo một schema đã đăng ký (bằng Avro, Protobuf, hoặc JSON Schema) trước khi được phép vào topic đó. Quan trọng hơn, registry còn thực thi các quy tắc tương thích schema (compatibility rules) khi evolution: một producer muốn thay đổi schema của một topic (thêm trường, đổi kiểu) phải làm theo cách được khai báo là tương thích với các consumer hiện có (ví dụ, tương thích BACKWARD yêu cầu schema mới vẫn phải đọc được bởi consumer đang dùng schema cũ, thường bằng cách yêu cầu trường mới phải có giá trị mặc định). Điều này biến schema drift, trong bối cảnh streaming, từ một thất bại runtime thầm lặng thành một lần ghi bị từ chối ngay tại thời điểm publish — chính xác cùng nguyên tắc “bắt lỗi tại điểm sớm nhất, rẻ nhất” từ validation lúc ingestion đã bàn ở trên, áp dụng riêng cho trường hợp streaming.

Governance như một thực hành Socio-Technical

Rất dễ để coi governance là một vấn đề công cụ — mua một catalog, cấu hình vài access control list, xong. Trong thực tế, governance hiệu quả mang tính socio-technical: công cụ là cần thiết nhưng không đủ, vì những vấn đề khó thực sự mang tính tổ chức. Ai chịu trách nhiệm khi dữ liệu của một bảng trở nên stale — có một owner được chỉ định rõ ràng, hay trách nhiệm bị pha loãng cho không ai cả? Chính sách truy cập nào áp dụng cho một cột chứa PII, và ai quyết định điều đó? Một “thuật ngữ glossary” như “active customer” có thực sự được thống nhất giữa Marketing, Sales, và Finance, hay mỗi team âm thầm duy trì một định nghĩa hơi khác nhau khiến ra ba con số khác nhau cho cùng một metric trên ba dashboard khác nhau?

Đây chính xác là vấn đề mang tính tổ chức mà phần thảo luận về data mesh trong 07 — Data Lakes & Modern Architectures đề cập từ góc độ kiến trúc: governance thành công khi nó federated (phân tán, liên bang) — các domain team sở hữu quality, tài liệu, và quy tắc truy cập cho dữ liệu họ tạo ra, vì họ là những người duy nhất có đủ context để làm điều đó tốt — thay vì là một chức năng trung tâm thuần túy, nơi một team governance cố hiểu và gác cổng (gatekeep) mọi dataset trong công ty và không thể tránh khỏi việc trở thành cả điểm nghẽn lẫn một con dấu cao su (rubber stamp, phê duyệt những thứ mà nó không có đủ context để thực sự đánh giá). Nguyên tắc thứ tư của data mesh theo Dehghani, federated computational governance, chính xác là điều này: các quy tắc toàn cục (một trường PII phải luôn được mask, mọi dataset phải có một owner được chỉ định, mọi bảng được publish phải pass quality gate) được thống nhất ở cấp trung tâm rồi được thực thi tự động, mang tính tính toán (computational), tại tầng platform — thay vì thủ công, bởi một con người review từng dataset trước khi nó được ship. Metadata là thứ khiến việc thực thi này khả thi ngay từ đầu: bạn không thể thực thi mang tính tính toán quy tắc “PII phải được mask” trừ khi có gì đó trong tầng metadata của bạn thực sự đánh dấu cột nào chứa PII ngay từ đầu. Đây là ý nghĩa mà quality, governance, và metadata là một hệ thống liên kết duy nhất chứ không phải ba sáng kiến riêng biệt: metadata là nền tảng, quality check là các test tự động dựa trên metadata, và governance là việc thực thi chính sách dựa trên metadata, tất cả đều dựa trên cùng một nền tảng cơ bản “chúng ta biết gì về dữ liệu này”. Việc kiểm thử toàn bộ hệ thống này đầu-cuối, bao gồm cả các pipeline thực thi những check này, được bàn ở 16 — Testing for Data Pipelines.

Best Practices

Bảng dưới đây cô đọng các chiều từ phần Kiến thức nền tảng thành check tự động cụ thể thường bắt được từng loại vi phạm trong thực tế — cách ánh xạ mà kỹ sư thực sự dùng khi quyết định triển khai cái gì trước.

Chiều data qualityVí dụ vi phạmCheck tự động điển hình
AccuracyTrường quốc gia của khách hàng mâu thuẫn với địa chỉ thanh toán / IP geolocationCheck đối chiếu chéo nguồn (cross-source reconciliation) với một tham chiếu độc lập (khó tự động hóa hơn; thường là một query audit định kỳ thay vì test theo từng dòng)
CompletenessIngestion upstream âm thầm mất dòng vào ngày lễ, bảng orders thiếu 3.000 dòngRow count check so với baseline lịch sử hoặc số dòng ở nguồn upstream
ConsistencyTổng chi tiêu trong orders mâu thuẫn với tổng của order_itemsCheck đối chiếu / aggregate liên bảng (dbt_utils.equality, hoặc một assertion SQL tùy chỉnh)
TimelinessDashboard “real-time” được nạp bởi một batch job refresh mỗi 6 giờFreshness check trên max(loaded_at) so với ngưỡng SLA (dbt source freshness của dbt, hoặc bộ giám sát freshness cấp catalog)
UniquenessDòng khách hàng trùng lặp từ một import bị retry không idempotentTest unique trên cột khóa tự nhiên hoặc surrogate key
ValidityCột discount_percent chứa giá trị 150Check range/accepted_values, hoặc ràng buộc JSON Schema / Avro schema được thực thi tại ingestion

Ngoài bảng ánh xạ đó, một số thực hành luôn phân biệt các team giữ được niềm tin của stakeholder với các team đánh mất nó:

Tài liệu tham khảo

Part of the Data Engineer Roadmap knowledge base.

Overview

Every other note in this knowledge base — ingestion (see 03 — Data Sources & Ingestion), pipelines (see 08 — ETL, ELT & Data Pipelines), warehouses and lakes (see 06 — Data Modeling & Warehousing and 07 — Data Lakes & Modern Architectures) — is, in the end, in service of one outcome: a person or a system makes a decision based on a number a pipeline produced. If that number is wrong, none of the engineering elegance upstream of it matters. A pipeline that runs on schedule, scales effortlessly, costs nothing to operate, and is a joy to read — but silently delivers the wrong revenue figure — is strictly worse than an ugly, slow, embarrassing pipeline that delivers the right one. This is the oldest aphorism in computing, “garbage in, garbage out,” but it is worth restating in its sharpest form for data engineering specifically: correctness is not one quality attribute among many; it is the product. Everything else — latency, cost, elegance — is a quality-of-life improvement on top of a product that has to be correct to exist at all.

What makes bad data uniquely dangerous, compared to almost any other class of software bug, is how asymmetrically it destroys trust. A crashed service is visible, gets paged, gets fixed, and is forgotten. A dashboard that quietly shows the wrong number for three weeks is often not visible — until a business stakeholder makes a decision on it, discovers the error some other way, and loses confidence not just in that one dashboard but in the entire data platform behind it. Once a VP has been burned by a wrong number, the rational response is to stop trusting any number from that system, including the 99% of dashboards that were always correct, and start asking an analyst to “double check it in Excel” before every meeting. That behavior — quietly re-deriving numbers by hand because nobody trusts the platform — is the single clearest symptom that an organization’s data trust has collapsed, and it is enormously expensive to rebuild once lost: trust is destroyed by one bad incident and rebuilt only by a long, boring track record of being right. This is why data quality, governance, and metadata management are not “nice to have” maturity-model line items to get to eventually — they are the mechanism by which a data platform earns and keeps the only thing that makes it useful at all: being believed.

This note covers the three disciplines that, together, keep data trustworthy at scale: data quality (does the data conform to what it should — is it accurate, complete, consistent, timely, unique, valid?), data governance (who owns this data, who may access it, and what rules apply to how it is produced and used?), and metadata management (what do we know about the data — its schema, its lineage, its meaning, its freshness — and can people find that out?). They are presented separately for clarity but are deeply intertwined in practice: you cannot govern data you cannot describe, you cannot automate quality checks without metadata about what “correct” means, and a data catalog is really just metadata management made searchable.

Fundamentals

The Dimensions of Data Quality

“Data quality” is not a single yes/no property; it decomposes into several independent dimensions, and a dataset can score well on one while failing badly on another (a table can be perfectly consistent — internally self-coherent — while being completely inaccurate, i.e., consistently wrong). Being precise about which dimension is broken is what turns a vague complaint (“the data looks off”) into a specific, testable, fixable claim.

DimensionDefinitionExample violation
AccuracyData correctly reflects the real-world value or event it representsA customer’s country field says “Canada” but their billing address and IP geolocation both say “USA”
CompletenessAll data that should be present, is present — no missing required fields or dropped recordsAn orders table is missing 3,000 rows for July 4th because an upstream ingestion job silently failed on a holiday with reduced traffic and nobody noticed the row count drop
ConsistencyThe same fact agrees across tables, systems, or fields — no contradictionsThe orders table shows a customer’s total lifetime spend as $4,200, but summing their individual order rows in the order_items table yields $3,850
TimelinessData is available and up to date within the window consumers needA “real-time” fraud dashboard is fed by a batch job that only refreshes every 6 hours, so it is reliably stale for the analysts relying on it to catch same-day fraud
UniquenessNo unintended duplicate records for the same real-world entityA CRM import runs twice due to a retry-without-idempotency bug, creating two rows for the same customer with two different customer IDs
ValidityData conforms to the syntactic and domain rules it’s supposed to (format, type, range, allowed values)A discount_percent column contains a value of 150, which is syntactically a valid integer but semantically impossible for a percentage discount

Two of these dimensions deserve a further note because they interact in a way that trips people up. Accuracy and validity are not the same thing: a value can be perfectly valid (right type, right format, within an allowed range) while still being inaccurate (the wrong value for that specific real-world fact) — a birth date of 1990-01-01 is a perfectly valid date, but it’s wrong if the person was actually born in 1985. Validity checks are cheap to automate because they only require knowing the rule (a regex, a type, a range); accuracy checks are expensive because they require an independent source of truth to compare against, which is exactly why most automated data quality tooling leans heavily on validity, completeness, and consistency checks, and treats true accuracy verification as a harder, often manual or statistical, problem.

Where Quality Checks Live in the Pipeline

Quality is not a single gate at the end of a pipeline; it is checked at multiple points, and where you check it changes what failure looks like:

This layered approach mirrors defense-in-depth in security: no single check is assumed sufficient, and catching an issue earlier is always cheaper than catching it later, because the “blast radius” of bad data grows with every downstream job that has already consumed it.

Key Concepts

Automated Data Quality Checks and Quality Gates

The practical mechanism for enforcing the dimensions above is to write automated data quality checks as an explicit, versioned part of the pipeline — not a manual spreadsheet review someone does when they remember to. The common check types are:

Two tools dominate how teams implement these checks in practice, and they solve the problem from different layers of the stack:

Great Expectations is a Python-based framework purpose-built for data quality. You define “Expectations” — declarative, human-readable assertions like “this column’s values must never be null” or “this column’s values must be between 0 and 100” — against a dataset, and Great Expectations validates the data, produces a pass/fail result, and can auto-generate human-readable HTML “Data Docs” reporting the outcome. It’s engine-agnostic (works against pandas DataFrames, SQL databases, Spark) and is typically run as an explicit step in an orchestrated pipeline (see 09 — Workflow Orchestration), often via its Airflow integration.

import great_expectations as gx

context = gx.get_context()
validator = context.sources.pandas_default.read_csv("orders.csv")

validator.expect_column_values_to_not_be_null("customer_id")
validator.expect_column_values_to_be_unique("order_id")
validator.expect_column_values_to_be_between("discount_percent", min_value=0, max_value=100)

results = validator.validate()
if not results.success:
    raise ValueError("Data quality checks failed — halting pipeline")

dbt tests solve the same problem but live natively inside the transformation layer, defined declaratively in YAML right alongside the models they check — which means the quality rule and the transformation it protects live in the same codebase, version-controlled and reviewed together. dbt ships four built-in generic tests out of the box (not_null, unique, accepted_values, relationships), and the community package dbt-utils and dbt-expectations extend this with distribution and statistical checks, effectively bringing Great-Expectations-style assertions into dbt’s YAML syntax.

# models/schema.yml
models:
  - name: orders
    columns:
      - name: order_id
        tests:
          - unique
          - not_null
      - name: customer_id
        tests:
          - not_null
          - relationships:
              to: ref('customers')
              field: customer_id
      - name: discount_percent
        tests:
          - dbt_utils.accepted_range:
              min_value: 0
              max_value: 100

Running dbt test executes every declared test as a SQL query against the built models and reports which ones failed — and critically, dbt build runs tests interleaved with model builds, so a failing test on an upstream model can stop its downstream dependents from being built with bad data underneath them at all.

This is the essence of a quality gate: a checkpoint, wired directly into the pipeline’s control flow, that halts data from progressing to the next stage — and therefore from ever reaching a dashboard or a downstream consumer — if it fails a defined check. A quality gate is a deliberate design choice to prefer a visible, loud failure (a failed pipeline run, a paged on-call engineer) over a silent, invisible one (a wrong number quietly shipped to a VP’s dashboard). This trade-off is almost always correct: a broken pipeline is annoying for a day; a wrong number believed for a month is a trust incident.

Schema Drift: The Silent Killer

Schema drift is what happens when an upstream source system changes its data shape — a column gets renamed, a type changes from integer to string, a field is dropped, a new required field appears — without any coordination with the teams consuming that data downstream. It is one of the most dangerous failure modes in data engineering precisely because it has two very different failure signatures, and the quiet one is far worse than the loud one:

Three concrete strategies mitigate schema drift:

Data Lineage

Data lineage is the record of where a piece of data came from and every transformation it passed through to reach its current form — effectively, a directed graph connecting every source table, transformation step, and downstream table or dashboard. Lineage answers two questions that are otherwise expensive or impossible to answer by inspection alone:

Lineage can be captured two ways, and the difference in reliability between them is stark:

The practical implication is unambiguous: prefer lineage that is derived from the actual execution artifacts (code, query logs, orchestrator metadata) over lineage that is authored by a human, because derived lineage cannot drift out of sync with the pipeline it describes, while authored lineage always eventually does.

Metadata Management and Data Catalogs

Metadata is data about data, and it splits cleanly into three categories that answer three different questions:

Metadata typeWhat it capturesExample
Technical metadataSchema, data types, table/column names, partitioning, storage formatorders.order_date is a TIMESTAMP column, partitioned by day, stored as Parquet
Business metadataHuman meaning, definitions, ownership, glossary terms”Active customer” is defined as: made a purchase in the trailing 90 days, per the Marketing team’s glossary entry, owned by the Growth domain
Operational metadataFreshness, run history, SLAs, data volume over time, quality check pass/fail historyThe orders table was last refreshed 14 minutes ago, its pipeline has a 99.2% success rate over the last 90 days, and its SLA is “fresh within 1 hour”

Any one of these categories alone is useful but incomplete: knowing a column’s type (technical) doesn’t tell you whether you’re allowed to use it in a customer-facing report (business/ownership); knowing a business definition doesn’t tell you whether the data behind it is actually fresh right now (operational). A data catalog — tools like DataHub, Amundsen, Atlan, or Collibra — exists to be the single searchable interface that ties all three together: an analyst searching for “customer churn” should find the right table, see its schema, see who owns it, see its business definition, see whether it’s fresh, and see whether it’s passing its quality checks, all in one lookup instead of five separate Slack messages to five different teams. This is what makes a catalog both a governance tool (it’s the enforcement surface for ownership and access policy) and a discoverability tool (it answers “what data do we have, and can I trust it?”) at the same time — the two purposes share the same underlying metadata, which is precisely why they’re usually built as one product rather than two.

Data Interoperability: Schemas and Schema Registries

Metadata’s most operationally critical form is the schema itself, and modern data systems increasingly externalize that schema into a shared, machine-readable contract rather than leaving it implicit in whatever code happens to produce or consume the data. Apache Parquet and Apache Avro both embed schema information directly alongside the data they store — Parquet as a columnar format optimized for analytical reads, Avro as a compact binary format with strong support for schema evolution, commonly used for row-oriented streaming data. JSON Schema provides the equivalent contract for JSON payloads, letting a producer and consumer agree on required fields, types, and constraints without either side needing to inspect the other’s code.

This matters most in streaming systems, where a producer and dozens of independent consumers may never coordinate deployments directly. A schema registry — most commonly the Confluent Schema Registry used with Apache Kafka — sits between producers and consumers and enforces that every message published to a topic conforms to a registered schema (in Avro, Protobuf, or JSON Schema) before it’s allowed onto the topic at all. Crucially, the registry also enforces schema compatibility rules on evolution: a producer wanting to change a topic’s schema (add a field, change a type) must do so in a way that’s declared compatible with existing consumers (e.g., BACKWARD compatibility requires new schemas to still be readable by consumers using the old schema, typically by requiring new fields to have defaults). This turns schema drift, in a streaming context, from a silent runtime failure into a rejected write at publish time — the exact same “catch it at the earliest, cheapest point” principle from the ingestion-time validation discussed above, applied specifically to the streaming case.

Governance as a Socio-Technical Practice

It’s tempting to treat governance as a tooling problem — buy a catalog, configure some access control lists, done. In practice, governance that works is socio-technical: the tools are necessary but not sufficient, because the actual hard problems are organizational. Who is accountable when a table’s data goes stale — is there a named owner, or does responsibility diffuse across nobody? What access policy applies to a column containing PII, and who decided that? Is a “glossary term” like “active customer” actually agreed upon across Marketing, Sales, and Finance, or does each team quietly maintain a slightly different definition that produces three different numbers for the same metric in three different dashboards?

This is exactly the organizational problem that 07 — Data Lakes & Modern Architectures’s discussion of data mesh addresses from the architecture side: governance succeeds when it is federated — domain teams own the quality, documentation, and access rules for the data they produce, because they’re the only ones with the context to do it well — rather than being a purely central function where one governance team tries to understand and gatekeep every dataset in the company and inevitably becomes both a bottleneck and a rubber stamp (approving things it doesn’t have the context to actually evaluate). Dehghani’s fourth data mesh principle, federated computational governance, is precisely this: global rules (a PII field must always be masked, every dataset must have a named owner, every published table must pass its quality gate) are agreed on centrally and then enforced automatically, computationally, at the platform layer — rather than manually, by a human reviewing every dataset before it ships. Metadata is what makes this enforcement possible at all: you cannot computationally enforce “PII must be masked” unless something in your metadata layer actually flags which columns contain PII in the first place. This is the sense in which quality, governance, and metadata are one connected system rather than three separate initiatives: metadata is the substrate, quality checks are metadata-driven automated tests, and governance is metadata-driven policy enforcement, all resting on the same underlying “what do we know about this data” foundation. Testing this entire system end-to-end, including the pipelines that enforce these checks, is covered in 16 — Testing for Data Pipelines.

Best Practices

The table below distills the dimensions from the Fundamentals section into the specific automated check that typically catches each kind of violation in practice — the mapping engineers actually reach for when deciding what to implement first.

Data quality dimensionExample violationTypical automated check
AccuracyCustomer country field disagrees with billing address / IP geolocationCross-source reconciliation check against an independent reference (harder to automate; often a periodic audit query rather than a per-row test)
CompletenessUpstream ingestion silently drops rows on a holiday, orders table missing 3,000 rowsRow count check against a historical baseline or upstream source count
ConsistencyTotal spend in orders disagrees with sum of order_itemsReconciliation / cross-table aggregate check (dbt_utils.equality, or a custom SQL assertion)
Timeliness”Real-time” dashboard fed by a batch job refreshing every 6 hoursFreshness check on max(loaded_at) against an SLA threshold (dbt’s dbt source freshness, or catalog-level freshness monitors)
UniquenessDuplicate customer rows from a non-idempotent retried importunique test on the natural or surrogate key column
Validitydiscount_percent contains 150Range/accepted_values check, or a JSON Schema / Avro schema constraint enforced at ingestion

Beyond that mapping, a handful of practices consistently separate teams that keep stakeholder trust from teams that lose it:

References