← Kỹ sư dữ liệu← Data Engineer
Kỹ sư dữ liệuData Engineer19 Th7, 2026Jul 19, 202622 phút đọc16 min read

Data Lake & Kiến trúc hiện đạiData Lakes & Modern Architectures

Thuộc bộ kiến thức Data Engineer Roadmap.

Tổng quan

Một data warehouse (xem 06 — Data Modeling & Warehousing) đòi hỏi bạn phải biết trước schema trước khi nạp bất kỳ dòng dữ liệu nào: bảng, cột, kiểu dữ liệu đều được định nghĩa từ đầu, và bất cứ thứ gì không khớp sẽ bị từ chối hoặc phải ép vào một giải pháp tình thế. Đảm bảo đó chính là điều khiến warehouse nhanh và đáng tin cậy với dữ liệu có cấu trúc, đã hiểu rõ, được truy vấn lặp đi lặp lại — nhưng nó lại không phù hợp với thực tế lộn xộn hơn mà hầu hết tổ chức phải đối mặt: JSON clickstream với các field thay đổi hàng tuần, dữ liệu telemetry từ cảm biến, file PDF, ảnh, dữ liệu xuất từ bên thứ ba theo bất kỳ định dạng nào nhà cung cấp muốn gửi, log mà chưa ai parse đầy đủ. Ép tất cả những thứ đó qua một schema cứng nhắc ngay lúc ingest sẽ hoặc âm thầm làm mất thông tin, hoặc chặn đứng pipeline cho tới khi ai đó cập nhật migration schema.

Data lake ra đời để loại bỏ yêu cầu định trước đó. Đây là một tầng lưu trữ — trước kia là HDFS, ngày nay hầu như luôn là object storage trên cloud như Amazon S3, Azure Data Lake Storage (ADLS), hoặc Google Cloud Storage (GCS) — chấp nhận dữ liệu ở định dạng gốc: có cấu trúc, bán cấu trúc, hoặc phi cấu trúc, với chi phí trên mỗi gigabyte thấp, không áp schema tại thời điểm ghi. Schema thay vào đó được áp dụng sau, tại thời điểm truy vấn, bởi bất kỳ công cụ nào đọc dữ liệu — mô hình này gọi là schema-on-read, đối lập trực tiếp với schema-on-write của warehouse. Điều này đảo ngược chi phí của sự linh hoạt: bạn gần như không mất gì để ingest bất cứ thứ gì, nhưng bạn trì hoãn (và đôi khi không bao giờ trả) chi phí để hiểu được ý nghĩa của nó.

Sự linh hoạt đó cũng chính là điểm yếu nổi tiếng nhất của data lake. Một lake không có cataloging, không có ownership, không có kiểm tra chất lượng, và không có tài liệu sẽ suy thoái thành thứ ngành gọi là data swamp — một kho mà dữ liệu về mặt kỹ thuật tồn tại nhưng không ai tin tưởng nó, không ai biết nó nghĩa là gì, và tìm ra thứ gì hữu ích tốn công sức hơn cả việc suy ra lại từ đầu. Toàn bộ hành trình “kiến trúc dữ liệu hiện đại” được đề cập trong bài này — lakehouse, open table format, data mesh, data fabric, thiết kế metadata-first — về cơ bản là câu trả lời chung của cả ngành cho một câu hỏi: làm sao giữ được lợi thế về chi phí và sự linh hoạt của một lake mà không để nó biến thành swamp?

Kiến thức nền tảng

Data Lake vs Data Warehouse: Đánh đổi cốt lõi

Khía cạnhData LakeData Warehouse
SchemaSchema-on-read (áp dụng lúc truy vấn)Schema-on-write (bắt buộc lúc nạp dữ liệu)
Loại dữ liệuCó cấu trúc, bán cấu trúc, phi cấu trúc (JSON, ảnh, log, Parquet, CSV, video)Chỉ có cấu trúc, đã được modeling thành bảng
Chi phí lưu trữThấp — object storage phổ thông, vài cent mỗi GB/thángCao hơn — thường gộp chung giá storage và compute
Đảm bảo transactionalTrước đây không có (chỉ là file thô); được bổ sung bởi table format (xem bên dưới)Mạnh — ACID theo thiết kế
Hiệu năng truy vấnBiến thiên; phụ thuộc file format, partitioning, engineĐược tối ưu — columnar storage, index, query planner
Người dùng điển hìnhData scientist, ML pipeline, khai phá ad hoc, lưu trữ archivalBI dashboard, analyst, báo cáo tài chính
Rủi ro governanceCao nếu không quản lý — vấn đề “data swamp”Thấp hơn — schema và access control được ép buộc theo cấu trúc
Điểm mạnh chínhChi phí, linh hoạt, không mất dữ liệu lúc ingestĐộ tin cậy, tốc độ truy vấn, sự tin tưởng của business

Không phía nào trong bảng này “lỗi thời” — một luồng event thô mà bạn có thể sẽ xử lý lại theo năm cách khác nhau trong hai năm tới thực sự thuộc về một lake; báo cáo doanh thu hàng tháng của team tài chính thực sự thuộc về một bảng warehouse được modeling nghiêm ngặt. Điều thay đổi cả ngành là nhận ra rằng hầu hết tổ chức đang xây dựng cả hai, và phải trả tiền hai lần.

Vì sao việc có hai bản sao dữ liệu trở thành vấn đề

Mô hình thống trị thập niên 2010 là: nạp dữ liệu thô vào lake, sau đó chạy các job ETL để copy, làm sạch, và nạp một phần dữ liệu đó vào warehouse phục vụ BI. Cách này vẫn hoạt động, nhưng nó tạo ra nỗi đau thực sự và ngày càng chồng chất:

Đây chính là vấn đề mà mô hình lakehouse được sinh ra để giải quyết.

Khái niệm chính

Lakehouse: Độ tin cậy của warehouse trên nền kinh tế của lake

Lakehouse là một kiến trúc lưu trữ dữ liệu một lần, dưới dạng file trong object storage rẻ (S3/ADLS/GCS), nhưng bổ sung một tầng metadata và transaction lên trên các file đó để mang lại các đảm bảo mà trước đây người ta cần warehouse mới có được: ACID transaction, ép buộc và tiến hóa schema, indexing, và time travel — trong khi nhiều compute engine khác nhau (Spark, Trino, Presto, Snowflake, Flink) có thể đọc và ghi cùng một tập file bên dưới đồng thời. Databricks là bên phổ biến hóa cả thuật ngữ lẫn kiến trúc này, và bài báo lakehouse của Databricks trình bày tại CIDR 2021 là luận điểm kỹ thuật kinh điển cho mô hình này.

Cơ chế giúp điều này khả thi là một open table format — một đặc tả và tập hợp các file metadata nằm cạnh các file columnar thuần túy (hầu như luôn là Apache Parquet) theo dõi tập file nào hiện đang cấu thành trạng thái hợp lệ của một bảng, schema là gì, và toàn bộ lịch sử thay đổi. Ba open table format thống trị là:

FormatNguồn gốcCơ chế nổi bật
Delta LakeDatabricks (mã nguồn mở từ 2019)Transaction log dạng JSON (_delta_log) ghi lại từng lần thêm/xóa file Parquet
Apache IcebergNetflix, nay là dự án ApacheCây metadata gồm manifest file và snapshot; được nhiều engine áp dụng mạnh (Snowflake, Trino, Spark, Flink)
Apache HudiUber, nay là dự án ApacheTimeline các commit; tối ưu cho upsert thường xuyên và ingestion kiểu incremental/CDC

Cả ba đều giải quyết về cơ bản cùng một vấn đề — gắn ngữ nghĩa ACID lên một thư mục file bất biến — với chi tiết triển khai khác nhau và thế mạnh hệ sinh thái khác nhau. Onehouse (do những người sáng lập Hudi thành lập) cũng đã đẩy mạnh các nỗ lực tương tác chéo để dữ liệu ghi bằng một format có thể được đọc qua tầng metadata của format khác, giảm chi phí thực tế nếu bạn “chọn sai.”

Vì sao open table format — và vì sao chữ “open” quan trọng

Trước khi table format ra đời, một file Parquet nằm trong S3 chỉ đơn thuần là một file: không có commit atomic đa file, không an toàn khi nhiều writer ghi đồng thời, không có cách nào tích hợp sẵn để biết “bảng này trông như thế nào một giờ trước.” Table format giải quyết điều đó, nhưng phần open trong “open table format” quan trọng không kém phần transactional, vì hai lý do:

Delta Lake chi tiết: Transaction log

Cơ chế của Delta Lake là một ví dụ cụ thể tốt để minh họa cách bất kỳ open table format nào đạt được ACID trên nền file thuần túy, nên đáng để đi sâu vào cụ thể.

Về mặt vật lý, một bảng Delta chỉ đơn giản là một thư mục chứa các file Parquet cộng với một thư mục con tên _delta_log/ chứa một chuỗi file JSON (000000.json, 000001.json, …), mỗi file ứng với một transaction đã commit. Mỗi entry JSON ghi lại một actionadd một file, remove một file, cập nhật schema, cập nhật metadata của bảng — thay vì ghi đè lên bất kỳ file dữ liệu nào tại chỗ. Đọc một bảng nghĩa là đọc log để dựng lại tập file “đang sống” hiện tại, rồi chỉ đọc các file Parquet đó; ghi nghĩa là thêm một entry JSON mới vào log mô tả những gì đã thay đổi. Vì commit là các entry log bất biến mới thay vì thay đổi tại chỗ, nhiều writer đồng thời có thể được đối soát bằng optimistic concurrency control, và mọi trạng thái trước đó của bảng vẫn có thể dựng lại được — nó chỉ là các entry log chưa bị garbage-collect.

Chính log đó cho phép time travel: truy vấn bảng ở trạng thái tại một phiên bản hoặc thời điểm trước đó chỉ đơn giản là replay log tới đúng điểm đó.

-- Truy vấn trạng thái hiện tại
SELECT * FROM sales;

-- Truy vấn tại một số phiên bản cụ thể
SELECT * FROM sales VERSION AS OF 12;

-- Truy vấn tại một thời điểm cụ thể
SELECT * FROM sales TIMESTAMP AS OF '2026-06-01';

-- Đưa bảng về lại một phiên bản trước đó
RESTORE TABLE sales TO VERSION AS OF 12;

Time travel hữu ích vượt xa việc thỏa mãn tò mò: tái tạo lại một run huấn luyện ML đúng với snapshot dữ liệu đã dùng tại thời điểm đó, audit “báo cáo này trông như thế nào trước khi được sửa,” và khôi phục sau một lần ghi lỗi (một DELETE sai hoặc một bug làm hỏng một batch) mà không cần restore từ backup. Bảo trì định kỳ (VACUUM của Delta, expiration snapshot của Iceberg) sẽ xóa các file cũ không còn được tham chiếu bởi bất kỳ phiên bản nào được giữ lại, đánh đổi độ sâu time-travel lấy chi phí lưu trữ.

Data Mesh: Phi tập trung hóa ownership, không chỉ storage

Tất cả những gì ở trên là câu trả lời kỹ thuật cho sự phân mảnh lake/warehouse. Data mesh, được Zhamak Dehghani giới thiệu trong bài viết năm 2019 “How to Move Beyond a Monolithic Data Lake to a Distributed Data Mesh” và mở rộng trong cuốn sách Data Mesh: Delivering Data-Driven Value at Scale, là câu trả lời cho một vấn đề tổ chức khác biệt: một team dữ liệu trung tâm duy nhất trở thành nút thắt cổ chai khi công ty phát triển.

Kiểu thất bại mà data mesh nhắm tới rất quen thuộc với bất kỳ ai từng làm việc trong một tổ chức lớn chỉ có một team data engineering trung tâm: mỗi domain (marketing, sales, logistics, fraud) gửi yêu cầu tới team trung tâm để ingest, modeling, và expose dữ liệu của họ; team trung tâm không có hiểu biết sâu về ngữ nghĩa dữ liệu của từng domain riêng lẻ; backlog tăng nhanh hơn tốc độ tuyển người; và các chuyên gia domain — những người thực sự hiểu dữ liệu — phải chờ đợi thay vì tự định hình nó. Câu trả lời của data mesh là đảo ngược ai sở hữu dữ liệu, chứ không chỉ là dữ liệu được lưu ở đâu.

Dehghani đóng khung data mesh quanh bốn nguyên tắc:

Nguyên tắcÝ nghĩaVấn đề nó giải quyết
Domain ownershipMỗi domain nghiệp vụ (không phải một team trung tâm) sở hữu dữ liệu mà nó tạo ra, từ đầu đến cuốiTeam trung tâm không còn cần hiểu biết sâu về mọi domain để modeling dữ liệu đúng
Data as a productDữ liệu của một domain được coi như một sản phẩm có owner, SLA, tài liệu, và cam kết chất lượng — team domain chịu trách nhiệm để consumer có thể tin tưởng và sử dụng nóNgăn các domain đổ dữ liệu thô, không tài liệu qua tường rồi gọi đó là “đã chia sẻ”
Self-serve data platformMột team platform trung tâm xây dựng hạ tầng tái sử dụng được (cấp phát storage, template pipeline, tích hợp catalog, access control) để các team domain dùng để xây dựng và phục vụ sản phẩm dữ liệu của chính họCác team domain không cần mỗi bên tự phát minh lại hạ tầng hay tuyển chuyên gia platform riêng
Federated computational governanceCác chuẩn toàn cục (tương tác, bảo mật, tuân thủ) được thống nhất giữa các domain và được ép buộc tự động/bằng máy tính, thay vì qua kiểm soát thủ công tập trungGiữ các domain tự chủ trong khi vẫn ngăn hỗn loạn — các chuẩn được chia sẻ dù ownership thì không

Lưu ý điều data mesh không quy định: đây không phải một công nghệ cụ thể, và nó không có nghĩa là phi tập trung hóa storage một cách vật lý thành hàng chục lake rời rạc không kết nối. Nhiều triển khai thực tế vẫn dùng một lakehouse hoặc object storage dùng chung bên dưới — thứ được phi tập trung hóa là trách nhiệm tổ chức trong việc sản xuất, tài liệu hóa, và chịu trách nhiệm cho dữ liệu của mỗi domain. Đây chính là lý do vấn đề tổ chức quan trọng không kém vấn đề kỹ thuật: mesh vừa là một cuộc tái cơ cấu tổ chức vừa là một kiến trúc, và việc áp dụng công nghệ (open table format, catalog self-serve) mà không có sự chuyển đổi tổ chức đi kèm (các team domain thực sự nhận ownership và được đầu tư để làm việc đó) thường thất bại.

Data Fabric: Một câu trả lời khác, thường bị nhầm với mesh

Data fabric là một khái niệm liên quan nhưng khác biệt, và hai thuật ngữ này liên tục bị nhầm lẫn với nhau, nên đáng để nói rõ ràng sự khác biệt:

Nói cách khác: một data fabric có thể được đặt lên trên một kiến trúc hoàn toàn tập trung, do một team duy nhất sở hữu, và vẫn mang lại giá trị (khả năng khám phá tốt hơn, lineage tự động, tích hợp chéo hệ thống dễ dàng hơn) mà không cần thay đổi tổ chức nào. Ngược lại, một data mesh sẽ vô nghĩa nếu không có sự chuyển đổi tổ chức về ownership, bất kể công cụ metadata nào nằm bên dưới nó. Hai khái niệm này không loại trừ lẫn nhau — một triển khai data mesh thường cần metadata chủ động mạnh và một catalog thực sự để federated governance và self-serve discovery hoạt động được, đó là chỗ hai khái niệm giao thoa và bị nhầm lẫn trong thực tế. Gartner phổ biến hóa “data fabric” chủ yếu như một danh mục vendor/kiến trúc; “data mesh” khởi nguồn như một đề xuất thiết kế tổ chức từ một practitioner — đây là một manh mối hữu ích để biết mỗi thuật ngữ thực sự đang mô tả góc nhìn nào.

Kiến trúc lấy metadata làm trung tâm

Cả federated governance của data mesh lẫn mô hình tích hợp của data fabric đều phụ thuộc vào cùng một sự chuyển dịch nền tảng: coi metadata — schema, lineage ở mức cột, ownership dữ liệu, tín hiệu về độ mới và chất lượng, chính sách truy cập, định nghĩa nghiệp vụ — như một sản phẩm hạng nhất theo đúng nghĩa của nó, chứ không phải thứ được gắn thêm vào giao diện catalog sau khi mọi thứ đã xong. Trong một thiết kế metadata-first, mỗi dataset được kỳ vọng phải công bố schema, owner, SLA, và tín hiệu chất lượng của nó như một phần điều kiện để được coi là “hoàn thành,” giống như một API được kỳ vọng phải công bố contract của nó. Đây chính là điều cho phép mô hình federated computational governance trong data mesh thực sự kiểm tra tuân thủ tự động (thay vì một người audit dò từng spreadsheet), và cũng là điều cho phép tự động hóa của một data fabric thực sự suy luận ra quan hệ và lineage giữa các hệ thống mà nó chưa từng được cấu hình cụ thể để biết. Chủ đề này được phát triển sâu hơn trong 14 — Data Quality, Governance & Metadata.

Thực tế lai (hybrid)

Rất ít tổ chức thực sự vận hành một data mesh thuần túy (hoàn toàn phi tập trung, không có team dữ liệu trung tâm nào) hoặc một kiến trúc hoàn toàn tập trung (một team, một warehouse, không có tự chủ domain nào) trong thực tế. Mô hình phổ biến trong thực tế là hybrid: một lakehouse hoặc cloud data platform dùng chung (xem 13 — Cloud Data Platforms) cung cấp hạ tầng self-serve và ép buộc governance nền tảng, trong khi các domain có ngữ cảnh cao cụ thể (ví dụ: team fraud, team supply-chain) được trao ownership đối với sản phẩm dữ liệu của chính họ và chịu trách nhiệm về chất lượng — mà không cần mọi domain trong công ty đều phải có riêng nhân sự data engineering chuyên trách. Điều này phản ánh một sự thật rộng hơn trong thiết kế hệ thống phân tán: tập trung hóa hoàn toàn và phi tập trung hóa hoàn toàn đều là hai thái cực dễ mô tả nhưng khó vận hành tốt; câu trả lời bền vững cho hầu hết tổ chức nằm ở đâu đó giữa hai thái cực, được định hình bởi quy mô team, độ phức tạp của domain, và mức đầu tư nền tảng trung tâm mà công ty sẵn sàng chi trả.

Best Practices

Tài liệu tham khảo

Part of the Data Engineer Roadmap knowledge base.

Overview

A data warehouse (see 06 — Data Modeling & Warehousing) demands that you know your schema before you load a single row: tables, columns, and types are defined up front, and anything that doesn’t fit gets rejected or forced into a workaround. That guarantee is exactly what makes warehouses fast and trustworthy for structured, well-understood, repeatedly-queried data — but it is a poor fit for the messier reality most organizations actually face: clickstream JSON with fields that change weekly, sensor telemetry, PDFs, images, third-party exports in whatever format the vendor felt like sending, logs nobody has fully parsed yet. Forcing all of that through a rigid schema at ingestion time either drops information silently or blocks the pipeline until someone updates a schema migration.

A data lake exists to remove that up-front requirement. It is a storage layer — historically HDFS, today almost always cloud object storage like Amazon S3, Azure Data Lake Storage (ADLS), or Google Cloud Storage (GCS) — that accepts data in its native format: structured, semi-structured, or unstructured, at low cost per gigabyte, with no schema enforced at write time. The schema is instead applied later, at query time, by whatever tool reads the data — a pattern called schema-on-read, in direct contrast to the warehouse’s schema-on-write. This inverts the cost of flexibility: you pay almost nothing to ingest anything, but you defer (and sometimes never pay off) the cost of making sense of it.

That flexibility is also the data lake’s best-known failure mode. A lake with no cataloging, no ownership, no quality checks, and no documentation degrades into what the industry calls a data swamp — a store where data technically exists but nobody trusts it, nobody knows what it means, and finding anything useful costs more effort than re-deriving it from scratch. The entire arc of “modern data architecture” covered in this note — lakehouses, open table formats, data mesh, data fabric, metadata-first design — is fundamentally an industry-wide response to one question: how do we keep the cost and flexibility advantages of a lake without it turning into a swamp?

Fundamentals

Data Lake vs Data Warehouse: The Core Trade-off

DimensionData LakeData Warehouse
SchemaSchema-on-read (applied at query time)Schema-on-write (enforced at load time)
Data typesStructured, semi-structured, unstructured (JSON, images, logs, Parquet, CSV, video)Structured only, modeled into tables
Storage costLow — commodity object storage, cents per GB/monthHigher — often couples storage and compute pricing
Transactional guaranteesHistorically none (plain files); added by table formats (see below)Strong — ACID by design
Query performanceVariable; depends on file format, partitioning, engineOptimized — columnar storage, indexes, query planner
Typical consumersData scientists, ML pipelines, ad hoc exploration, archivalBI dashboards, analysts, finance/reporting
Governance riskHigh if unmanaged — the “data swamp” problemLower — schema and access control enforced structurally
Primary strengthCost, flexibility, no data loss at ingestionReliability, query speed, business trust

Neither side of this table is “obsolete” — a raw event stream you might reprocess five different ways over the next two years genuinely belongs in a lake; a finance team’s monthly revenue report genuinely belongs in a rigorously modeled warehouse table. What changed the industry was realizing that most organizations were building both, and paying twice.

Why Two Copies of Data Became a Problem

The pattern that dominated the 2010s was: land raw data in a lake, then run ETL jobs that copy, clean, and load a subset of that data into a warehouse for BI consumption. This works, but it creates real, compounding pain:

This is the precise problem the lakehouse pattern was invented to solve.

Key Concepts

The Lakehouse: Warehouse Reliability on Lake Economics

A lakehouse is an architecture that stores data once, as files in cheap object storage (S3/ADLS/GCS), but adds a metadata and transaction layer on top of those files that gives you the guarantees people previously needed a warehouse for: ACID transactions, schema enforcement and evolution, indexing, and time travel — all while multiple different compute engines (Spark, Trino, Presto, Snowflake, Flink) can read and write the same underlying files concurrently. Databricks popularized both the term and the architecture, and the Databricks lakehouse paper presented at CIDR 2021 is the canonical technical case for it.

The mechanism that makes this possible is an open table format — a specification and set of metadata files sitting alongside plain columnar files (almost always Apache Parquet) that tracks which files currently constitute a table’s valid state, what the schema is, and the full history of changes. The three dominant open table formats are:

FormatOriginNotable mechanism
Delta LakeDatabricks (open-sourced 2019)JSON-based transaction log (_delta_log) recording every add/remove of a Parquet file
Apache IcebergNetflix, now Apache projectMetadata tree of manifest files and snapshots; strong multi-engine adoption (Snowflake, Trino, Spark, Flink)
Apache HudiUber, now Apache projectTimeline of commits; optimized for frequent upserts and incremental/CDC-style ingestion

All three solve essentially the same problem — bolting ACID semantics onto a directory of immutable files — with different implementation details and different ecosystem strengths. Onehouse (founded by Hudi’s original creators) has also pushed interoperability efforts so that data written in one format can be read via another’s metadata layer, reducing the practical cost of picking “wrong.”

Why Open Table Formats — and Why “Open” Matters

Before table formats existed, a Parquet file sitting in S3 was just a file: no atomic multi-file commits, no concurrent-writer safety, no built-in way to know “what did this table look like an hour ago.” Table formats solve that, but the open part of “open table format” matters just as much as the transactional part, for two reasons:

Delta Lake in Detail: The Transaction Log

Delta Lake’s mechanism is a good concrete illustration of how any open table format achieves ACID on top of plain files, so it’s worth walking through specifically.

A Delta table is, physically, just a directory of Parquet files plus a subdirectory called _delta_log/ containing a sequence of JSON files (000000.json, 000001.json, …), one per committed transaction. Each JSON entry records an actionadd a file, remove a file, update the schema, update table metadata — rather than rewriting any data file in place. Reading a table means reading the log to reconstruct the current set of “live” files, then reading only those Parquet files; writing means appending a new JSON entry to the log describing what changed. Because commits are new immutable log entries rather than in-place mutations, concurrent writers can be reconciled with optimistic concurrency control, and every prior state of the table is still reconstructable — it’s just log entries not yet garbage-collected.

That log is exactly what enables time travel: querying the table as it existed at a previous version or timestamp is simply a matter of replaying the log only up to that point.

-- Query the current state
SELECT * FROM sales;

-- Query as of a specific version number
SELECT * FROM sales VERSION AS OF 12;

-- Query as of a specific timestamp
SELECT * FROM sales TIMESTAMP AS OF '2026-06-01';

-- Roll a table back to a prior version
RESTORE TABLE sales TO VERSION AS OF 12;

Time travel is useful well beyond curiosity: reproducing an ML training run against the exact data snapshot used at the time, auditing “what did this report look like before the correction,” and recovering from a bad write (a botched DELETE or a bug that corrupted a batch) without restoring from a backup. Periodic maintenance (Delta’s VACUUM, Iceberg’s snapshot expiration) removes old files no longer referenced by any retained version, trading time-travel depth for storage cost.

Data Mesh: Decentralizing Ownership, Not Just Storage

Everything above is a technical answer to lake/warehouse fragmentation. Data mesh, introduced by Zhamak Dehghani in her 2019 article “How to Move Beyond a Monolithic Data Lake to a Distributed Data Mesh” and expanded in her book Data Mesh: Delivering Data-Driven Value at Scale, is an answer to a different, organizational problem: a single central data team becomes a bottleneck as a company grows.

The failure mode data mesh targets is familiar to anyone who has worked in a large organization with one central data engineering team: every domain (marketing, sales, logistics, fraud) submits a request to the central team to ingest their data, model it, and expose it; the central team has no deep context on any individual domain’s data semantics; the backlog grows faster than the team can hire; and domain experts who actually understand the data are stuck waiting instead of shaping it themselves. Data mesh’s answer is to invert who owns data, not just where it’s stored.

Dehghani frames data mesh around four principles:

PrincipleWhat it meansProblem it solves
Domain ownershipEach business domain (not a central team) owns the data it produces, end to endCentral team no longer needs deep expertise in every domain to model its data correctly
Data as a productA domain’s data is treated as a product with an owner, SLAs, documentation, and quality guarantees — the domain team is accountable for consumers being able to trust and use itPrevents domains from dumping raw, undocumented data over the wall and calling it “shared”
Self-serve data platformA central platform team builds reusable infrastructure (storage provisioning, pipeline templates, catalog integration, access control) that domain teams use to build and serve their own data productsDomain teams don’t each need to reinvent infrastructure or hire their own platform specialists
Federated computational governanceGlobal standards (interoperability, security, compliance) are agreed upon across domains and enforced automatically/computationally, rather than through central manual gatekeepingKeeps domains autonomous while still preventing chaos — standards are shared even though ownership isn’t

Note what data mesh does not prescribe: it is not a specific technology, and it does not mean literally decentralizing storage into dozens of unconnected lakes. Many real implementations still use a shared lakehouse or shared object storage underneath — what’s decentralized is the organizational responsibility for producing, documenting, and being accountable for each domain’s data. This is precisely why the organizational problem matters as much as the technical one: mesh is as much a re-org as it is an architecture, and adopting the technology (open table formats, self-serve catalogs) without the organizational shift (domain teams actually taking ownership and being funded to do so) tends to fail.

Data Fabric: A Different Answer, Often Confused with Mesh

Data fabric is a related but distinct concept, and the two terms get conflated constantly, so the distinction is worth stating plainly:

Put differently: a data fabric could be layered on top of a purely centralized, single-team-owned architecture and still deliver value (better discoverability, automated lineage, easier cross-system integration) with zero organizational change. A data mesh, conversely, is meaningless without an organizational shift in ownership, regardless of what metadata tooling sits underneath it. The two are not mutually exclusive — a data-mesh implementation typically needs strong active metadata and a catalog to make federated governance and self-serve discovery actually work, which is where the concepts overlap and get confused in practice. Gartner popularized “data fabric” largely as a vendor/architecture category; “data mesh” originated as an organizational design proposal from a practitioner, which is a useful tell for which lens each term is really describing.

Metadata-First Architecture

Both data mesh’s federated governance and data fabric’s integration model depend on the same underlying shift: treating metadata — schemas, column-level lineage, data ownership, freshness and quality signals, access policies, business definitions — as a first-class product in its own right, not an afterthought bolted onto a catalog UI after the fact. In a metadata-first design, every dataset is expected to publish its schema, its owner, its SLAs, and its quality signals as part of being considered “done,” in the same way an API is expected to publish its contract. This is what lets a federated computational governance model in data mesh actually check compliance automatically (rather than a human auditor spot-checking spreadsheets), and it’s what lets a data fabric’s automation actually infer relationships and lineage across systems it wasn’t specifically configured to know about. This topic is developed further in 14 — Data Quality, Governance & Metadata.

The Hybrid Reality

Very few organizations run a pure data mesh (fully decentralized, no central data team at all) or a purely centralized architecture (one team, one warehouse, no domain autonomy) in practice. The common real-world pattern is a hybrid: a shared lakehouse or shared cloud data platform (see 13 — Cloud Data Platforms) provides the self-serve infrastructure and enforces baseline governance, while specific high-context domains (e.g., a fraud team, a supply-chain team) are given ownership over their own data products and are accountable for their quality — without every single domain in the company needing its own dedicated data engineering staff. This mirrors a broader truth in distributed systems design: full centralization and full decentralization are both extremes that are easy to describe and hard to run well; the durable answer for most organizations sits somewhere in between, shaped by team size, domain complexity, and how much central platform investment the company is willing to fund.

Best Practices

References