Data Lake & Kiến trúc hiện đạiData Lakes & Modern Architectures
Thuộc bộ kiến thức Data Engineer Roadmap.
Tổng quan
Một data warehouse (xem 06 — Data Modeling & Warehousing) đòi hỏi bạn phải biết trước schema trước khi nạp bất kỳ dòng dữ liệu nào: bảng, cột, kiểu dữ liệu đều được định nghĩa từ đầu, và bất cứ thứ gì không khớp sẽ bị từ chối hoặc phải ép vào một giải pháp tình thế. Đảm bảo đó chính là điều khiến warehouse nhanh và đáng tin cậy với dữ liệu có cấu trúc, đã hiểu rõ, được truy vấn lặp đi lặp lại — nhưng nó lại không phù hợp với thực tế lộn xộn hơn mà hầu hết tổ chức phải đối mặt: JSON clickstream với các field thay đổi hàng tuần, dữ liệu telemetry từ cảm biến, file PDF, ảnh, dữ liệu xuất từ bên thứ ba theo bất kỳ định dạng nào nhà cung cấp muốn gửi, log mà chưa ai parse đầy đủ. Ép tất cả những thứ đó qua một schema cứng nhắc ngay lúc ingest sẽ hoặc âm thầm làm mất thông tin, hoặc chặn đứng pipeline cho tới khi ai đó cập nhật migration schema.
Data lake ra đời để loại bỏ yêu cầu định trước đó. Đây là một tầng lưu trữ — trước kia là HDFS, ngày nay hầu như luôn là object storage trên cloud như Amazon S3, Azure Data Lake Storage (ADLS), hoặc Google Cloud Storage (GCS) — chấp nhận dữ liệu ở định dạng gốc: có cấu trúc, bán cấu trúc, hoặc phi cấu trúc, với chi phí trên mỗi gigabyte thấp, không áp schema tại thời điểm ghi. Schema thay vào đó được áp dụng sau, tại thời điểm truy vấn, bởi bất kỳ công cụ nào đọc dữ liệu — mô hình này gọi là schema-on-read, đối lập trực tiếp với schema-on-write của warehouse. Điều này đảo ngược chi phí của sự linh hoạt: bạn gần như không mất gì để ingest bất cứ thứ gì, nhưng bạn trì hoãn (và đôi khi không bao giờ trả) chi phí để hiểu được ý nghĩa của nó.
Sự linh hoạt đó cũng chính là điểm yếu nổi tiếng nhất của data lake. Một lake không có cataloging, không có ownership, không có kiểm tra chất lượng, và không có tài liệu sẽ suy thoái thành thứ ngành gọi là data swamp — một kho mà dữ liệu về mặt kỹ thuật tồn tại nhưng không ai tin tưởng nó, không ai biết nó nghĩa là gì, và tìm ra thứ gì hữu ích tốn công sức hơn cả việc suy ra lại từ đầu. Toàn bộ hành trình “kiến trúc dữ liệu hiện đại” được đề cập trong bài này — lakehouse, open table format, data mesh, data fabric, thiết kế metadata-first — về cơ bản là câu trả lời chung của cả ngành cho một câu hỏi: làm sao giữ được lợi thế về chi phí và sự linh hoạt của một lake mà không để nó biến thành swamp?
Kiến thức nền tảng
Data Lake vs Data Warehouse: Đánh đổi cốt lõi
| Khía cạnh | Data Lake | Data Warehouse |
|---|---|---|
| Schema | Schema-on-read (áp dụng lúc truy vấn) | Schema-on-write (bắt buộc lúc nạp dữ liệu) |
| Loại dữ liệu | Có cấu trúc, bán cấu trúc, phi cấu trúc (JSON, ảnh, log, Parquet, CSV, video) | Chỉ có cấu trúc, đã được modeling thành bảng |
| Chi phí lưu trữ | Thấp — object storage phổ thông, vài cent mỗi GB/tháng | Cao hơn — thường gộp chung giá storage và compute |
| Đảm bảo transactional | Trước đây không có (chỉ là file thô); được bổ sung bởi table format (xem bên dưới) | Mạnh — ACID theo thiết kế |
| Hiệu năng truy vấn | Biến thiên; phụ thuộc file format, partitioning, engine | Được tối ưu — columnar storage, index, query planner |
| Người dùng điển hình | Data scientist, ML pipeline, khai phá ad hoc, lưu trữ archival | BI dashboard, analyst, báo cáo tài chính |
| Rủi ro governance | Cao nếu không quản lý — vấn đề “data swamp” | Thấp hơn — schema và access control được ép buộc theo cấu trúc |
| Điểm mạnh chính | Chi phí, linh hoạt, không mất dữ liệu lúc ingest | Độ tin cậy, tốc độ truy vấn, sự tin tưởng của business |
Không phía nào trong bảng này “lỗi thời” — một luồng event thô mà bạn có thể sẽ xử lý lại theo năm cách khác nhau trong hai năm tới thực sự thuộc về một lake; báo cáo doanh thu hàng tháng của team tài chính thực sự thuộc về một bảng warehouse được modeling nghiêm ngặt. Điều thay đổi cả ngành là nhận ra rằng hầu hết tổ chức đang xây dựng cả hai, và phải trả tiền hai lần.
Vì sao việc có hai bản sao dữ liệu trở thành vấn đề
Mô hình thống trị thập niên 2010 là: nạp dữ liệu thô vào lake, sau đó chạy các job ETL để copy, làm sạch, và nạp một phần dữ liệu đó vào warehouse phục vụ BI. Cách này vẫn hoạt động, nhưng nó tạo ra nỗi đau thực sự và ngày càng chồng chất:
- Trùng lặp và chi phí. Cùng một dữ liệu tồn tại vật lý hai lần, trong hai hệ thống lưu trữ khác nhau, thường với hai mô hình giá khác nhau.
- Sự cũ đi và trôi lệch (staleness/drift). Bản copy trong warehouse trễ hơn bản copy trong lake bằng đúng thời gian job ETL chạy; hai bản có thể không khớp nhau, và việc đối soát chúng trở thành một dự án riêng.
- Truy cập bị hạn chế vào sức mạnh thô của lake. Data scientist cần dữ liệu thô, chưa modeling để làm feature engineering hay ML thường không có được trải nghiệm tương đương, được quản trị, an toàn ACID trực tiếp trên lake — họ hoặc phải làm việc với các bản export cũ, hoặc làm việc trực tiếp trên file không được quản trị mà không có lưới an toàn nào của warehouse.
- Hai hệ thống cần bảo mật, giám sát, và trả lương kỹ sư vận hành. Mỗi hệ thống bổ sung là thêm bề mặt vận hành: access control list, chính sách backup, dashboard giám sát, và nhân sự chuyên trách cho từng hệ thống.
Đây chính là vấn đề mà mô hình lakehouse được sinh ra để giải quyết.
Khái niệm chính
Lakehouse: Độ tin cậy của warehouse trên nền kinh tế của lake
Lakehouse là một kiến trúc lưu trữ dữ liệu một lần, dưới dạng file trong object storage rẻ (S3/ADLS/GCS), nhưng bổ sung một tầng metadata và transaction lên trên các file đó để mang lại các đảm bảo mà trước đây người ta cần warehouse mới có được: ACID transaction, ép buộc và tiến hóa schema, indexing, và time travel — trong khi nhiều compute engine khác nhau (Spark, Trino, Presto, Snowflake, Flink) có thể đọc và ghi cùng một tập file bên dưới đồng thời. Databricks là bên phổ biến hóa cả thuật ngữ lẫn kiến trúc này, và bài báo lakehouse của Databricks trình bày tại CIDR 2021 là luận điểm kỹ thuật kinh điển cho mô hình này.
Cơ chế giúp điều này khả thi là một open table format — một đặc tả và tập hợp các file metadata nằm cạnh các file columnar thuần túy (hầu như luôn là Apache Parquet) theo dõi tập file nào hiện đang cấu thành trạng thái hợp lệ của một bảng, schema là gì, và toàn bộ lịch sử thay đổi. Ba open table format thống trị là:
| Format | Nguồn gốc | Cơ chế nổi bật |
|---|---|---|
| Delta Lake | Databricks (mã nguồn mở từ 2019) | Transaction log dạng JSON (_delta_log) ghi lại từng lần thêm/xóa file Parquet |
| Apache Iceberg | Netflix, nay là dự án Apache | Cây metadata gồm manifest file và snapshot; được nhiều engine áp dụng mạnh (Snowflake, Trino, Spark, Flink) |
| Apache Hudi | Uber, nay là dự án Apache | Timeline các commit; tối ưu cho upsert thường xuyên và ingestion kiểu incremental/CDC |
Cả ba đều giải quyết về cơ bản cùng một vấn đề — gắn ngữ nghĩa ACID lên một thư mục file bất biến — với chi tiết triển khai khác nhau và thế mạnh hệ sinh thái khác nhau. Onehouse (do những người sáng lập Hudi thành lập) cũng đã đẩy mạnh các nỗ lực tương tác chéo để dữ liệu ghi bằng một format có thể được đọc qua tầng metadata của format khác, giảm chi phí thực tế nếu bạn “chọn sai.”
Vì sao open table format — và vì sao chữ “open” quan trọng
Trước khi table format ra đời, một file Parquet nằm trong S3 chỉ đơn thuần là một file: không có commit atomic đa file, không an toàn khi nhiều writer ghi đồng thời, không có cách nào tích hợp sẵn để biết “bảng này trông như thế nào một giờ trước.” Table format giải quyết điều đó, nhưng phần open trong “open table format” quan trọng không kém phần transactional, vì hai lý do:
- Không bị vendor lock-in. Vì format là một đặc tả mở thay vì một storage engine độc quyền, bản thân dữ liệu có tính di động (portable). Bạn không bị khóa chặt vào compute của một vendor duy nhất để đọc chính dữ liệu của mình.
- Nhiều engine cùng truy cập một bản dữ liệu. Spark có thể ghi một bảng, Trino có thể truy vấn tương tác trên đó, Snowflake có thể đọc qua external table, và Flink có thể stream vào đó — tất cả trên cùng các file vật lý, không cần bước ETL nào để chuyển dữ liệu giữa các hệ thống chỉ để một engine khác chạm vào được. Đây chính là giải pháp kỹ thuật trực tiếp cho vấn đề “hai bản sao dữ liệu” nêu trên.
Delta Lake chi tiết: Transaction log
Cơ chế của Delta Lake là một ví dụ cụ thể tốt để minh họa cách bất kỳ open table format nào đạt được ACID trên nền file thuần túy, nên đáng để đi sâu vào cụ thể.
Về mặt vật lý, một bảng Delta chỉ đơn giản là một thư mục chứa các file Parquet cộng với một thư mục con tên _delta_log/ chứa một chuỗi file JSON (000000.json, 000001.json, …), mỗi file ứng với một transaction đã commit. Mỗi entry JSON ghi lại một action — add một file, remove một file, cập nhật schema, cập nhật metadata của bảng — thay vì ghi đè lên bất kỳ file dữ liệu nào tại chỗ. Đọc một bảng nghĩa là đọc log để dựng lại tập file “đang sống” hiện tại, rồi chỉ đọc các file Parquet đó; ghi nghĩa là thêm một entry JSON mới vào log mô tả những gì đã thay đổi. Vì commit là các entry log bất biến mới thay vì thay đổi tại chỗ, nhiều writer đồng thời có thể được đối soát bằng optimistic concurrency control, và mọi trạng thái trước đó của bảng vẫn có thể dựng lại được — nó chỉ là các entry log chưa bị garbage-collect.
Chính log đó cho phép time travel: truy vấn bảng ở trạng thái tại một phiên bản hoặc thời điểm trước đó chỉ đơn giản là replay log tới đúng điểm đó.
-- Truy vấn trạng thái hiện tại
SELECT * FROM sales;
-- Truy vấn tại một số phiên bản cụ thể
SELECT * FROM sales VERSION AS OF 12;
-- Truy vấn tại một thời điểm cụ thể
SELECT * FROM sales TIMESTAMP AS OF '2026-06-01';
-- Đưa bảng về lại một phiên bản trước đó
RESTORE TABLE sales TO VERSION AS OF 12;
Time travel hữu ích vượt xa việc thỏa mãn tò mò: tái tạo lại một run huấn luyện ML đúng với snapshot dữ liệu đã dùng tại thời điểm đó, audit “báo cáo này trông như thế nào trước khi được sửa,” và khôi phục sau một lần ghi lỗi (một DELETE sai hoặc một bug làm hỏng một batch) mà không cần restore từ backup. Bảo trì định kỳ (VACUUM của Delta, expiration snapshot của Iceberg) sẽ xóa các file cũ không còn được tham chiếu bởi bất kỳ phiên bản nào được giữ lại, đánh đổi độ sâu time-travel lấy chi phí lưu trữ.
Data Mesh: Phi tập trung hóa ownership, không chỉ storage
Tất cả những gì ở trên là câu trả lời kỹ thuật cho sự phân mảnh lake/warehouse. Data mesh, được Zhamak Dehghani giới thiệu trong bài viết năm 2019 “How to Move Beyond a Monolithic Data Lake to a Distributed Data Mesh” và mở rộng trong cuốn sách Data Mesh: Delivering Data-Driven Value at Scale, là câu trả lời cho một vấn đề tổ chức khác biệt: một team dữ liệu trung tâm duy nhất trở thành nút thắt cổ chai khi công ty phát triển.
Kiểu thất bại mà data mesh nhắm tới rất quen thuộc với bất kỳ ai từng làm việc trong một tổ chức lớn chỉ có một team data engineering trung tâm: mỗi domain (marketing, sales, logistics, fraud) gửi yêu cầu tới team trung tâm để ingest, modeling, và expose dữ liệu của họ; team trung tâm không có hiểu biết sâu về ngữ nghĩa dữ liệu của từng domain riêng lẻ; backlog tăng nhanh hơn tốc độ tuyển người; và các chuyên gia domain — những người thực sự hiểu dữ liệu — phải chờ đợi thay vì tự định hình nó. Câu trả lời của data mesh là đảo ngược ai sở hữu dữ liệu, chứ không chỉ là dữ liệu được lưu ở đâu.
Dehghani đóng khung data mesh quanh bốn nguyên tắc:
| Nguyên tắc | Ý nghĩa | Vấn đề nó giải quyết |
|---|---|---|
| Domain ownership | Mỗi domain nghiệp vụ (không phải một team trung tâm) sở hữu dữ liệu mà nó tạo ra, từ đầu đến cuối | Team trung tâm không còn cần hiểu biết sâu về mọi domain để modeling dữ liệu đúng |
| Data as a product | Dữ liệu của một domain được coi như một sản phẩm có owner, SLA, tài liệu, và cam kết chất lượng — team domain chịu trách nhiệm để consumer có thể tin tưởng và sử dụng nó | Ngăn các domain đổ dữ liệu thô, không tài liệu qua tường rồi gọi đó là “đã chia sẻ” |
| Self-serve data platform | Một team platform trung tâm xây dựng hạ tầng tái sử dụng được (cấp phát storage, template pipeline, tích hợp catalog, access control) để các team domain dùng để xây dựng và phục vụ sản phẩm dữ liệu của chính họ | Các team domain không cần mỗi bên tự phát minh lại hạ tầng hay tuyển chuyên gia platform riêng |
| Federated computational governance | Các chuẩn toàn cục (tương tác, bảo mật, tuân thủ) được thống nhất giữa các domain và được ép buộc tự động/bằng máy tính, thay vì qua kiểm soát thủ công tập trung | Giữ các domain tự chủ trong khi vẫn ngăn hỗn loạn — các chuẩn được chia sẻ dù ownership thì không |
Lưu ý điều data mesh không quy định: đây không phải một công nghệ cụ thể, và nó không có nghĩa là phi tập trung hóa storage một cách vật lý thành hàng chục lake rời rạc không kết nối. Nhiều triển khai thực tế vẫn dùng một lakehouse hoặc object storage dùng chung bên dưới — thứ được phi tập trung hóa là trách nhiệm tổ chức trong việc sản xuất, tài liệu hóa, và chịu trách nhiệm cho dữ liệu của mỗi domain. Đây chính là lý do vấn đề tổ chức quan trọng không kém vấn đề kỹ thuật: mesh vừa là một cuộc tái cơ cấu tổ chức vừa là một kiến trúc, và việc áp dụng công nghệ (open table format, catalog self-serve) mà không có sự chuyển đổi tổ chức đi kèm (các team domain thực sự nhận ownership và được đầu tư để làm việc đó) thường thất bại.
Data Fabric: Một câu trả lời khác, thường bị nhầm với mesh
Data fabric là một khái niệm liên quan nhưng khác biệt, và hai thuật ngữ này liên tục bị nhầm lẫn với nhau, nên đáng để nói rõ ràng sự khác biệt:
- Data mesh chủ yếu mang tính tổ chức — nó nói về ai sở hữu và chịu trách nhiệm cho dữ liệu (các team domain), được ép buộc qua federated governance.
- Data fabric chủ yếu mang tính kỹ thuật — nó là một kiến trúc sử dụng metadata chủ động, tự động (schema, lineage, mẫu sử dụng, quan hệ ngữ nghĩa) và ngày càng nhiều tự động hóa dựa trên AI/ML để khám phá, kết nối, và tích hợp dữ liệu vẫn nằm phân tán trên nhiều hệ thống hiện có — data warehouse, lake, database vận hành, công cụ SaaS — mà không nhất thiết phải di chuyển hay tập trung hóa nó, và không nhất thiết thay đổi ai sở hữu cái gì.
Nói cách khác: một data fabric có thể được đặt lên trên một kiến trúc hoàn toàn tập trung, do một team duy nhất sở hữu, và vẫn mang lại giá trị (khả năng khám phá tốt hơn, lineage tự động, tích hợp chéo hệ thống dễ dàng hơn) mà không cần thay đổi tổ chức nào. Ngược lại, một data mesh sẽ vô nghĩa nếu không có sự chuyển đổi tổ chức về ownership, bất kể công cụ metadata nào nằm bên dưới nó. Hai khái niệm này không loại trừ lẫn nhau — một triển khai data mesh thường cần metadata chủ động mạnh và một catalog thực sự để federated governance và self-serve discovery hoạt động được, đó là chỗ hai khái niệm giao thoa và bị nhầm lẫn trong thực tế. Gartner phổ biến hóa “data fabric” chủ yếu như một danh mục vendor/kiến trúc; “data mesh” khởi nguồn như một đề xuất thiết kế tổ chức từ một practitioner — đây là một manh mối hữu ích để biết mỗi thuật ngữ thực sự đang mô tả góc nhìn nào.
Kiến trúc lấy metadata làm trung tâm
Cả federated governance của data mesh lẫn mô hình tích hợp của data fabric đều phụ thuộc vào cùng một sự chuyển dịch nền tảng: coi metadata — schema, lineage ở mức cột, ownership dữ liệu, tín hiệu về độ mới và chất lượng, chính sách truy cập, định nghĩa nghiệp vụ — như một sản phẩm hạng nhất theo đúng nghĩa của nó, chứ không phải thứ được gắn thêm vào giao diện catalog sau khi mọi thứ đã xong. Trong một thiết kế metadata-first, mỗi dataset được kỳ vọng phải công bố schema, owner, SLA, và tín hiệu chất lượng của nó như một phần điều kiện để được coi là “hoàn thành,” giống như một API được kỳ vọng phải công bố contract của nó. Đây chính là điều cho phép mô hình federated computational governance trong data mesh thực sự kiểm tra tuân thủ tự động (thay vì một người audit dò từng spreadsheet), và cũng là điều cho phép tự động hóa của một data fabric thực sự suy luận ra quan hệ và lineage giữa các hệ thống mà nó chưa từng được cấu hình cụ thể để biết. Chủ đề này được phát triển sâu hơn trong 14 — Data Quality, Governance & Metadata.
Thực tế lai (hybrid)
Rất ít tổ chức thực sự vận hành một data mesh thuần túy (hoàn toàn phi tập trung, không có team dữ liệu trung tâm nào) hoặc một kiến trúc hoàn toàn tập trung (một team, một warehouse, không có tự chủ domain nào) trong thực tế. Mô hình phổ biến trong thực tế là hybrid: một lakehouse hoặc cloud data platform dùng chung (xem 13 — Cloud Data Platforms) cung cấp hạ tầng self-serve và ép buộc governance nền tảng, trong khi các domain có ngữ cảnh cao cụ thể (ví dụ: team fraud, team supply-chain) được trao ownership đối với sản phẩm dữ liệu của chính họ và chịu trách nhiệm về chất lượng — mà không cần mọi domain trong công ty đều phải có riêng nhân sự data engineering chuyên trách. Điều này phản ánh một sự thật rộng hơn trong thiết kế hệ thống phân tán: tập trung hóa hoàn toàn và phi tập trung hóa hoàn toàn đều là hai thái cực dễ mô tả nhưng khó vận hành tốt; câu trả lời bền vững cho hầu hết tổ chức nằm ở đâu đó giữa hai thái cực, được định hình bởi quy mô team, độ phức tạp của domain, và mức đầu tư nền tảng trung tâm mà công ty sẵn sàng chi trả.
Best Practices
- Đừng xây một lake mà không có catalog và mô hình ownership ngay từ ngày đầu. Yếu tố dự báo lớn nhất cho việc một lake trở thành swamp là trì hoãn governance “tới khi có nhiều dữ liệu hơn” — lúc đó backlog các dataset không tài liệu đã trở nên không thể quản lý.
- Mặc định dùng một open table format (Delta Lake, Iceberg, hoặc Hudi) thay vì file Parquet thô cho bất cứ thứ gì bạn sẽ truy vấn nhiều hơn một lần. Đảm bảo ACID và time-travel gần như miễn phí khi đã áp dụng, và việc gắn chúng ngược vào một lake file thô đã tồn tại tốn kém hơn nhiều so với xây dựng trên đó ngay từ đầu.
- Chọn open table format dựa trên engine và workload chủ đạo của bạn, không dựa trên trào lưu. Workload upsert/CDC nặng thường phù hợp với thiết kế của Hudi; khả năng tương tác đa engine rộng (đặc biệt ngoài hệ sinh thái Spark) thường phù hợp với Iceberg; công cụ Databricks-native sâu thì phù hợp với Delta Lake. Cả ba đang hội tụ về mặt tính năng, nên mức độ phù hợp hệ sinh thái quan trọng hơn so sánh tính năng thuần túy.
- Coi “data mesh” là một cam kết tổ chức, không phải một khoản mua công cụ. Áp dụng một platform self-serve mà không thực sự chuyển giao ownership và trách nhiệm cho các team domain sẽ tái tạo lại cùng một nút thắt cổ chai dưới cái tên mới.
- Đầu tư vào active metadata và một data catalog thực sự trước khi mở rộng domain ownership. Federated governance và self-serve discovery là bất khả thi nếu không ai tìm ra được dữ liệu gì tồn tại, ai sở hữu nó, hay nó có đáng tin hay không.
- Dùng mô hình medallion (bronze/silver/gold) bên trong lakehouse để tách các tầng thô, đã làm sạch, và sẵn sàng cho nghiệp vụ — xem 06 — Data Modeling & Warehousing — để sự linh hoạt schema-on-read ở tầng thô không rò rỉ vào các tầng mà người dùng nghiệp vụ thực sự truy vấn.
- Chỉ áp dụng data mesh đầy đủ cho những tổ chức có độ phức tạp domain và quy mô thực sự. Một startup mười người với một data engineer chẳng được lợi gì từ nghi thức federated governance; mô hình này chỉ xứng đáng với chi phí phát sinh ở quy mô nơi một team trung tâm đã thực sự chứng minh là nút thắt cổ chai.
- Định kỳ nén các file nhỏ và hết hạn các snapshot cũ. Table format tích lũy file nhỏ và các phiên bản lịch sử không còn được tham chiếu theo thời gian (đặc biệt dưới ingestion kiểu streaming/CDC); nếu không quản lý, điều này làm giảm hiệu năng truy vấn và tăng chi phí lưu trữ.
Tài liệu tham khảo
- Zhamak Dehghani, How to Move Beyond a Monolithic Data Lake to a Distributed Data Mesh (martinfowler.com, 2019)
- Zhamak Dehghani, Data Mesh: Delivering Data-Driven Value at Scale (O’Reilly, 2022)
- Armbrust et al., Lakehouse: A New Generation of Open Source Data Management (CIDR 2021, Databricks/UC Berkeley)
- Delta Lake — Documentation
- Apache Iceberg — Documentation
- Apache Hudi — Documentation
- Gartner — Data Fabric Architecture
- roadmap.sh — Data Engineer Roadmap
Part of the Data Engineer Roadmap knowledge base.
Overview
A data warehouse (see 06 — Data Modeling & Warehousing) demands that you know your schema before you load a single row: tables, columns, and types are defined up front, and anything that doesn’t fit gets rejected or forced into a workaround. That guarantee is exactly what makes warehouses fast and trustworthy for structured, well-understood, repeatedly-queried data — but it is a poor fit for the messier reality most organizations actually face: clickstream JSON with fields that change weekly, sensor telemetry, PDFs, images, third-party exports in whatever format the vendor felt like sending, logs nobody has fully parsed yet. Forcing all of that through a rigid schema at ingestion time either drops information silently or blocks the pipeline until someone updates a schema migration.
A data lake exists to remove that up-front requirement. It is a storage layer — historically HDFS, today almost always cloud object storage like Amazon S3, Azure Data Lake Storage (ADLS), or Google Cloud Storage (GCS) — that accepts data in its native format: structured, semi-structured, or unstructured, at low cost per gigabyte, with no schema enforced at write time. The schema is instead applied later, at query time, by whatever tool reads the data — a pattern called schema-on-read, in direct contrast to the warehouse’s schema-on-write. This inverts the cost of flexibility: you pay almost nothing to ingest anything, but you defer (and sometimes never pay off) the cost of making sense of it.
That flexibility is also the data lake’s best-known failure mode. A lake with no cataloging, no ownership, no quality checks, and no documentation degrades into what the industry calls a data swamp — a store where data technically exists but nobody trusts it, nobody knows what it means, and finding anything useful costs more effort than re-deriving it from scratch. The entire arc of “modern data architecture” covered in this note — lakehouses, open table formats, data mesh, data fabric, metadata-first design — is fundamentally an industry-wide response to one question: how do we keep the cost and flexibility advantages of a lake without it turning into a swamp?
Fundamentals
Data Lake vs Data Warehouse: The Core Trade-off
| Dimension | Data Lake | Data Warehouse |
|---|---|---|
| Schema | Schema-on-read (applied at query time) | Schema-on-write (enforced at load time) |
| Data types | Structured, semi-structured, unstructured (JSON, images, logs, Parquet, CSV, video) | Structured only, modeled into tables |
| Storage cost | Low — commodity object storage, cents per GB/month | Higher — often couples storage and compute pricing |
| Transactional guarantees | Historically none (plain files); added by table formats (see below) | Strong — ACID by design |
| Query performance | Variable; depends on file format, partitioning, engine | Optimized — columnar storage, indexes, query planner |
| Typical consumers | Data scientists, ML pipelines, ad hoc exploration, archival | BI dashboards, analysts, finance/reporting |
| Governance risk | High if unmanaged — the “data swamp” problem | Lower — schema and access control enforced structurally |
| Primary strength | Cost, flexibility, no data loss at ingestion | Reliability, query speed, business trust |
Neither side of this table is “obsolete” — a raw event stream you might reprocess five different ways over the next two years genuinely belongs in a lake; a finance team’s monthly revenue report genuinely belongs in a rigorously modeled warehouse table. What changed the industry was realizing that most organizations were building both, and paying twice.
Why Two Copies of Data Became a Problem
The pattern that dominated the 2010s was: land raw data in a lake, then run ETL jobs that copy, clean, and load a subset of that data into a warehouse for BI consumption. This works, but it creates real, compounding pain:
- Duplication and cost. The same data physically exists twice, in two different storage systems, often with two different pricing models.
- Staleness and drift. The warehouse copy lags the lake copy by however long the ETL job takes to run; the two can disagree, and reconciling them is its own project.
- Restricted access to the lake’s raw power. Data scientists who need the raw, unmodeled data for feature engineering or ML often can’t get an equivalent, governed, ACID-safe experience directly on the lake — they either work with stale exports or work directly on ungoverned files with none of the warehouse’s safety net.
- Two systems to secure, monitor, and pay engineers to operate. Every additional system is additional operational surface area: access control lists, backup policies, monitoring dashboards, and specialized staff for each.
This is the precise problem the lakehouse pattern was invented to solve.
Key Concepts
The Lakehouse: Warehouse Reliability on Lake Economics
A lakehouse is an architecture that stores data once, as files in cheap object storage (S3/ADLS/GCS), but adds a metadata and transaction layer on top of those files that gives you the guarantees people previously needed a warehouse for: ACID transactions, schema enforcement and evolution, indexing, and time travel — all while multiple different compute engines (Spark, Trino, Presto, Snowflake, Flink) can read and write the same underlying files concurrently. Databricks popularized both the term and the architecture, and the Databricks lakehouse paper presented at CIDR 2021 is the canonical technical case for it.
The mechanism that makes this possible is an open table format — a specification and set of metadata files sitting alongside plain columnar files (almost always Apache Parquet) that tracks which files currently constitute a table’s valid state, what the schema is, and the full history of changes. The three dominant open table formats are:
| Format | Origin | Notable mechanism |
|---|---|---|
| Delta Lake | Databricks (open-sourced 2019) | JSON-based transaction log (_delta_log) recording every add/remove of a Parquet file |
| Apache Iceberg | Netflix, now Apache project | Metadata tree of manifest files and snapshots; strong multi-engine adoption (Snowflake, Trino, Spark, Flink) |
| Apache Hudi | Uber, now Apache project | Timeline of commits; optimized for frequent upserts and incremental/CDC-style ingestion |
All three solve essentially the same problem — bolting ACID semantics onto a directory of immutable files — with different implementation details and different ecosystem strengths. Onehouse (founded by Hudi’s original creators) has also pushed interoperability efforts so that data written in one format can be read via another’s metadata layer, reducing the practical cost of picking “wrong.”
Why Open Table Formats — and Why “Open” Matters
Before table formats existed, a Parquet file sitting in S3 was just a file: no atomic multi-file commits, no concurrent-writer safety, no built-in way to know “what did this table look like an hour ago.” Table formats solve that, but the open part of “open table format” matters just as much as the transactional part, for two reasons:
- No vendor lock-in. Because the format is an open specification rather than a proprietary storage engine, the data itself is portable. You are not locked into a single vendor’s compute to read your own data.
- Multi-engine access to one copy of data. Spark can write a table, Trino can query it interactively, Snowflake can read it via external tables, and Flink can stream into it — all against the same physical files, with no ETL step required to move data between systems just so a different engine can touch it. This is the direct technical fix for the “two copies of data” problem described above.
Delta Lake in Detail: The Transaction Log
Delta Lake’s mechanism is a good concrete illustration of how any open table format achieves ACID on top of plain files, so it’s worth walking through specifically.
A Delta table is, physically, just a directory of Parquet files plus a subdirectory called _delta_log/ containing a sequence of JSON files (000000.json, 000001.json, …), one per committed transaction. Each JSON entry records an action — add a file, remove a file, update the schema, update table metadata — rather than rewriting any data file in place. Reading a table means reading the log to reconstruct the current set of “live” files, then reading only those Parquet files; writing means appending a new JSON entry to the log describing what changed. Because commits are new immutable log entries rather than in-place mutations, concurrent writers can be reconciled with optimistic concurrency control, and every prior state of the table is still reconstructable — it’s just log entries not yet garbage-collected.
That log is exactly what enables time travel: querying the table as it existed at a previous version or timestamp is simply a matter of replaying the log only up to that point.
-- Query the current state
SELECT * FROM sales;
-- Query as of a specific version number
SELECT * FROM sales VERSION AS OF 12;
-- Query as of a specific timestamp
SELECT * FROM sales TIMESTAMP AS OF '2026-06-01';
-- Roll a table back to a prior version
RESTORE TABLE sales TO VERSION AS OF 12;
Time travel is useful well beyond curiosity: reproducing an ML training run against the exact data snapshot used at the time, auditing “what did this report look like before the correction,” and recovering from a bad write (a botched DELETE or a bug that corrupted a batch) without restoring from a backup. Periodic maintenance (Delta’s VACUUM, Iceberg’s snapshot expiration) removes old files no longer referenced by any retained version, trading time-travel depth for storage cost.
Data Mesh: Decentralizing Ownership, Not Just Storage
Everything above is a technical answer to lake/warehouse fragmentation. Data mesh, introduced by Zhamak Dehghani in her 2019 article “How to Move Beyond a Monolithic Data Lake to a Distributed Data Mesh” and expanded in her book Data Mesh: Delivering Data-Driven Value at Scale, is an answer to a different, organizational problem: a single central data team becomes a bottleneck as a company grows.
The failure mode data mesh targets is familiar to anyone who has worked in a large organization with one central data engineering team: every domain (marketing, sales, logistics, fraud) submits a request to the central team to ingest their data, model it, and expose it; the central team has no deep context on any individual domain’s data semantics; the backlog grows faster than the team can hire; and domain experts who actually understand the data are stuck waiting instead of shaping it themselves. Data mesh’s answer is to invert who owns data, not just where it’s stored.
Dehghani frames data mesh around four principles:
| Principle | What it means | Problem it solves |
|---|---|---|
| Domain ownership | Each business domain (not a central team) owns the data it produces, end to end | Central team no longer needs deep expertise in every domain to model its data correctly |
| Data as a product | A domain’s data is treated as a product with an owner, SLAs, documentation, and quality guarantees — the domain team is accountable for consumers being able to trust and use it | Prevents domains from dumping raw, undocumented data over the wall and calling it “shared” |
| Self-serve data platform | A central platform team builds reusable infrastructure (storage provisioning, pipeline templates, catalog integration, access control) that domain teams use to build and serve their own data products | Domain teams don’t each need to reinvent infrastructure or hire their own platform specialists |
| Federated computational governance | Global standards (interoperability, security, compliance) are agreed upon across domains and enforced automatically/computationally, rather than through central manual gatekeeping | Keeps domains autonomous while still preventing chaos — standards are shared even though ownership isn’t |
Note what data mesh does not prescribe: it is not a specific technology, and it does not mean literally decentralizing storage into dozens of unconnected lakes. Many real implementations still use a shared lakehouse or shared object storage underneath — what’s decentralized is the organizational responsibility for producing, documenting, and being accountable for each domain’s data. This is precisely why the organizational problem matters as much as the technical one: mesh is as much a re-org as it is an architecture, and adopting the technology (open table formats, self-serve catalogs) without the organizational shift (domain teams actually taking ownership and being funded to do so) tends to fail.
Data Fabric: A Different Answer, Often Confused with Mesh
Data fabric is a related but distinct concept, and the two terms get conflated constantly, so the distinction is worth stating plainly:
- Data mesh is primarily organizational — it’s about who owns and is accountable for data (domain teams), enforced through federated governance.
- Data fabric is primarily technical — it’s an architecture that uses active, automated metadata (schemas, lineage, usage patterns, semantic relationships) and increasingly AI/ML-driven automation to discover, connect, and integrate data that stays distributed across many existing systems — data warehouses, lakes, operational databases, SaaS tools — without necessarily moving or centralizing it, and without necessarily changing who owns what.
Put differently: a data fabric could be layered on top of a purely centralized, single-team-owned architecture and still deliver value (better discoverability, automated lineage, easier cross-system integration) with zero organizational change. A data mesh, conversely, is meaningless without an organizational shift in ownership, regardless of what metadata tooling sits underneath it. The two are not mutually exclusive — a data-mesh implementation typically needs strong active metadata and a catalog to make federated governance and self-serve discovery actually work, which is where the concepts overlap and get confused in practice. Gartner popularized “data fabric” largely as a vendor/architecture category; “data mesh” originated as an organizational design proposal from a practitioner, which is a useful tell for which lens each term is really describing.
Metadata-First Architecture
Both data mesh’s federated governance and data fabric’s integration model depend on the same underlying shift: treating metadata — schemas, column-level lineage, data ownership, freshness and quality signals, access policies, business definitions — as a first-class product in its own right, not an afterthought bolted onto a catalog UI after the fact. In a metadata-first design, every dataset is expected to publish its schema, its owner, its SLAs, and its quality signals as part of being considered “done,” in the same way an API is expected to publish its contract. This is what lets a federated computational governance model in data mesh actually check compliance automatically (rather than a human auditor spot-checking spreadsheets), and it’s what lets a data fabric’s automation actually infer relationships and lineage across systems it wasn’t specifically configured to know about. This topic is developed further in 14 — Data Quality, Governance & Metadata.
The Hybrid Reality
Very few organizations run a pure data mesh (fully decentralized, no central data team at all) or a purely centralized architecture (one team, one warehouse, no domain autonomy) in practice. The common real-world pattern is a hybrid: a shared lakehouse or shared cloud data platform (see 13 — Cloud Data Platforms) provides the self-serve infrastructure and enforces baseline governance, while specific high-context domains (e.g., a fraud team, a supply-chain team) are given ownership over their own data products and are accountable for their quality — without every single domain in the company needing its own dedicated data engineering staff. This mirrors a broader truth in distributed systems design: full centralization and full decentralization are both extremes that are easy to describe and hard to run well; the durable answer for most organizations sits somewhere in between, shaped by team size, domain complexity, and how much central platform investment the company is willing to fund.
Best Practices
- Don’t build a lake without a catalog and ownership model from day one. The single biggest predictor of a lake becoming a swamp is postponing governance “until we have more data” — by then the backlog of undocumented datasets is unmanageable.
- Default to an open table format (Delta Lake, Iceberg, or Hudi) rather than raw Parquet files for anything you’ll query more than once. The ACID and time-travel guarantees are close to free once adopted, and retrofitting them onto an existing raw-file lake is much more expensive than building on them from the start.
- Pick an open table format based on your dominant engine and workload, not hype. Heavy upsert/CDC workloads often favor Hudi’s design; broad multi-engine interoperability (especially outside the Spark ecosystem) often favors Iceberg; deep Databricks-native tooling favors Delta Lake. All three are converging in capability, so ecosystem fit matters more than raw feature comparison.
- Treat “data mesh” as an organizational commitment, not a tooling purchase. Adopting a self-serve platform without genuinely transferring ownership and accountability to domain teams reproduces the same bottleneck under a new name.
- Invest in active metadata and a real data catalog before scaling out domain ownership. Federated governance and self-serve discovery are unworkable if nobody can find out what data exists, who owns it, or whether it’s trustworthy.
- Use the medallion pattern (bronze/silver/gold) inside your lakehouse to separate raw, cleaned, and business-ready layers — see 06 — Data Modeling & Warehousing — so schema-on-read flexibility at the raw layer doesn’t leak into the layers business users actually query.
- Reserve full data mesh adoption for organizations with real domain complexity and scale. A ten-person startup with one data engineer gains nothing from federated governance ceremony; the pattern earns its overhead at the size where a central team has demonstrably become a bottleneck.
- Periodically compact small files and expire old snapshots. Table formats accumulate small files and unreferenced historical versions over time (especially under streaming/CDC ingestion); left unmanaged this degrades query performance and inflates storage cost.
References
- Zhamak Dehghani, How to Move Beyond a Monolithic Data Lake to a Distributed Data Mesh (martinfowler.com, 2019)
- Zhamak Dehghani, Data Mesh: Delivering Data-Driven Value at Scale (O’Reilly, 2022)
- Armbrust et al., Lakehouse: A New Generation of Open Source Data Management (CIDR 2021, Databricks/UC Berkeley)
- Delta Lake — Documentation
- Apache Iceberg — Documentation
- Apache Hudi — Documentation
- Gartner — Data Fabric Architecture
- roadmap.sh — Data Engineer Roadmap