Giới thiệu về Data EngineeringIntroduction to Data Engineering
Thuộc bộ kiến thức Data Engineer Roadmap.
Tổng quan
Data engineering là ngành thiết kế, xây dựng và vận hành hạ tầng cùng các pipeline giúp di chuyển dữ liệu từ nơi nó được sinh ra đến nơi nó có thể được tin dùng — bởi analyst viết dashboard, bởi data scientist huấn luyện model, bởi các tính năng sản phẩm đọc dữ liệu từ API, và bởi ban lãnh đạo ra quyết định. Nếu bỏ qua các thuật ngữ hoa mỹ, công việc của một data engineer chỉ đơn giản là làm cho dữ liệu đến nơi một cách đáng tin cậy, ở định dạng dùng được, đúng thời điểm, và với chi phí mà doanh nghiệp chấp nhận được. Mọi thứ còn lại — công cụ cụ thể, nhà cung cấp cloud, framework orchestration — chỉ là chi tiết triển khai đặt trên nền tảng trách nhiệm cốt lõi đó.
Dữ liệu không tự sắp xếp bản thân nó. Nó được sinh ra từ những nguồn hỗn tạp, không đồng nhất: database ứng dụng, API bên thứ ba, cảm biến IoT, hệ thống theo dõi sự kiện người dùng, file spreadsheet, dữ liệu xuất từ các công cụ SaaS. Nếu để mặc, dữ liệu này sẽ bị phân mảnh (silo), thiếu nhất quán, và thường “chống đối” việc phân tích một cách chủ động (giá trị null, bản ghi trùng lặp, schema drift, timestamp lệch múi giờ). Data engineering tồn tại để hấp thụ sự hỗn loạn đó ngay ở tầng hạ tầng, để tất cả những người dùng phía sau — data scientist, analytics engineer, business analyst, hệ thống ML — có thể làm việc với dữ liệu sạch, được model hóa tốt và quản trị tốt, thay vì phải tự mình vật lộn với phần “ống nước”. Một bộ phận data engineering tốt gần như vô hình khi nó hoạt động đúng: dashboard tự làm mới đúng lịch, model luôn có feature mới, không ai bị gọi dậy lúc 3 giờ sáng vì một pipeline âm thầm làm rớt dữ liệu suốt hai tuần.
Bài viết này là điểm khởi đầu cho toàn bộ roadmap. Nó định nghĩa ngành, phân biệt data engineering với các vai trò liên quan, đi qua vòng đời (lifecycle) từ đầu đến cuối mà các chủ đề sau sẽ đào sâu từng phần, và liệt kê những kỹ năng mà công việc này thực sự đòi hỏi trong thực tế.
Kiến thức nền tảng
Data Engineer thực sự xây dựng những gì
Cụ thể, đầu ra công việc hàng ngày của một data engineer thường rơi vào một vài nhóm sau:
- Pipeline ingestion trích xuất dữ liệu từ hệ thống nguồn (database production qua CDC, API bên thứ ba, event stream, file được đẩy vào) và đưa nó vào một nơi lưu trữ bền vững.
- Hệ thống lưu trữ — data lake, data warehouse, lakehouse — được thiết kế với partitioning, định dạng file, và pattern truy cập phù hợp với cách dữ liệu sẽ thực sự được query.
- Logic transformation biến dữ liệu thô, mang hình dạng của nguồn, thành các bảng sạch, đã model hóa, sẵn sàng cho phân tích (đây là nơi SQL, dbt, và job Spark hoạt động).
- Orchestration sắp xếp trình tự tất cả những phần trên, retry khi thất bại, thực thi các phụ thuộc (dependency), và cảnh báo con người khi có sự cố.
- Quyết định về platform và tooling — chọn warehouse nào, orchestrator nào, hệ thống streaming nào, compute và storage được cấp phát và tính phí ra sao.
- Data quality, governance, và kiểm soát truy cập — đảm bảo dữ liệu chính xác, có tài liệu, dễ tìm kiếm (discoverable), và chỉ hiển thị cho những người được phép xem.
Không có phần nào trong số này hào nhoáng theo kiểu một tính năng hướng đến khách hàng, nhưng đây là nền tảng mà mọi thứ khác đứng trên đó. Một data scientist không thể xây dựng model tốt trên dữ liệu mà không ai tin tưởng được, và một analyst không thể đưa ra khuyến nghị có căn cứ từ những con số âm thầm mất dữ liệu trong quá trình load ban đêm.
Data Engineering vs Data Science vs Analytics Engineering
Ba vai trò này thường bị nhầm lẫn vì đều liên quan đến “dữ liệu”, nhưng chúng vận hành ở các tầng khác nhau trong stack và đòi hỏi trọng tâm kỹ năng khác nhau. Hiểu rõ ranh giới — và những chỗ chúng cố tình chồng lấn lên nhau — giúp bạn xác định phạm vi công việc của mình và hợp tác mà không giẫm chân lên trách nhiệm của người khác.
Một mô hình tư duy hữu ích: data engineer xây con đường, analytics engineer dựng biển báo và kẻ vạch làn, còn data scientist lái xe đến một nơi nào đó có ích. Data engineer chịu trách nhiệm để con đường tồn tại, được trải nhựa, và không tự nhiên đóng cửa lúc 2 giờ sáng; analytics engineer đảm bảo con đường được gắn nhãn đúng và dễ đi; data scientist dùng con đường đó để đến những nơi chưa ai từng đến.
| Khía cạnh | Data Engineer | Analytics Engineer | Data Scientist |
|---|---|---|---|
| Đầu ra chính | Pipeline và platform đáng tin cậy, có khả năng mở rộng | Data model sẵn sàng cho phân tích, có tài liệu | Model, thử nghiệm (experiment), insight |
| Công cụ cốt lõi | Python/Scala, Spark, Airflow/Dagster, cloud data platform, Kafka | SQL, dbt, semantic layer của BI | Python/R, pandas, scikit-learn, notebook |
| Mối quan tâm chính | Dữ liệu có đến đúng, đúng lúc, ở quy mô lớn, và bền vững không? | Dữ liệu có được model hóa, test, và đặt tên theo cách người dùng nghiệp vụ hiểu được không? | Dữ liệu nói lên điều gì, và ta có thể dự đoán được gì từ nó? |
| Làm việc gần nhất với | Hệ thống nguồn, hạ tầng, đội platform | Analyst, công cụ BI, stakeholder nghiệp vụ | Đội product, ML engineer, stakeholder nghiệp vụ |
| Đơn vị công việc điển hình | Một pipeline, job ingestion cho một bảng, một topic streaming | Một model dbt, một định nghĩa metric, một entry trong semantic layer | Một notebook, một model, một A/B test |
| Kiểu thất bại nếu thiếu vai trò này | Dữ liệu không đến, hoặc đến bị hỏng/trễ | Dữ liệu tồn tại nhưng mỗi team định nghĩa “revenue” một kiểu khác nhau | Model tồn tại nhưng được huấn luyện trên dữ liệu không đáng tin, chưa được model hóa |
Trong thực tế, các vai trò này nằm trên một dải phổ (spectrum) hơn là các ô hộp cứng nhắc. Analytics engineering xuất hiện vào khoảng giữa đến cuối những năm 2010 để lấp đầy khoảng trống mở ra giữa data engineer (ngày càng tập trung vào hạ tầng và ingestion dữ liệu thô ở quy mô lớn) và analyst (cần các bảng sạch, đã test, có tài liệu, nhưng trước đây thiếu các thực hành software engineering như version control và testing để xây dựng chúng một cách đáng tin cậy). Các công cụ như dbt khiến việc transform dựa trên SQL trở nên dễ tiếp cận đến mức analyst có thể “tốt nghiệp” lên tầng modeling này mà không cần học Spark hay Kafka. Các công ty nhỏ thường gộp cả ba vai trò vào một hoặc hai người; các tổ chức lớn hơn tách chúng thành các team riêng biệt với lịch trực on-call và roadmap của riêng mình.
Vì sao Data Engineering trở thành một ngành riêng biệt
Data engineering như một chức danh công việc được đặt tên riêng, khác biệt, khá non trẻ — nó chuyên nghiệp hóa trong suốt những năm 2010 và định hình rõ ràng vào những năm 2020. Trước đó, công việc này vẫn tồn tại nhưng bị rải rác dưới các chức danh khác: database administrator (DBA) tinh chỉnh database production và chạy các job ETL batch ban đêm; “data scientist” thường được kỳ vọng phải tự mình trích xuất, làm sạch, và dẫn dữ liệu của chính họ trước khi làm bất kỳ việc modeling thực sự nào; software engineer gắn thêm báo cáo analytics vào codebase ứng dụng như một việc phụ, nghĩ đến sau cùng.
Ba lực đã phá vỡ cách sắp xếp đó:
- Khối lượng, sự đa dạng, và tốc độ (volume, variety, velocity) đã vượt quá khả năng xử lý của các giải pháp một máy, một vai trò. Khi các công ty bắt đầu sinh ra dữ liệu từ clickstream trên web, ứng dụng mobile, thiết bị IoT, và hàng chục công cụ SaaS cùng lúc, một DBA duy nhất quản lý một database quan hệ không thể còn là toàn bộ câu chuyện dữ liệu. Lưu trữ và xử lý phân tán (Hadoop, sau đó là object storage trên cloud và warehouse MPP, rồi Spark) trở nên cần thiết chỉ để chứa và xử lý được dữ liệu.
- Mô hình “data scientist tự làm luôn ETL của mình” không mở rộng được (scale). Sẽ rất tốn kém và lãng phí nếu một người được đào tạo tiến sĩ về modeling phải dành 60–80% thời gian (một tỷ lệ thường được nhắc đến trong thực tế) để viết các script trích xuất dễ vỡ thay vì xây dựng model. Việc tách “làm cho dữ liệu có sẵn một cách đáng tin cậy” khỏi “làm điều gì đó có giá trị với dữ liệu” cho phép cả hai nhóm chuyên môn hóa và làm việc nhanh hơn ở phần việc tương ứng của mình.
- Các nền tảng cloud và “modern data stack” khiến việc chuyên môn hóa trở nên khả thi về mặt kinh tế. Warehouse được quản lý (BigQuery, Snowflake, Redshift), orchestration được quản lý, và compute trả tiền theo mức sử dụng nghĩa là một team nhỏ — thậm chí một người — có thể dựng lên hạ tầng mà trước đây cần cả một tổ chức ops riêng. Điều này hạ thấp rào cản đủ để data engineering trở thành một chuyên môn được định nghĩa rõ ràng, có nhu cầu cao, có thể dạy được, thay vì một tập hợp script chắp vá tùy tiện.
Kết quả là ngành như hiện nay: một chuyên môn vay mượn nhiều từ software engineering (version control, testing, CI/CD, code review) và hệ thống phân tán, áp dụng cụ thể cho bài toán đưa dữ liệu từ nơi nó sinh ra đến nơi nó tạo ra giá trị.
Khái niệm chính
Vòng đời Data Engineering (Data Engineering Lifecycle)
Joe Reis và Matt Housley, trong cuốn Fundamentals of Data Engineering, đóng khung ngành này xung quanh một vòng đời gồm năm giai đoạn tuần tự, cùng với một tập các “undercurrent” (dòng chảy ngầm) xuyên suốt chạm đến từng giai đoạn. Cách đóng khung này hữu ích vì nó cho bạn một tấm bản đồ: bất kỳ công cụ hay kỹ thuật nào bạn gặp đều có thể được đặt vào một giai đoạn, và bất kỳ vấn đề nào bạn gặp phải thường có thể truy về một undercurrent yếu (governance kém, thiếu kỷ luật orchestration, v.v.) hơn là một giai đoạn yếu.
| Giai đoạn | Điều gì xảy ra | Được đào sâu ở |
|---|---|---|
| Generation (Sinh dữ liệu) | Dữ liệu được sinh ra bởi hệ thống nguồn: database ứng dụng, cảm biến IoT, API bên thứ ba, công cụ SaaS, theo dõi sự kiện người dùng | 03 — Sources & Ingestion |
| Storage (Lưu trữ) | Dữ liệu được lưu trữ bền vững — trong database vận hành, data lake, warehouse, hoặc lakehouse — với các quyết định về định dạng, partitioning, và chi phí | 04–05 — Databases, 07 — Lakes & Modern Architectures |
| Ingestion (Nạp dữ liệu) | Dữ liệu được di chuyển từ hệ thống nguồn vào tầng lưu trữ/xử lý, theo kiểu batch hoặc streaming | 03 — Sources & Ingestion, 10–11 — Big Data & Streaming |
| Transformation (Biến đổi) | Dữ liệu thô được làm sạch, join, tổng hợp, và model hóa thành các dạng sẵn sàng cho phân tích | 06 — Modeling & Warehousing, 08 — ETL/ELT & Pipelines |
| Serving (Phục vụ) | Dữ liệu đã model hóa được đưa đến người dùng cuối: dashboard BI, SQL ad hoc, feature store cho ML, reverse ETL vào các công cụ SaaS | 19 — Analytics, BI & Data Serving |
Lưu ý rằng trong thực tế các giai đoạn này không hoàn toàn tuyến tính — serving có thể quay ngược lại nuôi generation (một gợi ý được hiển thị cho người dùng sinh ra một event clickstream mới), và transformation thường diễn ra một phần ngay trong ingestion (như trong các pipeline streaming tổng hợp dữ liệu ngay khi nó chảy qua). Vòng đời này là một mô hình tư duy hữu ích, không phải một sơ đồ pipeline cứng nhắc.
Generation
Việc sinh dữ liệu phần lớn nằm ngoài tầm kiểm soát của data engineer — bạn không sở hữu schema database ứng dụng mà team product thay đổi không báo trước, và bạn không kiểm soát được API bên thứ ba sẽ gửi gì cho bạn. Nhiệm vụ của data engineer ở giai đoạn này chủ yếu là hiểu các nguồn: schema của chúng, tần suất thay đổi, các cam kết về độ tin cậy, và khối lượng, để mọi thứ ở phía sau có thể được thiết kế dựa trên những kỳ vọng thực tế thay vì mong muốn viển vông.
Storage
Các quyết định về lưu trữ ảnh hưởng lan tỏa đến mọi thứ khác: định dạng file (Parquet so với CSV hay Avro), sơ đồ partitioning, dùng database vận hành hướng hàng (row-oriented) hay warehouse hướng cột (columnar), và cách bạn tách dữ liệu “nóng” được truy vấn thường xuyên khỏi dữ liệu “lạnh” lưu trữ dài hạn. Các quyết định lưu trữ tồi thường lộ ra vài tháng sau đó dưới dạng hóa đơn cloud tăng vọt hoặc các query mất mười phút thay vì mười giây.
Ingestion
Ingestion là về việc di chuyển dữ liệu đáng tin cậy từ A sang B. Ingestion theo batch (load đầy đủ/tăng dần vào ban đêm) dễ suy luận hơn nhưng gây ra độ trễ; ingestion streaming (qua Kafka, Kinesis, Pub/Sub) giảm độ trễ nhưng thêm vào độ phức tạp vận hành thực sự — thứ tự (ordering), ngữ nghĩa exactly-once, backpressure. Chọn giữa hai kiểu này là một sự đánh đổi thực sự, không phải chuyện cái này “hiện đại hơn” cái kia một cách phổ quát.
Transformation
Đây là nơi dữ liệu thô trở nên hữu ích: khử trùng lặp, ép kiểu (type casting), logic nghiệp vụ, join giữa các nguồn, và xây dựng các model dimensional hoặc bảng rộng (wide table) mà analyst và công cụ BI thực sự truy vấn. Đây cũng là nơi phần lớn công việc hàng ngày của data engineer và analytics engineer diễn ra, thường bằng SQL (thường qua dbt) hoặc Spark cho logic quy mô lớn hơn hoặc phức tạp hơn.
Serving
Dữ liệu mà không ai tiêu thụ thì đã tạo ra giá trị bằng không, bất kể pipeline xây dựng nó tinh vi đến đâu. Serving bao gồm công cụ BI, truy cập SQL ad hoc của analyst, feature store cho ML, embedded analytics trong sản phẩm, và reverse ETL (đẩy dữ liệu từ warehouse ngược lại vào các công cụ SaaS vận hành như Salesforce hay Braze). Công việc của một pipeline chưa xong khi “bảng đã tồn tại” — nó chỉ xong khi đúng người hoặc đúng hệ thống có thể sử dụng bảng đó một cách chính xác.
Undercurrents (Dòng chảy ngầm)
Reis và Housley mô tả năm undercurrent không thuộc về riêng giai đoạn nào mà chạy ngầm bên dưới và xuyên suốt tất cả các giai đoạn. Sự yếu kém ở một undercurrent thường gây ra kiểu thất bại không xuất hiện trong demo nhưng lộ ra sau ba tháng chạy production.
- Security (Bảo mật) — kiểm soát truy cập, mã hóa, và nguyên tắc least privilege áp dụng cho mọi giai đoạn, không phải gắn thêm vào lúc cuối.
- Data management / governance (Quản trị dữ liệu) — data quality, metadata, catalog, lineage, và tuân thủ (như GDPR) khiến dữ liệu đáng tin cậy và dễ tìm kiếm thay vì một hộp đen. Được đào sâu ở 14 — Data Quality, Governance & Metadata.
- DataOps — áp dụng các thực hành kiểu DevOps (tự động hóa, giám sát, xử lý sự cố, observability) cụ thể cho data pipeline, để các thất bại được phát hiện qua cảnh báo thay vì qua việc một analyst nhận ra dashboard trông có gì đó sai sai.
- Data architecture (Kiến trúc dữ liệu) — thiết kế có chủ đích, liên tục tiến hóa về cách storage, processing, và serving system khớp với nhau, cân bằng nhu cầu hiện tại với sự linh hoạt cho tương lai.
- Orchestration — điều phối phụ thuộc, retry, và lịch trình xuyên suốt mọi giai đoạn của vòng đời, thường qua các công cụ như Airflow, Dagster, hoặc Prefect. Được đào sâu ở 09 — Orchestration.
- Software engineering — version control, testing, CI/CD, và code review áp dụng cho pipeline và code transformation, vì một pipeline “chỉ chạy một script mà ai đó sửa tay” thì không thể mở rộng hay tồn tại được khi tác giả của nó rời team.
Kỹ năng và trách nhiệm của một Data Engineer
Bộ kỹ năng trải dài từ năng lực kỹ thuật cứng đến một kỹ năng mềm thực sự bị đánh giá thấp: hiểu được stakeholder thực sự cần gì từ dữ liệu trước khi xây dựng bất cứ thứ gì cho họ.
| Nhóm | Ví dụ | Vì sao quan trọng |
|---|---|---|
| Programming | Python, SQL (bắt buộc), Scala/Java (phổ biến ở các công ty dùng nhiều Spark) | SQL là ngôn ngữ hàng ngày của transformation; Python kết dính ingestion, orchestration, và scripting; Scala/Java xuất hiện trong code Spark hoặc Kafka nhạy cảm về hiệu năng |
| Hiểu biết về hệ thống phân tán | Cách Spark partition và shuffle dữ liệu, cách Kafka đảm bảo thứ tự trong một partition, các đánh đổi của CAP theorem | Pipeline thất bại ở production vì các lý do thuộc về hệ thống phân tán (skew, retry, thất bại một phần) mà không bao giờ xuất hiện khi chạy trên laptop với tập dữ liệu mẫu nhỏ |
| Cloud platform | AWS, GCP, hoặc Azure — dịch vụ storage, compute, và data service được quản lý của họ | Gần như toàn bộ hạ tầng dữ liệu hiện đại được cấp phát và tính phí trên cloud platform; hiểu mô hình chi phí quan trọng không kém hiểu API |
| Database internals | Indexing, query planning, normalization so với denormalization, OLTP so với OLAP | Chọn và tinh chỉnh storage đòi hỏi hiểu vì sao một query chậm, không chỉ biết là nó chậm |
| Công cụ orchestration | Airflow, Dagster, Prefect, dbt | Gần như mọi pipeline không tầm thường đều cần lập lịch, retry, và quản lý phụ thuộc; cron job viết tay nhanh chóng không mở rộng được |
| Data modeling | Dimensional modeling, normalization, slowly changing dimension | Quyết định liệu người dùng phía sau có thực sự trả lời được câu hỏi nghiệp vụ mà không cần SQL “anh hùng” hay không |
| Giao tiếp với stakeholder | Thu thập yêu cầu, dịch “chúng tôi cần báo cáo tốt hơn” thành schema và SLA cụ thể | Những pipeline tốn kém nhất là những pipeline mà ngay từ đầu không ai hỏi đúng câu; hiểu sai yêu cầu lãng phí nhiều thời gian hơn bất kỳ sai lầm kỹ thuật nào |
Dòng cuối cùng đáng được nhấn mạnh: một pipeline xuất sắc về mặt kỹ thuật nhưng trả lời sai câu hỏi, hoặc bỏ lỡ một SLA không ai viết ra nhưng ai cũng mặc định, vẫn là một thất bại bất kể code sạch đến đâu. Những data engineer giỏi dành thời gian thực sự để trò chuyện với analyst, scientist, và stakeholder nghiệp vụ đang tiêu thụ dữ liệu của họ — không chỉ với hạ tầng.
Bản đồ phần còn lại của Roadmap này
Bài giới thiệu này nằm ở đầu một chuỗi dài hơn. Nói một cách khái quát, roadmap đi từ tooling nền tảng, qua cơ chế lấy và lưu trữ dữ liệu, vào modeling và kiến trúc, rồi đến pipeline, big data, và các vấn đề platform, và cuối cùng đến các tầng tin cậy/vận hành và tầng giá trị giúp mọi thứ bền vững lâu dài:
- Nền tảng programming & tooling (02) — baseline Python, SQL, và command-line mà mọi thứ khác giả định bạn đã có.
- Sources & ingestion (03) — dữ liệu đến từ đâu và được di chuyển như thế nào.
- Databases (04–05) — nội bộ lưu trữ quan hệ và phi quan hệ.
- Modeling & warehousing (06) — biến dữ liệu thô thành cấu trúc sẵn sàng cho phân tích.
- Lakes & modern architectures (07) — lakehouse, kiến trúc medallion, các định dạng table mở.
- Pipelines & orchestration (08–09) — các pattern ETL/ELT và lập lịch chúng một cách đáng tin cậy.
- Big data & streaming (10–11) — xử lý phân tán và pipeline thời gian thực.
- Platform (12–13) — container/Kubernetes cho data workload và cloud data platform.
- Trust & ops (14–18) — data quality, governance, metadata, security, monitoring.
- Value delivery & tổng kết (19–20) — serving dữ liệu qua BI/analytics, và kết ở lộ trình sự nghiệp & best practices.
Mỗi chủ đề trong số đó sẽ đào sâu vào một lát cắt của vòng đời đã mô tả ở trên; nhiệm vụ của bài viết này chỉ là cung cấp khung tư duy giúp tất cả những phần còn lại khớp với nhau.
Best Practices
- Ưu tiên độ tin cậy trước sự khéo léo. Một pipeline “nhàm chán” chạy mỗi ngày mà không cần can thiệp tốt hơn một pipeline tinh vi nhưng đòi hỏi kỹ sư on-call phải trông chừng nó. Ưu tiên các thiết kế idempotent, có thể retry, được giám sát tốt hơn là những thiết kế thanh lịch nhưng dễ vỡ.
- Đối xử với code pipeline như phần mềm production. Version control, code review, automated testing, và CI/CD không phải là những thứ “thêm cho vui” đối với “chỉ là một script” — pipeline là hệ thống production và nên được phát triển với cùng mức độ kỷ luật như code hướng đến khách hàng.
- Thiết kế storage và schema theo cách dữ liệu sẽ được query, không chỉ theo cách nó đến. Dữ liệu mang hình dạng của nguồn hiếm khi là dữ liệu mang hình dạng để query; đầu tư vào modeling (dimensional model, partitioning hợp lý) sẽ đem lại lợi ích ngay lần đầu tiên một analyst cần câu trả lời nhanh.
- Mặc định ưu tiên managed service hơn là tự host hạ tầng. Tự vận hành cluster Kafka hay Spark của riêng bạn là một lựa chọn chính đáng, nhưng nó nên là một lựa chọn có chủ đích dựa trên nhu cầu thực sự, không phải điểm khởi đầu mặc định — managed service đánh đổi một phần chi phí và sự linh hoạt để giảm đáng kể gánh nặng vận hành.
- Xây dựng observability vào pipeline ngay từ đầu, chứ không phải sau sự cố mất dữ liệu âm thầm đầu tiên. Kiểm tra số lượng dòng, kiểm tra độ tươi (freshness) của dữ liệu, và cảnh báo thay đổi schema bắt được những thất bại mà chỉ dashboard thôi thì không thể.
- Trò chuyện với stakeholder trước khi viết code. Làm rõ dataset cần trả lời câu hỏi gì, “đúng” có nghĩa là gì đối với nó, và SLA về độ tươi thực sự cần thiết là gì — điều này ngăn việc xây dựng nhầm thứ (dù xuất sắc về mặt kỹ thuật).
- Tài liệu hóa data model và định nghĩa, đặc biệt là các metric dùng chung như “revenue” hay “active user”. Sự mơ hồ ở đây gây ra nhiều ma sát tổ chức hơn hầu hết mọi vấn đề kỹ thuật khác.
Tài liệu tham khảo
- Joe Reis & Matt Housley, Fundamentals of Data Engineering (O’Reilly, 2022)
- roadmap.sh — Data Engineer Roadmap
- dbt Labs Blog — What is Analytics Engineering?
- dbt Labs — The Analytics Engineering Guide
- Martin Kleppmann, Designing Data-Intensive Applications (O’Reilly, 2017)
- Google Cloud — What is a Data Engineer?
- AWS — What is Data Engineering?
Part of the Data Engineer Roadmap knowledge base.
Overview
Data engineering is the discipline of designing, building, and operating the infrastructure and pipelines that move data from where it is produced to where it can be trusted and used — by analysts writing dashboards, by data scientists training models, by product features that read from an API, and by executives making decisions. If you strip away the buzzwords, a data engineer’s job is to make data arrive reliably, in a usable shape, on time, and at a cost the business can afford. Everything else — the specific tools, the cloud provider, the orchestration framework — is implementation detail layered on top of that core responsibility.
Data does not organize itself. It is generated by messy, heterogeneous sources: application databases, third-party APIs, IoT sensors, event trackers, spreadsheets, SaaS exports. Left alone, this data is siloed, inconsistent, and often actively hostile to analysis (nulls, duplicate records, schema drift, clock-skewed timestamps). Data engineering exists to absorb that chaos at the infrastructure layer so that everyone downstream — data scientists, analytics engineers, business analysts, ML systems — can work with clean, well-modeled, well-governed data instead of fighting the plumbing themselves. A good data engineering function is largely invisible when it works: dashboards refresh on schedule, models get fresh features, nobody is paged at 3 a.m. because a pipeline silently dropped rows for two weeks.
This note is the entry point to the rest of the roadmap. It defines the discipline, distinguishes it from adjacent roles, walks through the end-to-end lifecycle that later topics dig into individually, and lists the skills the job actually demands in practice.
Fundamentals
What Data Engineers Actually Build
Concretely, a data engineer’s day-to-day output tends to fall into a few buckets:
- Ingestion pipelines that extract data from source systems (production databases via CDC, third-party APIs, event streams, file drops) and land it somewhere durable.
- Storage systems — data lakes, warehouses, lakehouses — designed with the right partitioning, file formats, and access patterns for how the data will actually be queried.
- Transformation logic that turns raw, source-shaped data into clean, modeled, analysis-ready tables (this is where SQL, dbt, and Spark jobs live).
- Orchestration that sequences all of the above, retries failures, enforces dependencies, and alerts humans when something breaks.
- Platform and tooling decisions — which warehouse, which orchestrator, which streaming system, how compute and storage are provisioned and paid for.
- Data quality, governance, and access control — making sure the data is correct, documented, discoverable, and only visible to those who should see it.
None of this is glamorous in the way a customer-facing feature is, but it is the foundation everything else stands on. A data scientist cannot build a good model on data nobody can trust, and an analyst cannot make a defensible recommendation from numbers that quietly drop rows during nightly loads.
Data Engineering vs Data Science vs Analytics Engineering
These three roles are frequently confused because they all touch “data,” but they operate at different layers of the stack and require different skill emphases. Understanding the boundaries — and where they deliberately overlap — helps you scope your own work and collaborate without stepping on someone else’s responsibilities.
A useful mental model: data engineers build the roads, analytics engineers put up the road signs and paint the lane markings, and data scientists drive somewhere useful. The data engineer is responsible for the road existing at all, being paved, and not randomly closing at 2 a.m.; the analytics engineer makes sure the road is labeled correctly and easy to navigate; the data scientist uses it to get somewhere no one has been before.
| Dimension | Data Engineer | Analytics Engineer | Data Scientist |
|---|---|---|---|
| Primary output | Reliable, scalable pipelines and platforms | Analysis-ready, documented data models | Models, experiments, insights |
| Core tools | Python/Scala, Spark, Airflow/Dagster, cloud data platforms, Kafka | SQL, dbt, BI semantic layers | Python/R, pandas, scikit-learn, notebooks |
| Primary concern | Is the data arriving correctly, on time, at scale, and durably? | Is the data modeled, tested, and named in a way business users understand? | What does the data tell us, and what can we predict from it? |
| Works closest with | Source systems, infrastructure, platform teams | Analysts, BI tools, business stakeholders | Product, ML engineers, business stakeholders |
| Typical unit of work | A pipeline, a table’s ingestion job, a streaming topic | A dbt model, a metric definition, a semantic layer entry | A notebook, a model, an A/B test |
| Failure mode if absent | No data arrives, or arrives corrupted/late | Data exists but every team defines “revenue” differently | Models exist but are trained on unreliable, unmodeled data |
In practice these roles sit on a spectrum rather than in hard boxes. Analytics engineering emerged in the mid-to-late 2010s specifically to fill the gap that opened up between data engineers (who were increasingly focused on infrastructure and raw ingestion at scale) and analysts (who needed clean, tested, documented tables but historically lacked software-engineering practices like version control and testing to build them reliably). Tools like dbt made SQL-based transformation approachable enough that analysts could “graduate” into this modeling layer without needing to learn Spark or Kafka. Smaller companies often collapse all three roles into one or two people; larger organizations split them into dedicated teams with their own on-call rotations and roadmaps.
Why Data Engineering Emerged as a Discipline
Data engineering as a named, distinct job title is relatively young — it professionalized through the 2010s and solidified in the 2020s. Before that, the work existed but was scattered across other titles: database administrators (DBAs) tuned production databases and ran nightly batch ETL jobs; “data scientists” were frequently expected to also extract, clean, and pipe their own data before doing any actual modeling; software engineers bolted analytics reporting onto application codebases as an afterthought.
Three forces broke that arrangement:
- Volume, variety, and velocity outgrew single-machine and single-role solutions. As companies started generating data from web clickstreams, mobile apps, IoT devices, and dozens of SaaS tools simultaneously, a single DBA managing one relational database could no longer be the whole data story. Distributed storage and processing (Hadoop, then cloud object storage and MPP warehouses, then Spark) became necessary just to hold and process the data at all.
- The “data scientist who also does their own ETL” model didn’t scale. It is expensive and wasteful for a PhD-trained modeler to spend 60–80% of their time (a commonly cited real-world split) writing brittle extraction scripts instead of building models. Separating “make the data reliably available” from “do something valuable with the data” let both groups specialize and get faster at their respective jobs.
- Cloud platforms and the “modern data stack” made specialization economically viable. Managed warehouses (BigQuery, Snowflake, Redshift), managed orchestration, and pay-as-you-go compute meant a small team — or even one person — could stand up infrastructure that previously required a dedicated ops org. This lowered the barrier enough that data engineering could become a well-defined, in-demand, teachable specialty rather than an ad hoc set of duct-taped scripts.
The result is the field as it exists today: a discipline that borrows heavily from software engineering (version control, testing, CI/CD, code review) and distributed systems, applied specifically to the problem of getting data from where it’s born to where it creates value.
Key Concepts
The Data Engineering Lifecycle
Joe Reis and Matt Housley, in Fundamentals of Data Engineering, frame the discipline around a lifecycle of five sequential stages, with a set of cross-cutting “undercurrents” that touch every one of them. This framing is useful because it gives you a map: any tool or technique you encounter can be placed at a stage, and any problem you hit can usually be traced to a weak undercurrent (poor governance, no orchestration discipline, etc.) rather than a weak stage.
| Stage | What happens | Deepened in |
|---|---|---|
| Generation | Data is produced by source systems: application databases, IoT sensors, third-party APIs, SaaS tools, user event tracking | 03 — Sources & Ingestion |
| Storage | Data is persisted durably — in operational databases, data lakes, warehouses, or lakehouses — with decisions about format, partitioning, and cost | 04–05 — Databases, 07 — Lakes & Modern Architectures |
| Ingestion | Data is moved from source systems into the storage/processing layer, in batch or streaming fashion | 03 — Sources & Ingestion, 10–11 — Big Data & Streaming |
| Transformation | Raw data is cleaned, joined, aggregated, and modeled into analysis-ready shapes | 06 — Modeling & Warehousing, 08 — ETL/ELT & Pipelines |
| Serving | Modeled data is exposed to its consumers: BI dashboards, ad hoc SQL, ML feature stores, reverse ETL into SaaS tools | 19 — Analytics, BI & Data Serving |
Notice that these stages are not strictly linear in practice — serving can feed back into generation (a recommendation shown to a user generates a new clickstream event), and transformation often happens partly during ingestion (as in streaming pipelines that aggregate on the fly). The lifecycle is a useful mental model, not a rigid pipeline diagram.
Generation
Data generation is largely outside the data engineer’s control — you don’t own the application database schema the product team changes without warning, and you don’t control what a third-party API decides to send you. The data engineer’s job at this stage is mostly to understand the sources: their schema, their change cadence, their reliability guarantees, and their volume, so that everything downstream can be designed around realistic expectations rather than wishful ones.
Storage
Storage decisions ripple through everything else: file format (Parquet vs. CSV vs. Avro), partitioning scheme, whether you use a row-oriented operational database or a columnar warehouse, and how you separate “hot” frequently-queried data from “cold” archival data. Bad storage decisions show up months later as runaway cloud bills or queries that take ten minutes instead of ten seconds.
Ingestion
Ingestion is about reliably getting data to move from A to B. Batch ingestion (nightly full/incremental loads) is simpler to reason about but introduces latency; streaming ingestion (via Kafka, Kinesis, Pub/Sub) reduces latency but adds real operational complexity — ordering, exactly-once semantics, backpressure. Choosing between them is a genuine trade-off, not a matter of one being universally “more modern” than the other.
Transformation
This is where raw data becomes useful: deduplication, type casting, business logic, joining across sources, and building the dimensional or wide-table models that analysts and BI tools actually query. This is also where a large share of a data engineer’s and analytics engineer’s day-to-day work happens, typically in SQL (often via dbt) or in Spark for larger-scale or more complex logic.
Serving
Data that nobody consumes has produced zero value, no matter how elegant the pipeline that built it. Serving covers BI tools, ad hoc analyst SQL access, ML feature stores, embedded analytics in products, and reverse ETL (pushing warehouse data back into operational SaaS tools like Salesforce or Braze). A pipeline’s job isn’t done at “the table exists” — it’s done when the right person or system can use that table correctly.
Undercurrents
Reis and Housley describe five undercurrents that don’t belong to any single stage but instead run underneath and through all of them. Weakness in an undercurrent tends to cause the kind of failure that doesn’t show up in a demo but shows up three months into production.
- Security — access control, encryption, and the principle of least privilege applied to every stage, not bolted on at the end.
- Data management / governance — data quality, metadata, cataloging, lineage, and compliance (things like GDPR) that make data trustworthy and discoverable rather than a black box. Deepened in 14 — Data Quality, Governance & Metadata.
- DataOps — applying DevOps-style practices (automation, monitoring, incident response, observability) to data pipelines specifically, so failures are caught by alerts rather than by an analyst noticing a dashboard looks wrong.
- Data architecture — the deliberate, evolving design of how storage, processing, and serving systems fit together, balancing today’s needs against tomorrow’s flexibility.
- Orchestration — coordinating dependencies, retries, and scheduling across every stage of the lifecycle, typically via tools like Airflow, Dagster, or Prefect. Deepened in 09 — Orchestration.
- Software engineering — version control, testing, CI/CD, and code review applied to pipelines and transformation code, because a pipeline that “just runs a script someone edits by hand” does not scale or survive its author leaving the team.
Skills and Responsibilities of a Data Engineer
The skill set spans hard technical ability and a genuinely underrated soft skill: understanding what stakeholders actually need before building anything for them.
| Category | Examples | Why it matters |
|---|---|---|
| Programming | Python, SQL (non-negotiable), Scala/Java (common in Spark-heavy shops) | SQL is the daily language of transformation; Python glues together ingestion, orchestration, and scripting; Scala/Java appear in performance-sensitive Spark or Kafka code |
| Distributed systems literacy | How Spark partitions and shuffles data, how Kafka guarantees ordering within a partition, CAP-theorem trade-offs | Pipelines fail in production for distributed-systems reasons (skew, retries, partial failures) that never show up on a laptop with a small sample dataset |
| Cloud platforms | AWS, GCP, or Azure — their storage, compute, and managed data services | Nearly all modern data infrastructure is provisioned and billed on cloud platforms; understanding cost models is as important as understanding APIs |
| Database internals | Indexing, query planning, normalization vs. denormalization, OLTP vs. OLAP | Choosing and tuning storage requires knowing why a query is slow, not just that it is |
| Orchestration tools | Airflow, Dagster, Prefect, dbt | Nearly every non-trivial pipeline needs scheduling, retries, and dependency management; hand-rolled cron jobs stop scaling quickly |
| Data modeling | Dimensional modeling, normalization, slowly changing dimensions | Determines whether downstream consumers can actually answer business questions without heroic SQL |
| Stakeholder communication | Requirements gathering, translating “we need better reporting” into concrete schemas and SLAs | The most expensive pipelines are the ones nobody asked for correctly the first time; misunderstanding requirements wastes more time than any technical mistake |
That last row deserves emphasis: a technically excellent pipeline that answers the wrong question, or that misses an SLA nobody wrote down but everyone assumed, is a failure regardless of how clean the code is. Strong data engineers spend real time in conversation with the analysts, scientists, and business stakeholders consuming their data — not just with the infrastructure.
A Map of the Rest of This Roadmap
This introduction sits at the top of a longer sequence. Roughly, the roadmap moves from foundational tooling, through the mechanics of getting and storing data, into modeling and architecture, then into pipelines, big data, and platform concerns, and finally into the trust/operations and value-delivery layers that make everything sustainable long-term:
- Programming & tooling foundations (02) — the Python, SQL, and command-line baseline everything else assumes.
- Sources & ingestion (03) — where data comes from and how it gets moved.
- Databases (04–05) — relational and non-relational storage internals.
- Modeling & warehousing (06) — turning raw data into analysis-ready structures.
- Lakes & modern architectures (07) — lakehouses, medallion architecture, open table formats.
- Pipelines & orchestration (08–09) — ETL/ELT patterns and scheduling them reliably.
- Big data & streaming (10–11) — distributed processing and real-time pipelines.
- Platform (12–13) — containers/Kubernetes for data workloads and cloud data platforms.
- Trust & ops (14–18) — data quality, governance, metadata, security, monitoring.
- Value delivery & wrap-up (19–20) — serving data through BI/analytics, then closing with career path and best practices.
Each of those topics goes deep into one slice of the lifecycle described above; this note’s job was only to give you the frame that makes the rest of them click together.
Best Practices
- Optimize for reliability before cleverness. A boring pipeline that runs every day without intervention beats a clever one that requires an on-call engineer to babysit it. Favor idempotent, retryable, well-monitored designs over elegant-but-fragile ones.
- Treat pipeline code like production software. Version control, code review, automated testing, and CI/CD are not optional extras for “just a script” — pipelines are production systems and should be developed with the same rigor as customer-facing code.
- Design storage and schemas for how data will be queried, not just how it arrives. Source-shaped data is rarely query-shaped data; investing in modeling (dimensional models, sensible partitioning) pays for itself the first time an analyst needs an answer quickly.
- Prefer managed services over self-hosted infrastructure by default. Running your own Kafka cluster or Spark cluster is a legitimate choice, but it should be a deliberate one made because of a real requirement, not the default starting point — managed services trade some cost and flexibility for dramatically less operational burden.
- Build observability into pipelines from day one, not after the first silent data-loss incident. Row counts, freshness checks, and schema-change alerts catch the failures that dashboards alone won’t.
- Talk to your stakeholders before writing code. Clarify what question a dataset needs to answer, what “correct” means for it, and what freshness SLA is actually required — this prevents building the wrong (if technically excellent) thing.
- Document data models and definitions, especially shared metrics like “revenue” or “active user.” Ambiguity here causes more organizational friction than almost any technical issue.
References
- Joe Reis & Matt Housley, Fundamentals of Data Engineering (O’Reilly, 2022)
- roadmap.sh — Data Engineer Roadmap
- dbt Labs Blog — What is Analytics Engineering?
- dbt Labs — The Analytics Engineering Guide
- Martin Kleppmann, Designing Data-Intensive Applications (O’Reilly, 2017)
- Google Cloud — What is a Data Engineer?
- AWS — What is Data Engineering?