← DevSecOps← DevSecOps
DevSecOpsDevSecOps19 Th7, 2026Jul 19, 202627 phút đọc20 min read

SIEM & Tự động hóa bảo mậtSIEM & Security Automation

Thuộc bộ kiến thức DevSecOps Roadmap.

Tổng quan

Monitoring & Logging đề cập đến việc từng hệ thống riêng lẻ phát ra metric, log, trace như thế nào. Điều đó cần thiết nhưng chưa đủ: một lần đăng nhập thất bại trên một máy chỉ là nhiễu; nhưng hàng trăm lần đăng nhập thất bại trên mười máy, tiếp theo là một lần đăng nhập thành công và một tài khoản admin mới được tạo — đó là một cuộc tấn công. Để nhìn ra pattern đó, cần gom log từ mọi hệ thống về một chỗ và đặt câu hỏi xuyên suốt nhiều nguồn — đây chính xác là việc mà SIEM (Security Information and Event Management) làm.

SIEM tổng hợp các sự kiện liên quan đến bảo mật từ khắp môi trường (server, thiết bị mạng, cloud API, ứng dụng, endpoint), chuẩn hóa chúng về một cấu trúc chung, correlate (tương quan) chúng với các rule và baseline hành vi, rồi trình bày kết quả dưới dạng alert và dashboard mà một analyst con người có thể hành động dựa trên đó. Nó là “hệ thần kinh” của một chương trình phát hiện tấn công (detection) — nơi biến “có gì đó xảy ra ở đâu đó” thành “đây là tín hiệu đã được ưu tiên cho thấy có điều xấu đang diễn ra”.

Phát hiện ra mối đe dọa mới chỉ là một nửa công việc. Khi một alert bật lên, ai đó (hoặc thứ gì đó) phải triage nó, quyết định nó có thật hay không, và hành động — vô hiệu hóa tài khoản, chặn một IP, cô lập một host. Làm việc này thủ công, mỗi lần, cho mọi alert, sẽ không scale được khi số lượng alert vượt quá vài chục mỗi ngày. SOAR (Security Orchestration, Automation, and Response) ra đời để lấp khoảng trống đó: các nền tảng SOAR mã hóa những phần lặp lại của quy trình phản ứng thành các playbook tự động, để analyst dành thời gian cho những quyết định cần phán đoán thay vì click lặp đi lặp lại.

Bài viết này bao quát cách SIEM và SOAR phối hợp với nhau — cách correlation rule biến sự kiện thô thành alert, cách framework MITRE ATT&CK được dùng để suy luận về độ bao phủ phát hiện (detection coverage), cách playbook của SOAR tự động hóa các phản ứng phổ biến, và cách alert đi qua quy trình triage rồi vào công việc điều tra/ngăn chặn sâu hơn được trình bày ở Incident Response and Digital Forensics.

Kiến thức nền tảng

SIEM khác gì so với log management thuần túy

Quản lý log tập trung (một cụm ELK dùng thuần để tìm kiếm log, hay một dịch vụ log aggregation) và một SIEM trông tương tự nhau ở bề mặt — cả hai đều nạp log từ nhiều nguồn vào một nơi có thể tìm kiếm — nhưng chúng giải quyết những vấn đề khác nhau.

Khía cạnhLog management thuần túySIEM
Mục đích chínhTìm kiếm, debug, xử lý sự cố vận hànhPhát hiện bảo mật, alerting, báo cáo compliance
Data modelThường thô/không có cấu trúc thống nhất, field đặc thù theo từng nguồnĐược chuẩn hóa về một schema chung (vd: “login failure” trông giống nhau bất kể đến từ hệ thống nào)
CorrelationThủ công — kỹ sư tự chạy query ad hocCó sẵn engine correlation liên tục đánh giá rule trên nhiều nguồn gần thời gian thực
AlertingThường thiếu hoặc gắn thêm sauTính năng cốt lõi — rule/analytics tự động sinh alert đã được ưu tiên
Enrichment ngữ cảnhHiếmPhổ biến — alert được enrich với threat intelligence, độ quan trọng của asset, điểm rủi ro của user
Yếu tố quyết định retentionNhu cầu vận hành (vài ngày đến vài tuần)Thường theo yêu cầu compliance (vài tháng đến vài năm, vd: PCI DSS yêu cầu 1 năm)
Người dùng điển hìnhDeveloper, SRESOC analyst, security engineer, auditor
Ví dụ công cụELK/OpenSearch thuần, Loki, CloudWatch LogsSplunk Enterprise Security, Elastic Security, Microsoft Sentinel, IBM QRadar

Trên thực tế, nhiều tổ chức xây cả hai trên cùng một nền tảng lưu trữ — ví dụ Elastic Stack có thể vừa dùng để tìm kiếm log thuần túy, vừa (khi thêm Security app, correlation rule, detection content) đóng vai trò SIEM. Điểm khác biệt không nằm ở storage engine mà ở việc có tồn tại lớp correlation/detection, case management, và workflow cho analyst được xây quanh nó hay không.

Các năng lực cốt lõi của SIEM

Correlation rule và use case

Một correlation rule (đôi khi gọi là “detection” hay “analytic”) diễn tả một pattern trên một hoặc nhiều sự kiện, mà khi kết hợp lại, cho thấy điều gì đó đáng để analyst chú ý. Rule có thể đơn giản chỉ khớp một sự kiện đơn lẻ, cho đến các chuỗi nhiều giai đoạn kéo dài từ vài phút đến vài ngày.

Cấu trúc điển hình của một correlation rule:

WHEN <loại sự kiện/nguồn> xảy ra
WHERE <điều kiện field khớp>
[AND correlate với <loại sự kiện khác>]
[WITHIN <khoảng thời gian>]
[GROUPED BY <entity, vd: user hoặc host>]
[THRESHOLD <số lượng/tần suất vượt ngưỡng>]
THEN sinh alert với <severity, ánh xạ MITRE ATT&CK, team được assign>

Các use case tiêu biểu, theo mức độ phức tạp tăng dần:

Use casePatternVí dụ logic rule
Brute-force loginSự kiện đơn tần suất cao10+ lần đăng nhập thất bại cho cùng tài khoản trong 5 phút
Impossible travelSự kiện correlate, enrich geoĐăng nhập thành công cho cùng user từ hai quốc gia cách nhau chưa đến 1 giờ, không khớp với thời gian di chuyển thực tế
Login thất bại → thành công → privilege escalationChuỗi nhiều giai đoạnĐăng nhập thất bại trên một tài khoản, sau đó thành công, rồi trong vòng 30 phút tài khoản đó được thêm vào group admin/đặc quyền
Chuẩn bị exfiltrate dữ liệuBất thường về khối lượng/hành viMột host bình thường chỉ gửi < 100 MB/ngày ra internet bỗng upload vài GB đến một đích bên ngoài lạ
Lateral movementCorrelation xuyên hostCùng một credential được dùng xác thực vào một số lượng host bất thường trong thời gian ngắn, đặc biệt với công cụ như PsExec/WMI
PowerShell độc hại / living-off-the-landPattern process + command-lineCommand-line PowerShell mã hóa/che giấu, hoặc certutil được dùng để tải file từ xa
Thay đổi IAM cloud đáng ngờAPI call hiếm/nghiêm trọngCreateAccessKey, AttachUserPolicy với AdministratorAccess, hoặc vô hiệu hóa CloudTrail — đặc biệt từ một identity hoặc vị trí bất thường
Kích hoạt lại tài khoản đã bị vô hiệu hóaCorrelation thay đổi trạng tháiTài khoản của một nhân viên đã nghỉ việc/bị vô hiệu hóa trước đó xác thực thành công

Một correlation rule tốt cân bằng giữa độ bao phủ phát hiệnnhiễu — quá rộng thì analyst chìm trong false positive (“alert fatigue”); quá hẹp thì tấn công thật có thể lọt qua chỉ bằng cách thay đổi một chi tiết mà rule giả định là cố định. Rule cần được tune liên tục dựa trên việc trong thực tế nó là true positive hay false positive.

MITRE ATT&CK như một tài liệu tham chiếu cho độ bao phủ phát hiện

MITRE ATT&CK là một knowledge base công khai, được duy trì và có cơ sở thực nghiệm, về tactic (mục tiêu của kẻ tấn công, vd: Privilege Escalation, Lateral Movement, Exfiltration) và technique (phương pháp cụ thể để đạt mục tiêu đó, vd: T1078 — Valid Accounts, T1059.001 — PowerShell), được rút ra từ các cuộc xâm nhập thực tế đã quan sát được.

Giá trị của nó với công việc SIEM/SOAR nằm ở việc là một ngôn ngữ chung và bản đồ độ bao phủ, không phải một checklist để triển khai máy móc từ trên xuống dưới:

ATT&CK mang tính mô tả (những gì kẻ tấn công đã được quan sát làm), không mang tính chỉ định (nó không nói bạn nên mua control nào) — các team vẫn phải tự quyết định, dựa trên risk profile của mình, khoảng trống nào quan trọng nhất cần đóng lại trước.

Khái niệm chính

SOAR: orchestration, automation, và response

Nền tảng SOAR nằm ở tầng phía sau (hoặc tích hợp với) SIEM và giải quyết ba vấn đề liên quan nhau:

Thuật ngữ playbook (đôi khi gọi là runbook) chỉ một workflow được định nghĩa rõ ràng, thường được mô hình hóa trực quan: điều kiện kích hoạt → các bước enrichment → điểm quyết định → hành động phản ứng → thông báo/ghi log. Playbook trải dài từ hoàn toàn tự động (không có con người trong vòng lặp, dùng cho các alert đã hiểu rõ, rủi ro thấp, độ tin cậy cao) đến bán tự động (playbook làm enrichment và đề xuất hành động, nhưng con người phải phê duyệt trước khi có gì mang tính phá hủy xảy ra) đến do con người dẫn dắt với trợ giúp tự động (playbook chỉ thu thập ngữ cảnh và bàn giao một case đã đầy đủ thông tin cho analyst).

Ví dụ playbook: tự động phản ứng với alert tài khoản bị compromise

Dưới đây là một playbook SOAR tiêu biểu cho một loại alert phổ biến, đã được hiểu rõ — một tài khoản user có dấu hiệu bị compromise (vd: đăng nhập impossible-travel theo sau bởi hoạt động bất thường). Ví dụ này minh họa pattern mà hầu hết playbook thực tế tuân theo: enrich, quyết định, hành động, ghi lại.

  1. Trigger (kích hoạt). Correlation rule của SIEM bật lên: đăng nhập “impossible travel” cho user jane.doe, theo sau trong vòng 10 phút bởi một rule chuyển tiếp mail mới được tạo trong hộp thư của cô ấy — một indicator account-takeover đã được biết đến rộng rãi.
  2. Enrich. Playbook tự động thu thập ngữ cảnh mà không cần chờ con người:
    • Truy vấn identity provider để lấy lịch sử đăng nhập gần đây và các session token hiện tại của tài khoản.
    • Truy vấn nền tảng EDR/endpoint xem IP nguồn hoặc thiết bị có từng xuất hiện trên bất kỳ thiết bị được quản lý nào trước đây không.
    • Kiểm tra IP nguồn với các feed threat intelligence (có phải là exit node VPN/proxy đã biết? Đã từng bị gắn cờ là độc hại chưa?).
    • Lấy vai trò/mức độ đặc quyền của user từ hệ thống IAM (đây là user thường hay admin — quyết định mức độ khẩn cấp và phạm vi ảnh hưởng).
  3. Decide (quyết định). Playbook đánh giá ngữ cảnh đã enrich theo một cây quyết định:
    • Nếu IP nguồn khớp với một indicator độc hại đã biết tài khoản có đặc quyền cao → tiến hành ngăn chặn tự động (độ tin cậy cao, tác động cao — tốc độ quan trọng hơn việc tránh một false positive hiếm gặp).
    • Nếu indicator mập mờ (vd: IP lạ nhưng chưa bị gắn cờ độc hại, user thường) → tạm dừng và chuyển cho analyst phê duyệt trước khi thực hiện hành động mang tính phá hủy.
  4. Contain (ngăn chặn — các hành động phản ứng tự động). Khi đã được kích hoạt (tự động hoặc qua phê duyệt của analyst):
    • Thu hồi mọi session token/refresh token đang hoạt động của tài khoản, buộc xác thực lại ở mọi nơi.
    • Vô hiệu hóa tài khoản (hoặc buộc đổi mật khẩu) trong identity provider.
    • Xóa hoặc vô hiệu hóa rule chuyển tiếp mail độc hại phát hiện được trong bước enrichment.
    • Chặn IP nguồn tại firewall/WAF/CDN trong một khoảng thời gian xác định.
    • Cô lập mọi endpoint liên quan đến session qua EDR (network containment, giữ máy khả dụng để phục vụ forensics thay vì tắt máy).
  5. Notify và document (thông báo và ghi lại). Playbook đồng thời:
    • Mở ticket/case trong hệ thống case management, điền sẵn mọi thông tin thu thập được ở bước 2, để analyst bắt đầu điều tra với đầy đủ ngữ cảnh thay vì từ con số không.
    • Thông báo cho chủ tài khoản và quản lý/IT support của họ qua kênh định sẵn (email, Slack, ticketing).
    • Ghi log có cấu trúc, có timestamp cho mọi hành động đã thực hiện (đã kiểm tra gì, quyết định gì, thay đổi gì) — đây vừa là audit trail phục vụ compliance, vừa là điểm khởi đầu cho một cuộc điều tra do con người dẫn dắt theo Incident Response and Digital Forensics nếu case được xác nhận là nghiêm trọng.
  6. Human review (con người xem xét lại). Analyst xem xét case, xác nhận việc ngăn chặn tự động có phù hợp hay không, và hoặc đóng case là đã giải quyết (nếu là false positive — gỡ block/kích hoạt lại tài khoản) hoặc escalate lên quy trình incident response đầy đủ (nếu là compromise được xác nhận, vd: kích hoạt quy trình từ Incident Response and Digital Forensics).

Một ví dụ phổ biến khác theo cùng khuôn mẫu cho alert “IP bị gắn cờ là traffic C2 độc hại”: enrich (host nào đã liên hệ với nó, khối lượng dữ liệu bao nhiêu, process nào khởi tạo kết nối) → decide (host hướng nội bộ hay hướng ra ngoài, độ tin cậy của nguồn threat intel) → act (chặn IP tại firewall perimeter/WAF, cô lập host khởi tạo qua EDR, kill process gây hại) → notify và document.

Tích hợp IDS/IPS với SIEM/SOAR

Network Security and Zero Trust trình bày cách các sensor IDS/IPS phát hiện và (với IPS) chặn traffic độc hại ở tầng mạng. Đứng riêng lẻ, alert IDS/IPS chỉ là một nguồn dữ liệu nhiễu khác — một IDS dựa trên signature có thể bật lên hàng nghìn lần mỗi ngày với traffic scan vô hại. Giá trị của chúng nhân lên khi được tích hợp vào pipeline SIEM/SOAR:

Quy trình triage alert: phát hiện → triage → điều tra → phản ứng

Ngay cả với tự động hóa mạnh, hầu hết các môi trường vẫn chuyển một phần đáng kể alert qua phán đoán của con người. Một quy trình chuẩn:

  1. Detection (phát hiện). Correlation rule của SIEM, anomaly từ UEBA, hoặc signature IDS/IPS bật lên và sinh alert kèm severity ban đầu và, lý tưởng nhất, ánh xạ technique ATT&CK.
  2. Triage. Một SOC analyst (hoặc một playbook SOAR tự động làm triage bước đầu) nhanh chóng đánh giá: Đây có phải là một pattern false positive đã biết? Nó có khớp với một điều kiện đang được suppress/kỳ vọng không (vd: một cuộc pentest đã lên lịch)? Phạm vi ảnh hưởng thế nào (asset nào, user nào, mức đặc quyền nào)? Triage kết thúc ở một trong số: bỏ qua (false positive, có ghi lại), escalate lên điều tra, hoặc chuyển thẳng đến playbook ngăn chặn tự động nếu đó là một case đã hiểu rõ và độ tin cậy cao.
  3. Investigation (điều tra). Với bất cứ thứ gì không bị bỏ qua hoặc tự động giải quyết hoàn toàn, analyst đào sâu hơn: pivot qua SIEM để dựng lại timeline, kéo thêm log không có trong alert ban đầu, kiểm tra các alert liên quan cho cùng user/host/IP, và hình thành giả thuyết về những gì thực sự xảy ra và mức độ lan rộng. Đây là lúc các kỹ thuật forensic sâu hơn trong Incident Response and Digital Forensics phát huy tác dụng cho bất cứ điều gì được xác nhận là sự cố thật.
  4. Response (phản ứng). Khi cuộc điều tra xác nhận hoạt động độc hại, phản ứng có thể theo một playbook tự động (cho các hành động ngăn chặn mô tả ở trên) hoặc quy trình incident response đầy đủ hơn (containment, eradication, recovery, rút kinh nghiệm) cho bất cứ điều gì đủ nghiêm trọng để được tuyên bố là một sự cố (incident).

Các chỉ số thường được theo dõi xuyên suốt quy trình này: Mean Time to Detect (MTTD), Mean Time to Triage, Mean Time to Respond/Remediate (MTTR), và tỷ lệ false positive trên mỗi rule — tất cả đều thúc đẩy việc tune liên tục correlation rule và playbook.

Bức tranh công cụ nhìn tổng quan

Danh mụcVí dụGhi chú
SIEM thương mạiSplunk Enterprise Security, IBM QRadar, Microsoft Sentinel, ExabeamCorrelation/UEBA trưởng thành, hỗ trợ vendor mạnh, licensing thường theo khối lượng dữ liệu nạp vào
SIEM xây trên nền open-sourceElastic Security (xây trên Elastic Stack)Mô hình open-core; phù hợp với team đã dùng Elasticsearch/OpenSearch cho log
SIEM cloud-nativeMicrosoft Sentinel, Google Chronicle/SecOps, AWS Security Hub + GuardDutyTích hợp sâu với control plane và hệ thống identity của nhà cung cấp cloud tương ứng
Bộ SIEM/detection mã nguồn mởWazuh, Security Onion, Graylog (kèm security content)Chi phí licensing thấp hơn, cần nhiều công sức tích hợp/bảo trì hơn, cộng đồng cung cấp rule set tích cực
SOAR thương mạiSplunk SOAR (Phantom), Palo Alto Cortex XSOAR, Microsoft Sentinel automation (Logic Apps/playbooks)Thư viện playbook và tích hợp (“apps”/“connectors”) có sẵn phong phú cho các công cụ phổ biến
SOAR mã nguồn mởShuffle, TheHive + Cortex (case management + automation)Phù hợp cho team nhỏ hơn hoặc muốn toàn quyền kiểm soát logic playbook

Việc chọn vendor ít quan trọng hơn việc tổ chức có thực sự viết và test detection content cùng playbook cho môi trường của riêng mình hay không — một SIEM chưa được cấu hình với rule mặc định phần lớn chỉ sinh ra nhiễu, và một SOAR chưa được cấu hình thì không tự động hóa được gì hữu ích.

SIEM vs. SOAR vs. EDR

Ba danh mục này thường bị nhầm lẫn vì chúng chồng lấn nhau trong một security stack hiện đại, nhưng mỗi cái có trách nhiệm chính riêng biệt.

Khía cạnhSIEMSOAREDR (Endpoint Detection and Response)
Trách nhiệm chínhTổng hợp và correlate sự kiện từ khắp môi trường; sinh alertĐiều phối các công cụ và tự động hóa workflow phản ứng khi đã có alertPhát hiện và phản ứng với mối đe dọa cụ thể trên endpoint (máy trạm, server)
Phạm vi quan sátRộng — mạng, cloud, identity, ứng dụng, endpoint (thường nạp cả telemetry EDR như một nguồn)Không phải nguồn dữ liệu — hành động trên/xuyên qua các công cụ khácHẹp nhưng sâu — process tree, file system, memory, registry trên một host cụ thể
Câu hỏi cốt lõi trả lời”Có điều gì đáng ngờ xảy ra, correlate xuyên hệ thống không?""Giờ đã biết rồi, hành động lặp lại nào nên xảy ra tiếp theo?""Chính xác điều gì đang xảy ra trên endpoint này, và có thể ngăn chặn ngay bây giờ không?”
Hành động điển hìnhAlerting, dashboard, báo cáo compliance, retention/search dài hạnThực thi playbook: enrichment, ticketing, chặn, thông báo, workflow xuyên công cụKill process độc hại, cách ly file, cô lập host khỏi mạng, rollback file bị ransomware mã hóa
Quan hệ với hai cái còn lạiThường là nguồn trigger cho playbook SOAR; nạp alert EDR như một inputThường gọi API EDR như một hành động phản ứng (vd: “cô lập host này”)Thường đẩy alert vào SIEM và expose API mà SOAR có thể gọi
Ví dụ công cụSplunk ES, Elastic Security, Microsoft SentinelSplunk SOAR, Cortex XSOAR, ShuffleCrowdStrike Falcon, Microsoft Defender for Endpoint, SentinelOne

Trong một stack trưởng thành, ba thành phần này được nối với nhau chứ không đứng độc lập: EDR phát hiện và có thể hành động cục bộ trên endpoint; alert của nó chảy vào SIEM để correlate với mọi thứ khác; một match của correlation SIEM kích hoạt playbook SOAR, playbook này có thể gọi ngược lại API EDR để cô lập host — vòng lặp khép kín xuyên suốt cả ba công cụ.

Best Practices

Tài liệu tham khảo

Part of the DevSecOps Roadmap knowledge base.

Overview

Monitoring and Logging covers how individual systems emit metrics, logs, and traces. That’s necessary but not sufficient: a single failed login on one host is noise; a hundred failed logins across ten hosts followed by a successful login and a new admin account being created is an attack. Seeing that pattern requires pulling logs from every system into one place and asking questions that span sources — which is exactly what a SIEM (Security Information and Event Management) system does.

A SIEM aggregates security-relevant events from across the environment (servers, network devices, cloud APIs, applications, endpoints), normalizes them into a common structure, correlates them against rules and behavioral baselines, and surfaces the results as alerts and dashboards that a human analyst can act on. It is the nervous system of a detection program — the place where “something happened over there” becomes “here is a prioritized signal that something bad is happening.”

Detecting a threat is only half the job. Once an alert fires, someone (or something) has to triage it, decide whether it’s real, and take action — disable an account, block an IP, isolate a host. Doing that by hand, every time, for every alert, does not scale once alert volume grows past a handful a day. SOAR (Security Orchestration, Automation, and Response) platforms exist to close that gap: they encode the repeatable parts of the response process into automated playbooks so analysts spend their time on judgment calls instead of repetitive clicks.

This note covers how SIEM and SOAR work together — how correlation rules turn raw events into alerts, how the MITRE ATT&CK framework is used to reason about detection coverage, how SOAR playbooks automate common responses, and how alerts flow through a triage workflow into the deeper investigation and containment work covered in Incident Response and Digital Forensics.

Fundamentals

SIEM vs. plain log management

Centralized log management (an ELK stack used purely for log search, or a log aggregation service) and a SIEM look similar on the surface — both ingest logs from many sources into one searchable place — but they solve different problems.

AspectPlain log managementSIEM
Primary purposeSearch, debugging, operational troubleshootingSecurity detection, alerting, and compliance reporting
Data modelOften raw/unstructured, source-specific fieldsNormalized into a common schema (e.g., a “login failure” looks the same regardless of source system)
CorrelationManual — an engineer runs ad hoc queriesBuilt-in correlation engine continuously evaluating rules across sources in near real time
AlertingUsually absent or bolted onCore feature — rules/analytics generate prioritized alerts automatically
Context enrichmentRareCommon — alerts enriched with threat intelligence, asset criticality, user risk score
Retention driverOperational needs (days to weeks)Often compliance-driven (months to years, e.g., PCI DSS requires 1 year)
Typical usersDevelopers, SREsSOC analysts, security engineers, auditors
Example toolsPlain ELK/OpenSearch, Loki, CloudWatch LogsSplunk Enterprise Security, Elastic Security, Microsoft Sentinel, IBM QRadar

In practice, many organizations build both on the same underlying storage — the Elastic Stack, for example, can serve as plain log search and, with the Security app, correlation rules, and detection content layered on top, function as a SIEM. The distinguishing feature is not the storage engine but whether there’s a correlation/detection layer, case management, and analyst workflow built around it.

Core SIEM capabilities

Correlation rules and use cases

A correlation rule (sometimes called a “detection” or “analytic”) expresses a pattern across one or more events that, taken together, indicates something worth an analyst’s attention. Rules range from simple single-event matches to multi-stage sequences spanning minutes or days.

Typical structure of a correlation rule:

WHEN <event type/source> occurs
WHERE <field conditions match>
[AND is correlated with <other event type>]
[WITHIN <time window>]
[GROUPED BY <entity, e.g. user or host>]
[THRESHOLD <count/frequency exceeded>]
THEN generate alert with <severity, MITRE ATT&CK mapping, assigned team>

Representative use cases, roughly in increasing complexity:

Use casePatternExample rule logic
Brute-force loginHigh-frequency single event10+ failed logins for the same account within 5 minutes
Impossible travelCorrelated event, geo enrichmentSuccessful logins for the same user from two countries less than 1 hour apart, geographically incompatible with travel time
Failed logins → success → privilege escalationMulti-stage sequenceFailed logins on an account, followed by a success, followed within 30 minutes by that account being added to an admin/privileged group
Data exfiltration stagingVolume/behavior anomalyA host that normally sends < 100 MB/day to the internet suddenly uploads several GB to an unfamiliar external destination
Lateral movementCross-host correlationThe same credential used to authenticate to an unusual number of distinct hosts within a short window, especially with tools like PsExec/WMI
Malicious PowerShell / living-off-the-landProcess + command-line patternEncoded/obfuscated PowerShell command line, or certutil used to download a remote file
Suspicious cloud IAM changeRare/critical API callCreateAccessKey, AttachUserPolicy with AdministratorAccess, or disabling CloudTrail — especially from a non-standard identity or location
New disabled-account reactivationState-change correlationA previously disabled or terminated employee’s account authenticates successfully

A good correlation rule balances detection coverage against noise — too broad and analysts drown in false positives (“alert fatigue”); too narrow and real attacks slip through by varying a detail the rule assumed was fixed. Rules are continuously tuned based on what turns out to be a true positive vs. a false positive in practice.

MITRE ATT&CK as a reference for detection coverage

MITRE ATT&CK is a publicly maintained, empirically-grounded knowledge base of adversary tactics (the attacker’s goal, e.g., Privilege Escalation, Lateral Movement, Exfiltration) and techniques (the specific method used to achieve that goal, e.g., T1078 — Valid Accounts, T1059.001 — PowerShell), derived from observed real-world intrusions.

Its value to SIEM/SOAR work is as a common vocabulary and coverage map, not as a checklist to blindly implement top to bottom:

ATT&CK is descriptive (what attackers have been observed doing), not prescriptive (it doesn’t tell you which controls to buy) — teams still have to decide, based on their own risk profile, which gaps matter most to close first.

Key Concepts

SOAR: orchestration, automation, and response

SOAR platforms sit downstream of (or integrated with) the SIEM and address three related problems:

The term playbook (sometimes runbook) refers to a defined, often visually-modeled workflow: trigger condition → enrichment steps → decision points → response actions → notification/documentation. Playbooks range from fully automated (no human in the loop, for well-understood, low-risk, high-confidence alerts) to semi-automated (the playbook does enrichment and proposes an action, but a human must approve before anything destructive happens) to human-driven with automated helpers (the playbook just gathers context and hands a fully-informed case to an analyst).

Example playbook: automatic response to a compromised-account alert

Below is a representative SOAR playbook for a common, well-understood alert type — a user account showing signs of compromise (e.g., impossible-travel login followed by unusual activity). This illustrates the pattern most real playbooks follow: enrich, decide, act, document.

  1. Trigger. SIEM correlation rule fires: “impossible travel” login for user jane.doe, followed within 10 minutes by a new mail forwarding rule created in her mailbox — a well-known account-takeover indicator.
  2. Enrich. The playbook automatically gathers context without waiting on a human:
    • Query the identity provider for the account’s recent sign-in history and current session tokens.
    • Query the EDR/endpoint platform for whether the source IP or device has been seen before on any managed device.
    • Check the source IP against threat intelligence feeds (known VPN/proxy exit node? Previously flagged as malicious?).
    • Pull the user’s role/privilege level from the IAM system (is this a standard user or an admin — changes the urgency and blast radius).
  3. Decide. The playbook evaluates the enriched context against a decision tree:
    • If the source IP matches a known-malicious indicator and the account has elevated privileges → proceed automatically to containment (high confidence, high impact — speed matters more than avoiding a rare false positive).
    • If indicators are ambiguous (e.g., unfamiliar but not flagged-malicious IP, standard user) → pause and route to an analyst for approval before taking destructive action.
  4. Contain (automated response actions). Once triggered (automatically or via analyst approval):
    • Revoke all active session tokens/refresh tokens for the account, forcing re-authentication everywhere.
    • Disable the account (or force a password reset) in the identity provider.
    • Delete or disable the malicious mail-forwarding rule found during enrichment.
    • Block the source IP at the firewall/WAF/CDN for a defined period.
    • Isolate any endpoint associated with the session via EDR (network containment, leaving the box available for forensics rather than shutting it down).
  5. Notify and document. The playbook simultaneously:
    • Opens a ticket/case in the case management system pre-populated with everything gathered in step 2, so an analyst starts the investigation with full context instead of from zero.
    • Notifies the account owner and their manager/IT support through a predefined channel (email, Slack, ticketing).
    • Writes a structured, timestamped log of every action taken (what was checked, what decision was made, what was changed) — this becomes both the audit trail for compliance and the starting point for a human-led investigation under Incident Response and Digital Forensics if the case proves serious.
  6. Human review. An analyst reviews the case, confirms whether the automated containment was appropriate, and either closes it as resolved (if it was a false positive — releases the block/re-enables the account) or escalates to full incident response (if it’s a confirmed compromise, e.g., invoking the process from Incident Response and Digital Forensics).

Another common example follows the same shape for an “IP flagged as malicious C2 traffic” alert: enrich (which hosts contacted it, what data volume, what process initiated the connection) → decide (internal-facing vs. external-facing host, confidence of the threat intel source) → act (block the IP at the perimeter firewall/WAF, isolate the initiating host via EDR, kill the offending process) → notify and document.

IDS/IPS integration with SIEM/SOAR

Network Security and Zero Trust covers how IDS/IPS sensors detect and (for IPS) block malicious traffic at the network layer. On their own, IDS/IPS alerts are just one more noisy data source — a signature-based IDS can fire thousands of times a day on benign scanning traffic. Their value multiplies once integrated into the SIEM/SOAR pipeline:

Alert triage workflow: detection → triage → investigation → response

Even with heavy automation, most environments still funnel a meaningful share of alerts through human judgment. A standard workflow:

  1. Detection. SIEM correlation rule, UEBA anomaly, or IDS/IPS signature fires and generates an alert with an initial severity and, ideally, an ATT&CK technique mapping.
  2. Triage. A SOC analyst (or a SOAR playbook doing first-pass triage automatically) quickly assesses: Is this a known false positive pattern? Does it match a currently-suppressed/expected condition (e.g., a scheduled pentest)? What’s the blast radius (which asset, which user, what privilege level)? Triage ends in one of: dismiss (false positive, documented), escalate to investigation, or route directly to an automated containment playbook if it’s a well-understood high-confidence case.
  3. Investigation. For anything that isn’t dismissed or fully auto-resolved, an analyst digs deeper: pivoting through the SIEM to reconstruct a timeline, pulling additional logs not included in the original alert, checking related alerts for the same user/host/IP, and forming a hypothesis about what actually happened and how far it went. This is where the deeper forensic techniques in Incident Response and Digital Forensics come in for anything confirmed as a real incident.
  4. Response. Once an investigation confirms malicious activity, the response follows either an automated playbook (for the containment actions described above) or the fuller incident response process (containment, eradication, recovery, lessons learned) for anything serious enough to be declared an incident.

Metrics commonly tracked across this workflow: Mean Time to Detect (MTTD), Mean Time to Triage, Mean Time to Respond/Remediate (MTTR), and the false positive rate per rule — all of which drive continuous tuning of correlation rules and playbooks.

Tooling landscape at a glance

CategoryExamplesNotes
Commercial SIEMSplunk Enterprise Security, IBM QRadar, Microsoft Sentinel, ExabeamMature correlation/UEBA, strong vendor support, licensing often by data volume ingested
SIEM built on open-source foundationsElastic Security (built on the Elastic Stack)Open-core model; strong for teams already running Elasticsearch/OpenSearch for logs
Cloud-native SIEMMicrosoft Sentinel, Google Chronicle/SecOps, AWS Security Hub + GuardDutyDeep integration with the respective cloud provider’s control plane and identity system
Open-source SIEM/detection stacksWazuh, Security Onion, Graylog (with security content)Lower licensing cost, more integration/maintenance effort, active community rule sets
Commercial SOARSplunk SOAR (Phantom), Palo Alto Cortex XSOAR, Microsoft Sentinel automation (Logic Apps/playbooks)Rich pre-built playbook libraries and integrations (“apps”/“connectors”) for common tools
Open-source SOARShuffle, TheHive + Cortex (case management + automation)Good fit for smaller teams or those wanting full control over playbook logic

Vendor choice matters less than whether the organization has actually written and tested detection content and playbooks for its own environment — an unconfigured SIEM with default rules produces mostly noise, and an unconfigured SOAR automates nothing useful.

SIEM vs. SOAR vs. EDR

These three categories are frequently confused because they overlap in a modern security stack, but each has a distinct primary responsibility.

AspectSIEMSOAREDR (Endpoint Detection and Response)
Primary responsibilityAggregate and correlate events from across the whole environment; generate alertsOrchestrate tools and automate the response workflow once an alert existsDetect and respond to threats specifically on endpoints (workstations, servers)
Scope of visibilityBroad — network, cloud, identity, application, endpoint (often ingesting EDR telemetry as one of its sources)Not a data source itself — acts on/across other toolsNarrow but deep — process trees, file system, memory, registry on a single host
Core question it answers”Did something suspicious happen, correlated across systems?""Now that we know, what repeatable actions should happen next?""What exactly is happening on this endpoint, and can we stop it right now?”
Typical actionsAlerting, dashboards, compliance reporting, long-term retention/searchExecuting playbooks: enrichment, ticketing, blocking, notification, cross-tool workflowsKilling a malicious process, quarantining a file, isolating a host from the network, rolling back ransomware-encrypted files
Relationship to the othersOften the trigger source for SOAR playbooks; ingests EDR alerts as one inputFrequently calls EDR APIs as a response action (e.g., “isolate this host”)Frequently feeds alerts into the SIEM and exposes an API SOAR can call
Example toolsSplunk ES, Elastic Security, Microsoft SentinelSplunk SOAR, Cortex XSOAR, ShuffleCrowdStrike Falcon, Microsoft Defender for Endpoint, SentinelOne

In a mature stack these three are wired together rather than standalone: EDR detects and can act locally on the endpoint; its alerts flow into the SIEM where they’re correlated with everything else; a SIEM correlation match triggers a SOAR playbook that may, in turn, call back into the EDR API to isolate the host — the loop closes across all three tools.

Best Practices

References