Giám sát & ObservabilityMonitoring & Observability
Thuộc bộ kiến thức Network Engineer Roadmap.
Tổng quan
Bạn không thể vận hành thứ mà bạn không nhìn thấy. Network monitoring và observability biến một “hộp đen” gồm cáp, switch, router, firewall và các link thành một hệ thống mà bạn có thể suy luận về sức khỏe, hiệu năng và hành vi gần như theo thời gian thực. Monitoring trả lời những câu hỏi bạn đã biết trước (“link có up không? CPU của core switch bao nhiêu? VLAN đó dùng bao nhiêu bandwidth?”). Observability đi xa hơn: mục tiêu là giúp bạn trả lời cả những câu hỏi bạn chưa lường trước, bằng cách phơi bày telemetry giàu thông tin, high-cardinality để bạn có thể cắt lát và tương quan (correlate) sau đó.
Tại sao phải đầu tư vào việc này?
- Phát hiện và chẩn đoán sự cố nhanh. Mean Time To Detect (MTTD) và Mean Time To Repair (MTTR) quyết định độ tin cậy mà người dùng cảm nhận. Telemetry tốt giảm cả hai.
- Hiểu hiệu năng và dung lượng (capacity). Latency tăng dần, buffer bloat, microburst, và link bão hòa đều vô hình cho đến khi được đo. Dữ liệu xu hướng (trend) là cơ sở cho capacity planning và biện minh ngân sách.
- Bảo mật và điều tra (forensics). Flow records và logs phơi bày exfiltration, scanning, DDoS, lateral movement, và cung cấp dấu vết kiểm toán sau sự cố.
- Chứng minh SLA / SLO. Hợp đồng với nhà cung cấp và service owner nội bộ cần bằng chứng khách quan, có lịch sử.
- Cấp dữ liệu cho automation. Closed-loop remediation và intent-based networking dựa trên một luồng trạng thái đáng tin cậy.
Có hai bộ môn kinh điển mà observability bao trùm:
| Bộ môn | Câu hỏi | Tín hiệu điển hình |
|---|---|---|
| Fault monitoring | Nó có up / có hỏng không? | interface up/down, SNMP trap, syslog error, reachability (ICMP/BFD) |
| Performance monitoring | Nó có đủ nhanh / khỏe không? | throughput, latency, jitter, loss, CPU/memory, queue drop, error |
Ghi chú này ánh xạ mô hình “ba trụ cột” tổng quát của observability sang networking, rồi đi qua các công nghệ telemetry cốt lõi — SNMP, flow export (NetFlow/sFlow/IPFIX), syslog, và streaming telemetry (gNMI) — cùng các stack lưu trữ và trực quan hóa chúng (Prometheus/Grafana, và các nền tảng thương mại như Datadog, Dynatrace). Cuối cùng là alerting, baselining, SLI/SLO, và capacity planning. Để chẩn đoán ở mức packet xem ./15-tools-and-troubleshooting.md; về các protocol bạn đang theo dõi xem ./06-core-protocols.md; và về hành vi queueing/marking mà các metric này phản ánh xem ./13-traffic-management-and-qos.md.
Kiến thức nền tảng
Ba trụ cột, điều chỉnh cho networking
Thế giới software-observability nói về metrics, logs, và traces. Networking có các tương đương trực tiếp cộng thêm một trụ cột riêng:
| Trụ cột | Thế giới software | Tương đương trong networking | Thu thập qua |
|---|---|---|---|
| Metrics | Counter/gauge theo thời gian | Interface counter, CPU/mem, latency, drop, số BGP peer | SNMP polling, streaming telemetry, synthetic probe |
| Logs | Dòng log ứng dụng | Syslog: link flap, thay đổi config, ACL deny, sự kiện auth | syslog (RFC 5424) |
| Traces | Span của request phân tán | Phân tích path/hop: traceroute, IP SLA, đo path chủ động | traceroute, IP SLA, TWAMP, synthetic test |
| Flows (riêng của networking) | — | Bản ghi theo từng phiên: ai nói với ai, bao nhiêu, port nào | NetFlow / sFlow / IPFIX |
Flows là trụ cột riêng của networking. Một flow record tóm tắt toàn bộ một “cuộc hội thoại” (5-tuple + số byte/packet + timing), rẻ hơn nhiều so với bắt từng packet nhưng giàu thông tin hơn nhiều so với một gauge bandwidth.
Push vs pull
- Pull (polling): collector định kỳ hỏi thiết bị về trạng thái. SNMP GET và Prometheus scraping là mô hình pull. Đơn giản, nhưng độ phân giải bị giới hạn bởi poll interval (thường 30–300 s) và collector phải biết từng target.
- Push (streaming/events): thiết bị gửi dữ liệu khi có thay đổi hoặc theo nhịp mịn. SNMP trap, syslog, NetFlow export, và gNMI dial-out là mô hình push. Độ trễ thấp hơn, phân giải cao hơn, tốt hơn khi scale — nhưng cần cấu hình thiết bị để nói với đúng collector.
Thiết kế hiện đại kết hợp cả hai: pull metric cho sức khỏe steady-state, push event/telemetry cho thay đổi và dữ liệu phân giải cao.
Cadence, cardinality, và retention
Ba “núm vặn” định hình mọi hệ thống monitoring:
- Cadence / resolution — tần suất lấy mẫu (1 s streaming vs 5 phút SNMP). Phân giải cao bắt được microburst nhưng tốn storage và tải.
- Cardinality — số lượng time series riêng biệt (per-interface × per-device × per-metric). Con số này bùng nổ rất nhanh; một chassis 400 port × 20 counter = 8000 series trên một thiết bị.
- Retention — giữ dữ liệu raw vs đã roll-up (downsample) trong bao lâu. Điển hình: raw vài ngày, rollup 5 phút vài tuần, rollup theo giờ hơn một năm.
Khái niệm chính
SNMP (Simple Network Management Protocol)
SNMP là “con ngựa thồ” nhiều thập kỷ để poll trạng thái thiết bị. Dù đã lâu đời, nó vẫn là mẫu số chung thấp nhất được gần như mọi thiết bị managed hỗ trợ.
Kiến trúc
- Manager (NMS): trạm polling/thu thập (ví dụ LibreNMS, Zabbix, PRTG, hoặc một SNMP exporter cấp dữ liệu cho Prometheus).
- Agent: phần mềm trên thiết bị managed, phơi bày dữ liệu và có thể phát notification.
- MIB (Management Information Base): schema dạng cây phân cấp của các object quản lý được. Nhà sản xuất công bố MIB (ví dụ
IF-MIB,BGP4-MIB,CISCO-PROCESS-MIB). - OID (Object Identifier): địa chỉ dạng dotted-decimal của một object trong cây MIB. Ví dụ:
1.3.6.1.2.1.2.2.1.10=ifInOctets(số byte nhận trên một interface).
Các operation
| PDU | Chiều | Mục đích |
|---|---|---|
GET / GETNEXT / GETBULK | Manager → Agent | Đọc một/kế tiếp/nhiều object (GETBULK hiệu quả khi walk bảng) |
SET | Manager → Agent | Ghi/đổi giá trị (hiếm khi bật — rủi ro bảo mật) |
TRAP | Agent → Manager | Notification tự phát, fire-and-forget (ví dụ linkDown) |
INFORM | Agent → Manager | Giống trap nhưng có acknowledge (đáng tin cậy) |
RESPONSE | Agent → Manager | Trả lời GET/SET |
Polling hoạt động bằng cách manager phát GET/GETBULK lên một danh sách OID theo lịch (ví dụ mỗi 60 s). Để tính rate (bps) bạn poll một counter hai lần rồi chia delta cho khoảng thời gian. Chú ý counter wrap: counter 32-bit (ifInOctets) wrap rất nhanh trên link 10G+ — luôn dùng counter high-capacity 64-bit (ifHCInOctets, OID 1.3.6.1.2.1.31.1.1.1.6).
Version & bảo mật
| Version | Bảo mật | Ghi chú |
|---|---|---|
| v1 | community string, plaintext | cũ, chỉ counter 32-bit — tránh dùng |
| v2c | community string, plaintext | phổ biến nhất thực tế; có GETBULK & counter 64-bit; không mã hóa |
| v3 | USM: authentication + privacy | theo user, hỗ trợ noAuthNoPriv / authNoPriv / authPriv; nên dùng authPriv (SHA + AES) |
Ưu tiên SNMPv3 authPriv ở mọi nơi thiết bị và NMS hỗ trợ; chỉ lùi về v2c trên management network cô lập.
Ví dụ — bật SNMP trên Cisco IOS (v2c read-only, rồi v3):
! SNMPv2c read-only, giới hạn bằng ACL chỉ cho NMS
access-list 99 permit host 10.0.0.10
snmp-server community R3adOnly-x8f2 RO 99
snmp-server location DC1-Row4-Rack7
snmp-server contact netops@example.com
snmp-server host 10.0.0.10 version 2c R3adOnly-x8f2
! SNMPv3 — group, view, user (authPriv, SHA + AES-128)
snmp-server view ALL iso included
snmp-server group MONgrp v3 priv read ALL
snmp-server user monitor MONgrp v3 auth sha S3cret-Auth-Pass priv aes 128 S3cret-Priv-Pass
snmp-server enable traps
snmp-server host 10.0.0.10 version 3 priv monitor
Truy vấn từ một collector Linux:
# v2c walk bảng interface
snmpwalk -v2c -c 'R3adOnly-x8f2' 10.0.0.1 IF-MIB::ifTable
# v3 authPriv GET một counter
snmpget -v3 -l authPriv -u monitor \
-a SHA -A 'S3cret-Auth-Pass' \
-x AES -X 'S3cret-Priv-Pass' \
10.0.0.1 IF-MIB::ifHCInOctets.2
Flow-based monitoring — NetFlow, sFlow, IPFIX
Trong khi SNMP cho biết bao nhiêu traffic đi qua một interface, flow monitoring cho biết traffic đó là gì: source/destination IP, port, protocol, ToS/DSCP, AS number, số byte/packet, và timing. Đây là nền tảng cho phân tích traffic, billing, quyết định peering, và phát hiện bảo mật (DDoS, scanning, exfiltration).
Flow export hoạt động thế nào (NetFlow/IPFIX)
- Router/switch duy trì một flow cache, key theo định nghĩa flow (kinh điển là 5-tuple: src IP, dst IP, src port, dst port, protocol; cộng ingress interface, ToS).
- Packet cập nhật counter của entry cache khớp.
- Một flow bị expire và export khi nó kết thúc (TCP FIN/RST), idle (idle timeout), sống quá lâu (active timeout, ví dụ 60 s), hoặc cache đầy.
- Các flow đã expire được đóng gói vào datagram UDP và gửi tới một flow collector (ví dụ nfdump/nfsen, Elastiflow, ntopng, Kentik).
Sampling. Trên link tốc độ cao, theo dõi mọi flow rất tốn kém, nên thiết bị lấy mẫu 1-trên-N packet (ví dụ 1:1000). Sampling là bắt buộc khi scale nhưng nghĩa là bạn phải nhân ngược số liệu lên và mất khả năng thấy các flow rất nhỏ.
Template. NetFlow v9 và IPFIX dựa trên template: exporter gửi trước một template mô tả các field, rồi các data record tham chiếu tới nó. Điều này giúp mở rộng linh hoạt (khác NetFlow v5 cố định). IPFIX còn hỗ trợ field độ dài thay đổi và element riêng của doanh nghiệp (ví dụ URL, application ID).
sFlow hoạt động khác: thay vì flow cache nó làm packet sampling — lấy N byte đầu (header) của mỗi packet thứ N và gửi mẫu đó ngay lập tức, cộng thêm counter interface định kỳ. Việc này stateless trên thiết bị (rẻ, thân thiện hardware, real-time) nhưng mang tính thống kê chứ không phải kế toán đầy đủ mọi cuộc hội thoại.
So sánh:
| Đặc điểm | NetFlow v5 | NetFlow v9 | IPFIX (v10) | sFlow |
|---|---|---|---|---|
| Nguồn gốc | Cisco | Cisco | Chuẩn IETF (RFC 7011) | sFlow.org / RFC 3176 |
| Mô hình | flow cache | flow cache | flow cache | packet + counter sampling |
| Field | cố định | dựa template | template, độ dài thay đổi, element doanh nghiệp | mẫu packet header |
| Sampling | tùy chọn | tùy chọn | tùy chọn | bản chất (luôn sampled) |
| Layer | L3/L4 | L3/L4 (+MPLS, IPv6) | L2–L7 | L2–L7 (raw header) |
| State trên thiết bị | có (cache) | có | có | không (stateless) |
| Real-time | trễ theo timeout | trễ theo timeout | trễ theo timeout | gần tức thời |
| Tính chuẩn | độc quyền | độc quyền (cơ sở của IPFIX) | mở, IETF | chuẩn đa nhà cung cấp |
Quy tắc ngón tay cái: dùng IPFIX/NetFlow v9 khi cần kế toán chính xác và field phong phú; dùng sFlow khi cần khả năng nhìn rẻ, real-time, dựa hardware trên nhiều port tốc độ cao.
Ví dụ — Flexible NetFlow (kiểu IPFIX) trên Cisco IOS:
flow record RECORD-v4
match ipv4 source address
match ipv4 destination address
match transport source-port
match transport destination-port
match ipv4 protocol
collect counter bytes
collect counter packets
collect interface input
flow exporter EXP-1
destination 10.0.0.20
transport udp 2055
export-protocol ipfix
template data timeout 60
flow monitor MON-1
record RECORD-v4
exporter EXP-1
cache timeout active 60
interface GigabitEthernet0/1
ip flow monitor MON-1 input
ip flow monitor MON-1 output
Syslog và quản lý log
Syslog (RFC 5424, cũ hơn là RFC 3164) là chuẩn để ghi log sự kiện thiết bị: link flap, thay đổi adjacency OSPF/BGP, thay đổi config, ACL/firewall deny, sự kiện authentication, cảnh báo hardware. Log là bản tường thuật giải thích tại sao một metric thay đổi.
Severity level (0 = tệ nhất). Một mnemonic hữu ích: Every Awesome Cisco Engineer Will Need Ice-cream Daily.
| Level | Từ khóa | Ý nghĩa |
|---|---|---|
| 0 | Emergency | hệ thống không dùng được |
| 1 | Alert | phải hành động ngay |
| 2 | Critical | tình huống critical |
| 3 | Error | tình huống error (ví dụ interface error) |
| 4 | Warning | tình huống warning |
| 5 | Notice | bình thường nhưng đáng chú ý (thay đổi config) |
| 6 | Informational | thông tin (link up) |
| 7 | Debug | thông điệp debug (nhiễu) |
Facility phân loại nguồn (ví dụ local0–local7, auth, kern). Một message kết hợp facility + severity thành một priority.
Gửi log tới một collector trung tâm — rsyslog/syslog-ng, một stack ELK/OpenSearch, Graylog, Loki, hoặc một SIEM (Splunk, Elastic Security) để correlate và lưu trữ. Tập trung hóa quan trọng vì một thiết bị crash sẽ mang theo log cục bộ của nó, và việc correlate xuyên thiết bị (một flap ở đây, một error ở kia) cần mọi thứ ở cùng một nơi.
Ví dụ — Cisco IOS gửi tới syslog server, có timestamp và lọc theo level:
service timestamps log datetime msec localtime show-timezone
logging host 10.0.0.30 transport udp port 514
logging trap informational ! gửi severity 6 trở xuống (tệ hơn)
logging source-interface Loopback0
logging buffered 64000 debugging
Luôn đồng bộ clock bằng NTP trước — log không có timestamp chính xác, nhất quán thì gần như vô dụng cho việc correlate.
Streaming telemetry (model-driven, gNMI)
SNMP polling có giới hạn thực sự: phân giải thô (không thể poll hàng nghìn OID mỗi giây), overhead cao (mỗi object một request/response), và cấu trúc MIB cứng nhắc. Model-driven streaming telemetry là câu trả lời hiện đại.
- Thiết bị push dữ liệu liên tục (subscribe một lần, stream mãi mãi) thay vì bị poll.
- Dữ liệu được cấu trúc theo YANG model (chuẩn mở như OpenConfig, hoặc model của nhà cung cấp) — cùng model dùng cho config, nên monitoring và configuration chia sẻ một schema.
- Transport thường là gNMI (gRPC Network Management Interface) trên HTTP/2 với TLS; encoding là Protobuf hoặc JSON. NETCONF/RESTCONF cũng có thể mang telemetry.
- Chế độ: dial-out (thiết bị kết nối tới collector) hoặc dial-in (collector subscribe tới thiết bị); sample (khoảng cố định, ví dụ 1 s) hoặc on-change (theo sự kiện, ví dụ thay đổi trạng thái BGP).
Kết quả: phân giải dưới giây, CPU thiết bị thấp hơn cho mỗi data point, và một data model nhất quán. Thu thập bằng một telemetry pipeline — Telegraf (với input gNMI/cisco_telemetry), Cisco Pipeline, hoặc gnmic — đổ vào một TSDB (InfluxDB, Prometheus) cho Grafana.
Ví dụ — subscription gNMI OpenConfig trên Cisco IOS-XR (path dial-in), cùng lệnh gnmic subscribe:
! Phía thiết bị (IOS-XR) — bật gRPC/gNMI
grpc
port 57400
tls-mutual
!
telemetry model-driven
sensor-group IFCOUNTERS
sensor-path openconfig-interfaces:interfaces/interface/state/counters
# Phía collector — subscribe với sample interval 1s
gnmic -a 10.0.0.1:57400 --skip-verify -u admin -p '***' \
subscribe --path "openconfig-interfaces:interfaces/interface/state/counters" \
--mode stream --stream-mode sample --sample-interval 1s
| Khía cạnh | SNMP polling | Streaming telemetry (gNMI) |
|---|---|---|
| Mô hình | pull / request-response | push / subscribe |
| Phân giải | thường 30 s – 5 phút | dưới giây, hoặc on-change |
| Data schema | MIB/OID | YANG (OpenConfig / vendor) |
| Overhead | cao (mỗi object) | thấp (stream gộp) |
| Encoding | BER | Protobuf / JSON |
| Bảo mật | v3 USM | gRPC + TLS |
| Phù hợp nhất | hỗ trợ thiết bị rộng, legacy | scale, phân giải cao, OS hiện đại |
Metrics & time-series stack — Prometheus + Grafana
Chuẩn thực tế open-source để lưu trữ và trực quan hóa metric.
- Prometheus — một TSDB kiểu pull, scrape các endpoint HTTP
/metrics, lưu time series đa chiều (tên metric + label), và cung cấp PromQL để query và alerting. - Exporter làm cầu nối các nguồn non-native vào Prometheus:
- snmp_exporter — Prometheus scrape nó qua HTTP; nó lại SNMP-poll thiết bị mạng dùng một
snmp.ymlđược sinh ra (dựng từ MIB bởigenerator). Đây là cách đưa dữ liệu SNMP của switch/router vào Prometheus. - node_exporter — metric mức host (CPU, disk, counter NIC) cho server/collector Linux.
- Khác: blackbox_exporter (probe ICMP/TCP/HTTP cho reachability & latency), gnmi qua Telegraf → remote_write.
- snmp_exporter — Prometheus scrape nó qua HTTP; nó lại SNMP-poll thiết bị mạng dùng một
- Grafana — dashboard trên Prometheus (và InfluxDB, Loki, …): panel time-series, heatmap, topology, tô màu theo threshold, và alerting riêng.
- Alertmanager — dedupe, group, silence, và route các alert của Prometheus tới email/Slack/PagerDuty.
Ví dụ — Prometheus scrape SNMP exporter cho hai switch:
# prometheus.yml
scrape_configs:
- job_name: 'snmp-switches'
static_configs:
- targets:
- 10.0.0.1 # core-sw-1
- 10.0.0.2 # core-sw-2
metrics_path: /snmp
params:
module: [if_mib] # module định nghĩa trong snmp.yml
relabel_configs:
- source_labels: [__address__]
target_label: __param_target
- source_labels: [__param_target]
target_label: instance
- target_label: __address__
replacement: 127.0.0.1:9116 # host:port của snmp_exporter
- job_name: 'node'
static_configs:
- targets: ['10.0.0.30:9100'] # node_exporter trên host collector
Ví dụ PromQL — throughput interface theo bit/giây và một error rate:
# Throughput vào (bps) mỗi interface, rate 5 phút
rate(ifHCInOctets[5m]) * 8
# Các interface có input-error rate đang tăng
rate(ifInErrors[5m]) > 0
Nền tảng observability thương mại
Stack open-source mạnh nhưng cần lắp ráp và vận hành. Các nền tảng SaaS thương mại đánh đổi chi phí lấy sự “chìa khóa trao tay”, khả năng correlate rộng, và hỗ trợ.
| Nền tảng | Điểm mạnh cho networking |
|---|---|
| Datadog | Hợp nhất infra + app + Network Performance Monitoring (NPM) và Network Device Monitoring (NDM/SNMP); nạp flow; correlate mạnh giữa app ↔ network; hợp với tổ chức thiên về cloud, DevOps. |
| Dynatrace | Root-cause dựa AI (Davis engine), topology tự động (Smartscape), network visibility hiểu ứng dụng; nghiêng về APM doanh nghiệp. |
| Khác | SolarWinds NPM, Cisco ThousandEyes (khả năng nhìn Internet/path, reachability SaaS), Kentik (phân tích flow ở quy mô lớn), LogicMonitor, Auvik. |
Chúng phù hợp nhất khi bạn coi trọng time-to-value, correlate xuyên miền (network ↔ application ↔ cloud), và không muốn tự vận hành pipeline; open source phù hợp khi bạn cần kiểm soát, model tùy biến, hoặc muốn tránh giá theo per-host/per-flow.
Packet analysis cho kiểm tra sâu
Metric, flow, và log cho biết rằng và đại khái cái gì; khi cần biết chính xác tại sao ở mức byte, hãy bắt packet. Wireshark/tshark và tcpdump decode các field protocol, phơi bày retransmission, frame lỗi, vấn đề TCP window, và TLS handshake. Đây là hình thức nhìn sâu nhất và tốn kém nhất — dùng có chọn lọc, kích hoạt bởi điều mà metric/flow đã báo. Được trình bày chi tiết trong ./15-tools-and-troubleshooting.md.
Alerting, baselining, SLI/SLO, và capacity planning
Baselining. Mạng có tính mùa vụ (giờ làm việc, backup ban đêm, cuối tháng). Một threshold tĩnh (“alert nếu >80% util”) tạo báo động giả và bỏ sót bất thường. Hãy thiết lập một baseline hành vi bình thường per-interface/per-hour, rồi alert khi lệch khỏi nó (anomaly detection). Baseline cũng định nghĩa thế nào là “khỏe” cho SLO.
Nguyên tắc alerting.
- Alert theo triệu chứng (tác động lên người dùng: link down, loss, latency) hơn là nguyên nhân; giữ dashboard nguyên nhân cho việc chẩn đoán.
- Làm cho mỗi alert actionable — nếu không có gì để làm thì đừng page.
- Dùng tầng severity: page cho khẩn cấp, ticket/notify cho phần còn lại.
- Thêm hysteresis / for-duration (
for: 5m) để tránh alert flapping. - Giảm nhiễu bằng grouping, deduplication, và suppression theo phụ thuộc (đừng page cho 40 thiết bị downstream khi link upstream mới là lỗi thật).
Ví dụ alert Prometheus:
groups:
- name: network
rules:
- alert: InterfaceDown
expr: ifOperStatus == 2 and ifAdminStatus == 1
for: 2m
labels: {severity: critical}
annotations:
summary: "Interface {{ $labels.ifName }} down trên {{ $labels.instance }}"
- alert: HighLinkUtilization
expr: (rate(ifHCInOctets[5m])*8 / ifHighSpeed / 1e6) > 0.85
for: 10m
labels: {severity: warning}
annotations:
summary: "Link >85% util trên {{ $labels.instance }} {{ $labels.ifName }}"
SLI / SLO cho mạng. Mượn thực hành SRE:
- SLI (indicator): một đại lượng đo được — ví dụ ”% cửa sổ 1 phút có packet loss < 0.1% trên WAN,” hoặc “p95 round-trip latency giữa DC-A và DC-B.”
- SLO (objective): mục tiêu — ví dụ “99.9% thời gian trong tháng, RTT ≤ 20 ms và loss ≤ 0.05%.”
- Error budget: phần thiếu hụt được phép (0.1% của tháng), dùng để cân bằng tốc độ thay đổi vs độ ổn định.
SLI mạng điển hình: availability (reachability qua synthetic probe/BFD), latency, jitter, loss, và throughput — đo bằng active probe (IP SLA, TWAMP, blackbox_exporter, ThousandEyes) chứ không chỉ counter thiết bị, để bạn đo path của người dùng.
Capacity planning. Dùng dữ liệu trend lưu dài để dự báo khi nào link, CPU, TCAM/FIB, hay NAT table sẽ cạn. Theo dõi utilization percentile 95 (metric billing của ngành cho transit), mô hình hóa tăng trưởng, và cấp phát trước khi vượt ~70–80% duy trì. Dữ liệu flow còn cho thấy ứng dụng hay peer nào thúc đẩy tăng trưởng, cấp thông tin cho quyết định peering và QoS (xem ./13-traffic-management-and-qos.md).
Bảng so sánh tool / kỹ thuật
| Kỹ thuật | Phù hợp nhất | Phân giải | Overhead | Ghi chú |
|---|---|---|---|---|
| SNMP polling | metric thiết bị rộng, health | 30 s–5 phút | trung bình | hỗ trợ phổ quát; dùng v3, counter 64-bit |
| Streaming telemetry (gNMI) | metric phân giải cao ở quy mô lớn | dưới giây / on-change | thấp | YANG/OpenConfig; chỉ OS hiện đại |
| NetFlow/IPFIX | kế toán traffic, bảo mật | theo flow-timeout | trung bình | ai-nói-với-ai, field phong phú |
| sFlow | real-time, nhiều port nhanh | gần tức thời (sampled) | thấp | stateless, thống kê |
| Syslog | sự kiện, config/audit, fault | theo sự kiện | thấp | cần NTP + store trung tâm |
| Synthetic probe (IP SLA/TWAMP/blackbox) | SLO theo path người dùng (latency/loss) | giây | thấp | chủ động, đo path chứ không phải box |
| Packet capture (Wireshark/tcpdump) | root-cause sâu | mỗi packet | cao | chỉ dùng có chọn lọc |
Best Practices
- Giám sát bốn golden signal cho link: utilization, error, discard/drop, và latency — cộng thêm availability. Error và discard là dấu hiệu cảnh báo sớm trước khi outage xảy ra.
- Dùng SNMPv3 authPriv trong production; giới hạn SNMP bằng ACL và một management VRF/network riêng; không bao giờ để community string
public/private. - Luôn dùng counter interface 64-bit (HC) trên link ≥1 Gbps để tránh counter wrap; tính rate từ delta, không dùng giá trị tuyệt đối.
- Đồng bộ thời gian mọi nơi bằng NTP trước tiên — metric, flow, log, và trace chỉ correlate được khi timestamp nhất quán.
- Tập trung hóa log và flow; lưu trữ theo nhu cầu compliance/forensic; downsample metric cho trend dài hạn.
- Ưu tiên streaming telemetry trên nền tảng đủ khả năng để có dữ liệu phân giải cao và on-change; giữ SNMP cho phần còn lại.
- Alert theo triệu chứng, làm alert actionable, thêm for-duration, và suppress alert phụ thuộc để chống mệt mỏi (alert fatigue).
- Baseline trước khi đặt threshold; threshold tĩnh trên traffic mùa vụ tạo nhiễu.
- Định nghĩa SLI/SLO mạng bằng active probe để đo trải nghiệm người dùng chứ không chỉ counter thiết bị, và theo dõi error budget.
- Capacity-plan theo trend percentile 95; hành động trước khi vượt ~70–80% util duy trì; để dữ liệu flow giải thích nguồn tăng trưởng.
- Bảo vệ chính mặt phẳng telemetry: mã hóa (SNMPv3/TLS/gRPC), authenticate collector, và coi dữ liệu flow/log là nhạy cảm (nó tiết lộ ai nói với ai).
- Instrument cả collector (node_exporter, self-monitoring) — một hệ thống monitoring bị “mù” còn tệ hơn không có.
- Giữ dashboard có mục đích: một overview per site/role để nhìn nhanh sức khỏe, dashboard sâu cho chẩn đoán; tránh “bức tường đồ thị” không ai đọc.
Tài liệu tham khảo
- RFC 3416 — Version 2 of the Protocol Operations for SNMP
- RFC 3414 / RFC 3410 — SNMPv3 User-based Security Model & Introduction
- RFC 7011 — Specification of the IP Flow Information Export (IPFIX) Protocol
- RFC 3954 — Cisco Systems NetFlow Services Export Version 9
- RFC 3176 — InMon sFlow: Traffic Monitoring using Packet Sampling
- RFC 5424 — The Syslog Protocol
- Prometheus documentation và snmp_exporter
- Grafana documentation · OpenConfig / gNMI
Part of the Network Engineer Roadmap knowledge base.
Overview
You cannot operate what you cannot see. Network monitoring and observability turn a black box of cables, switches, routers, firewalls, and links into a system whose health, performance, and behavior you can reason about in near real time. Monitoring answers the questions you already know to ask (“is the link up? what is the CPU on the core switch? how much bandwidth is that VLAN using?”). Observability goes further: it aims to let you answer questions you did not anticipate ahead of time, by exposing rich, high-cardinality telemetry that you can slice and correlate after the fact.
Why invest in this at all?
- Detect and diagnose faults fast. Mean Time To Detect (MTTD) and Mean Time To Repair (MTTR) dominate the user-perceived reliability of a network. Good telemetry cuts both.
- Understand performance and capacity. Latency creep, buffer bloat, microbursts, and link saturation are invisible until measured. Trend data drives capacity planning and budget justification.
- Security and forensics. Flow records and logs reveal exfiltration, scanning, DDoS, and lateral movement, and provide the audit trail after an incident.
- Prove SLAs / SLOs. Contracts with providers and internal service owners need objective, historical evidence.
- Feed automation. Closed-loop remediation and intent-based networking depend on a trustworthy stream of state.
There are two classic disciplines that observability spans:
| Discipline | Question | Typical signals |
|---|---|---|
| Fault monitoring | Is it up / broken? | interface up/down, SNMP traps, syslog errors, reachability (ICMP/BFD) |
| Performance monitoring | Is it fast / healthy enough? | throughput, latency, jitter, loss, CPU/memory, queue drops, errors |
This note maps the general “three pillars” of observability onto networking, then works through the core telemetry technologies — SNMP, flow export (NetFlow/sFlow/IPFIX), syslog, and streaming telemetry (gNMI) — and the stacks that store and visualize them (Prometheus/Grafana, and commercial platforms like Datadog and Dynatrace). It closes with alerting, baselining, SLIs/SLOs, and capacity planning. For hands-on packet-level diagnosis see ./15-tools-and-troubleshooting.md; for the protocols you are watching see ./06-core-protocols.md; and for the queueing/marking behavior these metrics reflect see ./13-traffic-management-and-qos.md.
Fundamentals
The three pillars, adapted to networking
The software-observability world talks about metrics, logs, and traces. Networking has direct analogs plus one extra pillar of its own:
| Pillar | Software world | Networking equivalent | Collected via |
|---|---|---|---|
| Metrics | Counters/gauges over time | Interface counters, CPU/mem, latency, drops, BGP peer count | SNMP polling, streaming telemetry, synthetic probes |
| Logs | App log lines | Syslog: link flaps, config changes, ACL denies, auth events | syslog (RFC 5424) |
| Traces | Distributed request spans | Path/hop analysis: traceroute, IP SLA, active path measurement | traceroute, IP SLA, TWAMP, synthetic tests |
| Flows (networking-specific) | — | Per-conversation records: who talked to whom, how much, on what ports | NetFlow / sFlow / IPFIX |
Flows are the pillar unique to networking. A single flow record summarizes an entire conversation (5-tuple + byte/packet counts + timing), which is far cheaper than capturing every packet but far richer than a bandwidth gauge.
Push vs pull
- Pull (polling): the collector periodically asks devices for state. SNMP GET and Prometheus scraping are pull models. Simple, but resolution is limited by poll interval (often 30–300 s) and the collector must know every target.
- Push (streaming/events): the device sends data when it changes or on a fine cadence. SNMP traps, syslog, NetFlow export, and gNMI dial-out are push models. Lower latency, higher resolution, better for scale — but needs the device to be configured to talk to the right collector.
Modern designs blend both: pull metrics for steady-state health, push events/telemetry for change and high-resolution data.
Cadence, cardinality, and retention
Three knobs shape every monitoring system:
- Cadence / resolution — how often you sample (1 s streaming vs 5 min SNMP). Higher resolution catches microbursts but costs storage and load.
- Cardinality — how many distinct time series (per-interface × per-device × per-metric). This explodes quickly; a chassis with 400 ports × 20 counters = 8000 series on one device.
- Retention — how long you keep raw vs rolled-up (downsampled) data. Typical: raw for days, 5-min rollups for weeks, hourly for a year+.
Key Concepts
SNMP (Simple Network Management Protocol)
SNMP is the decades-old workhorse for polling device state. Despite its age it remains the lowest-common-denominator supported by virtually every managed device.
Architecture
- Manager (NMS): the polling/collecting station (e.g., LibreNMS, Zabbix, PRTG, or an SNMP exporter feeding Prometheus).
- Agent: software on the managed device that exposes data and can emit notifications.
- MIB (Management Information Base): a hierarchical, tree-structured schema of manageable objects. Vendors publish MIBs (e.g.,
IF-MIB,BGP4-MIB,CISCO-PROCESS-MIB). - OID (Object Identifier): a dotted-decimal address of one object in the MIB tree. Example:
1.3.6.1.2.1.2.2.1.10=ifInOctets(bytes received on an interface).
Operations
| PDU | Direction | Purpose |
|---|---|---|
GET / GETNEXT / GETBULK | Manager → Agent | Read one/next/many objects (GETBULK is efficient for walking tables) |
SET | Manager → Agent | Write/change a value (rarely enabled — security risk) |
TRAP | Agent → Manager | Unsolicited, fire-and-forget notification (e.g., linkDown) |
INFORM | Agent → Manager | Like a trap but acknowledged (reliable) |
RESPONSE | Agent → Manager | Reply to GET/SET |
Polling works by the manager issuing GET/GETBULK against a list of OIDs on a schedule (e.g., every 60 s). To compute a rate (bps) you poll a counter twice and divide the delta by the interval. Watch for counter wrap: 32-bit counters (ifInOctets) wrap fast on 10G+ links — always use the 64-bit high-capacity counters (ifHCInOctets, OID 1.3.6.1.2.1.31.1.1.1.6).
Versions & security
| Version | Security | Notes |
|---|---|---|
| v1 | community string, plaintext | legacy, 32-bit counters only — avoid |
| v2c | community string, plaintext | most common in practice; GETBULK & 64-bit counters; no encryption |
| v3 | USM: authentication + privacy | user-based, supports noAuthNoPriv / authNoPriv / authPriv; use authPriv (SHA + AES) |
Prefer SNMPv3 authPriv wherever the device and NMS support it; fall back to v2c only on isolated management networks.
Example — enabling SNMP on Cisco IOS (v2c read-only, then v3):
! SNMPv2c read-only, restricted by ACL to the NMS
access-list 99 permit host 10.0.0.10
snmp-server community R3adOnly-x8f2 RO 99
snmp-server location DC1-Row4-Rack7
snmp-server contact netops@example.com
snmp-server host 10.0.0.10 version 2c R3adOnly-x8f2
! SNMPv3 — group, view, user (authPriv, SHA + AES-128)
snmp-server view ALL iso included
snmp-server group MONgrp v3 priv read ALL
snmp-server user monitor MONgrp v3 auth sha S3cret-Auth-Pass priv aes 128 S3cret-Priv-Pass
snmp-server enable traps
snmp-server host 10.0.0.10 version 3 priv monitor
Query it from a Linux collector:
# v2c walk of the interface table
snmpwalk -v2c -c 'R3adOnly-x8f2' 10.0.0.1 IF-MIB::ifTable
# v3 authPriv GET of a single counter
snmpget -v3 -l authPriv -u monitor \
-a SHA -A 'S3cret-Auth-Pass' \
-x AES -X 'S3cret-Priv-Pass' \
10.0.0.1 IF-MIB::ifHCInOctets.2
Flow-based monitoring — NetFlow, sFlow, IPFIX
Where SNMP tells you how much traffic crossed an interface, flow monitoring tells you what that traffic was: source/destination IPs, ports, protocol, ToS/DSCP, AS numbers, byte/packet counts, and timing. This is the foundation of traffic analysis, billing, peering decisions, and security detection (DDoS, scanning, exfiltration).
How flow export works (NetFlow/IPFIX)
- The router/switch maintains a flow cache, keyed by the flow definition (classically the 5-tuple: src IP, dst IP, src port, dst port, protocol; plus ingress interface, ToS).
- Packets update the matching cache entry’s counters.
- A flow is expired and exported when it ends (TCP FIN/RST), goes idle (idle timeout), lives too long (active timeout, e.g., 60 s), or the cache fills.
- Expired flows are packed into UDP datagrams and sent to a flow collector (e.g., nfdump/nfsen, Elastiflow, ntopng, Kentik).
Sampling. On high-speed links, tracking every flow is expensive, so devices sample 1-in-N packets (e.g., 1:1000). Sampling is essential at scale but means you must scale counts back up and lose visibility into very small flows.
Templates. NetFlow v9 and IPFIX are template-based: the exporter first sends a template describing the fields, then data records referencing it. This makes them extensible (unlike fixed NetFlow v5). IPFIX also supports variable-length fields and enterprise-specific elements (e.g., URLs, application IDs).
sFlow works differently: instead of a flow cache it does packet sampling — it grabs the first N bytes (header) of every Nth packet and ships that sample immediately, plus periodic interface counters. This is stateless on the device (cheap, hardware-friendly, real-time) but statistical rather than a full accounting of every conversation.
Comparison:
| Feature | NetFlow v5 | NetFlow v9 | IPFIX (v10) | sFlow |
|---|---|---|---|---|
| Origin | Cisco | Cisco | IETF standard (RFC 7011) | sFlow.org / RFC 3176 |
| Model | flow cache | flow cache | flow cache | packet + counter sampling |
| Fields | fixed | template-based | template, variable-length, enterprise elements | packet header samples |
| Sampling | optional | optional | optional | inherent (always sampled) |
| Layers | L3/L4 | L3/L4 (+MPLS, IPv6) | L2–L7 | L2–L7 (raw header) |
| State on device | yes (cache) | yes | yes | no (stateless) |
| Real-time | delayed by timeouts | delayed by timeouts | delayed by timeouts | near-instant |
| Standard | proprietary | proprietary (basis of IPFIX) | open IETF | multi-vendor standard |
Rule of thumb: IPFIX/NetFlow v9 when you want accurate accounting and rich fields; sFlow when you want cheap, real-time, hardware-based visibility across many high-speed ports.
Example — Flexible NetFlow (IPFIX-style) on Cisco IOS:
flow record RECORD-v4
match ipv4 source address
match ipv4 destination address
match transport source-port
match transport destination-port
match ipv4 protocol
collect counter bytes
collect counter packets
collect interface input
flow exporter EXP-1
destination 10.0.0.20
transport udp 2055
export-protocol ipfix
template data timeout 60
flow monitor MON-1
record RECORD-v4
exporter EXP-1
cache timeout active 60
interface GigabitEthernet0/1
ip flow monitor MON-1 input
ip flow monitor MON-1 output
Syslog and log management
Syslog (RFC 5424, older RFC 3164) is the standard for device event logging: link flaps, OSPF/BGP adjacency changes, config changes, ACL/firewall denies, authentication events, hardware alarms. Logs are the narrative record that explains why a metric moved.
Severity levels (0 = worst). A useful mnemonic: Every Awesome Cisco Engineer Will Need Ice-cream Daily.
| Level | Keyword | Meaning |
|---|---|---|
| 0 | Emergency | system unusable |
| 1 | Alert | act immediately |
| 2 | Critical | critical condition |
| 3 | Error | error condition (e.g., interface errors) |
| 4 | Warning | warning condition |
| 5 | Notice | normal but significant (config change) |
| 6 | Informational | informational (link up) |
| 7 | Debug | debug messages (noisy) |
Facilities categorize the source (e.g., local0–local7, auth, kern). A message combines facility + severity into a priority.
Ship logs to a central collector — rsyslog/syslog-ng, an ELK/OpenSearch stack, Graylog, Loki, or a SIEM (Splunk, Elastic Security) for correlation and retention. Centralization matters because a device that crashes takes its local logs with it, and cross-device correlation (a flap here, an error there) needs everything in one place.
Example — Cisco IOS to a syslog server, timestamped and level-filtered:
service timestamps log datetime msec localtime show-timezone
logging host 10.0.0.30 transport udp port 514
logging trap informational ! send severity 6 and worse
logging source-interface Loopback0
logging buffered 64000 debugging
Always sync clocks with NTP first — logs without accurate, consistent timestamps are nearly useless for correlation.
Streaming telemetry (model-driven, gNMI)
SNMP polling has real limits: coarse resolution (you can’t poll thousands of OIDs every second), high overhead (request/response per object), and rigid MIB structures. Model-driven streaming telemetry is the modern answer.
- The device pushes data continuously (subscribe once, stream forever) instead of being polled.
- Data is structured by YANG models (open standards like OpenConfig, or vendor models) — the same models used for config, so monitoring and configuration share one schema.
- Transport is typically gNMI (gRPC Network Management Interface) over HTTP/2 with TLS; encoding is Protobuf or JSON. NETCONF/RESTCONF can also carry telemetry.
- Modes: dial-out (device connects to collector) or dial-in (collector subscribes to device); sample (fixed interval, e.g., 1 s) or on-change (event-driven, e.g., a BGP state change).
Result: sub-second resolution, lower device CPU per data point, and a consistent data model. Collect with a telemetry pipeline — Telegraf (with the gNMI/cisco_telemetry inputs), Cisco’s Pipeline, or gnmic — landing into a TSDB (InfluxDB, Prometheus) for Grafana.
Example — OpenConfig gNMI subscription on Cisco IOS-XR (dial-in path), plus a gnmic subscribe:
! Device side (IOS-XR) — enable gRPC/gNMI
grpc
port 57400
tls-mutual
!
telemetry model-driven
sensor-group IFCOUNTERS
sensor-path openconfig-interfaces:interfaces/interface/state/counters
# Collector side — subscribe at 1s sample interval
gnmic -a 10.0.0.1:57400 --skip-verify -u admin -p '***' \
subscribe --path "openconfig-interfaces:interfaces/interface/state/counters" \
--mode stream --stream-mode sample --sample-interval 1s
| Aspect | SNMP polling | Streaming telemetry (gNMI) |
|---|---|---|
| Model | pull / request-response | push / subscribe |
| Resolution | 30 s – 5 min typical | sub-second, or on-change |
| Data schema | MIB/OID | YANG (OpenConfig / vendor) |
| Overhead | high (per-object) | low (batched stream) |
| Encoding | BER | Protobuf / JSON |
| Security | v3 USM | gRPC + TLS |
| Best for | broad device support, legacy | scale, high resolution, modern OS |
Metrics & time-series stack — Prometheus + Grafana
The open-source de-facto standard for storing and visualizing metrics.
- Prometheus — a pull-based TSDB that scrapes HTTP
/metricsendpoints, stores dimensional time series (metric name + labels), and provides PromQL for querying and alerting. - Exporters bridge non-native sources into Prometheus:
- snmp_exporter — Prometheus scrapes it over HTTP; it in turn SNMP-polls network devices using a generated
snmp.yml(built from MIBs by thegenerator). This is how you get switch/router SNMP data into Prometheus. - node_exporter — host-level metrics (CPU, disk, NIC counters) for Linux servers/collectors.
- Others: blackbox_exporter (ICMP/TCP/HTTP probes for reachability & latency), gnmi via Telegraf → remote_write.
- snmp_exporter — Prometheus scrapes it over HTTP; it in turn SNMP-polls network devices using a generated
- Grafana — dashboards over Prometheus (and InfluxDB, Loki, etc.): time-series panels, heatmaps, topology, threshold coloring, and its own alerting.
- Alertmanager — dedupes, groups, silences, and routes Prometheus alerts to email/Slack/PagerDuty.
Example — Prometheus scraping the SNMP exporter for two switches:
# prometheus.yml
scrape_configs:
- job_name: 'snmp-switches'
static_configs:
- targets:
- 10.0.0.1 # core-sw-1
- 10.0.0.2 # core-sw-2
metrics_path: /snmp
params:
module: [if_mib] # module defined in snmp.yml
relabel_configs:
- source_labels: [__address__]
target_label: __param_target
- source_labels: [__param_target]
target_label: instance
- target_label: __address__
replacement: 127.0.0.1:9116 # the snmp_exporter host:port
- job_name: 'node'
static_configs:
- targets: ['10.0.0.30:9100'] # collector host node_exporter
Example PromQL — interface throughput in bits/sec and an error rate:
# Ingress throughput (bps) per interface, 5-min rate
rate(ifHCInOctets[5m]) * 8
# Interfaces with a rising input-error rate
rate(ifInErrors[5m]) > 0
Commercial observability platforms
Open-source stacks are powerful but need assembly and operation. Commercial SaaS platforms trade cost for turnkey breadth, correlation, and support.
| Platform | Sweet spot for networking |
|---|---|
| Datadog | Unified infra + app + Network Performance Monitoring (NPM) and Network Device Monitoring (NDM/SNMP); flow ingestion; strong correlation across app ↔ network; good for cloud-heavy, DevOps-oriented orgs. |
| Dynatrace | AI-driven root-cause (Davis engine), automatic topology (Smartscape), application-aware network visibility; enterprise APM leaning. |
| Others | SolarWinds NPM, Cisco ThousandEyes (Internet/path visibility, SaaS reachability), Kentik (flow analytics at scale), LogicMonitor, Auvik. |
They fit best when you value time-to-value, cross-domain correlation (network ↔ application ↔ cloud), and don’t want to run the pipeline yourself; open source fits when you need control, custom models, or want to avoid per-host/per-flow pricing.
Packet analysis for deep inspection
Metrics, flows, and logs tell you that and roughly what; when you need to know exactly why at the byte level, capture packets. Wireshark/tshark and tcpdump decode protocol fields, expose retransmissions, malformed frames, TCP window issues, and TLS handshakes. This is the deepest and most expensive form of visibility — use it surgically, triggered by what your metrics/flows flagged. Covered in depth in ./15-tools-and-troubleshooting.md.
Alerting, baselining, SLIs/SLOs, and capacity planning
Baselining. Networks are seasonal (business hours, backups at night, month-end). A static threshold (“alert if >80% util”) produces false alarms and misses anomalies. Establish a baseline of normal per-interface/per-hour behavior, then alert on deviation from it (anomaly detection). Baselines also define what “healthy” means for SLOs.
Alerting principles.
- Alert on symptoms (user-facing impact: link down, loss, latency) more than causes; keep cause dashboards for diagnosis.
- Make every alert actionable — if there’s nothing to do, it shouldn’t page.
- Use severity tiers: page for urgent, ticket/notify for the rest.
- Add hysteresis / for-duration (
for: 5m) to avoid flapping alerts. - Reduce noise via grouping, deduplication, and dependency suppression (don’t page for 40 downstream devices when their upstream link is the real fault).
Example Prometheus alert:
groups:
- name: network
rules:
- alert: InterfaceDown
expr: ifOperStatus == 2 and ifAdminStatus == 1
for: 2m
labels: {severity: critical}
annotations:
summary: "Interface {{ $labels.ifName }} down on {{ $labels.instance }}"
- alert: HighLinkUtilization
expr: (rate(ifHCInOctets[5m])*8 / ifHighSpeed / 1e6) > 0.85
for: 10m
labels: {severity: warning}
annotations:
summary: "Link >85% utilized on {{ $labels.instance }} {{ $labels.ifName }}"
SLIs / SLOs for networks. Borrow SRE practice:
- SLI (indicator): a measured quantity — e.g., ”% of 1-min windows with packet loss < 0.1% on the WAN,” or “p95 round-trip latency between DC-A and DC-B.”
- SLO (objective): the target — e.g., “99.9% of the month, RTT ≤ 20 ms and loss ≤ 0.05%.”
- Error budget: the allowed shortfall (0.1% of the month), used to balance change velocity vs stability.
Typical network SLIs: availability (reachability via synthetic probes/BFD), latency, jitter, loss, and throughput — measured with active probes (IP SLA, TWAMP, blackbox_exporter, ThousandEyes) rather than only device counters, so you measure the user’s path.
Capacity planning. Use long-retention trend data to forecast when links, CPU, TCAM/FIB, or NAT tables will exhaust. Track the 95th-percentile utilization (the industry billing metric for transit), model growth, and provision before you cross ~70–80% sustained. Flow data additionally shows which applications or peers drive growth, informing peering and QoS decisions (see ./13-traffic-management-and-qos.md).
Tool / technique comparison
| Technique | Best for | Resolution | Overhead | Notes |
|---|---|---|---|---|
| SNMP polling | broad device metrics, health | 30 s–5 min | medium | universal support; use v3, 64-bit counters |
| Streaming telemetry (gNMI) | high-res metrics at scale | sub-second / on-change | low | YANG/OpenConfig; modern OS only |
| NetFlow/IPFIX | traffic accounting, security | flow-timeout | medium | who-talks-to-whom, rich fields |
| sFlow | real-time, many fast ports | near-instant (sampled) | low | stateless, statistical |
| Syslog | events, config/audit, faults | event-driven | low | needs NTP + central store |
| Synthetic probes (IP SLA/TWAMP/blackbox) | user-path SLOs (latency/loss) | seconds | low | active, measures the path not the box |
| Packet capture (Wireshark/tcpdump) | deep root-cause | per-packet | high | surgical use only |
Best Practices
- Monitor the four golden signals for links: utilization, errors, discards/drops, and latency — plus availability. Errors and discards are early warning signs that precede outages.
- Use SNMPv3 authPriv on production; restrict SNMP with ACLs and a dedicated management VRF/network; never leave
public/privatecommunity strings. - Always use 64-bit (HC) interface counters on ≥1 Gbps links to avoid counter wrap; compute rates from deltas, not absolute values.
- Sync time everywhere with NTP before anything else — metrics, flows, logs, and traces are only correlatable with consistent timestamps.
- Centralize logs and flows; retain per compliance/forensic need; downsample metrics for long-term trends.
- Prefer streaming telemetry on capable platforms for high-resolution and on-change data; keep SNMP for coverage of everything else.
- Alert on symptoms, make alerts actionable, add for-durations, and suppress dependent alerts to fight fatigue.
- Baseline before you threshold; static thresholds on seasonal traffic create noise.
- Define network SLIs/SLOs with active probes so you measure the user’s experience, not just device counters, and track error budgets.
- Capacity-plan on 95th-percentile trends; act before ~70–80% sustained utilization; let flow data explain the growth.
- Secure the telemetry plane itself: encrypt (SNMPv3/TLS/gRPC), authenticate collectors, and treat flow/log data as sensitive (it reveals who talks to whom).
- Instrument the collector too (node_exporter, self-monitoring) — a blind monitoring system is worse than none.
- Keep dashboards purposeful: an overview per site/role for at-a-glance health, deep dashboards for diagnosis; avoid “wall of graphs” nobody reads.
References
- RFC 3416 — Version 2 of the Protocol Operations for SNMP
- RFC 3414 / RFC 3410 — SNMPv3 User-based Security Model & Introduction
- RFC 7011 — Specification of the IP Flow Information Export (IPFIX) Protocol
- RFC 3954 — Cisco Systems NetFlow Services Export Version 9
- RFC 3176 — InMon sFlow: Traffic Monitoring using Packet Sampling
- RFC 5424 — The Syslog Protocol
- Prometheus documentation and snmp_exporter
- Grafana documentation · OpenConfig / gNMI