3.4 KiB
3.4 KiB
01 - Architecture Overview
This documentation suite serves as a complete playbook to build a production-grade monitoring pipeline from scratch.
The Architecture Stack
Our monitoring infrastructure is built on a PULL-based architecture combined with PUSH-based alerting.
- Node Exporters (Remote Nodes): Tiny agents installed on every server/desktop. They sit on a port (usually
9100for CPU/RAM,9835for GPUs) and simply expose raw hardware metrics. - Prometheus (Central Server): The brain. It actively PULLS (scrapes) data from all the remote Node Exporters every 15 seconds. It also evaluates Alert Rules (e.g., "Is a server offline?").
- Grafana (Central Server): The visualizer. It connects to the Prometheus database to draw beautiful, real-time dashboards.
- Alertmanager (Central Server): The router. When Prometheus detects an anomaly, it PUSHES an alert to Alertmanager. Alertmanager groups, delays, and routes the alert to external systems like Email and Rocket.Chat.
graph TD
subgraph Remote Nodes
N1[Node Exporter 1]
N2[Node Exporter 2]
G1[GPU Exporter]
end
subgraph Central Server
P[Prometheus]
A[Alertmanager]
G[Grafana]
end
subgraph Notifications
E[Email]
R[Rocket.Chat]
end
P -- "Pulls metrics (15s)" --> N1
P -- "Pulls metrics (15s)" --> N2
P -- "Pulls metrics (15s)" --> G1
G -- "Queries DB" --> P
P -- "Pushes alerts" --> A
A -- "Routes alert" --> E
A -- "Routes alert" --> R
Setup Sequence
To deploy this infrastructure from scratch, you must follow the steps in this order:
- Setup the Central Server (
02_central_server_setup.md) - Setup the Remote Nodes (
03_remote_node_setup.md) - Review the Day-to-Day Operations (
04_operations_and_debugging.md)
Understanding Alert Timing
Timing configurations exist in both Prometheus and Alertmanager to ensure you receive timely alerts without being spammed by false positives.
The Chronological Flow
- T=0s: A remote server suddenly loses power.
- T=15s (Prometheus
scrape_interval): The security guard checks the shop. Prometheus attempts its regular 15s scrape, gets a connection error, and marks the server state asPENDING. - T=25s (Prometheus
for: 10s): The guard waits to be sure. After being down for a full 10 seconds, Prometheus promotes the alert fromPENDINGtoFIRING. - T=35s (Alertmanager
group_wait: 10s): The boss waits for more news. Alertmanager receives the firing alert, but waits 10 seconds to see if any other servers crash at the same time so it can bundle them into a single message. It then sends the first notification to Rocket.Chat/Email.
The Follow-up Rules (Alertmanager)
group_interval: 1m: If a brand new server crashes immediately after that first notification was sent, Alertmanager waits 1 minute before sending an update to avoid spamming the channel.repeat_interval: 5m: If nobody fixes the server and it remains offline, Alertmanager will re-send the exact same notification every 5 minutes until it is resolved. (Note:repeat_intervalmust always be larger thangroup_interval!)resolve_timeout: 5m: When the server finally comes back online, Alertmanager waits 5 minutes before sending the greenRESOLVEDnotification. This ensures the server is completely stable and not just flickering on and off.