63 lines
3.4 KiB
Markdown
63 lines
3.4 KiB
Markdown
# 01 - Architecture Overview
|
|
|
|
This documentation suite serves as a complete playbook to build a production-grade monitoring pipeline from scratch.
|
|
|
|
## The Architecture Stack
|
|
Our monitoring infrastructure is built on a **PULL-based architecture** combined with **PUSH-based alerting**.
|
|
|
|
1. **Node Exporters (Remote Nodes):** Tiny agents installed on every server/desktop. They sit on a port (usually `9100` for CPU/RAM, `9835` for GPUs) and simply expose raw hardware metrics.
|
|
2. **Prometheus (Central Server):** The brain. It actively **PULLS** (scrapes) data from all the remote Node Exporters every 15 seconds. It also evaluates Alert Rules (e.g., "Is a server offline?").
|
|
3. **Grafana (Central Server):** The visualizer. It connects to the Prometheus database to draw beautiful, real-time dashboards.
|
|
4. **Alertmanager (Central Server):** The router. When Prometheus detects an anomaly, it **PUSHES** an alert to Alertmanager. Alertmanager groups, delays, and routes the alert to external systems like Email and Rocket.Chat.
|
|
|
|
```mermaid
|
|
graph TD
|
|
subgraph Remote Nodes
|
|
N1[Node Exporter 1]
|
|
N2[Node Exporter 2]
|
|
G1[GPU Exporter]
|
|
end
|
|
|
|
subgraph Central Server
|
|
P[Prometheus]
|
|
A[Alertmanager]
|
|
G[Grafana]
|
|
end
|
|
|
|
subgraph Notifications
|
|
E[Email]
|
|
R[Rocket.Chat]
|
|
end
|
|
|
|
P -- "Pulls metrics (15s)" --> N1
|
|
P -- "Pulls metrics (15s)" --> N2
|
|
P -- "Pulls metrics (15s)" --> G1
|
|
G -- "Queries DB" --> P
|
|
P -- "Pushes alerts" --> A
|
|
A -- "Routes alert" --> E
|
|
A -- "Routes alert" --> R
|
|
```
|
|
|
|
## Setup Sequence
|
|
To deploy this infrastructure from scratch, you must follow the steps in this order:
|
|
1. Setup the Central Server (`02_central_server_setup.md`)
|
|
2. Setup the Remote Nodes (`03_remote_node_setup.md`)
|
|
3. Review the Day-to-Day Operations (`04_operations_and_debugging.md`)
|
|
|
|
---
|
|
|
|
## Understanding Alert Timing
|
|
Timing configurations exist in both Prometheus and Alertmanager to ensure you receive timely alerts without being spammed by false positives.
|
|
|
|
### The Chronological Flow
|
|
|
|
- **T=0s:** A remote server suddenly loses power.
|
|
- **T=15s (Prometheus `scrape_interval`):** The security guard checks the shop. Prometheus attempts its regular 15s scrape, gets a connection error, and marks the server state as `PENDING`.
|
|
- **T=25s (Prometheus `for: 10s`):** The guard waits to be sure. After being down for a full 10 seconds, Prometheus promotes the alert from `PENDING` to `FIRING`.
|
|
- **T=35s (Alertmanager `group_wait: 10s`):** The boss waits for more news. Alertmanager receives the firing alert, but waits 10 seconds to see if any *other* servers crash at the same time so it can bundle them into a single message. It then sends the first notification to Rocket.Chat/Email.
|
|
|
|
### The Follow-up Rules (Alertmanager)
|
|
- **`group_interval: 1m`:** If a brand *new* server crashes immediately after that first notification was sent, Alertmanager waits 1 minute before sending an update to avoid spamming the channel.
|
|
- **`repeat_interval: 5m`:** If nobody fixes the server and it remains offline, Alertmanager will re-send the exact same notification every 5 minutes until it is resolved. *(Note: `repeat_interval` must always be larger than `group_interval`!)*
|
|
- **`resolve_timeout: 5m`:** When the server finally comes back online, Alertmanager waits 5 minutes before sending the green `RESOLVED` notification. This ensures the server is completely stable and not just flickering on and off.
|