406 lines
11 KiB
Markdown
406 lines
11 KiB
Markdown
# Central Monitoring System — Complete Setup
|
|
|
|
## Prometheus + Grafana + Alertmanager + systemd Process Management
|
|
|
|
Run every command below on your **central Linux server** (native Linux, not WSL).
|
|
|
|
---
|
|
|
|
## Prerequisites
|
|
|
|
- Ubuntu 22.04 or 24.04 LTS (native, always-on machine)
|
|
- You have already installed Tailscale and MeshCentral (see `08_central_server_production_setup.md`)
|
|
- Your Tailscale IP is noted (run `tailscale ip -4` to get it)
|
|
- You are running all commands as your normal user (use `sudo` where shown)
|
|
|
|
---
|
|
|
|
## Part 1: Prometheus
|
|
|
|
### Step 1 — Download and Install
|
|
|
|
```bash
|
|
cd /tmp
|
|
PROM_VERSION="2.51.2"
|
|
wget https://github.com/prometheus/prometheus/releases/download/v${PROM_VERSION}/prometheus-${PROM_VERSION}.linux-amd64.tar.gz
|
|
tar xvf prometheus-${PROM_VERSION}.linux-amd64.tar.gz
|
|
cd prometheus-${PROM_VERSION}.linux-amd64
|
|
|
|
sudo mv prometheus promtool /usr/local/bin/
|
|
sudo mkdir -p /etc/prometheus /var/lib/prometheus
|
|
sudo mv prometheus.yml consoles console_libraries /etc/prometheus/
|
|
```
|
|
|
|
Verify:
|
|
|
|
```bash
|
|
prometheus --version
|
|
# Should print: prometheus, version 2.51.2
|
|
```
|
|
|
|
### Step 2 — Write the Config File
|
|
|
|
```bash
|
|
sudo tee /etc/prometheus/prometheus.yml << 'EOF'
|
|
global:
|
|
scrape_interval: 15s
|
|
evaluation_interval: 15s
|
|
|
|
alerting:
|
|
alertmanagers:
|
|
- static_configs:
|
|
- targets: ['localhost:9093']
|
|
|
|
rule_files:
|
|
- "/etc/prometheus/rules.yml"
|
|
|
|
scrape_configs:
|
|
- job_name: 'prometheus'
|
|
static_configs:
|
|
- targets: ['localhost:9090']
|
|
|
|
# Add each client machine's Tailscale IP here after connecting them
|
|
# - job_name: 'node_exporter'
|
|
# static_configs:
|
|
# - targets:
|
|
# - '100.x.x.x:9100' # machine name
|
|
# - '100.x.x.x:9100' # machine name
|
|
EOF
|
|
```
|
|
|
|
### Step 3 — Write Alert Rules
|
|
|
|
```bash
|
|
sudo tee /etc/prometheus/rules.yml << 'EOF'
|
|
groups:
|
|
- name: server_alerts
|
|
rules:
|
|
- alert: ServerDown
|
|
expr: up == 0
|
|
for: 1m
|
|
labels:
|
|
severity: critical
|
|
annotations:
|
|
summary: "Server {{ $labels.instance }} is DOWN"
|
|
description: "{{ $labels.instance }} has been unreachable for over 1 minute."
|
|
|
|
- alert: HighCPU
|
|
expr: 100 - (avg by(instance)(rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100) > 85
|
|
for: 5m
|
|
labels:
|
|
severity: warning
|
|
annotations:
|
|
summary: "High CPU on {{ $labels.instance }}"
|
|
description: "CPU usage is above 85% for 5 minutes."
|
|
|
|
- alert: HighRAM
|
|
expr: (1 - (node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes)) * 100 > 90
|
|
for: 5m
|
|
labels:
|
|
severity: warning
|
|
annotations:
|
|
summary: "High RAM on {{ $labels.instance }}"
|
|
description: "RAM usage is above 90% for 5 minutes."
|
|
|
|
- alert: LowDisk
|
|
expr: (node_filesystem_avail_bytes{mountpoint="/"} / node_filesystem_size_bytes{mountpoint="/"}) * 100 < 15
|
|
for: 5m
|
|
labels:
|
|
severity: warning
|
|
annotations:
|
|
summary: "Low disk on {{ $labels.instance }}"
|
|
description: "Root disk has less than 15% free space."
|
|
|
|
- alert: HighGPUTemp
|
|
expr: nvidia_smi_temperature_gpu > 85
|
|
for: 5m
|
|
labels:
|
|
severity: critical
|
|
annotations:
|
|
summary: "GPU overheating on {{ $labels.instance }}"
|
|
description: "GPU temperature is above 85°C."
|
|
EOF
|
|
```
|
|
|
|
### Step 4 — Create systemd Service
|
|
|
|
```bash
|
|
sudo tee /etc/systemd/system/prometheus.service << 'EOF'
|
|
[Unit]
|
|
Description=Prometheus Monitoring
|
|
Documentation=https://prometheus.io/docs/
|
|
After=network.target
|
|
|
|
[Service]
|
|
User=root
|
|
ExecStart=/usr/local/bin/prometheus \
|
|
--config.file=/etc/prometheus/prometheus.yml \
|
|
--storage.tsdb.path=/var/lib/prometheus \
|
|
--storage.tsdb.retention.time=30d \
|
|
--storage.tsdb.retention.size=10GB \
|
|
--web.listen-address=0.0.0.0:9090 \
|
|
--web.enable-lifecycle
|
|
Restart=always
|
|
RestartSec=5
|
|
|
|
[Install]
|
|
WantedBy=multi-user.target
|
|
EOF
|
|
|
|
sudo systemctl daemon-reload
|
|
sudo systemctl enable prometheus
|
|
sudo systemctl start prometheus
|
|
sudo systemctl status prometheus
|
|
```
|
|
|
|
Expected: `active (running)`
|
|
|
|
Verify in browser: `http://localhost:9090`
|
|
|
|
### Step 5 — Adding a New Client to Prometheus
|
|
|
|
After connecting a client machine via Tailscale and deploying Node Exporter:
|
|
|
|
```bash
|
|
sudo nano /etc/prometheus/prometheus.yml
|
|
```
|
|
|
|
Add the machine under `scrape_configs`:
|
|
|
|
```yaml
|
|
- job_name: "node_exporter"
|
|
static_configs:
|
|
- targets:
|
|
- "100.x.x.x:9100" # machine hostname
|
|
```
|
|
|
|
Reload without restarting (since `--web.enable-lifecycle` is set):
|
|
|
|
```bash
|
|
curl -X POST http://localhost:9090/-/reload
|
|
```
|
|
|
|
Check `http://localhost:9090/targets` — new machine should appear green.
|
|
|
|
---
|
|
|
|
## Part 2: Alertmanager
|
|
|
|
### Step 1 — Download and Install
|
|
|
|
```bash
|
|
cd /tmp
|
|
AM_VERSION="0.27.0"
|
|
wget https://github.com/prometheus/alertmanager/releases/download/v${AM_VERSION}/alertmanager-${AM_VERSION}.linux-amd64.tar.gz
|
|
tar xvf alertmanager-${AM_VERSION}.linux-amd64.tar.gz
|
|
cd alertmanager-${AM_VERSION}.linux-amd64
|
|
|
|
sudo mv alertmanager amtool /usr/local/bin/
|
|
sudo mkdir -p /etc/alertmanager /var/lib/alertmanager
|
|
```
|
|
|
|
### Step 2 — Configure Alertmanager
|
|
|
|
```bash
|
|
sudo tee /etc/alertmanager/alertmanager.yml << 'EOF'
|
|
global:
|
|
resolve_timeout: 5m
|
|
smtp_smarthost: "smtp.gmail.com:587"
|
|
smtp_from: "no-reply@seekright.com"
|
|
smtp_auth_username: "no-reply@seekright.com"
|
|
smtp_auth_password: "clzfywewriqxlvqv"
|
|
|
|
route:
|
|
receiver: "fallback-do-nothing"
|
|
group_wait: 10s # ← ADD THIS
|
|
group_interval: 1m # ← ADD THIS (must be ≤ repeat_interval)
|
|
repeat_interval: 1m # now this actually works
|
|
routes:
|
|
- receiver: "rocket-chat-alerts"
|
|
continue: true
|
|
- receiver: "email-alerts"
|
|
continue: true
|
|
|
|
receivers:
|
|
- name: "fallback-do-nothing"
|
|
|
|
- name: "rocket-chat-alerts"
|
|
webhook_configs: # Changed from slack_configs
|
|
- url: "https://text.seekright.com/hooks/6a05a268a88367bc6e0fe318/aNjZzmLy2wB3R4BR7ZtoLtAZYyfNyfLAoQMszmhfBDveKrB>
|
|
send_resolved: true
|
|
- name: "email-alerts"
|
|
EOF
|
|
```
|
|
|
|
Replace:
|
|
|
|
- `YOUR_GMAIL@gmail.com` — your Gmail address
|
|
- `YOUR_GMAIL_APP_PASSWORD` — Gmail App Password (not your login password — create one at myaccount.google.com → Security → App Passwords)
|
|
- `YOUR_ROCKETCHAT_WEBHOOK_URL` — your Rocket.Chat incoming webhook URL
|
|
|
|
### Step 3 — Create systemd Service
|
|
|
|
```bash
|
|
sudo tee /etc/systemd/system/alertmanager.service << 'EOF'
|
|
[Unit]
|
|
Description=Alertmanager
|
|
Documentation=https://prometheus.io/docs/alerting/alertmanager/
|
|
After=network.target
|
|
|
|
[Service]
|
|
User=root
|
|
ExecStart=/usr/local/bin/alertmanager \
|
|
--config.file=/etc/alertmanager/alertmanager.yml \
|
|
--storage.path=/var/lib/alertmanager \
|
|
--web.listen-address=0.0.0.0:9093
|
|
Restart=always
|
|
RestartSec=5
|
|
|
|
[Install]
|
|
WantedBy=multi-user.target
|
|
EOF
|
|
|
|
sudo systemctl daemon-reload
|
|
sudo systemctl enable alertmanager
|
|
sudo systemctl start alertmanager
|
|
sudo systemctl status alertmanager
|
|
```
|
|
|
|
Verify in browser: `http://localhost:9093`
|
|
|
|
---
|
|
|
|
## Part 3: Grafana
|
|
|
|
### Step 1 — Install via apt (Official Repo)
|
|
|
|
```bash
|
|
sudo apt-get install -y apt-transport-https software-properties-common wget
|
|
sudo mkdir -p /etc/apt/keyrings/
|
|
wget -q -O - https://apt.grafana.com/gpg.key | gpg --dearmor | sudo tee /etc/apt/keyrings/grafana.gpg > /dev/null
|
|
echo "deb [signed-by=/etc/apt/keyrings/grafana.gpg] https://apt.grafana.com stable main" | sudo tee /etc/apt/sources.list.d/grafana.list
|
|
sudo apt-get update
|
|
sudo apt-get install -y grafana
|
|
|
|
sudo systemctl daemon-reload
|
|
sudo systemctl enable grafana-server
|
|
sudo systemctl start grafana-server
|
|
sudo systemctl status grafana-server
|
|
```
|
|
|
|
Verify in browser: `http://localhost:3000`
|
|
Default login: `admin` / `admin` — change password immediately on first login.
|
|
|
|
### Step 2 — Connect Grafana to Prometheus
|
|
|
|
1. Go to **Connections → Data Sources → Add data source**
|
|
2. Choose **Prometheus**
|
|
3. Set URL: `http://localhost:9090`
|
|
4. Click **Save & Test** — must show "Data source is working"
|
|
|
|
### Step 3 — Import Dashboards
|
|
|
|
Go to **Dashboards → Import** and import these IDs one by one:
|
|
|
|
| Dashboard ID | What It Shows |
|
|
| ------------ | ------------------------------------------------------------- |
|
|
| `1860` | Node Exporter Full — CPU, RAM, Disk, Network for all machines |
|
|
| `14574` | NVIDIA GPU Metrics — GPU temp, VRAM, utilization |
|
|
|
|
---
|
|
|
|
## Part 4: Process Management Reference
|
|
|
|
### Check all service statuses at once
|
|
|
|
```bash
|
|
sudo systemctl status prometheus alertmanager grafana-server
|
|
```
|
|
|
|
### Start / Stop / Restart any service
|
|
|
|
```bash
|
|
sudo systemctl start prometheus
|
|
sudo systemctl stop prometheus
|
|
sudo systemctl restart prometheus
|
|
|
|
sudo systemctl start alertmanager
|
|
sudo systemctl stop alertmanager
|
|
sudo systemctl restart alertmanager
|
|
|
|
sudo systemctl start grafana-server
|
|
sudo systemctl stop grafana-server
|
|
sudo systemctl restart grafana-server
|
|
```
|
|
|
|
### Check live logs for any service
|
|
|
|
```bash
|
|
sudo journalctl -u prometheus -f
|
|
sudo journalctl -u alertmanager -f
|
|
sudo journalctl -u grafana-server -f
|
|
```
|
|
|
|
### Reload Prometheus config without restarting
|
|
|
|
```bash
|
|
curl -X POST http://localhost:9090/-/reload
|
|
# Only works because we set --web.enable-lifecycle
|
|
```
|
|
|
|
### Check all listening ports
|
|
|
|
```bash
|
|
ss -tlnp | grep -E "9090|9093|3000|4430"
|
|
```
|
|
|
|
### Check all enabled services (confirms they survive reboot)
|
|
|
|
```bash
|
|
sudo systemctl is-enabled prometheus alertmanager grafana-server
|
|
# All three must print: enabled
|
|
```
|
|
|
|
---
|
|
|
|
## Part 5: Final Verification Checklist
|
|
|
|
Run this after complete setup to confirm everything is working:
|
|
|
|
```bash
|
|
echo "=== Service Status ===" && \
|
|
sudo systemctl is-active prometheus alertmanager grafana-server && \
|
|
echo "=== Prometheus Targets ===" && \
|
|
curl -s http://localhost:9090/api/v1/targets | python3 -c "import sys,json; [print(t['labels']['job'], t['health']) for t in json.load(sys.stdin)['data']['activeTargets']]" && \
|
|
echo "=== Ports Listening ===" && \
|
|
ss -tlnp | grep -E "9090|9093|3000"
|
|
```
|
|
|
|
Expected output:
|
|
|
|
```
|
|
=== Service Status ===
|
|
active
|
|
active
|
|
active
|
|
=== Prometheus Targets ===
|
|
prometheus up
|
|
node_exporter up
|
|
=== Ports Listening ===
|
|
LISTEN *:9090 prometheus
|
|
LISTEN *:9093 alertmanager
|
|
LISTEN *:3000 grafana
|
|
```
|
|
|
|
---
|
|
|
|
## Quick Reference
|
|
|
|
| Service | Port | Config File | Log Command |
|
|
| ------------ | ---- | -------------------------------------------- | --------------------------------- |
|
|
| Prometheus | 9090 | `/etc/prometheus/prometheus.yml` | `journalctl -u prometheus -f` |
|
|
| Alertmanager | 9093 | `/etc/alertmanager/alertmanager.yml` | `journalctl -u alertmanager -f` |
|
|
| Grafana | 3000 | `/etc/grafana/grafana.ini` | `journalctl -u grafana-server -f` |
|
|
| MeshCentral | 4430 | `~/meshcentral/meshcentral-data/config.json` | `journalctl -u meshcentral -f` |
|
|
| Tailscale | — | — | `tailscale status` |
|