initial commit

This commit is contained in:
kawsik
2026-06-02 10:37:57 +05:30
commit b4acb8d1df
49 changed files with 15350 additions and 0 deletions

405
09_central_system_setup.md Normal file
View File

@@ -0,0 +1,405 @@
# Central Monitoring System — Complete Setup
## Prometheus + Grafana + Alertmanager + systemd Process Management
Run every command below on your **central Linux server** (native Linux, not WSL).
---
## Prerequisites
- Ubuntu 22.04 or 24.04 LTS (native, always-on machine)
- You have already installed Tailscale and MeshCentral (see `08_central_server_production_setup.md`)
- Your Tailscale IP is noted (run `tailscale ip -4` to get it)
- You are running all commands as your normal user (use `sudo` where shown)
---
## Part 1: Prometheus
### Step 1 — Download and Install
```bash
cd /tmp
PROM_VERSION="2.51.2"
wget https://github.com/prometheus/prometheus/releases/download/v${PROM_VERSION}/prometheus-${PROM_VERSION}.linux-amd64.tar.gz
tar xvf prometheus-${PROM_VERSION}.linux-amd64.tar.gz
cd prometheus-${PROM_VERSION}.linux-amd64
sudo mv prometheus promtool /usr/local/bin/
sudo mkdir -p /etc/prometheus /var/lib/prometheus
sudo mv prometheus.yml consoles console_libraries /etc/prometheus/
```
Verify:
```bash
prometheus --version
# Should print: prometheus, version 2.51.2
```
### Step 2 — Write the Config File
```bash
sudo tee /etc/prometheus/prometheus.yml << 'EOF'
global:
scrape_interval: 15s
evaluation_interval: 15s
alerting:
alertmanagers:
- static_configs:
- targets: ['localhost:9093']
rule_files:
- "/etc/prometheus/rules.yml"
scrape_configs:
- job_name: 'prometheus'
static_configs:
- targets: ['localhost:9090']
# Add each client machine's Tailscale IP here after connecting them
# - job_name: 'node_exporter'
# static_configs:
# - targets:
# - '100.x.x.x:9100' # machine name
# - '100.x.x.x:9100' # machine name
EOF
```
### Step 3 — Write Alert Rules
```bash
sudo tee /etc/prometheus/rules.yml << 'EOF'
groups:
- name: server_alerts
rules:
- alert: ServerDown
expr: up == 0
for: 1m
labels:
severity: critical
annotations:
summary: "Server {{ $labels.instance }} is DOWN"
description: "{{ $labels.instance }} has been unreachable for over 1 minute."
- alert: HighCPU
expr: 100 - (avg by(instance)(rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100) > 85
for: 5m
labels:
severity: warning
annotations:
summary: "High CPU on {{ $labels.instance }}"
description: "CPU usage is above 85% for 5 minutes."
- alert: HighRAM
expr: (1 - (node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes)) * 100 > 90
for: 5m
labels:
severity: warning
annotations:
summary: "High RAM on {{ $labels.instance }}"
description: "RAM usage is above 90% for 5 minutes."
- alert: LowDisk
expr: (node_filesystem_avail_bytes{mountpoint="/"} / node_filesystem_size_bytes{mountpoint="/"}) * 100 < 15
for: 5m
labels:
severity: warning
annotations:
summary: "Low disk on {{ $labels.instance }}"
description: "Root disk has less than 15% free space."
- alert: HighGPUTemp
expr: nvidia_smi_temperature_gpu > 85
for: 5m
labels:
severity: critical
annotations:
summary: "GPU overheating on {{ $labels.instance }}"
description: "GPU temperature is above 85°C."
EOF
```
### Step 4 — Create systemd Service
```bash
sudo tee /etc/systemd/system/prometheus.service << 'EOF'
[Unit]
Description=Prometheus Monitoring
Documentation=https://prometheus.io/docs/
After=network.target
[Service]
User=root
ExecStart=/usr/local/bin/prometheus \
--config.file=/etc/prometheus/prometheus.yml \
--storage.tsdb.path=/var/lib/prometheus \
--storage.tsdb.retention.time=30d \
--storage.tsdb.retention.size=10GB \
--web.listen-address=0.0.0.0:9090 \
--web.enable-lifecycle
Restart=always
RestartSec=5
[Install]
WantedBy=multi-user.target
EOF
sudo systemctl daemon-reload
sudo systemctl enable prometheus
sudo systemctl start prometheus
sudo systemctl status prometheus
```
Expected: `active (running)`
Verify in browser: `http://localhost:9090`
### Step 5 — Adding a New Client to Prometheus
After connecting a client machine via Tailscale and deploying Node Exporter:
```bash
sudo nano /etc/prometheus/prometheus.yml
```
Add the machine under `scrape_configs`:
```yaml
- job_name: "node_exporter"
static_configs:
- targets:
- "100.x.x.x:9100" # machine hostname
```
Reload without restarting (since `--web.enable-lifecycle` is set):
```bash
curl -X POST http://localhost:9090/-/reload
```
Check `http://localhost:9090/targets` — new machine should appear green.
---
## Part 2: Alertmanager
### Step 1 — Download and Install
```bash
cd /tmp
AM_VERSION="0.27.0"
wget https://github.com/prometheus/alertmanager/releases/download/v${AM_VERSION}/alertmanager-${AM_VERSION}.linux-amd64.tar.gz
tar xvf alertmanager-${AM_VERSION}.linux-amd64.tar.gz
cd alertmanager-${AM_VERSION}.linux-amd64
sudo mv alertmanager amtool /usr/local/bin/
sudo mkdir -p /etc/alertmanager /var/lib/alertmanager
```
### Step 2 — Configure Alertmanager
```bash
sudo tee /etc/alertmanager/alertmanager.yml << 'EOF'
global:
resolve_timeout: 5m
smtp_smarthost: "smtp.gmail.com:587"
smtp_from: "no-reply@seekright.com"
smtp_auth_username: "no-reply@seekright.com"
smtp_auth_password: "clzfywewriqxlvqv"
route:
receiver: "fallback-do-nothing"
group_wait: 10s # ← ADD THIS
group_interval: 1m # ← ADD THIS (must be ≤ repeat_interval)
repeat_interval: 1m # now this actually works
routes:
- receiver: "rocket-chat-alerts"
continue: true
- receiver: "email-alerts"
continue: true
receivers:
- name: "fallback-do-nothing"
- name: "rocket-chat-alerts"
webhook_configs: # Changed from slack_configs
- url: "https://text.seekright.com/hooks/6a05a268a88367bc6e0fe318/aNjZzmLy2wB3R4BR7ZtoLtAZYyfNyfLAoQMszmhfBDveKrB>
send_resolved: true
- name: "email-alerts"
EOF
```
Replace:
- `YOUR_GMAIL@gmail.com` — your Gmail address
- `YOUR_GMAIL_APP_PASSWORD` — Gmail App Password (not your login password — create one at myaccount.google.com → Security → App Passwords)
- `YOUR_ROCKETCHAT_WEBHOOK_URL` — your Rocket.Chat incoming webhook URL
### Step 3 — Create systemd Service
```bash
sudo tee /etc/systemd/system/alertmanager.service << 'EOF'
[Unit]
Description=Alertmanager
Documentation=https://prometheus.io/docs/alerting/alertmanager/
After=network.target
[Service]
User=root
ExecStart=/usr/local/bin/alertmanager \
--config.file=/etc/alertmanager/alertmanager.yml \
--storage.path=/var/lib/alertmanager \
--web.listen-address=0.0.0.0:9093
Restart=always
RestartSec=5
[Install]
WantedBy=multi-user.target
EOF
sudo systemctl daemon-reload
sudo systemctl enable alertmanager
sudo systemctl start alertmanager
sudo systemctl status alertmanager
```
Verify in browser: `http://localhost:9093`
---
## Part 3: Grafana
### Step 1 — Install via apt (Official Repo)
```bash
sudo apt-get install -y apt-transport-https software-properties-common wget
sudo mkdir -p /etc/apt/keyrings/
wget -q -O - https://apt.grafana.com/gpg.key | gpg --dearmor | sudo tee /etc/apt/keyrings/grafana.gpg > /dev/null
echo "deb [signed-by=/etc/apt/keyrings/grafana.gpg] https://apt.grafana.com stable main" | sudo tee /etc/apt/sources.list.d/grafana.list
sudo apt-get update
sudo apt-get install -y grafana
sudo systemctl daemon-reload
sudo systemctl enable grafana-server
sudo systemctl start grafana-server
sudo systemctl status grafana-server
```
Verify in browser: `http://localhost:3000`
Default login: `admin` / `admin` — change password immediately on first login.
### Step 2 — Connect Grafana to Prometheus
1. Go to **Connections → Data Sources → Add data source**
2. Choose **Prometheus**
3. Set URL: `http://localhost:9090`
4. Click **Save & Test** — must show "Data source is working"
### Step 3 — Import Dashboards
Go to **Dashboards → Import** and import these IDs one by one:
| Dashboard ID | What It Shows |
| ------------ | ------------------------------------------------------------- |
| `1860` | Node Exporter Full — CPU, RAM, Disk, Network for all machines |
| `14574` | NVIDIA GPU Metrics — GPU temp, VRAM, utilization |
---
## Part 4: Process Management Reference
### Check all service statuses at once
```bash
sudo systemctl status prometheus alertmanager grafana-server
```
### Start / Stop / Restart any service
```bash
sudo systemctl start prometheus
sudo systemctl stop prometheus
sudo systemctl restart prometheus
sudo systemctl start alertmanager
sudo systemctl stop alertmanager
sudo systemctl restart alertmanager
sudo systemctl start grafana-server
sudo systemctl stop grafana-server
sudo systemctl restart grafana-server
```
### Check live logs for any service
```bash
sudo journalctl -u prometheus -f
sudo journalctl -u alertmanager -f
sudo journalctl -u grafana-server -f
```
### Reload Prometheus config without restarting
```bash
curl -X POST http://localhost:9090/-/reload
# Only works because we set --web.enable-lifecycle
```
### Check all listening ports
```bash
ss -tlnp | grep -E "9090|9093|3000|4430"
```
### Check all enabled services (confirms they survive reboot)
```bash
sudo systemctl is-enabled prometheus alertmanager grafana-server
# All three must print: enabled
```
---
## Part 5: Final Verification Checklist
Run this after complete setup to confirm everything is working:
```bash
echo "=== Service Status ===" && \
sudo systemctl is-active prometheus alertmanager grafana-server && \
echo "=== Prometheus Targets ===" && \
curl -s http://localhost:9090/api/v1/targets | python3 -c "import sys,json; [print(t['labels']['job'], t['health']) for t in json.load(sys.stdin)['data']['activeTargets']]" && \
echo "=== Ports Listening ===" && \
ss -tlnp | grep -E "9090|9093|3000"
```
Expected output:
```
=== Service Status ===
active
active
active
=== Prometheus Targets ===
prometheus up
node_exporter up
=== Ports Listening ===
LISTEN *:9090 prometheus
LISTEN *:9093 alertmanager
LISTEN *:3000 grafana
```
---
## Quick Reference
| Service | Port | Config File | Log Command |
| ------------ | ---- | -------------------------------------------- | --------------------------------- |
| Prometheus | 9090 | `/etc/prometheus/prometheus.yml` | `journalctl -u prometheus -f` |
| Alertmanager | 9093 | `/etc/alertmanager/alertmanager.yml` | `journalctl -u alertmanager -f` |
| Grafana | 3000 | `/etc/grafana/grafana.ini` | `journalctl -u grafana-server -f` |
| MeshCentral | 4430 | `~/meshcentral/meshcentral-data/config.json` | `journalctl -u meshcentral -f` |
| Tailscale | — | — | `tailscale status` |