Files
RMM-UI-BE/instructions.md
2026-06-02 10:37:57 +05:30

630 lines
16 KiB
Markdown

# PROMETHEUS SETUP INSTRUCTIONS FROM SCRATCH
_Follow these instructions sequentially to build a production-grade Prometheus and Node Exporter setup from scratch._
---
## PART 1: INSTALLING THE PROMETHEUS CENTRAL SERVER
**1. Download Prometheus**
```bash
wget https://github.com/prometheus/prometheus/releases/download/v2.51.2/prometheus-2.51.2.linux-amd64.tar.gz
```
**2. Extract Prometheus**
```bash
tar -xvf prometheus-2.51.2.linux-amd64.tar.gz
```
**3. Create a restricted system user for security**
```bash
sudo useradd --no-create-home --shell /bin/false prometheus
```
**4. Create the Configuration and Data folders**
```bash
sudo mkdir /etc/prometheus
sudo mkdir /var/lib/prometheus
```
**5. Hand ownership of those folders to the new user**
```bash
sudo chown prometheus:prometheus /etc/prometheus
sudo chown prometheus:prometheus /var/lib/prometheus
```
**6. Copy the main engine and spell-checker to the secure Linux bin folder**
```bash
sudo cp prometheus-2.51.2.linux-amd64/prometheus /usr/local/bin/
sudo cp prometheus-2.51.2.linux-amd64/promtool /usr/local/bin/
```
**7. Hand ownership of the executables to the new user**
```bash
sudo chown prometheus:prometheus /usr/local/bin/prometheus
sudo chown prometheus:prometheus /usr/local/bin/promtool
```
**8. Create the main configuration file**
```bash
sudo nano /etc/prometheus/prometheus.yml
```
_Paste this exact code into the file and save:_
```yaml
global:
scrape_interval: 15s
evaluation_interval: 15s
scrape_configs:
- job_name: "prometheus_self_monitor"
static_configs:
- targets: ["localhost:9090"]
- job_name: "laptop_server_1"
static_configs:
- targets: ["localhost:9100"]
```
_Verify the YAML syntax is perfect so you don't crash the server:_
```bash
promtool check config /etc/prometheus/prometheus.yml
```
**9. Create the background service file**
```bash
sudo nano /etc/systemd/system/prometheus.service
```
_Paste this exact code into the file and save:_
```ini
[Unit]
Description=Prometheus
Wants=network-online.target
After=network-online.target
[Service]
User=prometheus
Group=prometheus
Type=simple
ExecStart=/usr/local/bin/prometheus \
--config.file /etc/prometheus/prometheus.yml \
--storage.tsdb.path /var/lib/prometheus/
[Install]
WantedBy=multi-user.target
```
**10. Turn Prometheus on forever**
```bash
sudo systemctl daemon-reload
sudo systemctl start prometheus
sudo systemctl enable prometheus
```
---
## PART 2: INSTALLING NODE EXPORTER (For hardware data)
**1. Download Node Exporter**
```bash
wget https://github.com/prometheus/node_exporter/releases/download/v1.7.0/node_exporter-1.7.0.linux-amd64.tar.gz
```
**2. Extract Node Exporter**
```bash
tar -xvf node_exporter-1.7.0.linux-amd64.tar.gz
```
**3. Create a restricted system user for security**
```bash
sudo useradd --no-create-home --shell /bin/false node_exporter
```
**4. Copy the main engine to the secure Linux bin folder**
```bash
sudo cp node_exporter-1.7.0.linux-amd64/node_exporter /usr/local/bin/
```
**5. Hand ownership of the executable to the new user**
```bash
sudo chown node_exporter:node_exporter /usr/local/bin/node_exporter
```
**6. Create the background service file**
```bash
sudo nano /etc/systemd/system/node_exporter.service
```
_Paste this exact code into the file and save:_
```ini
[Unit]
Description=Node Exporter
Wants=network-online.target
After=network-online.target
[Service]
User=node_exporter
Group=node_exporter
Type=simple
ExecStart=/usr/local/bin/node_exporter
[Install]
WantedBy=multi-user.target
```
**7. Turn Node Exporter on forever**
```bash
sudo systemctl daemon-reload
sudo systemctl start node_exporter
sudo systemctl enable node_exporter
```
_(You can verify everything is working by visiting `http://localhost:9090/targets` in your browser. Both targets should say UP)._
---
## PART 2.5: INSTALLING NVIDIA GPU EXPORTER (Only for servers with GPUs)
_Note: The server must already have Nvidia drivers installed so the `nvidia-smi` command works in the terminal._
**1. Download the Exporter**
```bash
wget https://github.com/utkuozdemir/nvidia_gpu_exporter/releases/download/v1.4.1/nvidia_gpu_exporter_1.4.1_linux_x86_64.tar.gz
```
**2. Extract it**
```bash
tar -xvf nvidia_gpu_exporter_1.4.1_linux_x86_64.tar.gz
```
**3. Create a restricted system user**
```bash
sudo useradd --no-create-home --shell /bin/false nvidia_exporter
```
**4. Copy the engine to the secure bin folder**
```bash
sudo cp nvidia_gpu_exporter /usr/local/bin/
```
**5. Hand ownership to the new user**
```bash
sudo chown nvidia_exporter:nvidia_exporter /usr/local/bin/nvidia_gpu_exporter
```
**6. Create the background service file**
```bash
sudo nano /etc/systemd/system/nvidia_gpu_exporter.service
```
_Paste this exact code into the file and save:_
```ini
[Unit]
Description=Nvidia GPU Exporter
Wants=network-online.target
After=network-online.target
[Service]
User=nvidia_exporter
Group=nvidia_exporter
Type=simple
ExecStart=/usr/local/bin/nvidia_gpu_exporter
[Install]
WantedBy=multi-user.target
```
**7. Turn it on forever**
```bash
sudo systemctl daemon-reload
sudo systemctl start nvidia_gpu_exporter
sudo systemctl enable nvidia_gpu_exporter
```
_(Once running, the GPU Exporter listens on port `9835`. You must go back to the Central Server and add a new job in `prometheus.yml` targeting `[SERVER_IP]:9835`)._
---
## PART 3: ESSENTIAL PROMQL QUERIES
_Use these queries in the Prometheus search bar or Grafana to monitor hardware._
**1. RAM: Percentage Available**
```promql
(node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes) * 100
```
**2. CPU: Percentage Actively Used**
_(Node Exporter tracks how long the CPU is "idle". We subtract the idle speed from 100% to get the active usage)._
```promql
100 - (avg by (instance) (rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100)
```
**3. DISK (SSD/HDD): Percentage Free Space (Root Drive)**
```promql
(node_filesystem_avail_bytes{mountpoint="/"} / node_filesystem_size_bytes{mountpoint="/"}) * 100
```
**4. NETWORK: Total Upload Speed in Megabytes/second (MB/s)**
```promql
rate(node_network_transmit_bytes_total[5m]) / 1024 / 1024
```
**5. HARDWARE: Server Uptime in Days**
```promql
(time() - node_boot_time_seconds) / 86400
```
**6. SERVERS: Is a server down? (1 = UP, 0 = DOWN)**
```promql
up
```
---
## PART 4: INSTALLING GRAFANA (THE DASHBOARD UI)
_Grafana runs on port 3000 by default. Use this to visualize Prometheus data._
**1. Download the Grafana security key:**
```bash
sudo apt-get install -y apt-transport-https software-properties-common wget
sudo mkdir -p /etc/apt/keyrings/
wget -q -O - https://apt.grafana.com/gpg.key | gpg --dearmor | sudo tee /etc/apt/keyrings/grafana.gpg > /dev/null
```
**2. Add the official Grafana repository to Linux:**
```bash
echo "deb [signed-by=/etc/apt/keyrings/grafana.gpg] https://apt.grafana.com stable main" | sudo tee -a /etc/apt/sources.list.d/grafana.list
```
**3. Install Grafana and start the service:**
```bash
sudo apt-get update
sudo apt-get install grafana
sudo systemctl daemon-reload
sudo systemctl start grafana-server
sudo systemctl enable grafana-server
```
---
## PART 5: GRAFANA DASHBOARD TEMPLATE IDs & DEBUGGING
When importing templates from `grafana.com/dashboards`, use these exact IDs:
- **Node Exporter Full (CPU/RAM/Network):** `1860`
- **Nvidia GPU Exporter (Temperature/VRAM):** `14574`
### How to Fix "Blank" or "No Data" Dashboards
1. **Variables Mismatch:** Look at the dropdown menus at the very top of the dashboard. If they say "N/A", click them and select your specific `job_name` (e.g. `sasi_gpu` instead of `nvidia_gpu_exporter`).
2. **PromQL Mismatch:** If the graph is still blank, click the `...` next to the graph title and hit **Edit**. Look at the raw PromQL code. Sometimes community dashboards hardcode filters like `{job="gpu"}`. Simply delete that strict filter so your specific data can flow through!
---
### 1. Browser Blocks HTTP Exporters
If your exporter is running (e.g., `curl http://localhost:9100/metrics` works in the terminal) but the browser says "Connection Refused", make sure your browser didn't automatically force `https://`. Exporters only run on plain `http://` by default!
### 2. Nvidia GPU Exporter Panic `[ms]` Error
If you use an outdated version of `nvidia_gpu_exporter` (like `v1.2.1`), the engine will instantly crash with a `panic: descriptor Desc{...}` error because newer Nvidia drivers output a unit called `[ms]`. Prometheus strictly forbids spaces and brackets in metric names!
**The Fix:** Always use `v1.4.1` or newer, as the developer hardcoded a fix for this exact regex bug.
_(Note: Node Exporter does NOT monitor GPUs. To monitor GPU temperatures and VRAM, you must install a separate exporter called `nvidia_smi_exporter`)._
---
## PART 6: INSTALLING ALERTMANAGER (For Slack/Email Notifications)
**1. Download and Extract Alertmanager**
```bash
wget https://github.com/prometheus/alertmanager/releases/download/v0.27.0/alertmanager-0.27.0.linux-amd64.tar.gz
tar -xvf alertmanager-0.27.0.linux-amd64.tar.gz
```
**2. Create a restricted user and folders**
```bash
sudo useradd --no-create-home --shell /bin/false alertmanager
sudo mkdir /etc/alertmanager
sudo mkdir /var/lib/alertmanager
sudo chown alertmanager:alertmanager /var/lib/alertmanager
```
**3. Move the binary and hand over ownership**
```bash
sudo cp alertmanager-0.27.0.linux-amd64/alertmanager /usr/local/bin/
sudo chown alertmanager:alertmanager /usr/local/bin/alertmanager
```
**4. Create the Configuration File**
```bash
sudo nano /etc/alertmanager/alertmanager.yml
```
_Paste this code to route alerts to both Rocket.Chat and Gmail simultaneously:_
```yaml
global:
resolve_timeout: 5m
smtp_smarthost: "smtp.gmail.com:587"
smtp_from: "knmkaushik@gmail.com"
smtp_auth_username: "knmkaushik@gmail.com"
smtp_auth_password: "YOUR_APP_PASSWORD"
route:
receiver: "fallback-do-nothing"
group_wait: 10s
group_interval: 1m
repeat_interval: 5m
routes:
- receiver: "rocket-chat-alerts"
continue: true
- receiver: "email-alerts"
continue: true
receivers:
- name: "fallback-do-nothing"
- name: "rocket-chat-alerts"
webhook_configs:
- url: "YOUR_ROCKETCHAT_WEBHOOK_URL"
send_resolved: true
- name: "email-alerts"
email_configs:
- to: "kawsik97@gmail.com"
send_resolved: true
```
**5. Hand over ownership of the config file**
```bash
sudo chown alertmanager:alertmanager /etc/alertmanager/alertmanager.yml
```
**6. Create the background service file**
```bash
sudo nano /etc/systemd/system/alertmanager.service
```
_Paste this exact code into the file and save:_
```ini
[Unit]
Description=Alertmanager
Wants=network-online.target
After=network-online.target
[Service]
User=alertmanager
Group=alertmanager
Type=simple
ExecStart=/usr/local/bin/alertmanager --config.file=/etc/alertmanager/alertmanager.yml --storage.path=/var/lib/alertmanager
[Install]
WantedBy=multi-user.target
```
**7. Turn Alertmanager on forever**
```bash
sudo systemctl daemon-reload
sudo systemctl start alertmanager
sudo systemctl enable alertmanager
```
**8. Configure Rocket.Chat Webhook Script**
Rocket.Chat natively only understands `{"text": "..."}` payloads, but Alertmanager sends structured JSON. To fix this, you must add a transformation script in the Rocket.Chat admin panel:
1. Go to your Rocket.Chat Admin Panel -> **Integrations** -> **Incoming**.
2. Create or Edit the incoming webhook for your channel (e.g., `#devops-alerts`).
3. Toggle **Script Enabled** to **ON**.
4. Paste the following into the Script box:
```javascript
class Script {
process_incoming_request({ request }) {
var content = request.content;
var status = content.status ? content.status.toUpperCase() : "UNKNOWN";
var alertName = "Unknown";
var instance = "";
var summary = "";
if (content.alerts && content.alerts[0]) {
var a = content.alerts[0];
if (a.labels) {
alertName = a.labels.alertname || "Unknown";
instance = a.labels.instance || "";
}
if (a.annotations) {
summary = a.annotations.summary || "";
}
}
var icon = status === "FIRING" ? ":red_circle:" : ":white_check_mark:";
var msg = icon + " *" + status + "* — " + alertName;
if (instance) {
msg = msg + " (" + instance + ")";
}
if (summary) {
msg = msg + "\n" + summary;
}
return {
content: {
text: msg,
},
};
}
}
```
5. Make sure the "Script Sandbox" is set to "Secure Sandbox".
6. Click **Save** and use the generated Webhook URL in your `alertmanager.yml`.
---
## PART 7: LINKING PROMETHEUS & ALERTMANAGER
Prometheus must be told that Alertmanager exists, and where the "Rules" are stored.
**1. Open the main Prometheus config**
```bash
sudo nano /etc/prometheus/prometheus.yml
```
**2. Update the `alerting` and `rule_files` blocks at the top:**
```yaml
alerting:
alertmanagers:
- static_configs:
- targets: ["localhost:9093"]
rule_files:
- "rules.yml"
```
_(Save and exit)._
**3. Create the Rules file (Where the Alert Logic lives):**
```bash
sudo nano /etc/prometheus/rules.yml
```
**4. Paste this exact rule and save:**
_(This logic says: If any server goes offline for more than 1 minute, fire a CRITICAL alert)._
```yaml
groups:
- name: hardware_alerts
rules:
- alert: ServerDown
expr: up == 0
for: 10s
labels:
severity: critical
annotations:
summary: "Server {{ $labels.instance }} is {{ if eq $value 0.0 }}DOWN{{ else }}UP{{ end }}!"
description: "Server {{ $labels.instance }} has been unreachable for more than 10 seconds."
```
**5. Hand over ownership and restart Prometheus:**
```bash
sudo chown prometheus:prometheus /etc/prometheus/rules.yml
sudo systemctl restart prometheus
```
---
## PART 8: MANAGING & DEBUGGING SERVICES (Crash, Reboot, & Maintenance)
Since we configured all components (Prometheus, Node Exporter, Grafana, Alertmanager) as Linux `systemd` services and used the `enable` command, **they will start automatically every time you boot or reboot your server.**
If you ever experience a crash, need to reboot, or want to test if services are running, use the following commands.
### 1. Check the Status of All Services at Once
If you just rebooted your server and want to ensure everything turned on properly, run this:
```bash
sudo systemctl status prometheus node_exporter grafana-server alertmanager
```
_Tip: If the output locks your terminal and shows `(END)` at the bottom, just press the `q` key on your keyboard to quit the reader mode and return to your prompt._
### 2. Restarting a Service After a Crash or Configuration Change
If you edit a configuration file (like `prometheus.yml` or `rules.yml`), the changes won't apply until you restart the service. If a service ever crashes and you want to force it to start fresh, use `restart`:
```bash
# Restart Prometheus
sudo systemctl restart prometheus
# Restart Alertmanager
sudo systemctl restart alertmanager
# Restart Grafana
sudo systemctl restart grafana-server
```
### 3. Turning Services On / Off Manually
If you need to intentionally stop a service to perform maintenance:
```bash
sudo systemctl stop prometheus
```
To start it back up:
```bash
sudo systemctl start prometheus
```
### 4. Viewing Error Logs (Debugging)
If a service says it is "failed" or "inactive" when you check its status, you can view its exact crash logs by checking the Linux `journalctl` (the master log book).
To see the last 50 logs for Prometheus (useful for finding YAML syntax errors):
```bash
sudo journalctl -u prometheus -n 50 --no-pager
```
To stream the logs live in real-time (press `Ctrl+C` to stop):
```bash
sudo journalctl -u prometheus -f
```