initial commit
This commit is contained in:
629
instructions.md
Normal file
629
instructions.md
Normal file
@@ -0,0 +1,629 @@
|
||||
# PROMETHEUS SETUP INSTRUCTIONS FROM SCRATCH
|
||||
|
||||
_Follow these instructions sequentially to build a production-grade Prometheus and Node Exporter setup from scratch._
|
||||
|
||||
---
|
||||
|
||||
## PART 1: INSTALLING THE PROMETHEUS CENTRAL SERVER
|
||||
|
||||
**1. Download Prometheus**
|
||||
|
||||
```bash
|
||||
wget https://github.com/prometheus/prometheus/releases/download/v2.51.2/prometheus-2.51.2.linux-amd64.tar.gz
|
||||
```
|
||||
|
||||
**2. Extract Prometheus**
|
||||
|
||||
```bash
|
||||
tar -xvf prometheus-2.51.2.linux-amd64.tar.gz
|
||||
```
|
||||
|
||||
**3. Create a restricted system user for security**
|
||||
|
||||
```bash
|
||||
sudo useradd --no-create-home --shell /bin/false prometheus
|
||||
```
|
||||
|
||||
**4. Create the Configuration and Data folders**
|
||||
|
||||
```bash
|
||||
sudo mkdir /etc/prometheus
|
||||
sudo mkdir /var/lib/prometheus
|
||||
```
|
||||
|
||||
**5. Hand ownership of those folders to the new user**
|
||||
|
||||
```bash
|
||||
sudo chown prometheus:prometheus /etc/prometheus
|
||||
sudo chown prometheus:prometheus /var/lib/prometheus
|
||||
```
|
||||
|
||||
**6. Copy the main engine and spell-checker to the secure Linux bin folder**
|
||||
|
||||
```bash
|
||||
sudo cp prometheus-2.51.2.linux-amd64/prometheus /usr/local/bin/
|
||||
sudo cp prometheus-2.51.2.linux-amd64/promtool /usr/local/bin/
|
||||
```
|
||||
|
||||
**7. Hand ownership of the executables to the new user**
|
||||
|
||||
```bash
|
||||
sudo chown prometheus:prometheus /usr/local/bin/prometheus
|
||||
sudo chown prometheus:prometheus /usr/local/bin/promtool
|
||||
```
|
||||
|
||||
**8. Create the main configuration file**
|
||||
|
||||
```bash
|
||||
sudo nano /etc/prometheus/prometheus.yml
|
||||
```
|
||||
|
||||
_Paste this exact code into the file and save:_
|
||||
|
||||
```yaml
|
||||
global:
|
||||
scrape_interval: 15s
|
||||
evaluation_interval: 15s
|
||||
|
||||
scrape_configs:
|
||||
- job_name: "prometheus_self_monitor"
|
||||
static_configs:
|
||||
- targets: ["localhost:9090"]
|
||||
|
||||
- job_name: "laptop_server_1"
|
||||
static_configs:
|
||||
- targets: ["localhost:9100"]
|
||||
```
|
||||
|
||||
_Verify the YAML syntax is perfect so you don't crash the server:_
|
||||
|
||||
```bash
|
||||
promtool check config /etc/prometheus/prometheus.yml
|
||||
```
|
||||
|
||||
**9. Create the background service file**
|
||||
|
||||
```bash
|
||||
sudo nano /etc/systemd/system/prometheus.service
|
||||
```
|
||||
|
||||
_Paste this exact code into the file and save:_
|
||||
|
||||
```ini
|
||||
[Unit]
|
||||
Description=Prometheus
|
||||
Wants=network-online.target
|
||||
After=network-online.target
|
||||
|
||||
[Service]
|
||||
User=prometheus
|
||||
Group=prometheus
|
||||
Type=simple
|
||||
ExecStart=/usr/local/bin/prometheus \
|
||||
--config.file /etc/prometheus/prometheus.yml \
|
||||
--storage.tsdb.path /var/lib/prometheus/
|
||||
|
||||
[Install]
|
||||
WantedBy=multi-user.target
|
||||
```
|
||||
|
||||
**10. Turn Prometheus on forever**
|
||||
|
||||
```bash
|
||||
sudo systemctl daemon-reload
|
||||
sudo systemctl start prometheus
|
||||
sudo systemctl enable prometheus
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## PART 2: INSTALLING NODE EXPORTER (For hardware data)
|
||||
|
||||
**1. Download Node Exporter**
|
||||
|
||||
```bash
|
||||
wget https://github.com/prometheus/node_exporter/releases/download/v1.7.0/node_exporter-1.7.0.linux-amd64.tar.gz
|
||||
```
|
||||
|
||||
**2. Extract Node Exporter**
|
||||
|
||||
```bash
|
||||
tar -xvf node_exporter-1.7.0.linux-amd64.tar.gz
|
||||
```
|
||||
|
||||
**3. Create a restricted system user for security**
|
||||
|
||||
```bash
|
||||
sudo useradd --no-create-home --shell /bin/false node_exporter
|
||||
```
|
||||
|
||||
**4. Copy the main engine to the secure Linux bin folder**
|
||||
|
||||
```bash
|
||||
sudo cp node_exporter-1.7.0.linux-amd64/node_exporter /usr/local/bin/
|
||||
```
|
||||
|
||||
**5. Hand ownership of the executable to the new user**
|
||||
|
||||
```bash
|
||||
sudo chown node_exporter:node_exporter /usr/local/bin/node_exporter
|
||||
```
|
||||
|
||||
**6. Create the background service file**
|
||||
|
||||
```bash
|
||||
sudo nano /etc/systemd/system/node_exporter.service
|
||||
```
|
||||
|
||||
_Paste this exact code into the file and save:_
|
||||
|
||||
```ini
|
||||
[Unit]
|
||||
Description=Node Exporter
|
||||
Wants=network-online.target
|
||||
After=network-online.target
|
||||
|
||||
[Service]
|
||||
User=node_exporter
|
||||
Group=node_exporter
|
||||
Type=simple
|
||||
ExecStart=/usr/local/bin/node_exporter
|
||||
|
||||
[Install]
|
||||
WantedBy=multi-user.target
|
||||
```
|
||||
|
||||
**7. Turn Node Exporter on forever**
|
||||
|
||||
```bash
|
||||
sudo systemctl daemon-reload
|
||||
sudo systemctl start node_exporter
|
||||
sudo systemctl enable node_exporter
|
||||
```
|
||||
|
||||
_(You can verify everything is working by visiting `http://localhost:9090/targets` in your browser. Both targets should say UP)._
|
||||
|
||||
---
|
||||
|
||||
## PART 2.5: INSTALLING NVIDIA GPU EXPORTER (Only for servers with GPUs)
|
||||
|
||||
_Note: The server must already have Nvidia drivers installed so the `nvidia-smi` command works in the terminal._
|
||||
|
||||
**1. Download the Exporter**
|
||||
|
||||
```bash
|
||||
wget https://github.com/utkuozdemir/nvidia_gpu_exporter/releases/download/v1.4.1/nvidia_gpu_exporter_1.4.1_linux_x86_64.tar.gz
|
||||
```
|
||||
|
||||
**2. Extract it**
|
||||
|
||||
```bash
|
||||
tar -xvf nvidia_gpu_exporter_1.4.1_linux_x86_64.tar.gz
|
||||
```
|
||||
|
||||
**3. Create a restricted system user**
|
||||
|
||||
```bash
|
||||
sudo useradd --no-create-home --shell /bin/false nvidia_exporter
|
||||
```
|
||||
|
||||
**4. Copy the engine to the secure bin folder**
|
||||
|
||||
```bash
|
||||
sudo cp nvidia_gpu_exporter /usr/local/bin/
|
||||
```
|
||||
|
||||
**5. Hand ownership to the new user**
|
||||
|
||||
```bash
|
||||
sudo chown nvidia_exporter:nvidia_exporter /usr/local/bin/nvidia_gpu_exporter
|
||||
```
|
||||
|
||||
**6. Create the background service file**
|
||||
|
||||
```bash
|
||||
sudo nano /etc/systemd/system/nvidia_gpu_exporter.service
|
||||
```
|
||||
|
||||
_Paste this exact code into the file and save:_
|
||||
|
||||
```ini
|
||||
[Unit]
|
||||
Description=Nvidia GPU Exporter
|
||||
Wants=network-online.target
|
||||
After=network-online.target
|
||||
|
||||
[Service]
|
||||
User=nvidia_exporter
|
||||
Group=nvidia_exporter
|
||||
Type=simple
|
||||
ExecStart=/usr/local/bin/nvidia_gpu_exporter
|
||||
|
||||
[Install]
|
||||
WantedBy=multi-user.target
|
||||
```
|
||||
|
||||
**7. Turn it on forever**
|
||||
|
||||
```bash
|
||||
sudo systemctl daemon-reload
|
||||
sudo systemctl start nvidia_gpu_exporter
|
||||
sudo systemctl enable nvidia_gpu_exporter
|
||||
```
|
||||
|
||||
_(Once running, the GPU Exporter listens on port `9835`. You must go back to the Central Server and add a new job in `prometheus.yml` targeting `[SERVER_IP]:9835`)._
|
||||
|
||||
---
|
||||
|
||||
## PART 3: ESSENTIAL PROMQL QUERIES
|
||||
|
||||
_Use these queries in the Prometheus search bar or Grafana to monitor hardware._
|
||||
|
||||
**1. RAM: Percentage Available**
|
||||
|
||||
```promql
|
||||
(node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes) * 100
|
||||
```
|
||||
|
||||
**2. CPU: Percentage Actively Used**
|
||||
_(Node Exporter tracks how long the CPU is "idle". We subtract the idle speed from 100% to get the active usage)._
|
||||
|
||||
```promql
|
||||
100 - (avg by (instance) (rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100)
|
||||
```
|
||||
|
||||
**3. DISK (SSD/HDD): Percentage Free Space (Root Drive)**
|
||||
|
||||
```promql
|
||||
(node_filesystem_avail_bytes{mountpoint="/"} / node_filesystem_size_bytes{mountpoint="/"}) * 100
|
||||
```
|
||||
|
||||
**4. NETWORK: Total Upload Speed in Megabytes/second (MB/s)**
|
||||
|
||||
```promql
|
||||
rate(node_network_transmit_bytes_total[5m]) / 1024 / 1024
|
||||
```
|
||||
|
||||
**5. HARDWARE: Server Uptime in Days**
|
||||
|
||||
```promql
|
||||
(time() - node_boot_time_seconds) / 86400
|
||||
```
|
||||
|
||||
**6. SERVERS: Is a server down? (1 = UP, 0 = DOWN)**
|
||||
|
||||
```promql
|
||||
up
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## PART 4: INSTALLING GRAFANA (THE DASHBOARD UI)
|
||||
|
||||
_Grafana runs on port 3000 by default. Use this to visualize Prometheus data._
|
||||
|
||||
**1. Download the Grafana security key:**
|
||||
|
||||
```bash
|
||||
sudo apt-get install -y apt-transport-https software-properties-common wget
|
||||
sudo mkdir -p /etc/apt/keyrings/
|
||||
wget -q -O - https://apt.grafana.com/gpg.key | gpg --dearmor | sudo tee /etc/apt/keyrings/grafana.gpg > /dev/null
|
||||
```
|
||||
|
||||
**2. Add the official Grafana repository to Linux:**
|
||||
|
||||
```bash
|
||||
echo "deb [signed-by=/etc/apt/keyrings/grafana.gpg] https://apt.grafana.com stable main" | sudo tee -a /etc/apt/sources.list.d/grafana.list
|
||||
```
|
||||
|
||||
**3. Install Grafana and start the service:**
|
||||
|
||||
```bash
|
||||
sudo apt-get update
|
||||
sudo apt-get install grafana
|
||||
sudo systemctl daemon-reload
|
||||
sudo systemctl start grafana-server
|
||||
sudo systemctl enable grafana-server
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## PART 5: GRAFANA DASHBOARD TEMPLATE IDs & DEBUGGING
|
||||
|
||||
When importing templates from `grafana.com/dashboards`, use these exact IDs:
|
||||
|
||||
- **Node Exporter Full (CPU/RAM/Network):** `1860`
|
||||
- **Nvidia GPU Exporter (Temperature/VRAM):** `14574`
|
||||
|
||||
### How to Fix "Blank" or "No Data" Dashboards
|
||||
|
||||
1. **Variables Mismatch:** Look at the dropdown menus at the very top of the dashboard. If they say "N/A", click them and select your specific `job_name` (e.g. `sasi_gpu` instead of `nvidia_gpu_exporter`).
|
||||
2. **PromQL Mismatch:** If the graph is still blank, click the `...` next to the graph title and hit **Edit**. Look at the raw PromQL code. Sometimes community dashboards hardcode filters like `{job="gpu"}`. Simply delete that strict filter so your specific data can flow through!
|
||||
|
||||
---
|
||||
|
||||
### 1. Browser Blocks HTTP Exporters
|
||||
|
||||
If your exporter is running (e.g., `curl http://localhost:9100/metrics` works in the terminal) but the browser says "Connection Refused", make sure your browser didn't automatically force `https://`. Exporters only run on plain `http://` by default!
|
||||
|
||||
### 2. Nvidia GPU Exporter Panic `[ms]` Error
|
||||
|
||||
If you use an outdated version of `nvidia_gpu_exporter` (like `v1.2.1`), the engine will instantly crash with a `panic: descriptor Desc{...}` error because newer Nvidia drivers output a unit called `[ms]`. Prometheus strictly forbids spaces and brackets in metric names!
|
||||
**The Fix:** Always use `v1.4.1` or newer, as the developer hardcoded a fix for this exact regex bug.
|
||||
|
||||
_(Note: Node Exporter does NOT monitor GPUs. To monitor GPU temperatures and VRAM, you must install a separate exporter called `nvidia_smi_exporter`)._
|
||||
|
||||
---
|
||||
|
||||
## PART 6: INSTALLING ALERTMANAGER (For Slack/Email Notifications)
|
||||
|
||||
**1. Download and Extract Alertmanager**
|
||||
|
||||
```bash
|
||||
wget https://github.com/prometheus/alertmanager/releases/download/v0.27.0/alertmanager-0.27.0.linux-amd64.tar.gz
|
||||
tar -xvf alertmanager-0.27.0.linux-amd64.tar.gz
|
||||
```
|
||||
|
||||
**2. Create a restricted user and folders**
|
||||
|
||||
```bash
|
||||
sudo useradd --no-create-home --shell /bin/false alertmanager
|
||||
sudo mkdir /etc/alertmanager
|
||||
sudo mkdir /var/lib/alertmanager
|
||||
sudo chown alertmanager:alertmanager /var/lib/alertmanager
|
||||
```
|
||||
|
||||
**3. Move the binary and hand over ownership**
|
||||
|
||||
```bash
|
||||
sudo cp alertmanager-0.27.0.linux-amd64/alertmanager /usr/local/bin/
|
||||
sudo chown alertmanager:alertmanager /usr/local/bin/alertmanager
|
||||
```
|
||||
|
||||
**4. Create the Configuration File**
|
||||
|
||||
```bash
|
||||
sudo nano /etc/alertmanager/alertmanager.yml
|
||||
```
|
||||
|
||||
_Paste this code to route alerts to both Rocket.Chat and Gmail simultaneously:_
|
||||
|
||||
```yaml
|
||||
global:
|
||||
resolve_timeout: 5m
|
||||
smtp_smarthost: "smtp.gmail.com:587"
|
||||
smtp_from: "knmkaushik@gmail.com"
|
||||
smtp_auth_username: "knmkaushik@gmail.com"
|
||||
smtp_auth_password: "YOUR_APP_PASSWORD"
|
||||
|
||||
route:
|
||||
receiver: "fallback-do-nothing"
|
||||
group_wait: 10s
|
||||
group_interval: 1m
|
||||
repeat_interval: 5m
|
||||
routes:
|
||||
- receiver: "rocket-chat-alerts"
|
||||
continue: true
|
||||
- receiver: "email-alerts"
|
||||
continue: true
|
||||
|
||||
receivers:
|
||||
- name: "fallback-do-nothing"
|
||||
|
||||
- name: "rocket-chat-alerts"
|
||||
webhook_configs:
|
||||
- url: "YOUR_ROCKETCHAT_WEBHOOK_URL"
|
||||
send_resolved: true
|
||||
|
||||
- name: "email-alerts"
|
||||
email_configs:
|
||||
- to: "kawsik97@gmail.com"
|
||||
send_resolved: true
|
||||
```
|
||||
|
||||
**5. Hand over ownership of the config file**
|
||||
|
||||
```bash
|
||||
sudo chown alertmanager:alertmanager /etc/alertmanager/alertmanager.yml
|
||||
```
|
||||
|
||||
**6. Create the background service file**
|
||||
|
||||
```bash
|
||||
sudo nano /etc/systemd/system/alertmanager.service
|
||||
```
|
||||
|
||||
_Paste this exact code into the file and save:_
|
||||
|
||||
```ini
|
||||
[Unit]
|
||||
Description=Alertmanager
|
||||
Wants=network-online.target
|
||||
After=network-online.target
|
||||
|
||||
[Service]
|
||||
User=alertmanager
|
||||
Group=alertmanager
|
||||
Type=simple
|
||||
ExecStart=/usr/local/bin/alertmanager --config.file=/etc/alertmanager/alertmanager.yml --storage.path=/var/lib/alertmanager
|
||||
|
||||
[Install]
|
||||
WantedBy=multi-user.target
|
||||
```
|
||||
|
||||
**7. Turn Alertmanager on forever**
|
||||
|
||||
```bash
|
||||
sudo systemctl daemon-reload
|
||||
sudo systemctl start alertmanager
|
||||
sudo systemctl enable alertmanager
|
||||
```
|
||||
|
||||
**8. Configure Rocket.Chat Webhook Script**
|
||||
|
||||
Rocket.Chat natively only understands `{"text": "..."}` payloads, but Alertmanager sends structured JSON. To fix this, you must add a transformation script in the Rocket.Chat admin panel:
|
||||
|
||||
1. Go to your Rocket.Chat Admin Panel -> **Integrations** -> **Incoming**.
|
||||
2. Create or Edit the incoming webhook for your channel (e.g., `#devops-alerts`).
|
||||
3. Toggle **Script Enabled** to **ON**.
|
||||
4. Paste the following into the Script box:
|
||||
|
||||
```javascript
|
||||
class Script {
|
||||
process_incoming_request({ request }) {
|
||||
var content = request.content;
|
||||
var status = content.status ? content.status.toUpperCase() : "UNKNOWN";
|
||||
var alertName = "Unknown";
|
||||
var instance = "";
|
||||
var summary = "";
|
||||
|
||||
if (content.alerts && content.alerts[0]) {
|
||||
var a = content.alerts[0];
|
||||
if (a.labels) {
|
||||
alertName = a.labels.alertname || "Unknown";
|
||||
instance = a.labels.instance || "";
|
||||
}
|
||||
if (a.annotations) {
|
||||
summary = a.annotations.summary || "";
|
||||
}
|
||||
}
|
||||
|
||||
var icon = status === "FIRING" ? ":red_circle:" : ":white_check_mark:";
|
||||
var msg = icon + " *" + status + "* — " + alertName;
|
||||
if (instance) {
|
||||
msg = msg + " (" + instance + ")";
|
||||
}
|
||||
if (summary) {
|
||||
msg = msg + "\n" + summary;
|
||||
}
|
||||
|
||||
return {
|
||||
content: {
|
||||
text: msg,
|
||||
},
|
||||
};
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
5. Make sure the "Script Sandbox" is set to "Secure Sandbox".
|
||||
6. Click **Save** and use the generated Webhook URL in your `alertmanager.yml`.
|
||||
|
||||
---
|
||||
|
||||
## PART 7: LINKING PROMETHEUS & ALERTMANAGER
|
||||
|
||||
Prometheus must be told that Alertmanager exists, and where the "Rules" are stored.
|
||||
|
||||
**1. Open the main Prometheus config**
|
||||
|
||||
```bash
|
||||
sudo nano /etc/prometheus/prometheus.yml
|
||||
```
|
||||
|
||||
**2. Update the `alerting` and `rule_files` blocks at the top:**
|
||||
|
||||
```yaml
|
||||
alerting:
|
||||
alertmanagers:
|
||||
- static_configs:
|
||||
- targets: ["localhost:9093"]
|
||||
|
||||
rule_files:
|
||||
- "rules.yml"
|
||||
```
|
||||
|
||||
_(Save and exit)._
|
||||
|
||||
**3. Create the Rules file (Where the Alert Logic lives):**
|
||||
|
||||
```bash
|
||||
sudo nano /etc/prometheus/rules.yml
|
||||
```
|
||||
|
||||
**4. Paste this exact rule and save:**
|
||||
_(This logic says: If any server goes offline for more than 1 minute, fire a CRITICAL alert)._
|
||||
|
||||
```yaml
|
||||
groups:
|
||||
- name: hardware_alerts
|
||||
rules:
|
||||
- alert: ServerDown
|
||||
expr: up == 0
|
||||
for: 10s
|
||||
labels:
|
||||
severity: critical
|
||||
annotations:
|
||||
summary: "Server {{ $labels.instance }} is {{ if eq $value 0.0 }}DOWN{{ else }}UP{{ end }}!"
|
||||
description: "Server {{ $labels.instance }} has been unreachable for more than 10 seconds."
|
||||
```
|
||||
|
||||
**5. Hand over ownership and restart Prometheus:**
|
||||
|
||||
```bash
|
||||
sudo chown prometheus:prometheus /etc/prometheus/rules.yml
|
||||
sudo systemctl restart prometheus
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## PART 8: MANAGING & DEBUGGING SERVICES (Crash, Reboot, & Maintenance)
|
||||
|
||||
Since we configured all components (Prometheus, Node Exporter, Grafana, Alertmanager) as Linux `systemd` services and used the `enable` command, **they will start automatically every time you boot or reboot your server.**
|
||||
|
||||
If you ever experience a crash, need to reboot, or want to test if services are running, use the following commands.
|
||||
|
||||
### 1. Check the Status of All Services at Once
|
||||
|
||||
If you just rebooted your server and want to ensure everything turned on properly, run this:
|
||||
|
||||
```bash
|
||||
sudo systemctl status prometheus node_exporter grafana-server alertmanager
|
||||
```
|
||||
|
||||
_Tip: If the output locks your terminal and shows `(END)` at the bottom, just press the `q` key on your keyboard to quit the reader mode and return to your prompt._
|
||||
|
||||
### 2. Restarting a Service After a Crash or Configuration Change
|
||||
|
||||
If you edit a configuration file (like `prometheus.yml` or `rules.yml`), the changes won't apply until you restart the service. If a service ever crashes and you want to force it to start fresh, use `restart`:
|
||||
|
||||
```bash
|
||||
# Restart Prometheus
|
||||
sudo systemctl restart prometheus
|
||||
|
||||
# Restart Alertmanager
|
||||
sudo systemctl restart alertmanager
|
||||
|
||||
# Restart Grafana
|
||||
sudo systemctl restart grafana-server
|
||||
```
|
||||
|
||||
### 3. Turning Services On / Off Manually
|
||||
|
||||
If you need to intentionally stop a service to perform maintenance:
|
||||
|
||||
```bash
|
||||
sudo systemctl stop prometheus
|
||||
```
|
||||
|
||||
To start it back up:
|
||||
|
||||
```bash
|
||||
sudo systemctl start prometheus
|
||||
```
|
||||
|
||||
### 4. Viewing Error Logs (Debugging)
|
||||
|
||||
If a service says it is "failed" or "inactive" when you check its status, you can view its exact crash logs by checking the Linux `journalctl` (the master log book).
|
||||
|
||||
To see the last 50 logs for Prometheus (useful for finding YAML syntax errors):
|
||||
|
||||
```bash
|
||||
sudo journalctl -u prometheus -n 50 --no-pager
|
||||
```
|
||||
|
||||
To stream the logs live in real-time (press `Ctrl+C` to stop):
|
||||
|
||||
```bash
|
||||
sudo journalctl -u prometheus -f
|
||||
```
|
||||
Reference in New Issue
Block a user