# PROMETHEUS SETUP INSTRUCTIONS FROM SCRATCH _Follow these instructions sequentially to build a production-grade Prometheus and Node Exporter setup from scratch._ --- ## PART 1: INSTALLING THE PROMETHEUS CENTRAL SERVER **1. Download Prometheus** ```bash wget https://github.com/prometheus/prometheus/releases/download/v2.51.2/prometheus-2.51.2.linux-amd64.tar.gz ``` **2. Extract Prometheus** ```bash tar -xvf prometheus-2.51.2.linux-amd64.tar.gz ``` **3. Create a restricted system user for security** ```bash sudo useradd --no-create-home --shell /bin/false prometheus ``` **4. Create the Configuration and Data folders** ```bash sudo mkdir /etc/prometheus sudo mkdir /var/lib/prometheus ``` **5. Hand ownership of those folders to the new user** ```bash sudo chown prometheus:prometheus /etc/prometheus sudo chown prometheus:prometheus /var/lib/prometheus ``` **6. Copy the main engine and spell-checker to the secure Linux bin folder** ```bash sudo cp prometheus-2.51.2.linux-amd64/prometheus /usr/local/bin/ sudo cp prometheus-2.51.2.linux-amd64/promtool /usr/local/bin/ ``` **7. Hand ownership of the executables to the new user** ```bash sudo chown prometheus:prometheus /usr/local/bin/prometheus sudo chown prometheus:prometheus /usr/local/bin/promtool ``` **8. Create the main configuration file** ```bash sudo nano /etc/prometheus/prometheus.yml ``` _Paste this exact code into the file and save:_ ```yaml global: scrape_interval: 15s evaluation_interval: 15s scrape_configs: - job_name: "prometheus_self_monitor" static_configs: - targets: ["localhost:9090"] - job_name: "laptop_server_1" static_configs: - targets: ["localhost:9100"] ``` _Verify the YAML syntax is perfect so you don't crash the server:_ ```bash promtool check config /etc/prometheus/prometheus.yml ``` **9. Create the background service file** ```bash sudo nano /etc/systemd/system/prometheus.service ``` _Paste this exact code into the file and save:_ ```ini [Unit] Description=Prometheus Wants=network-online.target After=network-online.target [Service] User=prometheus Group=prometheus Type=simple ExecStart=/usr/local/bin/prometheus \ --config.file /etc/prometheus/prometheus.yml \ --storage.tsdb.path /var/lib/prometheus/ [Install] WantedBy=multi-user.target ``` **10. Turn Prometheus on forever** ```bash sudo systemctl daemon-reload sudo systemctl start prometheus sudo systemctl enable prometheus ``` --- ## PART 2: INSTALLING NODE EXPORTER (For hardware data) **1. Download Node Exporter** ```bash wget https://github.com/prometheus/node_exporter/releases/download/v1.7.0/node_exporter-1.7.0.linux-amd64.tar.gz ``` **2. Extract Node Exporter** ```bash tar -xvf node_exporter-1.7.0.linux-amd64.tar.gz ``` **3. Create a restricted system user for security** ```bash sudo useradd --no-create-home --shell /bin/false node_exporter ``` **4. Copy the main engine to the secure Linux bin folder** ```bash sudo cp node_exporter-1.7.0.linux-amd64/node_exporter /usr/local/bin/ ``` **5. Hand ownership of the executable to the new user** ```bash sudo chown node_exporter:node_exporter /usr/local/bin/node_exporter ``` **6. Create the background service file** ```bash sudo nano /etc/systemd/system/node_exporter.service ``` _Paste this exact code into the file and save:_ ```ini [Unit] Description=Node Exporter Wants=network-online.target After=network-online.target [Service] User=node_exporter Group=node_exporter Type=simple ExecStart=/usr/local/bin/node_exporter [Install] WantedBy=multi-user.target ``` **7. Turn Node Exporter on forever** ```bash sudo systemctl daemon-reload sudo systemctl start node_exporter sudo systemctl enable node_exporter ``` _(You can verify everything is working by visiting `http://localhost:9090/targets` in your browser. Both targets should say UP)._ --- ## PART 2.5: INSTALLING NVIDIA GPU EXPORTER (Only for servers with GPUs) _Note: The server must already have Nvidia drivers installed so the `nvidia-smi` command works in the terminal._ **1. Download the Exporter** ```bash wget https://github.com/utkuozdemir/nvidia_gpu_exporter/releases/download/v1.4.1/nvidia_gpu_exporter_1.4.1_linux_x86_64.tar.gz ``` **2. Extract it** ```bash tar -xvf nvidia_gpu_exporter_1.4.1_linux_x86_64.tar.gz ``` **3. Create a restricted system user** ```bash sudo useradd --no-create-home --shell /bin/false nvidia_exporter ``` **4. Copy the engine to the secure bin folder** ```bash sudo cp nvidia_gpu_exporter /usr/local/bin/ ``` **5. Hand ownership to the new user** ```bash sudo chown nvidia_exporter:nvidia_exporter /usr/local/bin/nvidia_gpu_exporter ``` **6. Create the background service file** ```bash sudo nano /etc/systemd/system/nvidia_gpu_exporter.service ``` _Paste this exact code into the file and save:_ ```ini [Unit] Description=Nvidia GPU Exporter Wants=network-online.target After=network-online.target [Service] User=nvidia_exporter Group=nvidia_exporter Type=simple ExecStart=/usr/local/bin/nvidia_gpu_exporter [Install] WantedBy=multi-user.target ``` **7. Turn it on forever** ```bash sudo systemctl daemon-reload sudo systemctl start nvidia_gpu_exporter sudo systemctl enable nvidia_gpu_exporter ``` _(Once running, the GPU Exporter listens on port `9835`. You must go back to the Central Server and add a new job in `prometheus.yml` targeting `[SERVER_IP]:9835`)._ --- ## PART 3: ESSENTIAL PROMQL QUERIES _Use these queries in the Prometheus search bar or Grafana to monitor hardware._ **1. RAM: Percentage Available** ```promql (node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes) * 100 ``` **2. CPU: Percentage Actively Used** _(Node Exporter tracks how long the CPU is "idle". We subtract the idle speed from 100% to get the active usage)._ ```promql 100 - (avg by (instance) (rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100) ``` **3. DISK (SSD/HDD): Percentage Free Space (Root Drive)** ```promql (node_filesystem_avail_bytes{mountpoint="/"} / node_filesystem_size_bytes{mountpoint="/"}) * 100 ``` **4. NETWORK: Total Upload Speed in Megabytes/second (MB/s)** ```promql rate(node_network_transmit_bytes_total[5m]) / 1024 / 1024 ``` **5. HARDWARE: Server Uptime in Days** ```promql (time() - node_boot_time_seconds) / 86400 ``` **6. SERVERS: Is a server down? (1 = UP, 0 = DOWN)** ```promql up ``` --- ## PART 4: INSTALLING GRAFANA (THE DASHBOARD UI) _Grafana runs on port 3000 by default. Use this to visualize Prometheus data._ **1. Download the Grafana security key:** ```bash sudo apt-get install -y apt-transport-https software-properties-common wget sudo mkdir -p /etc/apt/keyrings/ wget -q -O - https://apt.grafana.com/gpg.key | gpg --dearmor | sudo tee /etc/apt/keyrings/grafana.gpg > /dev/null ``` **2. Add the official Grafana repository to Linux:** ```bash echo "deb [signed-by=/etc/apt/keyrings/grafana.gpg] https://apt.grafana.com stable main" | sudo tee -a /etc/apt/sources.list.d/grafana.list ``` **3. Install Grafana and start the service:** ```bash sudo apt-get update sudo apt-get install grafana sudo systemctl daemon-reload sudo systemctl start grafana-server sudo systemctl enable grafana-server ``` --- ## PART 5: GRAFANA DASHBOARD TEMPLATE IDs & DEBUGGING When importing templates from `grafana.com/dashboards`, use these exact IDs: - **Node Exporter Full (CPU/RAM/Network):** `1860` - **Nvidia GPU Exporter (Temperature/VRAM):** `14574` ### How to Fix "Blank" or "No Data" Dashboards 1. **Variables Mismatch:** Look at the dropdown menus at the very top of the dashboard. If they say "N/A", click them and select your specific `job_name` (e.g. `sasi_gpu` instead of `nvidia_gpu_exporter`). 2. **PromQL Mismatch:** If the graph is still blank, click the `...` next to the graph title and hit **Edit**. Look at the raw PromQL code. Sometimes community dashboards hardcode filters like `{job="gpu"}`. Simply delete that strict filter so your specific data can flow through! --- ### 1. Browser Blocks HTTP Exporters If your exporter is running (e.g., `curl http://localhost:9100/metrics` works in the terminal) but the browser says "Connection Refused", make sure your browser didn't automatically force `https://`. Exporters only run on plain `http://` by default! ### 2. Nvidia GPU Exporter Panic `[ms]` Error If you use an outdated version of `nvidia_gpu_exporter` (like `v1.2.1`), the engine will instantly crash with a `panic: descriptor Desc{...}` error because newer Nvidia drivers output a unit called `[ms]`. Prometheus strictly forbids spaces and brackets in metric names! **The Fix:** Always use `v1.4.1` or newer, as the developer hardcoded a fix for this exact regex bug. _(Note: Node Exporter does NOT monitor GPUs. To monitor GPU temperatures and VRAM, you must install a separate exporter called `nvidia_smi_exporter`)._ --- ## PART 6: INSTALLING ALERTMANAGER (For Slack/Email Notifications) **1. Download and Extract Alertmanager** ```bash wget https://github.com/prometheus/alertmanager/releases/download/v0.27.0/alertmanager-0.27.0.linux-amd64.tar.gz tar -xvf alertmanager-0.27.0.linux-amd64.tar.gz ``` **2. Create a restricted user and folders** ```bash sudo useradd --no-create-home --shell /bin/false alertmanager sudo mkdir /etc/alertmanager sudo mkdir /var/lib/alertmanager sudo chown alertmanager:alertmanager /var/lib/alertmanager ``` **3. Move the binary and hand over ownership** ```bash sudo cp alertmanager-0.27.0.linux-amd64/alertmanager /usr/local/bin/ sudo chown alertmanager:alertmanager /usr/local/bin/alertmanager ``` **4. Create the Configuration File** ```bash sudo nano /etc/alertmanager/alertmanager.yml ``` _Paste this code to route alerts to both Rocket.Chat and Gmail simultaneously:_ ```yaml global: resolve_timeout: 5m smtp_smarthost: "smtp.gmail.com:587" smtp_from: "knmkaushik@gmail.com" smtp_auth_username: "knmkaushik@gmail.com" smtp_auth_password: "YOUR_APP_PASSWORD" route: receiver: "fallback-do-nothing" group_wait: 10s group_interval: 1m repeat_interval: 5m routes: - receiver: "rocket-chat-alerts" continue: true - receiver: "email-alerts" continue: true receivers: - name: "fallback-do-nothing" - name: "rocket-chat-alerts" webhook_configs: - url: "YOUR_ROCKETCHAT_WEBHOOK_URL" send_resolved: true - name: "email-alerts" email_configs: - to: "kawsik97@gmail.com" send_resolved: true ``` **5. Hand over ownership of the config file** ```bash sudo chown alertmanager:alertmanager /etc/alertmanager/alertmanager.yml ``` **6. Create the background service file** ```bash sudo nano /etc/systemd/system/alertmanager.service ``` _Paste this exact code into the file and save:_ ```ini [Unit] Description=Alertmanager Wants=network-online.target After=network-online.target [Service] User=alertmanager Group=alertmanager Type=simple ExecStart=/usr/local/bin/alertmanager --config.file=/etc/alertmanager/alertmanager.yml --storage.path=/var/lib/alertmanager [Install] WantedBy=multi-user.target ``` **7. Turn Alertmanager on forever** ```bash sudo systemctl daemon-reload sudo systemctl start alertmanager sudo systemctl enable alertmanager ``` **8. Configure Rocket.Chat Webhook Script** Rocket.Chat natively only understands `{"text": "..."}` payloads, but Alertmanager sends structured JSON. To fix this, you must add a transformation script in the Rocket.Chat admin panel: 1. Go to your Rocket.Chat Admin Panel -> **Integrations** -> **Incoming**. 2. Create or Edit the incoming webhook for your channel (e.g., `#devops-alerts`). 3. Toggle **Script Enabled** to **ON**. 4. Paste the following into the Script box: ```javascript class Script { process_incoming_request({ request }) { var content = request.content; var status = content.status ? content.status.toUpperCase() : "UNKNOWN"; var alertName = "Unknown"; var instance = ""; var summary = ""; if (content.alerts && content.alerts[0]) { var a = content.alerts[0]; if (a.labels) { alertName = a.labels.alertname || "Unknown"; instance = a.labels.instance || ""; } if (a.annotations) { summary = a.annotations.summary || ""; } } var icon = status === "FIRING" ? ":red_circle:" : ":white_check_mark:"; var msg = icon + " *" + status + "* — " + alertName; if (instance) { msg = msg + " (" + instance + ")"; } if (summary) { msg = msg + "\n" + summary; } return { content: { text: msg, }, }; } } ``` 5. Make sure the "Script Sandbox" is set to "Secure Sandbox". 6. Click **Save** and use the generated Webhook URL in your `alertmanager.yml`. --- ## PART 7: LINKING PROMETHEUS & ALERTMANAGER Prometheus must be told that Alertmanager exists, and where the "Rules" are stored. **1. Open the main Prometheus config** ```bash sudo nano /etc/prometheus/prometheus.yml ``` **2. Update the `alerting` and `rule_files` blocks at the top:** ```yaml alerting: alertmanagers: - static_configs: - targets: ["localhost:9093"] rule_files: - "rules.yml" ``` _(Save and exit)._ **3. Create the Rules file (Where the Alert Logic lives):** ```bash sudo nano /etc/prometheus/rules.yml ``` **4. Paste this exact rule and save:** _(This logic says: If any server goes offline for more than 1 minute, fire a CRITICAL alert)._ ```yaml groups: - name: hardware_alerts rules: - alert: ServerDown expr: up == 0 for: 10s labels: severity: critical annotations: summary: "Server {{ $labels.instance }} is {{ if eq $value 0.0 }}DOWN{{ else }}UP{{ end }}!" description: "Server {{ $labels.instance }} has been unreachable for more than 10 seconds." ``` **5. Hand over ownership and restart Prometheus:** ```bash sudo chown prometheus:prometheus /etc/prometheus/rules.yml sudo systemctl restart prometheus ``` --- ## PART 8: MANAGING & DEBUGGING SERVICES (Crash, Reboot, & Maintenance) Since we configured all components (Prometheus, Node Exporter, Grafana, Alertmanager) as Linux `systemd` services and used the `enable` command, **they will start automatically every time you boot or reboot your server.** If you ever experience a crash, need to reboot, or want to test if services are running, use the following commands. ### 1. Check the Status of All Services at Once If you just rebooted your server and want to ensure everything turned on properly, run this: ```bash sudo systemctl status prometheus node_exporter grafana-server alertmanager ``` _Tip: If the output locks your terminal and shows `(END)` at the bottom, just press the `q` key on your keyboard to quit the reader mode and return to your prompt._ ### 2. Restarting a Service After a Crash or Configuration Change If you edit a configuration file (like `prometheus.yml` or `rules.yml`), the changes won't apply until you restart the service. If a service ever crashes and you want to force it to start fresh, use `restart`: ```bash # Restart Prometheus sudo systemctl restart prometheus # Restart Alertmanager sudo systemctl restart alertmanager # Restart Grafana sudo systemctl restart grafana-server ``` ### 3. Turning Services On / Off Manually If you need to intentionally stop a service to perform maintenance: ```bash sudo systemctl stop prometheus ``` To start it back up: ```bash sudo systemctl start prometheus ``` ### 4. Viewing Error Logs (Debugging) If a service says it is "failed" or "inactive" when you check its status, you can view its exact crash logs by checking the Linux `journalctl` (the master log book). To see the last 50 logs for Prometheus (useful for finding YAML syntax errors): ```bash sudo journalctl -u prometheus -n 50 --no-pager ``` To stream the logs live in real-time (press `Ctrl+C` to stop): ```bash sudo journalctl -u prometheus -f ```