How to Set Up Grafana + Prometheus Monitoring on Proxmox LXC 2026
Home Lab

How to Set Up Grafana + Prometheus Monitoring on Proxmox LXC 2026

Ricardo Gil
October 5, 2026
13 min read
#proxmox #self-hosting #lxc #grafana #prometheus
πŸ›’

Products in This Post

Affiliate links

As an Amazon Associate I earn from qualifying purchases at no extra cost to you.

For the first year of running Proxmox on my Beelink, I was essentially flying blind. My monitoring strategy was SSH-ing into containers, running htop, checking df -h, and hoping for the best. Then one evening I noticed Navidrome had been unresponsive for hours β€” the LXC had quietly hit its memory ceiling and the OOM killer had been at work. I found out because my wife couldn't play music. Not great.

That was the push I needed to build a real monitoring stack. I'd been putting it off because it felt like overkill for a single-node homelab. I was wrong. After setting up Prometheus and Grafana in a dedicated LXC container, I can see CPU utilization trends across six weeks, catch disk saturation before it becomes a problem, and get an email alert when any service goes down β€” all from a single dashboard. It's changed how I run my homelab completely.

In this guide I'll walk through exactly how I set up Prometheus and Grafana inside a single Proxmox LXC, with Node Exporter running on the Proxmox host and my other containers. By the end you'll have live dashboards showing CPU, memory, disk I/O, and network throughput for every node in your homelab. Before we start, if you're new to the LXC workflow, my Proxmox LXC complete self-hosting guide covers the fundamentals of container creation and networking.

How the Stack Works (and Why Prometheus + Grafana)

I've already written about Beszel vs Grafana + Prometheus if you want the full comparison. The short version: Beszel is quicker to set up and perfectly fine for basic host monitoring. But Prometheus + Grafana wins on every dimension that matters for a serious homelab: persistent time-series storage, a rich PromQL query language, thousands of community dashboards, and a proper alerting pipeline via Alertmanager. Once I had custom alert rules firing on disk space and service downtime, I never looked back.

The architecture is clean and lightweight. A single "monitoring" LXC hosts both Prometheus (metrics collector and time-series database) and Grafana (the visualization and alerting frontend). Node Exporter β€” a minimal metrics agent that exposes system stats over HTTP on port 9100 β€” runs on the Proxmox host and on any LXC containers you want to observe. Prometheus scrapes each Node Exporter endpoint every 15 seconds, stores the data in its local TSDB, and Grafana queries that data to render dashboards and evaluate alert rules.

Resource footprint: on my Beelink SER5 setup with around 12 scrape targets, the monitoring LXC uses about 800MB of RAM (Prometheus is the heavier process), 2 vCPUs, and the TSDB grows at roughly 1.5GB per month with a 15-second scrape interval. I allocate 30GB of disk and run 90-day retention, which leaves plenty of headroom. Even on a tight homelab build, this is very manageable.

Step 1: Create the Monitoring LXC in Proxmox

I use a Debian 12 template for the monitoring container β€” same base I use for most of my LXC workloads. Download the template from the Proxmox shell if you don't have it already, then create the container. Adjust the VMID, IP, and storage pool for your environment:

bash
# From the Proxmox host shell
pveam update
pveam download local debian-12-standard_12.7-1_amd64.tar.zst

# Create the monitoring LXC (adjust IDs and storage to your setup)
pct create 200 local:vztmpl/debian-12-standard_12.7-1_amd64.tar.zst \
  --hostname monitoring \
  --cores 2 \
  --memory 2048 \
  --swap 512 \
  --rootfs local-lvm:30 \
  --net0 name=eth0,bridge=vmbr0,ip=dhcp \
  --unprivileged 1 \
  --start 1

Once the container boots, log in via the Proxmox console (pct enter 200) and run the usual first-run setup:

bash
apt update && apt upgrade -y
apt install -y curl wget gnupg2 software-properties-common apt-transport-https

One thing I always do immediately: set a static IP reservation in my router's DHCP settings for this container's MAC address. The monitoring LXC's IP gets hardcoded into every Prometheus scrape config β€” if it changes after a DHCP lease renewal, your scrape targets break. Alternatively you can set a static IP in the LXC's /etc/network/interfaces. Either way, lock it down before you continue.

Step 2: Install Prometheus

I install Prometheus from the official binary releases rather than the Debian repos. The apt packages tend to lag by a few minor versions, and the binary install is just as clean with a proper systemd unit. At the time of writing, Prometheus 2.54.x is current stable.

bash
# Create dedicated system user
useradd --no-create-home --shell /bin/false prometheus

# Create config and data directories
mkdir -p /etc/prometheus /var/lib/prometheus

# Download and extract the binary release
PROM_VERSION="2.54.1"
wget -q https://github.com/prometheus/prometheus/releases/download/v${PROM_VERSION}/prometheus-${PROM_VERSION}.linux-amd64.tar.gz
tar xzf prometheus-${PROM_VERSION}.linux-amd64.tar.gz
cd prometheus-${PROM_VERSION}.linux-amd64

# Install binaries and console templates
cp prometheus promtool /usr/local/bin/
cp -r consoles/ console_libraries/ /etc/prometheus/

# Fix ownership
chown -R prometheus:prometheus /etc/prometheus /var/lib/prometheus
chown prometheus:prometheus /usr/local/bin/prometheus /usr/local/bin/promtool

Now create the base configuration file. This is nearly identical to the config I run on my Beelink today β€” replace the IPs with your actual Proxmox host and LXC addresses:

yaml
# /etc/prometheus/prometheus.yml
global:
  scrape_interval: 15s
  evaluation_interval: 15s

alerting:
  alertmanagers:
    - static_configs:
        - targets: []

rule_files:
  - "/etc/prometheus/alerts/*.yml"

scrape_configs:
  - job_name: "prometheus"
    static_configs:
      - targets: ["localhost:9090"]

  # The Proxmox host itself
  - job_name: "proxmox-host"
    static_configs:
      - targets: ["192.168.1.10:9100"]
        labels:
          instance: "beelink-proxmox"

  # All LXC containers as one job (easier to manage)
  - job_name: "lxc-containers"
    static_configs:
      - targets:
          - "192.168.1.101:9100"   # jellyfin
          - "192.168.1.102:9100"   # vaultwarden
          - "192.168.1.103:9100"   # navidrome
          - "192.168.1.104:9100"   # n8n
          - "192.168.1.105:9100"   # ollama

Create the systemd service unit, then enable and start Prometheus:

bash
cat <<EOF > /etc/systemd/system/prometheus.service
[Unit]
Description=Prometheus Monitoring
Wants=network-online.target
After=network-online.target

[Service]
User=prometheus
Group=prometheus
Type=simple
ExecStart=/usr/local/bin/prometheus \
    --config.file /etc/prometheus/prometheus.yml \
    --storage.tsdb.path /var/lib/prometheus/ \
    --web.console.templates=/etc/prometheus/consoles \
    --web.console.libraries=/etc/prometheus/console_libraries \
    --storage.tsdb.retention.time=90d \
    --web.enable-lifecycle

[Install]
WantedBy=multi-user.target
EOF

systemctl daemon-reload
systemctl enable --now prometheus
systemctl status prometheus

Prometheus starts up and binds to port 9090. You can immediately access the expression browser at http://<monitoring-lxc-ip>:9090 β€” it's sparse but useful for querying raw metrics and checking scrape status. The Targets page (/targets) is where I check every time I add a new scrape target.

Step 3: Install Node Exporter on Every Host You Want to Monitor

Node Exporter is the lightweight agent that exposes system metrics over HTTP on port 9100. You run it on the Proxmox host itself and on each LXC container you want visibility into. The binary is about 20MB and uses almost no resources in steady state. Here's the install script I've templated so I can drop it into any new LXC in under a minute:

bash
#!/bin/bash
# Run this on the Proxmox HOST and on each LXC you want to monitor

NODE_VERSION="1.8.2"

useradd --no-create-home --shell /bin/false node_exporter

wget -q https://github.com/prometheus/node_exporter/releases/download/v${NODE_VERSION}/node_exporter-${NODE_VERSION}.linux-amd64.tar.gz
tar xzf node_exporter-${NODE_VERSION}.linux-amd64.tar.gz
cp node_exporter-${NODE_VERSION}.linux-amd64/node_exporter /usr/local/bin/
chown node_exporter:node_exporter /usr/local/bin/node_exporter

cat <<EOF > /etc/systemd/system/node_exporter.service
[Unit]
Description=Prometheus Node Exporter
After=network.target

[Service]
User=node_exporter
Group=node_exporter
Type=simple
ExecStart=/usr/local/bin/node_exporter

[Install]
WantedBy=multi-user.target
EOF

systemctl daemon-reload
systemctl enable --now node_exporter
echo "Node Exporter is running on port 9100"

A quick note on running Node Exporter inside unprivileged LXC containers: since they share the Proxmox host's kernel, a few collectors (hardware sensors, NUMA topology) won't return meaningful data. That's expected and harmless β€” the metrics that matter for homelab monitoring (CPU, memory, disk I/O, filesystem usage, network bytes in/out) all work correctly. After install, verify with curl http://localhost:9100/metrics | grep ^node_cpu and you should see a wall of CPU counters.

Run this script on the Proxmox host via a root SSH session or the PVE shell, and then on each LXC you care about. I also run it on my Uptime Kuma container β€” between Node Exporter metrics and Uptime Kuma's service health checks, I have both infrastructure metrics and application availability covered from two angles, which has saved me more than once.

Step 4: Verify Prometheus Is Scraping Successfully

Once Node Exporter is running on your targets, update prometheus.yml with their IPs and send Prometheus a reload signal (no restart needed, thanks to --web.enable-lifecycle):

bash
curl -X POST http://localhost:9090/-/reload

Then navigate to http://<monitoring-ip>:9090/targets. Each configured target should show a green "UP" state and the time since the last successful scrape. If a target shows "DOWN", the usual culprits are: firewall rules inside the target LXC blocking port 9100 inbound, or a wrong IP in your scrape config. On my setup I don't run ufw inside individual LXCs, so it just works. If you do, open port 9100 from the monitoring container's IP.

You can test connectivity manually from the monitoring LXC before adding a target to the config:

bash
# Test that you can reach a target's Node Exporter
curl http://192.168.1.101:9100/metrics | head -5

Once all targets are UP, try your first PromQL query in the expression browser. node_memory_MemAvailable_bytes will return available RAM across all scraped hosts. 100 - (avg by (instance) (rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100) gives you CPU utilization percentage per instance. These confirm end-to-end that scraping is working correctly before you set up Grafana.

Step 5: Install Grafana

I install Grafana OSS from the official Grafana apt repository. This keeps it manageable via apt upgrade and avoids the stale package problem that affects distro repos:

bash
# Add the official Grafana GPG key and repo
wget -q -O - https://apt.grafana.com/gpg.key | gpg --dearmor | tee /usr/share/keyrings/grafana.gpg > /dev/null
echo "deb [signed-by=/usr/share/keyrings/grafana.gpg] https://apt.grafana.com stable main" | tee /etc/apt/sources.list.d/grafana.list

apt update && apt install -y grafana

systemctl daemon-reload
systemctl enable --now grafana-server
systemctl status grafana-server

Grafana starts on port 3000. Hit http://<monitoring-ip>:3000 in a browser and log in with admin / admin β€” you'll be prompted to set a new password on first login. Do it, and while you're in the UI, navigate to Administration β†’ Plugins and update any bundled plugins that show available updates. Fresh installs sometimes ship with slightly outdated panel plugins.

If you're exposing Grafana through Traefik with HTTPS (which I do β€” see my Traefik reverse proxy guide for the full setup), update /etc/grafana/grafana.ini with your domain before you start setting up dashboards. Set domain = monitoring.yourdomain.com and root_url = https://monitoring.yourdomain.com in the [server] section, then restart the service. Getting this right early avoids broken redirect loops later.

Step 6: Add Prometheus as a Data Source and Import Dashboards

In the Grafana UI, go to Connections β†’ Data sources β†’ Add data source β†’ Prometheus. Since Grafana and Prometheus run in the same LXC, set the URL to http://localhost:9090. Leave authentication empty and click Save & test β€” you should see a green "Data source is working" banner. If not, confirm Prometheus is running with systemctl status prometheus and that it's bound to all interfaces (ss -tlnp | grep 9090).

Now import the Node Exporter community dashboard β€” one of the most downloaded dashboards in Grafana's catalog and the first thing I install on any new Prometheus setup. Go to Dashboards β†’ Import, enter ID 1860 ("Node Exporter Full"), select your Prometheus data source, and click Import. You'll see an immediately useful dashboard with CPU usage by mode, memory breakdown, disk I/O in bytes/sec, network throughput, filesystem utilization, and load average β€” all pulling live data from your running targets. Use the top Instance dropdown to switch between your Proxmox host and each LXC.

Two other dashboards I add right away: ID 7362 ("Node Exporter Full β€” Kubernetes Full") for a clean multi-host overview when you have more than 4 nodes, and ID 11074 for per-instance drill-downs. You can save your own modified copies of any imported dashboard β€” I've built a homelab-specific variant with panels for Proxmox ZFS arc size, total VMs/LXC counts (pulled from the Proxmox VE exporter, a separate optional install), and my WD Black NVMe temperature pulled from smartmontools. Once you understand PromQL basics, dashboard customization gets addictive.

Step 7: Production Tips After 6 Months of Running This Stack

After running Prometheus + Grafana on my Beelink for about six months now, here are the things I'd go back and tell myself from day one.

Set up at least one alert rule. The whole point of monitoring is knowing when something breaks before someone else tells you. I use a simple alerts file with two rules: one that fires if any scrape target goes dark for more than 2 minutes, and one that warns when free disk on any host drops below 10%. These alone have caught two near-outages (a runaway log file eating my n8n container's disk, and a Jellyfin transcoder cache growing unbounded). Here's the exact config in /etc/prometheus/alerts/node_alerts.yml:

yaml
groups:
  - name: node_alerts
    rules:
      - alert: InstanceDown
        expr: up == 0
        for: 2m
        labels:
          severity: critical
        annotations:
          summary: "Instance {{ $labels.instance }} down"
          description: "{{ $labels.instance }} has been unreachable for 2+ minutes."

      - alert: DiskSpaceLow
        expr: (node_filesystem_avail_bytes{mountpoint="/"} / node_filesystem_size_bytes{mountpoint="/"}) * 100 < 10
        for: 5m
        labels:
          severity: warning
        annotations:
          summary: "Low disk on {{ $labels.instance }}"
          description: "Less than 10% free disk on {{ $labels.instance }} (mountpoint: {{ $labels.mountpoint }})"

Tune retention based on your actual use. The default is 15 days, which is too short for spotting seasonal trends or comparing week-over-week traffic. I run 90-day retention (--storage.tsdb.retention.time=90d). With ~12 targets at 15s intervals, 90 days of data uses about 5.5GB on my NVMe β€” completely reasonable. If you're on tight storage, 30 days is the minimum I'd recommend.

Split your scrape jobs by service type. I break my prometheus.yml into separate jobs: proxmox-host, lxc-core (Vaultwarden, Navidrome, etc.), lxc-ai (Ollama, Open WebUI), lxc-media (Jellyfin, arr stack). This makes dashboard filtering cleaner, lets you set different scrape intervals per group (I scrape the AI stack every 30s since Ollama's metrics don't change that fast), and makes alert rules easier to scope. Starting with one big "all LXCs" job works fine, but you'll want to split it once you have more than 8 targets.

Back up the monitoring LXC. Proxmox Backup Server handles LXC snapshots cleanly even while Prometheus is writing to its TSDB. I back up the monitoring container weekly. Losing months of historical trend data is genuinely painful β€” don't learn this the hard way when PBS makes it trivial to avoid.

Hardware That Runs This Comfortably

Adding the monitoring LXC to my Beelink costs about 400MB of RAM overhead on top of everything else I'm running. On a 16GB machine like mine, that's nothing. But if you're starting out on a tighter budget, the two things I'd think about are total RAM and NVMe speed (Prometheus is write-heavy to its TSDB).

For anyone just building their first Proxmox homelab and looking for something that can run monitoring plus 8–10 other services comfortably, the GMKtec G3 N100 Mini PC (~$189) is my current recommendation at the entry level β€” 16GB RAM, 1TB NVMe, and about 10W idle power draw. The Intel N100 handles Prometheus + Grafana fine alongside a full self-hosted stack.

If you also want to run Ollama with a decent model (7B+), you'll want more headroom. That's what I run on the Beelink SER5 with Ryzen 5 5500U (~$299). The extra cores handle Prometheus's TSDB compaction cycles without any latency spikes on other containers. One upgrade I'd make if I were starting over is maxing out RAM from day one β€” if your mini PC uses SODIMM slots, Crucial 32GB DDR4 SODIMM (~$59) lets you grow your LXC fleet without watching memory pressure creep up on your dashboards.

Want this kind of setup for your business?

I deploy private, self-hosted infrastructure on Proxmox: AI, automation and internal tools on hardware you own, documented and handed off. See the Private AI Stack β†’

Not sure what you need? Book a free discovery call.

Ricardo Gil is a full-stack developer (.NET/Angular) with 6+ years of experience. He runs a personal homelab on a Beelink mini PC with Proxmox VE, self-hosting everything from Jellyfin to local LLMs with Ollama. More about Ricardo β†’
πŸ“¬Weekly Newsletter

Get the best home lab & AI content

No spam. One email per week. Unsubscribe anytime.

Share this article