Auto Draft

Grafana + Prometheus on Docker: The Complete Homelab Monitoring Stack Guide

Why Every Homelab Needs a Monitoring Stack

Running a homelab without proper monitoring is a bit like driving with your dashboard blacked out — everything feels fine until it suddenly isn’t. You wake up to find your Plex server running at 98% CPU, your NAS out of disk space, or a container that quietly died overnight. A solid monitoring stack changes all of that. With Grafana and Prometheus, you get real-time dashboards, historical metrics, and alert capabilities that rival what enterprise teams use — all running locally on your own hardware, no cloud subscription required.

In this guide, I’ll walk you through deploying a complete Grafana + Prometheus monitoring stack using Docker Compose. By the end, you’ll be scraping metrics from your host machine, your containers, and any other services on your network — and visualizing them in beautiful, customizable dashboards you can actually act on.

If you’re new to Docker, I’d strongly recommend starting with our beginner’s guide to Docker before diving in here. We’ll be using Docker Compose heavily throughout this tutorial.


How Grafana and Prometheus Work Together

Before we dive into configuration files, it’s worth understanding the architecture. These two tools have very different jobs and understanding the split helps when you’re debugging or extending the stack:

  • Prometheus is a time-series database and scraping engine. It polls configured targets at regular intervals (default is every 15 seconds), pulls metrics in a specific text format, and stores them with timestamps. It also evaluates alerting rules and fires alerts when conditions are met.
  • Grafana is a visualization and dashboarding frontend. It connects to Prometheus (and many other data sources) and lets you build panels and dashboards from PromQL queries. Grafana doesn’t store metrics itself — it queries Prometheus on demand.

The “exporters” are the third piece of the puzzle. Since most software doesn’t natively expose Prometheus-format metrics, exporters act as translators. Node Exporter exposes Linux host metrics (CPU, RAM, disk I/O, network). cAdvisor (Container Advisor) exposes Docker container-level metrics. You run these as containers or services, and Prometheus scrapes them at the configured interval.

Here’s the data flow in plain terms:

Node Exporter  ─┐
cAdvisor       ─┼──► Prometheus (scrape & store) ──► Grafana (query & visualize)
Your App        ─┘
                        ↓
                   Alertmanager (optional)

Prerequisites

  • A Linux host — Debian, Ubuntu, Fedora, or any systemd-based distro works fine
  • Docker Engine 24+ and Docker Compose v2 installed
  • At least 2 GB RAM available for the stack (4 GB+ recommended for long-term storage)
  • Basic familiarity with YAML syntax and the Linux command line

I’m running this on a Proxmox VM with 4 vCPUs and 8 GB RAM, but a Raspberry Pi 4 (4 GB model), any mini-PC, or a spare laptop running Ubuntu will work equally well. The stack is lightweight — Prometheus typically uses around 300–500 MB RAM for a small homelab deployment.


Setting Up the Directory Structure

First, create a clean, organized directory layout for the monitoring stack. Keeping configuration files in version control (or at least backed up) makes the whole setup reproducible:

mkdir -p ~/monitoring/{prometheus/rules,grafana/{data,dashboards,provisioning/{datasources,dashboards}}}
cd ~/monitoring

Your target directory structure should look like this:

monitoring/
├── docker-compose.yml
├── prometheus/
│   ├── prometheus.yml
│   └── rules/
│       └── alerts.yml
└── grafana/
    ├── data/
    ├── dashboards/
    └── provisioning/
        ├── datasources/
        │   └── prometheus.yml
        └── dashboards/
            └── dashboards.yml

Prometheus Configuration

Create the Prometheus config at ~/monitoring/prometheus/prometheus.yml:

global:
  scrape_interval: 15s
  evaluation_interval: 15s

alerting:
  alertmanagers:
    - static_configs:
        - targets: []

rule_files:
  - "/etc/prometheus/rules/*.yml"

scrape_configs:
  - job_name: "prometheus"
    static_configs:
      - targets: ["localhost:9090"]

  - job_name: "node_exporter"
    static_configs:
      - targets: ["node_exporter:9100"]

  - job_name: "cadvisor"
    static_configs:
      - targets: ["cadvisor:8080"]

This config tells Prometheus to scrape itself, Node Exporter, and cAdvisor every 15 seconds. We’re using Docker service names as hostnames — since everything runs on the same Docker bridge network, Docker’s internal DNS resolves these automatically.

Need to add other services? Traefik, for example, exposes Prometheus metrics on port 8082 at /metrics. Just add another scrape job:

  - job_name: "traefik"
    static_configs:
      - targets: ["traefik:8082"]

Grafana Provisioning Configuration

Grafana supports automatic provisioning of data sources and dashboards through YAML config files. This means you never have to configure the data source through the UI after a fresh deploy — it’s all declarative and reproducible.

Create ~/monitoring/grafana/provisioning/datasources/prometheus.yml:

apiVersion: 1

datasources:
  - name: Prometheus
    type: prometheus
    access: proxy
    url: http://prometheus:9090
    isDefault: true
    editable: false

Create ~/monitoring/grafana/provisioning/dashboards/dashboards.yml:

apiVersion: 1

providers:
  - name: "Default"
    orgId: 1
    folder: ""
    type: file
    disableDeletion: false
    updateIntervalSeconds: 30
    options:
      path: /var/lib/grafana/dashboards

This tells Grafana to automatically load any .json dashboard files from /var/lib/grafana/dashboards, checking for updates every 30 seconds. We’ll mount our local dashboards directory there in the Compose file.


The Docker Compose File

Now for the main event. Create ~/monitoring/docker-compose.yml:

version: "3.8"

networks:
  monitoring:
    driver: bridge

volumes:
  prometheus_data: {}
  grafana_data: {}

services:
  prometheus:
    image: prom/prometheus:v2.53.0
    container_name: prometheus
    restart: unless-stopped
    volumes:
      - ./prometheus/prometheus.yml:/etc/prometheus/prometheus.yml:ro
      - ./prometheus/rules:/etc/prometheus/rules:ro
      - prometheus_data:/prometheus
    command:
      - "--config.file=/etc/prometheus/prometheus.yml"
      - "--storage.tsdb.path=/prometheus"
      - "--web.console.libraries=/usr/share/prometheus/console_libraries"
      - "--web.console.templates=/usr/share/prometheus/consoles"
      - "--storage.tsdb.retention.time=30d"
      - "--web.enable-lifecycle"
    ports:
      - "9090:9090"
    networks:
      - monitoring

  node_exporter:
    image: prom/node-exporter:v1.8.1
    container_name: node_exporter
    restart: unless-stopped
    volumes:
      - /proc:/host/proc:ro
      - /sys:/host/sys:ro
      - /:/rootfs:ro
    command:
      - "--path.procfs=/host/proc"
      - "--path.rootfs=/rootfs"
      - "--path.sysfs=/host/sys"
      - "--collector.filesystem.mount-points-exclude=^/(sys|proc|dev|host|etc)($$|/)"
    ports:
      - "9100:9100"
    networks:
      - monitoring

  cadvisor:
    image: gcr.io/cadvisor/cadvisor:v0.49.1
    container_name: cadvisor
    restart: unless-stopped
    privileged: true
    devices:
      - /dev/kmsg:/dev/kmsg
    volumes:
      - /:/rootfs:ro
      - /var/run:/var/run:ro
      - /sys:/sys:ro
      - /var/lib/docker:/var/lib/docker:ro
      - /cgroup:/cgroup:ro
    ports:
      - "8080:8080"
    networks:
      - monitoring

  grafana:
    image: grafana/grafana-oss:11.1.0
    container_name: grafana
    restart: unless-stopped
    depends_on:
      - prometheus
    volumes:
      - grafana_data:/var/lib/grafana
      - ./grafana/provisioning:/etc/grafana/provisioning:ro
      - ./grafana/dashboards:/var/lib/grafana/dashboards:ro
    environment:
      - GF_SECURITY_ADMIN_USER=admin
      - GF_SECURITY_ADMIN_PASSWORD=changeme
      - GF_USERS_ALLOW_SIGN_UP=false
      - GF_SERVER_DOMAIN=localhost
    ports:
      - "3000:3000"
    networks:
      - monitoring

A few notes on specific settings:

  • --storage.tsdb.retention.time=30d keeps 30 days of metrics history. On a small homelab, 30 days of Node Exporter data typically uses 1–3 GB of disk. Adjust to your available storage.
  • --web.enable-lifecycle allows hot-reloading the Prometheus config with an HTTP POST to /-/reload — no container restart needed when you update prometheus.yml.
  • cAdvisor requires privileged: true to access the container runtime’s cgroup data. This is expected and standard for cAdvisor deployments.
  • Change GF_SECURITY_ADMIN_PASSWORD to something secure before deploying.

Deploying the Stack

With all files in place, bring the stack up:

cd ~/monitoring
docker compose up -d

Watch the startup logs to confirm everything initializes cleanly:

docker compose logs -f

After 30–60 seconds, verify all containers are running:

docker compose ps

NAME             IMAGE                                    STATUS         PORTS
cadvisor         gcr.io/cadvisor/cadvisor:v0.49.1         Up 2 minutes   0.0.0.0:8080->8080/tcp
grafana          grafana/grafana-oss:11.1.0               Up 2 minutes   0.0.0.0:3000->3000/tcp
node_exporter    prom/node-exporter:v1.8.1                Up 2 minutes   0.0.0.0:9100->9100/tcp
prometheus       prom/prometheus:v2.53.0                  Up 2 minutes   0.0.0.0:9090->9090/tcp

Open http://YOUR_HOST_IP:9090 and navigate to Status → Targets. All three targets — prometheus, node_exporter, and cadvisor — should show UP in green.

Then open http://YOUR_HOST_IP:3000 for Grafana. Log in with admin and your password. The Prometheus data source should already appear under Configuration → Data Sources — fully configured and connected, no manual steps needed.


Importing Community Dashboards

One of the best things about the Grafana ecosystem is the community dashboard library at grafana.com/dashboards. Instead of building dashboards from scratch, you can import fully-featured ones in about 30 seconds.

Node Exporter Full (Dashboard ID: 1860) is the gold standard for Linux host metrics. It provides per-core CPU breakdown, memory usage with cache/buffer details, disk I/O throughput, network statistics, and system load — all in a single dashboard.

To import it:

  1. In Grafana, click the + menu → Import
  2. Type 1860 in the “Import via grafana.com” field and click Load
  3. Select Prometheus as the data source
  4. Click Import

For container metrics, cAdvisor exporter (Dashboard ID: 14282) gives you a per-container view of CPU throttling, memory usage, and network I/O. It’s particularly useful for spotting memory leaks or containers that are consuming unexpectedly high resources.

To make dashboards persistent across container rebuilds and manageable as code, export them as JSON (Dashboard settings → JSON Model) and save the files to ~/monitoring/grafana/dashboards/. Grafana’s provisioning config will auto-load them on every startup.


Essential PromQL Queries

PromQL (Prometheus Query Language) is how you retrieve and transform metrics. The Grafana “Explore” view is a great place to test queries interactively. Here are the ones I use most often:

CPU usage percentage (averaged across all cores):

100 - (avg(irate(node_cpu_seconds_total{mode="idle"}[5m])) * 100)

Available memory in GB:

node_memory_MemAvailable_bytes / 1024 / 1024 / 1024

Disk usage percentage for the root filesystem:

100 - ((node_filesystem_avail_bytes{mountpoint="/"} / node_filesystem_size_bytes{mountpoint="/"}) * 100)

Per-container CPU usage (percent):

sum(rate(container_cpu_usage_seconds_total{name!=""}[5m])) by (name) * 100

Per-container memory usage (MB):

container_memory_usage_bytes{name!=""} / 1024 / 1024

Network received bytes per second on eth0:

irate(node_network_receive_bytes_total{device="eth0"}[5m])

These form the backbone of most homelab dashboards. Combine them with Grafana’s visualization types (time series, gauge, stat panel) to build exactly the overview you need.


Setting Up Alerting Rules

Prometheus evaluates alerting rules at the evaluation_interval (also 15 seconds by default) and fires alerts when conditions are met. Create ~/monitoring/prometheus/rules/alerts.yml:

groups:
  - name: homelab_alerts
    rules:
      - alert: HighCPUUsage
        expr: 100 - (avg(irate(node_cpu_seconds_total{mode="idle"}[5m])) * 100) > 85
        for: 5m
        labels:
          severity: warning
        annotations:
          summary: "High CPU on {{ $labels.instance }}"
          description: "CPU has been above 85% for 5 minutes."

      - alert: LowDiskSpace
        expr: (node_filesystem_avail_bytes{mountpoint="/"} / node_filesystem_size_bytes{mountpoint="/"}) * 100 < 15
        for: 2m
        labels:
          severity: critical
        annotations:
          summary: "Low disk on {{ $labels.instance }}"
          description: "Root filesystem has less than 15% free space."

      - alert: HighMemoryUsage
        expr: (1 - (node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes)) * 100 > 90
        for: 5m
        labels:
          severity: warning
        annotations:
          summary: "High memory usage on {{ $labels.instance }}"
          description: "Memory usage is above 90% for 5 minutes."

      - alert: ScrapTargetDown
        expr: up == 0
        for: 1m
        labels:
          severity: critical
        annotations:
          summary: "Scrape target {{ $labels.job }} is down"
          description: "Prometheus cannot reach {{ $labels.instance }}."

The ScrapTargetDown alert fires when any monitored endpoint stops responding — this covers node_exporter going offline, cAdvisor crashing, or any other scrape failure. It’s your catch-all for “something stopped reporting.”

After updating your rules file, hot-reload Prometheus without a container restart:

curl -X POST http://localhost:9090/-/reload

Firing alerts show up under the Alerts tab in the Prometheus UI. To route them to Slack, email, or PagerDuty, you’d add Alertmanager to the stack — but the rules themselves are already in place and evaluated regardless.


Expanding to More Targets

The real power of this stack comes from scraping everything on your network. Some common homelab additions:

Blackbox Exporter probes HTTP endpoints, ICMP ping, DNS, and TCP ports from the outside. Add it to your Compose file:

  blackbox:
    image: prom/blackbox-exporter:v0.25.0
    container_name: blackbox
    restart: unless-stopped
    ports:
      - "9115:9115"
    networks:
      - monitoring

Then in prometheus.yml, add an HTTP probe job:

  - job_name: "blackbox_http"
    metrics_path: /probe
    params:
      module: [http_2xx]
    static_configs:
      - targets:
          - https://your-public-site.com
          - http://192.168.1.10:32400/web/index.html
    relabel_configs:
      - source_labels: [__address__]
        target_label: __param_target
      - source_labels: [__param_target]
        target_label: instance
      - target_label: __address__
        replacement: blackbox:9115

This gives you synthetic uptime monitoring similar to Uptime Kuma, but feeding into Grafana dashboards with response time history. The two tools complement each other well — Uptime Kuma for clean public status pages and on-call alerting, Prometheus + Grafana for deep metrics correlation.

You can also pull sensor data from Home Assistant directly into Grafana. Enable the built-in Prometheus integration in Home Assistant’s configuration.yaml:

prometheus:

Then add a scrape job in prometheus.yml pointing to http://homeassistant:8123/api/prometheus with a valid long-lived access token in the Authorization header. Every sensor, binary sensor, and entity state becomes a queryable metric in Grafana — temperature readings, power consumption, motion events, all of it.


Persistent Storage and Backups

The Compose file uses named Docker volumes for both Prometheus and Grafana data. To find the physical paths on disk:

docker volume inspect monitoring_grafana_data
# → "Mountpoint": "/var/lib/docker/volumes/monitoring_grafana_data/_data"

For regular backups, a simple daily cron job works well. Grafana’s data volume holds your dashboards, users, and settings — the most important things to protect:

0 2 * * * docker run --rm \
  -v monitoring_grafana_data:/source:ro \
  -v /mnt/backup/grafana:/dest \
  alpine tar czf /dest/grafana-$(date +\%Y\%m\%d).tar.gz -C /source .

Prometheus metric data is less critical — if lost, Prometheus simply resumes collecting from the current moment. But if you need historical continuity, the same backup approach applies to the monitoring_prometheus_data volume.


Security Considerations

By default, Prometheus and the exporters have no authentication. For a LAN-only homelab this is acceptable, but if you’re exposing any of these ports externally (via port forwarding or Tailscale), take these precautions:

  • Proxy Prometheus and cAdvisor behind Nginx Proxy Manager or Traefik with HTTP basic authentication
  • Change the Grafana admin password immediately and set GF_USERS_ALLOW_SIGN_UP=false (already in our Compose file)
  • Consider enabling Grafana’s built-in TLS or terminating SSL at the reverse proxy layer
  • Bind Node Exporter to 127.0.0.1:9100 if it doesn’t need to be reachable from other hosts on the network

Wrapping Up

A Grafana + Prometheus monitoring stack is one of the highest-value additions you can make to a homelab. Once it’s running, you’ll start noticing patterns you never knew existed — memory leaks in long-running containers, disk I/O spikes correlating with specific cron jobs, network saturation at odd hours. It turns your homelab from a black box into something you actually understand deeply.

The stack we built here is the same architecture that powers observability at companies running thousands of servers. You’re running it at homelab scale, but the tools and skills transfer directly. Start with the Node Exporter Full dashboard, get comfortable writing PromQL queries, then expand — scrape your router if it supports it, add the Blackbox Exporter for external probes, pull in Home Assistant sensor data. Before long you’ll have a monitoring setup that most sysadmins would envy.

Enjoying this post?

Get more guides like this delivered straight to your inbox. No spam, just tech and trails.