Scientyfic World

How to Monitor Docker Containers with Prometheus, Grafana and Loki

When a container is slow, restarts or runs out of memory, you want answers in seconds: which container, since when, and what else changed. This guide builds a small monitoring...

Share:

Get an AI summary of this article

monitor and log docker containers

When a container is slow, restarts or runs out of memory, you want answers in seconds: which container, since when, and what else changed. This guide builds a small monitoring stack that gives you exactly that: cAdvisor measures every container, Prometheus stores the numbers and fires alerts, Loki stores logs, Grafana Alloy ships the logs, and Grafana shows it all in one place. The whole thing is one Compose file plus four small config files.

Updated October 2026: I rewrote this 2024 post. The original had several steps that fail in practice: it told Grafana to use http://localhost:9090 for Prometheus and Prometheus to scrape localhost:8080 for cAdvisor, but inside a container localhost is the container itself, so neither connection works. It also installed the Loki Docker logging driver (which Grafana’s documentation now steers people away from in favour of Alloy), repeated the cAdvisor section twice, used an obsolete version: line and “scaled” Prometheus to two replicas, which does not give you high availability. This version fixes all of that, and I ran every config file through the matching tool’s own validator.

What Each Piece Does

ComponentJobTalks to
cAdvisorMeasures CPU, memory, network and disk for each container on the hostDocker and the kernel; scraped by Prometheus
PrometheusScrapes and stores metrics; evaluates alert rulescAdvisor (pull)
LokiStores and indexes logsReceives logs from Alloy
Grafana AlloyCollects container logs from Docker and forwards themDocker socket, then Loki
GrafanaDashboards, exploration and alert viewsPrometheus and Loki as data sources

Everything runs on one Docker Compose network, so containers reach each other by service name (prometheus, loki, cadvisor). That is the key to the localhost mistake above: use service names between containers, and localhost only for things on the same machine.

Prerequisites

  • A Linux host with Docker Engine and the Compose plugin. cAdvisor reads container data from the host’s cgroups and Docker directories, so this stack is designed for Linux servers. On Docker Desktop (Mac or Windows) containers run inside a VM, and the cAdvisor documentation does not cover that case, so expect gaps; use a Linux VM or server for real monitoring.
  • About 1 GB of free memory for the monitoring containers themselves, and disk for metrics and logs. The stack below keeps 15 days of metrics and 7 days of logs.
  • Basic comfort with YAML. Create a project folder, for example monitoring/, and put every file below in it.

Step 1: Prometheus Configuration

global:
  scrape_interval: 15s
  evaluation_interval: 15s

rule_files:
  - /etc/prometheus/alerts.yml

scrape_configs:
  - job_name: prometheus
    static_configs:
      - targets: ["localhost:9090"]

  - job_name: cadvisor
    static_configs:
      - targets: ["cadvisor:8080"]
  • cadvisor:8080 uses the Compose service name. Prometheus pulls metrics from each target every 15 seconds.
  • The prometheus job lets Prometheus monitor itself (inside its own container, localhost:9090 is correct).
  • rule_files points to the alert rules in the next step.

Step 2: Alert Rules You Can Test

Dashboards are for looking; alerts are for the moments you are not looking. These three rules cover the common container problems:

groups:
  - name: containers
    rules:
      - alert: ContainerHighMemory
        expr: |
          container_memory_working_set_bytes{name!=""}
            / (container_spec_memory_limit_bytes{name!=""} > 0) > 0.90
        for: 5m
        labels:
          severity: warning
        annotations:
          summary: "{{ $labels.name }} is using over 90% of its memory limit"

      - alert: ContainerHighCpu
        expr: sum by (name) (rate(container_cpu_usage_seconds_total{name!=""}[5m])) > 0.9
        for: 10m
        labels:
          severity: warning
        annotations:
          summary: "{{ $labels.name }} has used close to a full CPU core for 10 minutes"

      - alert: ContainerGone
        expr: time() - container_last_seen{name!=""} > 120
        for: 1m
        labels:
          severity: critical
        annotations:
          summary: "{{ $labels.name }} has not been seen for more than 2 minutes"
  • ContainerHighMemory compares memory use to the container’s configured limit. The > 0 part skips containers with no limit, which cAdvisor reports as 0 and would otherwise cause a divide-by-zero.
  • ContainerHighCpu fires when a container averages close to a full core for ten minutes.
  • ContainerGone fires when cAdvisor has not seen a container for two minutes. Be aware that it also fires for containers you stop on purpose, and that cAdvisor forgets a container a short while after it disappears, so treat it as a hint and not a complete “is it running” check.
  • The for: clause means the condition must hold for that long before the alert fires, which avoids alerting on brief spikes. Prometheus’ own alerting practices guide recommends alerting on symptoms that need a human, not on every metric.

Prometheus can unit-test alert rules. This test feeds in made-up series and checks that only the container above 90% of its limit alerts, while a container with plenty of headroom and one with no limit do not (see the unit testing docs):

rule_files:
  - alerts.yml
evaluation_interval: 1m
tests:
  - interval: 1m
    input_series:
      - series: 'container_memory_working_set_bytes{name="web"}'
        values: "950000000+0x10"
      - series: 'container_spec_memory_limit_bytes{name="web"}'
        values: "1000000000+0x10"
      - series: 'container_memory_working_set_bytes{name="db"}'
        values: "100000000+0x10"
      - series: 'container_spec_memory_limit_bytes{name="db"}'
        values: "1000000000+0x10"
      - series: 'container_memory_working_set_bytes{name="nolimit"}'
        values: "900000000+0x10"
      - series: 'container_spec_memory_limit_bytes{name="nolimit"}'
        values: "0+0x10"
    alert_rule_test:
      - eval_time: 8m
        alertname: ContainerHighMemory
        exp_alerts:
          - exp_labels:
              severity: warning
              name: web
            exp_annotations:
              summary: "web is using over 90% of its memory limit"
promtool check rules alerts.yml
promtool test rules alerts_test.yml
# SUCCESS

I ran both commands with Prometheus’ promtool and they pass. Putting a test like this in your repository means that a future edit to the rule is checked before it can silently stop alerting. Note that these alerts only appear in Prometheus’ Alerts page until you connect an Alertmanager (or use Grafana’s own alerting, below) to send notifications.

Step 3: Loki and Alloy for Logs

Metrics tell you that something is wrong; logs tell you why. Loki stores them:

auth_enabled: false

server:
  http_listen_port: 3100

common:
  instance_addr: 127.0.0.1
  path_prefix: /loki
  storage:
    filesystem:
      chunks_directory: /loki/chunks
      rules_directory: /loki/rules
  replication_factor: 1
  ring:
    kvstore:
      store: inmemory

schema_config:
  configs:
    - from: 2024-01-01
      store: tsdb
      object_store: filesystem
      schema: v13
      index:
        prefix: index_
        period: 24h

limits_config:
  retention_period: 168h

compactor:
  working_directory: /loki/compactor
  retention_enabled: true
  delete_request_store: filesystem

This is a single-process setup with data on local disk and 7-day retention (the compactor deletes older logs). It is suitable for one host; larger setups use object storage. Loki’s own -verify-config check accepts this file.

To collect logs, the original post installed the Loki Docker logging driver and edited /etc/docker/daemon.json, which needs a daemon restart. Grafana’s Loki driver documentation describes the driver as not carrying the same support commitment as Grafana Alloy and recommends Alloy’s Docker components to avoid known problems such as daemon deadlocks. Alloy reads logs through the Docker API instead, with no change to the daemon:

discovery.docker "containers" {
	host = "unix:///var/run/docker.sock"
}

discovery.relabel "containers" {
	targets = discovery.docker.containers.targets

	rule {
		source_labels = ["__meta_docker_container_name"]
		regex         = "/(.*)"
		target_label  = "container"
	}
}

loki.source.docker "containers" {
	host       = "unix:///var/run/docker.sock"
	targets    = discovery.relabel.containers.output
	forward_to = [loki.write.local.receiver]
}

loki.write "local" {
	endpoint {
		url = "http://loki:3100/loki/api/v1/push"
	}
}

The first block discovers running containers, the second turns Docker’s container name (which starts with a slash) into a clean container label, the third reads their logs (see the loki.source.docker docs) and the last sends them to Loki. I validated this file with alloy validate.

A security note: mounting /var/run/docker.sock into a container gives that container the ability to control Docker, which is effectively root on the host. It is mounted read-only here, but a read-only socket can still be used to read sensitive information. Run only the official image, pin its version, and do not expose Alloy’s port.

Step 4: Grafana Provisioning

Instead of clicking through menus to add data sources, define them in code with Grafana’s provisioning so a rebuilt container comes back configured. Save this as provisioning/datasources/datasources.yml:

apiVersion: 1

datasources:
  - name: Prometheus
    uid: prometheus
    type: prometheus
    access: proxy
    url: http://prometheus:9090
    isDefault: true

  - name: Loki
    uid: loki
    type: loki
    access: proxy
    url: http://loki:3100

Note the URLs: http://prometheus:9090 and http://loki:3100, using service names. This single fix removes the most common “data source connected but no data” problem from the older version of this tutorial.

Step 5: The Compose File

services:
  prometheus:
    image: prom/prometheus:v3.15.0
    command:
      - --config.file=/etc/prometheus/prometheus.yml
      - --storage.tsdb.retention.time=15d
    volumes:
      - ./prometheus.yml:/etc/prometheus/prometheus.yml:ro
      - ./alerts.yml:/etc/prometheus/alerts.yml:ro
      - prometheus_data:/prometheus
    ports:
      - "127.0.0.1:9090:9090"
    restart: unless-stopped

  cadvisor:
    image: ghcr.io/google/cadvisor:v0.60.6
    volumes:
      - /:/rootfs:ro
      - /var/run:/var/run:rw
      - /sys:/sys:ro
      - /var/lib/docker/:/var/lib/docker:ro
    restart: unless-stopped
    # no published port: Prometheus reaches it as cadvisor:8080 on the Compose network

  loki:
    image: grafana/loki:3.7.8
    command: -config.file=/etc/loki/config.yaml
    volumes:
      - ./loki-config.yaml:/etc/loki/config.yaml:ro
      - loki_data:/loki
    restart: unless-stopped

  alloy:
    image: grafana/alloy:v1.20.1
    command:
      - run
      - --server.http.listen-addr=0.0.0.0:12345
      - --storage.path=/var/lib/alloy/data
      - /etc/alloy/config.alloy
    volumes:
      - ./config.alloy:/etc/alloy/config.alloy:ro
      - /var/run/docker.sock:/var/run/docker.sock:ro
    depends_on:
      - loki
    restart: unless-stopped

  grafana:
    image: grafana/grafana:13.2.3
    environment:
      GF_SECURITY_ADMIN_PASSWORD__FILE: /run/secrets/grafana_admin_password
    secrets:
      - grafana_admin_password
    volumes:
      - ./provisioning:/etc/grafana/provisioning:ro
      - grafana_data:/var/lib/grafana
    ports:
      - "127.0.0.1:3000:3000"
    depends_on:
      - prometheus
      - loki
    restart: unless-stopped

secrets:
  grafana_admin_password:
    file: ./grafana_admin_password.txt

volumes:
  prometheus_data:
  loki_data:
  grafana_data:
  • Only Prometheus and Grafana publish ports, and only on 127.0.0.1. cAdvisor, Loki and Alloy have no published ports, because the other containers reach them over the Compose network. Metrics and logs can contain sensitive information, so do not expose these services to the internet; reach Grafana through an SSH tunnel or a reverse proxy with HTTPS and authentication (see our guide to Docker ports).
  • The Grafana admin password comes from a secret file via the __FILE suffix, not from an environment variable in the Compose file. Create it before starting: openssl rand -base64 24 > grafana_admin_password.txt, keep it out of version control, and read it when you first log in as admin.
  • Versions are pinned (I confirmed each tag exists on its registry in October 2026). Avoid latest for monitoring; an unexpected upgrade should not change your alerts. Update the tags deliberately and read the release notes.
  • Named volumes keep metrics, logs and Grafana data across container re-creation.
  • cAdvisor’s mounts (/, /var/run, /sys and /var/lib/docker) follow the example in the cAdvisor running guide. Its image is hosted at ghcr.io/google/cadvisor. Its docs mention extra flags (such as --privileged and --pid=host) only for process-level metrics, and --userns=host if your daemon uses user namespaces; this stack does not need them for container metrics.
  • There is no version: line, and no container_name, so you can run several projects on one host without name clashes.
openssl rand -base64 24 > grafana_admin_password.txt
docker compose up -d
docker compose ps

What I tested and what I did not: I validated each config with its own tool (promtool for Prometheus and the alert rules and their unit test, loki -verify-config, alloy validate) and checked that the Compose file parses and that all image tags exist. I could not run Docker on the machine where I wrote this post, so I did not start the stack; the first run on your Linux host is the real test. If something is off, docker compose logs SERVICE is the first place to look.

Step 6: Check That It Works

  1. Open http://localhost:9090/targets on the host. The cadvisor and prometheus targets should show UP. If cadvisor is down, check the logs and the service name in prometheus.yml.
  2. In the Prometheus query box, run container_memory_working_set_bytes. You should get one series per container.
  3. Open http://localhost:3000 and sign in as admin. Under Connections, check that both data sources exist and use Save & test to confirm they connect.
  4. Open Explore, choose Loki and run {container="prometheus"}. You should see that container’s log lines.

Step 7: Build Useful Dashboards

You can build panels yourself or import ready-made ones: in Grafana, go to Dashboards, then New, then Import dashboard, and paste a dashboard URL or ID from Grafana’s dashboard catalogue (see the import documentation). Search the catalogue for “cAdvisor” or “Docker” and check a dashboard’s last update date and reviews before you rely on it. If you build your own, these queries (all of which I syntax-checked with promtool) cover the essentials:

PanelQuery
CPU cores used per containersum by (name) (rate(container_cpu_usage_seconds_total{name!=""}[5m]))
Memory in usecontainer_memory_working_set_bytes{name!=""}
Memory as a share of the limitcontainer_memory_working_set_bytes{name!=""} / (container_spec_memory_limit_bytes{name!=""} > 0)
Network receive ratesum by (name) (rate(container_network_receive_bytes_total{name!=""}[5m]))
Network transmit ratesum by (name) (rate(container_network_transmit_bytes_total{name!=""}[5m]))
Errors in logs (Loki){container="web"} |= "error"

The name!="" filter hides the host-level series that cAdvisor also reports. Use a Grafana variable on the name label to get a dropdown of containers, and put the log panel next to the CPU and memory panels so a spike and its log lines appear together.

Alerting and Notifications

Prometheus evaluates the rules above, but sending an email or chat message requires an Alertmanager or Grafana’s built-in alerting. Grafana’s alerting documentation covers contact points (email, Slack, webhooks) and notification policies. A sound approach: start with a few alerts that always mean “a human should look” (the three here), send them to one channel, and tune thresholds for a couple of weeks before adding more. Too many alerts teach people to ignore all of them.

Running It in Production

  • Monitor the monitor. Alert if Prometheus targets go down (the built-in up metric is 0), and make sure the disk holding metrics and logs cannot fill up. Retention settings (15 days of metrics, 7 days of logs here) bound the growth.
  • Back up Grafana’s volume (or keep dashboards as JSON in version control), because dashboards you build in the UI live in its database.
  • Do not “scale” Prometheus by running two copies. Two instances with separate storage each hold their own data; high availability needs a different design (two independent instances scraping the same targets, plus a query layer or remote storage).
  • Secure access. Put Grafana behind HTTPS and real authentication, and if you ever expose the Docker daemon remotely, follow Docker’s guide to protecting access rather than opening its port.
  • Consider the daemon’s own metrics. Docker Engine can expose Prometheus metrics about the engine itself (see the Docker documentation), a useful addition to cAdvisor’s per-container data.
  • Know when to stop self-hosting. For a handful of containers this stack is cheap and educational; if you run many hosts, compare it with a managed service so you are not maintaining monitoring instead of your product.

Conclusion

Docker monitoring comes down to a clear pipeline: cAdvisor measures, Prometheus stores and alerts, Alloy ships logs to Loki, and Grafana shows metrics and logs side by side. The details that decide whether it works are using service names instead of localhost, not publishing ports you do not need, pinning versions, and testing your alert rules. Start with the stack above on a Linux host, confirm the targets are UP, and add alerts one at a time.

Why can’t Grafana connect to Prometheus at localhost:9090?

Inside the Grafana container, localhost means the Grafana container itself, not your host or the Prometheus container. Use the Compose service name instead, for example http://prometheus:9090, as long as both containers are on the same Compose network.

Do I need cAdvisor to monitor Docker containers with Prometheus?

It is the most common way to get per-container CPU, memory, network and disk metrics. Docker Engine can also expose its own metrics, but those describe the engine, not each container’s resource use. Best results come from running cAdvisor on a Linux host.

Should I use the Loki Docker logging driver or Grafana Alloy?

Grafana recommends Alloy. Its documentation says the Docker driver does not carry the same support commitment, and it points to Alloy’s Docker components to avoid issues such as daemon deadlocks. Alloy also needs no change to the Docker daemon configuration.

Is it safe to mount the Docker socket into a monitoring container?

It is a real risk, because access to the socket is effectively root access to the host. Mount it read-only, use only the official pinned image, keep the container off public networks and avoid exposing its ports.

How do I get alert notifications from this stack?

Prometheus evaluates the alert rules, but you need an Alertmanager or Grafana alerting with a contact point such as email or Slack to send notifications. Start with a few alerts that always need human action, and test your rules with promtool test rules.

Snehasish Konger
Developed @scientyficworld.org | Technical writer @Nected | Content Developer
Connect with Snehasish Konger

On This page

Take a Pause with Intervals

A Sunday letter on building, writing, and thinking deeper as a developer — short, honest, and worth your time.

Snehasish Konger profile photo

"Hey there — I'm Snehasish. Hope this post saved you some head-scratching time! I've spent years turning technical chaos into clarity, and I'm here to be your guide through the maze of modern tech. Stick around for more lightbulb moments — we're just getting started."

Related Posts