cAdvisor on the Flatcar Docker hosts, for per-container metrics (DWA-65)
Find a file
M. G. (Michael) de Bruin a095bf359a Add the role: cAdvisor on the Flatcar Docker hosts
bix, mara and inara run Docker outside Kubernetes and export nothing
about it -- no container_* or engine_daemon_* series exist for them, so
"what runs there, and is it still running" could only be answered by
logging in (DWA-65).

A systemd unit wrapping docker run, because Flatcar has no package
manager and no compose, and systemd already supervises everything else
on these hosts. The image is pinned by digest; the tag sits beside it
for readability only.

Deliberately does not touch daemon.json. Docker's own metrics would add
engine_daemon_container_states_containers, the one signal that tells
stopped from absent -- but enabling it restarts Docker and every
container on the host. That is a maintenance window, not a side effect
of deploying a monitoring role.

Defaults chosen against that: port 9101 rather than cAdvisor's 8080,
bound to the host's primary address rather than 0.0.0.0, container
labels off (a compose stack can carry a dozen, and label cardinality is
how a metrics store grows without anyone deciding to), and --docker_only
because the raw cgroup hierarchy roughly triples the series.

The cost worth knowing: cAdvisor runs --privileged, which is what
upstream documents for Docker. On a host facing the DMZ that is a real
trade, mitigated by the single bind address and by mercury passing
nothing inbound -- checked against its config rather than assumed.
2026-09-21 16:53:38 +00:00
defaults Add the role: cAdvisor on the Flatcar Docker hosts 2026-09-21 16:53:38 +00:00
handlers Add the role: cAdvisor on the Flatcar Docker hosts 2026-09-21 16:53:38 +00:00
meta Add the role: cAdvisor on the Flatcar Docker hosts 2026-09-21 16:53:38 +00:00
tasks Add the role: cAdvisor on the Flatcar Docker hosts 2026-09-21 16:53:38 +00:00
templates Add the role: cAdvisor on the Flatcar Docker hosts 2026-09-21 16:53:38 +00:00
tests Add the role: cAdvisor on the Flatcar Docker hosts 2026-09-21 16:53:38 +00:00
.ansible-lint Add the role: cAdvisor on the Flatcar Docker hosts 2026-09-21 16:53:38 +00:00
.gitignore Add the role: cAdvisor on the Flatcar Docker hosts 2026-09-21 16:53:38 +00:00
.yamllint Add the role: cAdvisor on the Flatcar Docker hosts 2026-09-21 16:53:38 +00:00
mise.toml Add the role: cAdvisor on the Flatcar Docker hosts 2026-09-21 16:53:38 +00:00
README.md Add the role: cAdvisor on the Flatcar Docker hosts 2026-09-21 16:53:38 +00:00

ansible_role_container_metrics

cAdvisor on the Flatcar Docker hosts — bix, mara and inara — so the containers they run stop being invisible (DWA-65).

Those hosts run Docker outside Kubernetes. Nothing exported container_* or engine_daemon_* for them, so "what is running there, and is it still running" could only be answered by logging in.

Why a systemd unit around docker run

Flatcar has no package manager, and no compose by default. systemd already supervises everything else on these hosts, so a unit is the mechanism with the fewest new moving parts. The image is pinned by digest; the tag sits beside it for readability only, so a moved tag cannot change what runs.

What it deliberately does not do

It does not touch daemon.json. Docker's own metrics (metrics-addr) would add engine_daemon_container_states_containers, which is the one signal that distinguishes stopped from absent. Enabling it requires restarting Docker, and that restarts every container on the host — a maintenance window, not a side effect of deploying a monitoring role.

That gap matters when writing alerts: a container that stops being reported by cAdvisor produces silence, not a zero. Alerting on "memory high" or "restarted" will never fire for a container that simply vanished. Until daemon metrics exist, the positive signal has to come from an expected-containers list and absent().

Variables

variable default notes
container_metrics_port 9101 not cAdvisor's 8080, which applications often want
container_metrics_bind primary IPv4 not 0.0.0.0; bix faces the DMZ
container_metrics_store_labels false container labels become Prometheus labels; a compose stack can carry a dozen
container_metrics_docker_only true otherwise the raw cgroup hierarchy roughly triples the series

The cost worth knowing

cAdvisor runs --privileged, which is what upstream documents for Docker: it reads cgroups, the Docker socket and block device state. On a host that faces the DMZ that is a real trade. It is mitigated by binding the port to one address, and by mercury passing nothing inbound to these hosts — checked against its config, not assumed.