- Jinja 100%
bix, mara and inara run Docker outside Kubernetes and export nothing about it -- no container_* or engine_daemon_* series exist for them, so "what runs there, and is it still running" could only be answered by logging in (DWA-65). A systemd unit wrapping docker run, because Flatcar has no package manager and no compose, and systemd already supervises everything else on these hosts. The image is pinned by digest; the tag sits beside it for readability only. Deliberately does not touch daemon.json. Docker's own metrics would add engine_daemon_container_states_containers, the one signal that tells stopped from absent -- but enabling it restarts Docker and every container on the host. That is a maintenance window, not a side effect of deploying a monitoring role. Defaults chosen against that: port 9101 rather than cAdvisor's 8080, bound to the host's primary address rather than 0.0.0.0, container labels off (a compose stack can carry a dozen, and label cardinality is how a metrics store grows without anyone deciding to), and --docker_only because the raw cgroup hierarchy roughly triples the series. The cost worth knowing: cAdvisor runs --privileged, which is what upstream documents for Docker. On a host facing the DMZ that is a real trade, mitigated by the single bind address and by mercury passing nothing inbound -- checked against its config rather than assumed. |
||
|---|---|---|
| defaults | ||
| handlers | ||
| meta | ||
| tasks | ||
| templates | ||
| tests | ||
| .ansible-lint | ||
| .gitignore | ||
| .yamllint | ||
| mise.toml | ||
| README.md | ||
ansible_role_container_metrics
cAdvisor on the Flatcar Docker hosts — bix, mara and inara — so the containers they run stop being invisible (DWA-65).
Those hosts run Docker outside Kubernetes. Nothing exported container_* or
engine_daemon_* for them, so "what is running there, and is it still running"
could only be answered by logging in.
Why a systemd unit around docker run
Flatcar has no package manager, and no compose by default. systemd already supervises everything else on these hosts, so a unit is the mechanism with the fewest new moving parts. The image is pinned by digest; the tag sits beside it for readability only, so a moved tag cannot change what runs.
What it deliberately does not do
It does not touch daemon.json. Docker's own metrics (metrics-addr) would
add engine_daemon_container_states_containers, which is the one signal that
distinguishes stopped from absent. Enabling it requires restarting Docker,
and that restarts every container on the host — a maintenance window, not a
side effect of deploying a monitoring role.
That gap matters when writing alerts: a container that stops being reported by
cAdvisor produces silence, not a zero. Alerting on "memory high" or
"restarted" will never fire for a container that simply vanished. Until daemon
metrics exist, the positive signal has to come from an expected-containers list
and absent().
Variables
| variable | default | notes |
|---|---|---|
container_metrics_port |
9101 |
not cAdvisor's 8080, which applications often want |
container_metrics_bind |
primary IPv4 | not 0.0.0.0; bix faces the DMZ |
container_metrics_store_labels |
false |
container labels become Prometheus labels; a compose stack can carry a dozen |
container_metrics_docker_only |
true |
otherwise the raw cgroup hierarchy roughly triples the series |
The cost worth knowing
cAdvisor runs --privileged, which is what upstream documents for Docker: it
reads cgroups, the Docker socket and block device state. On a host that faces
the DMZ that is a real trade. It is mitigated by binding the port to one
address, and by mercury passing nothing inbound to these hosts — checked
against its config, not assumed.