Add the role: cAdvisor on the Flatcar Docker hosts #1

Merged
mgbruin merged 1 commit from add-cadvisor into main 2026-09-21 16:55:41 +00:00
Owner

First half of DWA-65. bix, mara and inara run Docker outside Kubernetes and export nothing about it — no container_* or engine_daemon_* series exist for them, so "what runs there, and is it still running" can only be answered by logging in.

Shape

A systemd unit wrapping docker run. Flatcar has no package manager and no compose by default, and systemd already supervises everything else on these hosts — so this adds the fewest new moving parts. The image is pinned by digest (cAdvisor v0.55.1); the tag sits beside it for readability only, so a moved tag cannot change what runs.

Defaults, and why they are not cAdvisor's

variable default why
container_metrics_port 9101 not 8080, which applications on a container host often want
container_metrics_bind primary IPv4 not 0.0.0.0; bix faces the DMZ
container_metrics_store_labels false container labels become Prometheus labels, and a compose stack can carry a dozen
container_metrics_docker_only true the raw cgroup hierarchy roughly triples the series

What it deliberately does not do

It does not touch daemon.json. Docker's own metrics would add engine_daemon_container_states_containers — the one signal that distinguishes stopped from absent. Enabling it restarts Docker and every container on the host: a maintenance window, not a side effect of deploying a monitoring role.

That gap shapes the alerting that follows. A container that stops being reported by cAdvisor produces silence, not a zero, so "memory high" or "restarted" alerts will never fire for one that simply vanished. Until daemon metrics exist, the positive signal must come from an expected-containers list and absent().

The cost worth knowing

cAdvisor runs --privileged, which is what upstream documents for the Docker setup: it reads cgroups, the Docker socket and block device state. On a host facing the DMZ that is a genuine trade. Mitigations: the port binds to one address, and mercury passes nothing inbound to these hosts — checked against its config backup (WAN rules pass only lysithea's mail/DNS/SSH, inara 80/443, plex and wireguard), not assumed.

Verified

  • ansible-lint at the production profile: passed, 0 failures
  • yamllint: clean
  • mise run check:templates: the unit carries ansible_managed and role_name, and none of the per-run template variables that break idempotence
  • ansible-playbook --syntax-check: passes

⚠️ river's pinned ansible-lint and yamllint are broken — both pipx installs fail with ModuleNotFoundError, in this repo and in ansible_role_monitoring alike, so mise run lint cannot currently pass on river for any role. I ran the same pinned versions through uvx instead. Worth fixing separately.

Note on main

The repository was created empty, so main holds a placeholder README committed directly — repository initialisation has to start somewhere. Everything substantive is in this PR.

Not in this PR

The play targeting the flatcar group, and the Prometheus scrape config, alerts and dashboard on the fennec side. Those follow once this is tagged.

First half of DWA-65. bix, mara and inara run Docker outside Kubernetes and export **nothing** about it — no `container_*` or `engine_daemon_*` series exist for them, so "what runs there, and is it still running" can only be answered by logging in. ## Shape A systemd unit wrapping `docker run`. Flatcar has no package manager and no compose by default, and systemd already supervises everything else on these hosts — so this adds the fewest new moving parts. The image is pinned **by digest** (cAdvisor v0.55.1); the tag sits beside it for readability only, so a moved tag cannot change what runs. ## Defaults, and why they are not cAdvisor's | variable | default | why | |---|---|---| | `container_metrics_port` | `9101` | not 8080, which applications on a container host often want | | `container_metrics_bind` | primary IPv4 | not `0.0.0.0`; bix faces the DMZ | | `container_metrics_store_labels` | `false` | container labels become Prometheus labels, and a compose stack can carry a dozen | | `container_metrics_docker_only` | `true` | the raw cgroup hierarchy roughly triples the series | ## What it deliberately does not do **It does not touch `daemon.json`.** Docker's own metrics would add `engine_daemon_container_states_containers` — the one signal that distinguishes *stopped* from *absent*. Enabling it restarts Docker and every container on the host: a maintenance window, not a side effect of deploying a monitoring role. That gap shapes the alerting that follows. A container that stops being reported by cAdvisor produces **silence, not a zero**, so "memory high" or "restarted" alerts will never fire for one that simply vanished. Until daemon metrics exist, the positive signal must come from an expected-containers list and `absent()`. ## The cost worth knowing cAdvisor runs `--privileged`, which is what upstream documents for the Docker setup: it reads cgroups, the Docker socket and block device state. On a host facing the DMZ that is a genuine trade. Mitigations: the port binds to one address, and mercury passes nothing inbound to these hosts — checked against its config backup (WAN rules pass only lysithea's mail/DNS/SSH, inara 80/443, plex and wireguard), not assumed. ## Verified - `ansible-lint` at the **production** profile: passed, 0 failures - `yamllint`: clean - `mise run check:templates`: the unit carries `ansible_managed` and `role_name`, and none of the per-run template variables that break idempotence - `ansible-playbook --syntax-check`: passes ⚠️ **river's pinned `ansible-lint` and `yamllint` are broken** — both pipx installs fail with `ModuleNotFoundError`, in this repo and in `ansible_role_monitoring` alike, so `mise run lint` cannot currently pass on river for any role. I ran the same pinned versions through `uvx` instead. Worth fixing separately. ## Note on main The repository was created empty, so `main` holds a placeholder README committed directly — repository initialisation has to start somewhere. Everything substantive is in this PR. ## Not in this PR The play targeting the `flatcar` group, and the Prometheus scrape config, alerts and dashboard on the fennec side. Those follow once this is tagged.
mgbruin force-pushed add-cadvisor from 905a486370 to a095bf359a 2026-09-21 16:53:39 +00:00 Compare
Sign in to join this conversation.
No reviewers
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
mgbruin/ansible_role_container_metrics!1
No description provided.