Add the role: cAdvisor on the Flatcar Docker hosts #1
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "add-cadvisor"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
First half of DWA-65. bix, mara and inara run Docker outside Kubernetes and export nothing about it — no
container_*orengine_daemon_*series exist for them, so "what runs there, and is it still running" can only be answered by logging in.Shape
A systemd unit wrapping
docker run. Flatcar has no package manager and no compose by default, and systemd already supervises everything else on these hosts — so this adds the fewest new moving parts. The image is pinned by digest (cAdvisor v0.55.1); the tag sits beside it for readability only, so a moved tag cannot change what runs.Defaults, and why they are not cAdvisor's
container_metrics_port9101container_metrics_bind0.0.0.0; bix faces the DMZcontainer_metrics_store_labelsfalsecontainer_metrics_docker_onlytrueWhat it deliberately does not do
It does not touch
daemon.json. Docker's own metrics would addengine_daemon_container_states_containers— the one signal that distinguishes stopped from absent. Enabling it restarts Docker and every container on the host: a maintenance window, not a side effect of deploying a monitoring role.That gap shapes the alerting that follows. A container that stops being reported by cAdvisor produces silence, not a zero, so "memory high" or "restarted" alerts will never fire for one that simply vanished. Until daemon metrics exist, the positive signal must come from an expected-containers list and
absent().The cost worth knowing
cAdvisor runs
--privileged, which is what upstream documents for the Docker setup: it reads cgroups, the Docker socket and block device state. On a host facing the DMZ that is a genuine trade. Mitigations: the port binds to one address, and mercury passes nothing inbound to these hosts — checked against its config backup (WAN rules pass only lysithea's mail/DNS/SSH, inara 80/443, plex and wireguard), not assumed.Verified
ansible-lintat the production profile: passed, 0 failuresyamllint: cleanmise run check:templates: the unit carriesansible_managedandrole_name, and none of the per-run template variables that break idempotenceansible-playbook --syntax-check: passes⚠️ river's pinned
ansible-lintandyamllintare broken — both pipx installs fail withModuleNotFoundError, in this repo and inansible_role_monitoringalike, somise run lintcannot currently pass on river for any role. I ran the same pinned versions throughuvxinstead. Worth fixing separately.Note on main
The repository was created empty, so
mainholds a placeholder README committed directly — repository initialisation has to start somewhere. Everything substantive is in this PR.Not in this PR
The play targeting the
flatcargroup, and the Prometheus scrape config, alerts and dashboard on the fennec side. Those follow once this is tagged.905a486370toa095bf359a