Part 13: Running It Calmly
Part 5 ended with a problem it did not solve. The metrics stack was running when the disk filled to 97%, the graph was correct, and nobody looked at it. Graphs answer questions you think to ask. This post covers the layer that tells you things you did not ask about.
Three pieces do that work: searchable logs through Loki, a report-only update watcher, and one morning message that collects the rest. A fourth idea holds them together — deciding which work happens at which time of day.
Logs: Loki and Alloy
Metrics tell you a container is using memory. Logs tell you what it was doing. Until this point the answer to “why did that restart?” was docker logs over SSH, one container at a time, with nothing retained after a container was recreated.
Loki stores the logs and Grafana queries them, so the existing dashboard stays the single interface. Grafana Alloy collects them: it discovers containers through the Docker API, reads their log streams, and ships them with a small set of labels.
That label set is the design decision worth copying:
host=forge-docker
container
compose_project
compose_service
Four labels, deliberately. Loki indexes labels, not log bodies, so every label value multiplies the index. Request IDs, filenames, user IDs and URLs stay in the log body where a query can still search them. Promote them to labels and the index grows faster than the logs do. Retention is 14 days on the local filesystem, which suits one host.
Alloy’s Docker access is read-only on the socket path, with a caveat recorded in the build notes: mounting a socket :ro protects the file, not the API behind it. The collector has no update or command configuration, so it can only read. A restricted socket proxy is the stronger version of that boundary, and it remains a later hardening pass.
What Broke: The Capability That Was Not There
The first Alloy deployment went into a crash loop, and the reason is a good lesson in hardening by habit.
The container config started from the usual secure baseline: drop every Linux capability, run unprivileged. Alloy then failed to traverse its own /var/lib/alloy directory, which the image creates owned by the alloy user.
The subtlety is that cap_drop: ALL strips DAC_OVERRIDE, and DAC_OVERRIDE is the capability that lets root ignore file permission checks. Running as root is normally enough to read any path — that privilege is a capability, and dropping all of them takes it away. The result is a root process that cannot enter a directory it owns, which reads as a nonsense permission error.
The fix grants exactly that one capability back:
alloy:
image: grafana/alloy:v1.18.1
# Root plus DAC_OVERRIDE so Alloy can traverse the image's
# alloy-owned directories, write its data volume and read the
# Docker socket; all other capabilities stay dropped.
user: "0:0"
cap_add:
- DAC_OVERRIDE
volumes:
- ./alloy/config.alloy:/etc/alloy/config.alloy:ro
- /var/run/docker.sock:/var/run/docker.sock:ro
- alloy-data:/var/lib/alloy/data
The comment is in the tracked Compose file, not just in my memory, because a future reader who deletes that cap_add to “tighten security” reproduces the crash loop exactly.
One more deployment finding from the same session: a stale pve-exporter container refused to be adopted by the new stack and had to be removed once with docker rm -f. It was stateless, so the next run recreated it.
Update Watching: WUD, and Why It Only Reports
Part 9 covered Renovate proposing dependency updates as pull requests. WUD — What’s Up Docker — answers a different question: of the containers running right now, which have newer images published?
The first rollout is deliberately report-only. WUD has no trigger configured that can change a workload — no Docker action, no Compose action, no command, no webhook that mutates anything. It can see, and it cannot act.
The first scan proved why that restraint was right. Unconstrained tag searches proposed a PostgreSQL beta, a 32-bit Redis image, rootless image variants, and Immich release candidates. Every one of those is technically “a newer tag”. None was an upgrade I wanted. The fix was per-container tag filters, declared as labels:
labels:
- 'wud.tag.include=^v1\.\d+\.\d+$$'
That pattern accepts v1.18.1 and rejects release candidates, betas and architecture variants. The doubled $$ is Compose escaping a literal $ so the regex anchor survives variable interpolation.
WUD and Renovate disagree regularly, and the build notes explain why in a way worth repeating. WUD reports the instant a newer tag appears in a registry. Renovate only proposes what it will actually stage as a pull request, after a three-day minimum release age and a memory of pull requests you closed without merging. So an image can look days old in WUD while Renovate waits on a point release cut yesterday. The authoritative view for applications is Renovate’s dependency dashboard. WUD’s view is the running runtime.
The combination closed the loop from Part 9: Renovate proposes, a human merges, the allowlisted workflow deploys, and WUD confirms the running version changed.
One Morning Message
Five sources of truth is four too many to check before coffee. The maintenance digest is an n8n workflow that collects them into a single Telegram message each morning:
- WUD update notifications, received by a dedicated webhook and stored for the next digest
- Open Renovate pull requests in the approved repository allowlist
- Open GitHub pull requests from an explicit allowlist
- Failed CI runs
- Platform health and backup freshness from the existing status checks
Like WUD, the digest is read-only by design. The GitHub token is fine-grained, restricted to named repositories, and granted only metadata, contents, pull-request and actions read access. The workflow must not approve, merge, close, comment on, label or modify anything. Its job is to tell me what happened, and the original links stay the source of truth even when a summary is generated.
The WUD webhook gets a boundary of its own: a long bearer secret, requests accepted from the Docker network, and a single job of validating and recording the payload. It returns promptly and never calls the Docker or Portainer APIs.
This is the piece that answers Part 5’s complaint. The disk-usage graph still exists, and now a daily message arrives whether or not I think to open Grafana.
Quiet Hours
The last idea has no software at all. The build notes split platform work into two lanes:
- Quiet hours: long storage copies, checksum verification, media imports, backups, SMART tests, and controlled disk reuse.
- Daytime: dashboards, monitoring, read-only research, maintenance reports, workflow drafting, and anything that does not compete for storage I/O.
The reason is the shared ZFS mirror from Part 3. A 44GB photo import and a backup verification pass are both reasonable jobs that become unreasonable when they run at the same time as a music stream. Scheduling them into a lane costs nothing and removes a whole category of “why is everything slow” investigation.
The scheduled jobs follow the same split. The image prune from Part 4 runs on a Sunday morning. WUD scans at 06:00, before the digest is built. The Forgejo backup runs every six hours, including through the night.
What This Stage Delivered
The platform now reports on itself. Logs are searchable across every container from one query box, with a label set that will not explode. Update information arrives as advice rather than as automatic change. A single morning message covers health, backups, updates and pull requests, and it arrives without anyone opening a dashboard.
What is still open is alerting on the layer that does the alerting: Loki has retention but no alert for unexpectedly rapid volume growth, and its filesystem backend will not delete data to free space. Retention plus a disk alert are both required, and only the first exists.
Next in the Series
Part 14 covers the services that deliberately get no public hostname at all. Pi-hole on its own LXC, Home Assistant as a VM, Time Machine over Samba, and a bridge that puts the music library into a Sonos system from the previous decade. Part 8 explained how to publish a service safely. The next post is about deciding not to.