Docker monitors
Monitor a Docker host and every container running on it.
Overview
A Docker monitor watches a Docker host. Every container on that host is monitored, so you add one monitor per host, not one per container. The monitor reports how many containers are running. Each container gets its own detail page with resource charts and, if you turn it on, recent log lines.
Checkmate reaches the host in one of three ways: the local Docker socket, a remote daemon over TLS or a Capture endpoint.
The monitor's up/down status reflects whether the Docker host answers, not the state of its containers. A host with every container stopped is still up if the daemon responds. Container problems surface through container alerts, which put the monitor into the breached state.
Creating a Docker monitor
- Navigate to Docker in the sidebar
- Click Create a monitor
- Select Docker as the monitor type
- Enter the Docker host
- Set the check interval
- Click Create monitor
Connecting to a host
The Docker host field takes one of three values.
Local socket — unix:///var/run/docker.sock, or a bare absolute path like /var/run/docker.sock, reads the daemon on the machine the Checkmate server runs on. If Checkmate is itself running in Docker, mount the host's socket into the container, otherwise the monitor cannot see anything:
volumes:
- /var/run/docker.sock:/var/run/docker.sock:roRead-only is enough. Checkmate only reads: it pings the daemon, lists and inspects containers, then reads stats and logs.
The shipped docker/docker-compose.yaml does not mount the Docker socket. If you run Checkmate from that file and want a local-socket monitor, add the volume above yourself.
Remote daemon over TLS — tcp://host:2376 or https://host:2376 with client certificates. Port 2376 is assumed when you leave it off. Plain unencrypted TCP is not possible: a tcp:// host is always upgraded to TLS and requires a client key. See TLS credentials below.
Capture endpoint — a Capture URL ending in /metrics/docker, with its authorization secret. Use this when Capture already runs on the machine.
SSH hosts (ssh://) are not supported.
TLS credentials
For a remote daemon, paste the PEM files you would pass to docker --tlsverify:
- CA certificate — the authority that signed the daemon's certificate. It can be a multi-certificate bundle. Optional when Ignore TLS/SSL errors is on.
- Client certificate — the certificate the daemon checks
- Client key — required when you create the monitor. Leaving it blank when editing keeps the stored key, which is never shown again.
Checkmate rejects the key if it does not match the certificate.
Ignore TLS/SSL errors skips verifying the daemon's certificate, which makes the CA certificate optional. The client certificate and key are still required, because the daemon still verifies the client.
Passphrase-protected keys are rejected. Remove the passphrase before pasting the key.
ENCRYPTION_KEY is required
Client keys are encrypted before they are stored, so the server needs ENCRYPTION_KEY set. Generate one with openssl rand -base64 32 and set the same value for both the API and the worker, then restart.
Without it, saving a TLS monitor fails. If the key is removed or rotated away after the fact, the monitor's checks fail with "Docker TLS key could not be decrypted" rather than a validation error, so the monitor goes down with no obvious cause.
Container alerts
Two opt-in settings decide when containers affect the host's status:
- Alert when a container is stopped
- Alert when a container is unhealthy
While any container breaches one of the enabled rules, the host is marked breached and notifications fire on the monitor's channels. The alert message names the container and the reason, including the exit code where there is one.
A container counts as stopped when Docker reports it as dead, or as exited with a non-zero exit code. Containers that exit cleanly with code 0 are treated as finished one-shot jobs, not failures.
A container counts as unhealthy when its Docker healthcheck reports unhealthy. Containers with no healthcheck defined never trigger this rule.
Capture does not report exit codes. On a Capture host, every stopped container counts as stopped, including one-shot containers that exited cleanly.
Container logs
Collect container logs stores the most recent output from every container on the host on each check. Logs are off by default.
- Up to 200 lines per container are read on each check
- Lines longer than 4096 bytes are truncated, with a marker showing where
- Logs are kept for 7 days, then deleted automatically. The period is fixed.
- Both
stdoutandstderrare captured
Only output written after you turn collection on is stored, so the Logs tab stays empty until the next check runs. Where a check returns a full 200 lines, Checkmate marks a gap, because more lines were likely written between checks than it could read. A chatty container on a long interval shows gaps often.
Log collection costs database space. A single container can write 200 lines of 4 KB per check, for every container on the host. Leave it off for hosts with chatty containers. Only administrators can view container logs.
Log collection is unavailable on Capture hosts.
Container states
Docker reports one of seven states: created, running, paused, restarting, removing, exited and dead. Health, when a container defines a healthcheck, is healthy, unhealthy, starting or none.
CPU and memory are only collected for running containers, so stopped containers report zero.
Container history is keyed by container name, not ID. Renaming a container starts its history over. Recreating a container under an old name continues the old one.