Docker+machine: refresh a liveness label on tracked machines
What does this MR do?
docker+machine managers stamp a runner_manager_heartbeat label (unix seconds) on every machine they track. Two new settings under [runners.machine]:
HeartbeatInterval: seconds between refreshes per machine. 0 disables the feature (default).[machine.heartbeat] concurrency: process-global bound on parallel label writes (default 5), in the global[machine]section likeshutdown_drain.
Beats ride the existing updateMachines pass. Each machine gets a random initial offset so writes spread over the interval instead of bursting against the cloud provider's operation quota. An attempt counts as a beat whether it succeeds or fails, so unreachable machines retry next interval instead of hot-looping. Machines in creating or removing states are skipped.
The label write itself is the new docker-machine update-labels command from gitlab-org/ci-cd/docker-machine!190 (merged).
Why is this MR needed?
When a manager dies uncleanly, the VMs it tracked keep running with nothing to remove them. In INC-13578 (GitLab.com, 2026-08-27) crash-looping managers orphaned ~8,000 VMs and exhausted two GCP quotas. External cleanup can't use instance age to find orphans because machine reuse (MaxBuilds) makes multi-day lifetimes legitimate. A liveness label measures the property directly: a stale runner_manager_heartbeat means no live manager tracks the instance, whatever the cause.
Design and corrective-action context: gitlab-com/gl-infra/production-engineering#29652 (closed)
What's the best way to test this MR?
Unit tests cover due/not-due, disabled interval, failure handling, and first-beat spreading. Validated end to end in a production-shape GKE sandbox with real GCE VMs: labels appeared on all machines spread over the configured interval, refreshed on schedule, a machine removed from the store had its label freeze while the tracked fleet kept advancing, and the replacement machine was labeled correctly. Label writes took about a second each.
With HeartbeatInterval unset the feature is inert, and on a docker-machine binary without update-labels each attempt logs a warning once per interval per machine.