Docker+machine: refresh a liveness label on tracked machines

What does this MR do?

docker+machine managers stamp a runner_manager_heartbeat label (unix seconds) on every machine they track. Two new settings under [runners.machine]:

  • HeartbeatInterval: seconds between refreshes per machine. 0 disables the feature (default).
  • [machine.heartbeat] concurrency: process-global bound on parallel label writes (default 5), in the global [machine] section like shutdown_drain.

Beats ride the existing updateMachines pass. Each machine gets a random initial offset so writes spread over the interval instead of bursting against the cloud provider's operation quota. An attempt counts as a beat whether it succeeds or fails, so unreachable machines retry next interval instead of hot-looping. Machines in creating or removing states are skipped.

The label write itself is the new docker-machine update-labels command from gitlab-org/ci-cd/docker-machine!190 (merged).

Why is this MR needed?

When a manager dies uncleanly, the VMs it tracked keep running with nothing to remove them. In INC-13578 (GitLab.com, 2026-08-27) crash-looping managers orphaned ~8,000 VMs and exhausted two GCP quotas. External cleanup can't use instance age to find orphans because machine reuse (MaxBuilds) makes multi-day lifetimes legitimate. A liveness label measures the property directly: a stale runner_manager_heartbeat means no live manager tracks the instance, whatever the cause.

Design and corrective-action context: gitlab-com/gl-infra/production-engineering#29652 (closed)

What's the best way to test this MR?

Unit tests cover due/not-due, disabled interval, failure handling, and first-beat spreading. Validated end to end in a production-shape GKE sandbox with real GCE VMs: labels appeared on all machines spread over the configured interval, refreshed on schedule, a machine removed from the store had its label freeze while the tracked fleet kept advancing, and the replacement machine was labeled correctly. Label writes took about a second each.

With HeartbeatInterval unset the feature is inert, and on a docker-machine binary without update-labels each attempt logs a warning once per interval per machine.

Edited by Igor

Merge request reports

Loading
Loading