deprioritize recently stocked-out Google flex selections

NOTE THAT THIS FORK IS MAINTAINED FOR CRITICAL BUG FIXES AFFECTING RUNNING COSTS ONLY. NO OTHER CONTRIBUTIONS WILL BE ACCEPTED.

What critical bug is this MR fixing?

The Google BulkInsert path retries configured placement selections independently for every machine create. A recognized stockout correctly advances the current create to the next selection, but the driver forgets that result when the command exits. Concurrent and subsequent creates therefore retry the same recently constrained first selection and repeatedly pay its operation latency before reaching a viable fallback.

This MR adds an optional, short-lived placement cooldown shared by Docker Machine command processes on the same host. After an explicit stockout-class failure, subsequent creates temporarily try that exact selection after non-cooling alternatives, so it remains available if those alternatives fail. After cooldown expiry, one process receives a priority recovery probe at the selection's configured position; placement success clears the cooldown, while another stockout reopens it.

The feature is disabled by default. With the default configuration the selection order and create path are unchanged.

How does this change help reduce cost of usage? What scale of cost reduction is it?

The change reduces repeated provider operations that are unlikely to produce a VM while a placement class is constrained. It is intended to improve useful capacity produced per creation-slot minute, reducing paid and operational overhead associated with delayed replenishment and prolonged queueing.

The production effect is not claimed in advance. A released binary will be canaried against unchanged managers. Keep the change only if it improves ready VM yield and intent-to-ready latency without increasing quota errors, create failures, or cleanup anomalies.

In what scenarios is this change usable with GitLab Runner's docker+machine executor?

The feature applies only when the Google driver uses regional BulkInsert with multiple --google-flex-selection values and receives one of the driver's existing recognized capacity-error codes or API reasons.

Recommended initial canary configuration:

--google-flex-stockout-cooldown=120s
--google-flex-stockout-probe-lease=5m

--google-flex-stockout-cooldown=0s disables the feature and preserves current behavior. The cooldown is configurable rather than hard-coded; 120 seconds is the initial canary value. The probe lease must cover the configured provider-operation timeout so another process cannot claim the same priority probe while the first provider call may still be active.

The state is manager-local. It is stored at the global Docker Machine store level so independent create command processes can share it. Missing, corrupt, expired, unknown-version, or temporarily locked state fails open to the operator-configured selection order.

Behavior

  • Only the existing conservative stockout classification opens a cooldown.
  • Quota, authentication, configuration, and generic backend errors remain fatal and do not update placement health.
  • Operator order is preserved among eligible selections.
  • If all selections are cooling, the driver preserves current behavior and makes one configured ladder pass for that create.
  • If another selection is healthy, one process may priority-probe one expired class at its configured position. Other processes keep it behind healthy alternatives during the lease, but retain it as fallback if those alternatives fail.
  • A successful BulkInsert placement closes the cooldown even if later SSH or provisioning fails; post-placement failure is not a placement stockout.
  • The cross-process lock is bounded and is never held during a provider call.
  • State writes use atomic replacement and contain no credentials, metadata, customer identity, or provider response payloads.

Compatibility and scope

  • Direct Instances.Insert is unchanged.
  • BulkInsert zone selection, per-selection disk configuration, post-create setup, resolved placement fields, and cleanup are unchanged.
  • Existing configurations remain unchanged because the new cooldown defaults to disabled.
  • This is intentionally a narrow change in GitLab's existing maintained Docker Machine repository. It does not introduce another fork, a service, fleet-shared state, dynamic scoring, speculative creates, or batching.

Testing

Targeted tests cover disabled behavior, stockout ordering, non-stockout errors, all-cooling fallback, a single concurrent probe lease, placement recovery, post-placement failures, missing/corrupt/oversized/unknown-version state, future-version preservation, lock contention, atomic concurrent readers/writers, no lost updates, bounded state, normalized placement keys, absent StorePath, BulkInsert loop integration, and Windows cross-compilation.

Validation:

go test -race ./drivers/google
go vet ./drivers/google
go build ./...
GOOS=windows GOARCH=amd64 go test -c ./drivers/google

Full make validate still reports existing drivers/vmwarefusion vet failures unrelated to this change. The changed package passes vet.

Rollout

The code has no effect until a released binary is pinned and the cooldown is enabled.

  1. Publish an immutable Docker Machine artifact with the feature disabled by default.
  2. Pin it on a low-throughput manager cohort and enable the cooldown for lifecycle, state, cleanup, and rollback safety.
  3. Enable it on one bounded higher-volume cohort with unchanged siblings as controls.
  4. Compare ready VMs per creation-slot minute, intent-to-ready p95, no_free_executor, create outcomes, quota/throttle errors, and orphan cleanup.
  5. Disable the option or restore the prior binary pin on any regression.

The confidential internal tracking issue contains fleet-specific evidence and detailed canary thresholds. This public MR intentionally omits production topology, capacity values, customer impact, and live provider-availability details.

Follow-up excluded from this MR

BulkInsert microbatching is a separate experiment. It requires per-instance partial-success, idempotency, late-success, restart, and cleanup semantics and is not bundled with this cooldown change.

Edited by Kam Kyrala

Merge request reports

Loading
Loading