Add a separate COS readiness URL egress check

What

Split the Google COS readiness gate into two independent opt-ins:

  • --google-cos-docker-network-readiness-gate (existing) now waits for cloud-init only.
  • --google-cos-docker-network-readiness-url (new) runs the container egress probe against the given URL, and only when set.

Why

!183 (merged) bundled two fixes under one flag: the cloud-init wait (workers racing the NVIDIA driver install) and a container egress probe (workers where egress was dead despite healthy-looking Docker). The probe fetched the metadata server on :80 from a container, which is exactly what the worker's own DOCKER-USER rule drops to keep job containers away from service-account credentials. So the probe could never pass on a correctly-firewalled worker, and every docker-machine create failed. On the K8s GPU shard that meant every manager restart-looped and served no jobs. Diagnosis: production-engineering#29697.

The probe has to test real egress, not just rule presence (a static iptables -C was the thing !183 (merged) moved away from, since rules can be present while egress is dead). Testing real egress means reaching a destination the firewall permits. Two options don't work: the metadata server is blocked by design, and a DNS probe to it hits the :53 ACCEPT in DOCKER-USER before Docker's own egress rule, so it passes even when that rule is missing. A configurable URL goes off-bridge through MASQUERADE, falls through to the docker0 FORWARD egress rule, and isn't the blocked address. For the GPU shard the target is https://gitlab.com, which a worker must reach to run a job anyway.

Behaviour

  • Success means an HTTP response came back. wget -S output is matched for an HTTP/ line, so a 301 or 503 from a reachable target still passes; only a blackout fails. This keeps a transient non-2xx on the target from fail-closing provisioning fleet-wide.
  • The one repair-restart and fail-closed cleanup from !183 (merged) are unchanged.
  • The URL is validated driver-side (http/https, host present, no shell metacharacters) before it reaches the worker shell command.

Rollout

The K8s GPU managers are drained to zero in argocd/apps!3364 pending this. Once a new tag ships and the manager image picks it up, the shard re-enables the gate for the cloud-init wait, sets the readiness URL to https://gitlab.com, and scales back up.

Rollback

Unset the readiness URL. The gate flag stays on independently for the cloud-init wait.

Note for the Chef fleet

Chef pins its own docker-machine version in roles/runners-manager.json, so this k8s change doesn't touch it. When Chef is bumped to a version with this change, set --google-cos-docker-network-readiness-url=https://gitlab.com in the same bump. The Chef GPU role turns the gate on but sets no URL, and without the URL the egress probe stops running. Bump and URL together, so the check isn't lost.

Edited by Igor

Merge request reports

Loading
Loading