fix(gitlab-runner): drain GPU K8s managers before the readiness-gate fix
What
Scale the saas-linux-medium-amd64-gpu-standard K8s runner-manager deployments on runner-managers-gprd-1 back to zero replicas.
Why
After the soft-start (!3356 (merged)) the 9 GPU managers restart-loop. boot_verify fails on every worker because the google-cos-docker-network-readiness-gate probe requests the metadata server on :80, which the worker DOCKER-USER firewall rule drops by design. The check can never pass on a correctly-firewalled worker, so every machine creation fails and no manager reaches Ready. Full diagnosis in confidential production-engineering#29697.
Why zero and not a gate revert
The gate flag enables two things in docker-machine: the cloud-init wait (the actual fix for the GPU-driver race in production-engineering#29677) and the broken bridge-network probe. Disabling the gate while the shard is at replicas: 3 would re-expose the #29677 race on newly provisioned workers. Draining to zero first removes any worker that could hit that race, and Chef carries GPU capacity meanwhile.
This clears the way to fix the probe in docker-machine (revert the :80 metadata target to the static NAT/FORWARD rules check) and re-enable the gate without provisioning workers against a broken check.
Rollback
Set replicas back to a positive value. This is a drain, not a config change.