google_cos: wait on cloud-init.target instead of cloud-init status --wait
What
--google-cos-docker-network-readiness-gate ran cloud-init status --wait over SSH before configuring Docker. On stock COS that command never exits 0, and SSHCommand fails on any non-zero exit. The gate now waits with systemctl start cloud-init.target and then runs cloud-init status once: exit 0 and 2 pass, exit 1 (errors recorded by cloud-init, e.g. a failed runcmd) and 124 (5 min timeout) fail the create.
Why --wait does not work on stock COS
--wait polls get_status_details() every 0.25 s and returns as soon as the status is anything other than running. Inside that function, cloud-init ≥ 24.1 has a check that runs while cloud-init is still in progress:
if running_status == RunningStatus.RUNNING and uses_systemd() and systemd_failed(wait=wait):
running_status = RunningStatus.DONE
condition_status = ConditionStatus.ERROR
description = "Failed due to systemd unit failure"systemd_failed() looks at the four cloud-init services and returns true for any that is not UnitFileState=enabled or static. The idea is that a disabled stage will never run, so stop waiting for it. On COS the services are symlinked into multi-user.target.wants, but their [Install] sections say WantedBy=cloud-init.target, so systemd reports them as disabled even though they run:
Id=cloud-init-local.service ActiveState=activating UnitFileState=disabled
Id=cloud-init.service ActiveState=inactive UnitFileState=disabled
Id=cloud-config.service ActiveState=inactive UnitFileState=disabled
Id=cloud-final.service ActiveState=inactive UnitFileState=disabledSo the first poll flips the status to error - done and --wait exits 1 within a second. docker-machine reaches the VM about 10 s after boot, so every stock create failed (last_update: 00:00:10, Machine creation failed ... time=58s in the sandbox). There is no stage after which --wait starts working: the check applies until result.json exists, and the units stay "disabled" for the whole boot. The check is unchanged on cloud-init main.
Once cloud-init is done, status exits 2 (degraded done) instead:
recoverable_errors:
WARNING:
- Getting data from <class 'cloudinit.sources.DataSourceGCE.DataSourceGCELocal'> failed File ".../cloudinit/sources/DataSourceGCE.py", line 131, in _get_data
UnboundLocalError: cannot access local variable 'ret' where it is not associated with a valueCOS has no dhclient or udhcpc, so DataSourceGCELocal skips every NIC in init-local and crashes on an unassigned variable. cloud-init falls back to DataSourceGCE and applies user-data. This happens before user-data is read, and the image's datasource_list: [GCE, NoCloud, None] cannot exclude GCELocal. The crash is fixed in cloud-init 25.1, but the no-lease path still logs a WARNING there, so exit 2 is to be expected on COS for the foreseeable future.
Both were checked on bare cos-125-lts and cos-stable (cos-121) VMs with no user-data, both cloud-init 24.4.1. The fleet image (COS 109, cloud-init 23.2.1) predates the systemd check and exits 0, which is why the gate has worked on gprd.
Why cloud-init.target
cloud-init documents it as the synchronization point for "all of cloud-init's initial system configuration tasks have completed" and tells units to order themselves with After=cloud-init.target. systemctl start on a target blocks until it is reached and returns immediately if it already was. On COS the target is After=multi-user.target, which is not reached until its oneshot wants, cloud-final included, have exited. It is reached even when cloud-final failed (Wants=, not Requires=), which is why cloud-init status still runs afterwards.
| image | cloud-init | systemctl start cloud-init.target |
cloud-init status after |
|---|---|---|---|
cos-125-lts, runcmd: sleep 150 |
24.4.1 | 0 after 149 s, cloud-final active |
2 |
| fleet COS 109, same | 23.2.1 | 0 after 153 s | 0 |
cos-125-lts, failing runcmd |
24.4.1 | 0 after 60 s, cloud-final failed |
1, scripts-user ... Runparts: 1 failures (runcmd) |
Why it matters for the GPU shard
systemctl start gpu-driver.service in runcmd blocks until the oneshot finishes (cloud-final.service ended the same second as gpu-driver.service on both images: fleet 9.7 s / 8.5 s, stock 60.7 s / 59.7 s), so the gate is what keeps a job off a VM whose driver install and remount,exec have not run yet. The sandbox GPU experiments did not set it, which is where the canary exit 126/127 failures came from. With this change the gate is on for both, and boot-verify and a real nvidia-smi job pass on the fleet image and on stock COS 125.
Rename
The flag is now --google-cos-wait-for-cloud-init, which is what it does. The old name described the incident that motivated it and has nothing to do with the network. --google-cos-docker-network-readiness-gate still works and logs a deprecation warning. The instance metadata key becomes gitlab-wait-for-cloud-init. It is written and read by the same binary and a machine is only provisioned once, so nothing depends on the old key. The gprd GPU shard config and the chef role can move to the new name after the release.