google_cos: wait on cloud-init.target instead of cloud-init status --wait

What

--google-cos-docker-network-readiness-gate ran cloud-init status --wait over SSH before configuring Docker. On stock COS that command never exits 0, and SSHCommand fails on any non-zero exit. The gate now waits with systemctl start cloud-init.target and then runs cloud-init status once: exit 0 and 2 pass, exit 1 (errors recorded by cloud-init, e.g. a failed runcmd) and 124 (5 min timeout) fail the create.

Why --wait does not work on stock COS

--wait polls get_status_details() every 0.25 s and returns as soon as the status is anything other than running. Inside that function, cloud-init ≥ 24.1 has a check that runs while cloud-init is still in progress:

if running_status == RunningStatus.RUNNING and uses_systemd() and systemd_failed(wait=wait):
    running_status = RunningStatus.DONE
    condition_status = ConditionStatus.ERROR
    description = "Failed due to systemd unit failure"

systemd_failed() looks at the four cloud-init services and returns true for any that is not UnitFileState=enabled or static. The idea is that a disabled stage will never run, so stop waiting for it. On COS the services are symlinked into multi-user.target.wants, but their [Install] sections say WantedBy=cloud-init.target, so systemd reports them as disabled even though they run:

Id=cloud-init-local.service  ActiveState=activating  UnitFileState=disabled
Id=cloud-init.service        ActiveState=inactive    UnitFileState=disabled
Id=cloud-config.service      ActiveState=inactive    UnitFileState=disabled
Id=cloud-final.service       ActiveState=inactive    UnitFileState=disabled

So the first poll flips the status to error - done and --wait exits 1 within a second. docker-machine reaches the VM about 10 s after boot, so every stock create failed (last_update: 00:00:10, Machine creation failed ... time=58s in the sandbox). There is no stage after which --wait starts working: the check applies until result.json exists, and the units stay "disabled" for the whole boot. The check is unchanged on cloud-init main.

Once cloud-init is done, status exits 2 (degraded done) instead:

recoverable_errors:
WARNING:
	- Getting data from <class 'cloudinit.sources.DataSourceGCE.DataSourceGCELocal'> failed
  File ".../cloudinit/sources/DataSourceGCE.py", line 131, in _get_data
UnboundLocalError: cannot access local variable 'ret' where it is not associated with a value

COS has no dhclient or udhcpc, so DataSourceGCELocal skips every NIC in init-local and crashes on an unassigned variable. cloud-init falls back to DataSourceGCE and applies user-data. This happens before user-data is read, and the image's datasource_list: [GCE, NoCloud, None] cannot exclude GCELocal. The crash is fixed in cloud-init 25.1, but the no-lease path still logs a WARNING there, so exit 2 is to be expected on COS for the foreseeable future.

Both were checked on bare cos-125-lts and cos-stable (cos-121) VMs with no user-data, both cloud-init 24.4.1. The fleet image (COS 109, cloud-init 23.2.1) predates the systemd check and exits 0, which is why the gate has worked on gprd.

Why cloud-init.target

cloud-init documents it as the synchronization point for "all of cloud-init's initial system configuration tasks have completed" and tells units to order themselves with After=cloud-init.target. systemctl start on a target blocks until it is reached and returns immediately if it already was. On COS the target is After=multi-user.target, which is not reached until its oneshot wants, cloud-final included, have exited. It is reached even when cloud-final failed (Wants=, not Requires=), which is why cloud-init status still runs afterwards.

image cloud-init systemctl start cloud-init.target cloud-init status after
cos-125-lts, runcmd: sleep 150 24.4.1 0 after 149 s, cloud-final active 2
fleet COS 109, same 23.2.1 0 after 153 s 0
cos-125-lts, failing runcmd 24.4.1 0 after 60 s, cloud-final failed 1, scripts-user ... Runparts: 1 failures (runcmd)

Why it matters for the GPU shard

systemctl start gpu-driver.service in runcmd blocks until the oneshot finishes (cloud-final.service ended the same second as gpu-driver.service on both images: fleet 9.7 s / 8.5 s, stock 60.7 s / 59.7 s), so the gate is what keeps a job off a VM whose driver install and remount,exec have not run yet. The sandbox GPU experiments did not set it, which is where the canary exit 126/127 failures came from. With this change the gate is on for both, and boot-verify and a real nvidia-smi job pass on the fleet image and on stock COS 125.

Rename

The flag is now --google-cos-wait-for-cloud-init, which is what it does. The old name described the incident that motivated it and has nothing to do with the network. --google-cos-docker-network-readiness-gate still works and logs a deprecation warning. The instance metadata key becomes gitlab-wait-for-cloud-init. It is written and read by the same binary and a machine is only provisioned once, so nothing depends on the old key. The gprd GPU shard config and the chef role can move to the new name after the release.

Edited by Igor

Merge request reports

Loading
Loading