Recover the instance zone for stop, start, and inspect in bulk mode

What does this MR do?

Extends the empty-zone recovery that delete already has to stop, start, and instance lookups.

Why was this MR needed?

A failed bulkInsert can leave a machine with no resolved zone. Delete recovers via AggregatedList since !175 (merged), but stop, start, and inspect still called the API with an empty zone and failed with:

googleapi: Error 400: Invalid value for field 'zone': ''. Must be a match of regex ...

gitlab-runner hits the stop path when draining a machine whose create never resolved a zone. As with delete, a machine that was never placed now yields a not-found error so callers treat it as gone and reap local state.

API traffic

Healthy machines pay a field check, and bulk-mode creates now make one call fewer: the pre-create existence check never found anything there (no zone is resolved before placement, so it always 400ed and passed) and is skipped. A machine with no recorded zone previously made one guaranteed-400 call per operation and callers retried forever. It now makes a single AggregatedList that either resolves the zone, which is written back to the machine config so nothing rediscovers it, or returns not-found so the retrying stops.

What's the best way to test this MR?

Unit tests drive all three operations through both recovery outcomes against a fake compute API: the recovered zone is targeted, and a never-placed instance returns not-found.

Edited by Igor

Merge request reports

Loading
Loading