Synchronize config reload across engine services

Closes #788 (closed).

Hot reload wrote config structs in place with no lock in provision, zfs, cloning, retrieval, platform and embeddedui, racing every handler that read them. Two more aliasing bugs in the same area: CanStartRefresh was a lock-free check-then-act, so two concurrent full refreshes both passed it, and SnapshotList handed out the internal slice that removeSnapshotFromList shifts in place.

Four commits, one per task group.

Retrieval state and the refresh slot (task 22)

State keeps its fields unexported behind accessors that take its mutex. TryStartRefresh checks the preconditions and claims the slot under one lock; FullRefresh calls it and releases with a defer that covers every early return, and reports a failed claim as an error instead of nil.

CanStartRefresh stays a read-only precondition check. Making it the claim would have let the handler precheck consume the slot, so POST /full-refresh would answer 200 and do nothing.

The claim is a dedicated flag rather than a transition to Refreshing, which the issue suggested. Setting the status inside the claim makes RefreshData bail out with pool is still busy, so the refresh would never run.

Locked reload (tasks 23a-23d)

Each service holds its config as a value behind cfgMu, swapped as a whole under the write lock and read through an accessor. platform.Service.Reload copies fields one by one instead of *s = *newService, which go vet copylocks rejects once the struct carries a mutex, and the API client moved behind Client() so a reload cannot swap it mid-request.

Two consequences worth review:

  1. An AppConfig handed to a clone now carries its own copy of the database config. A reload no longer rewrites the credentials of a clone that is already being provisioned.

  2. Cloning's Reload used to write the global config through the pointer it shared with retrieval and the HTTP server. That aliased write raced with every reader of the same memory, and it was also the only path by which a reloaded global: section reached them. The propagation is now explicit: Retrieval.Reload takes the global config and Server.SetGlobal replaces it under the server's own lock, the way SetRetention already does. This touches internal/srv/server.go and cmd/database-lab/main.go, which are outside the file list in the issue.

Snapshot list (task 24)

SnapshotList returns slices.Clone of its list. The caller in createSnapshot that copied the slice itself before sorting no longer has to.

Compatibility

Nothing under pkg/ or api/ changed, no YAML tag moved, and no endpoint, status code or response body changed. Three runtime behavior changes: a skipped refresh now logs where it was silent before, an in-flight clone keeps its database config across a reload, and global-config propagation runs through explicit calls. provision.New and cloning.NewBase now dereference their config arguments, so nil panics where it was tolerated.

Merge request reports

Loading
Loading