Code quality plan, MR F: reload synchronization

Phase F of the code quality action plan (https://gitlab.com/postgres-ai/database-lab/-/blob/docs/code-quality-action-plan/docs/plans/20260913-code-quality-assessment-action-plan.md). One MR, four commits.

Hot reload writes config structs in place with no lock in provision, zfs, cloning, retrieval, platform, embeddedui. Two more aliasing bugs in the same area. Confirmed on origin/master 8a42a6ef.

Findings

  • every Reload writes fields while readers may be mid-request
  • CanStartRefresh is a lock-free check-then-act on State.Status; two concurrent full refreshes both pass
  • SnapshotList returns the internal slice; removeSnapshotFromList shifts it in place under the lock while callers iterate

Tasks

  • 22. State accessors under its mutex; keep CanStartRefresh read-only for the handler precheck at routes.go:1293; add exported TryStartRefresh called only inside FullRefresh, which releases the slot on every early return. Turning CanStartRefresh itself into a claim would make the handler consume the slot and FullRefresh return nil, so POST /full-refresh would answer 200 and do nothing.
  • 23a. cfgMu in provision and zfs
  • 23b. cfgMu in cloning
  • 23c. cfgMu in retrieval
  • 23d. cfgMu in platform and embeddedui; platform.Service.Reload must copy fields instead of *s = *newService, otherwise go vet copylocks fails
  • 24. SnapshotList returns slices.Clone(m.snapshots)

Every task ships a -race test that fails before the fix. Pattern to follow: Server.configMu (internal/srv/server.go:56).