Code quality plan, MR F: reload synchronization
Phase F of the code quality action plan (https://gitlab.com/postgres-ai/database-lab/-/blob/docs/code-quality-action-plan/docs/plans/20260913-code-quality-assessment-action-plan.md). One MR, four commits.
Hot reload writes config structs in place with no lock in provision, zfs, cloning, retrieval, platform, embeddedui. Two more aliasing bugs in the same area. Confirmed on origin/master 8a42a6ef.
Findings
- every
Reloadwrites fields while readers may be mid-request CanStartRefreshis a lock-free check-then-act onState.Status; two concurrent full refreshes both passSnapshotListreturns the internal slice;removeSnapshotFromListshifts it in place under the lock while callers iterate
Tasks
- 22.
Stateaccessors under its mutex; keepCanStartRefreshread-only for the handler precheck atroutes.go:1293; add exportedTryStartRefreshcalled only insideFullRefresh, which releases the slot on every early return. TurningCanStartRefreshitself into a claim would make the handler consume the slot andFullRefreshreturn nil, soPOST /full-refreshwould answer 200 and do nothing. - 23a.
cfgMuinprovisionandzfs - 23b.
cfgMuincloning - 23c.
cfgMuinretrieval - 23d.
cfgMuinplatformandembeddedui;platform.Service.Reloadmust copy fields instead of*s = *newService, otherwisego vetcopylocks fails - 24.
SnapshotListreturnsslices.Clone(m.snapshots)
Every task ships a -race test that fails before the fix. Pattern to follow: Server.configMu (internal/srv/server.go:56).