docs: Fix Geo registry notification secret and failover guidance
What does this MR do?
Fixes two documentation defects on the Geo container registry pages, both found while reproducing a self-managed customer's Geo registry migration on a two-site lab.
The replication page documents a configuration key that does not exist
geo/replication/container_registry.md tells operators to set registry['notification_secret'], in two places: once in a note for external registries, and once as a step for GitLab HA web nodes.
That key does not exist in omnibus-gitlab. An exhaustive search of the repository for notification_secret finds no consumer under the registry namespace at all: every hit is the gitlab_rails['registry_notification_secret'] attribute, its derivation in gitlab/libraries/registry.rb, the templates that render it, and their specs. The registry namespace in gitlab.rb is a Gitlab::ConfigMash with auto-vivification during from_file, so an unknown key inside it is accepted without complaint and then ignored.
On a single-node primary running the bundled registry nobody notices, because the Linux package derives the secret automatically. The two cases the page is specifically addressing are exactly the two where that derivation does not happen, so both end up with notification_secret empty, every registry notification rejected 401, and container images that never replicate on their own.
Measured on 19.2.1, on a node with no registry['notifications'] block:
| Step | gitlab.yml notification_secret: |
|---|---|
| Baseline | empty |
After registry['notification_secret'] = '<value>', reconfigure rc=0 |
still empty |
After gitlab_rails['registry_notification_secret'] = '<value>', reconfigure rc=0 |
the value |
The correct key is already documented on the non-Geo page, at administration/packages/container_registry.md, which is how the two pages came to disagree.
This MR also:
- Documents that the automatic derivation is gated on the notification endpoint being named exactly
geo_event.gitlab/libraries/registry.rbrunsregistry_notification_secret ||= endpoint['headers']['Authorization'].last if endpoint['name'] == 'geo_event', so any other name skips the derivation and GitLab is left with no secret unlessgitlab_rails['registry_notification_secret']is set explicitly, which is exactly what this MR's own fix requires on HA web nodes and for external registries. The page's own example happens to usegeo_event, so an operator who copies it verbatim is fine by luck, but the endpoint name reads as a free-form label. The derived secret is never written togitlab-secrets.json, so it is re-derived on every reconfigure: a misnamed endpoint with no explicit setting is not a one-time mistake that a later run repairs, it stays broken for the life of the site. - Adds an empty-secret check to the "Registry events logs response status 401 Unauthorized unaccepted" section. That section currently says to make the headers match "as should be done during step Configure primary site", which an operator on the HA or external-registry path cannot do, because the step they followed set nothing.
The section is named after an error string, so it is worth recording that a misnamed endpoint produces that exact string. Renaming the endpoint and changing nothing else, on 19.2.1:
| Endpoint name | GitLab's derived secret | Registry log | API result |
|---|---|---|---|
registry_event |
empty | response status 401 Unauthorized unaccepted, five times |
401 |
geo_event |
present, 34 characters | nothing | 200 |
Five rejected delivery attempts were logged, with the endpoint at the page's example maxretries of 5. I am reporting the count I observed rather than deriving it from the setting, because the two do not obviously agree: the registry builds its retry policy with backoff.WithMaxRetries(b, 5), which would permit an initial attempt plus five retries. Either reading leads to the same outcome, which is the part that matters here: after the last attempt the event is dropped, the push itself succeeds, and no Geo event is created, so replication does not start for that push. While the secret stays empty, no later push starts it either.
The page already documents recovery once the secret is fixed, under "Repository not resynced after downtime", and the MR now links the two together rather than implying the loss is permanent.
The MR also points readers at api_json.log. The notification endpoint is part of the REST API, so its requests, including the rejections, are logged in api_json.log and never in production_json.log, which is where an administrator looks first. That is a small addition but it is the difference between finding the problem and concluding there isn't one.
The planned failover page predates object storage and the metadata database
geo/disaster_recovery/planned_failover.md already has a "Container registry" section, so the problem is not that the registry is unmentioned. The problem is that everything it offers describes a deployment shape that a modern site no longer has:
rsync --archive --perms --delete root@<geo-primary>:/var/opt/gitlab/gitlab-rails/shared/registry/. /var/opt/gitlab/gitlab-rails/shared/registryOn a site using object storage that directory is empty, so the command exits 0 and copies nothing. On a site using the metadata database the catalog is not in storage at all. The section is gated on "If you use local storage", which is accurate and leaves every other reader with no guidance. Meanwhile ## Preflight checks, the list an administrator works through before a maintenance window, does not mention the registry.
This MR splits the guidance by backend, and adds:
- That the metadata database is not replicated between sites and needs no promotion, unlike the Geo tracking database. This is good news that currently reaches nobody, and it changes how a maintenance window is budgeted.
- The two role-dependent settings that do not carry across a promotion: the promoted site needs its own
registry['notifications']block, and any remaining secondary needsgeo_registry_replication_primary_api_urlrepointed. - A registry preflight check: compare catalogs between sites, and confirm the active metadata backend. It also warns against relying on Geo's container repository sync status alone, because tag-level sync failures are logged rather than recorded against the registry.
The catalog comparison spells out how to get a token, which is worth the four extra lines. Reading /v2/_catalog needs the registry:catalog:* scope, which GitLab issues only to an administrator, so the ordinary docker login flow does not produce a usable token and a bare curl returns 401. The added step generates one from /jwt/auth and then reads the catalog with it. It also documents the pagination, which matters more than it looks: the catalog returns 100 repositories by default, n caps at 1000, and asking for more than that returns an error rather than a page. A site with more repositories than the page size otherwise compares clean while differing beyond the first page. Read from registry/handlers/catalog.go at v4.40.2-gitlab, the version the lab ran: defaultMaximumReturnedEntries = 100, maximumReturnEntriesUpperLimit = 1000, and writeCatalogResponse sets a Link header when more entries remain.
What promotion does to the registry, observed
I ran the promotion on a two-site 19.2.1 pair with both sites on the metadata database.
Promotion does not touch the registry. It was never restarted: same process id across the promotion, with its database configuration, storage lockfile and catalog all unchanged. That is worth stating plainly on the page, because "not replicated and needs no promotion" reads like a caveat when it is in fact the good news.
The missing notifications block is silent. With no endpoint configured the registry has nothing to notify, so unlike the misnamed-endpoint case there is no 401 and no log line at all, on either side. Measured as an A/B on the promoted site:
| As promotion leaves it | After adding the block | |
|---|---|---|
| Events reaching GitLab | none attempted | accepted |
| Errors logged, either side | none | none |
| Tag present on the other site | no | yes, and it also caught up the tag missed earlier |
| Geo container repository state | synced, last-synced timestamp predates the push |
synced, timestamp advanced |
The replication page already notes that "the Geo admin UI might still report 100% replication because the sync status is based on the registry entry state, not on content verification", so that half is not new. What is new is that a promotion puts a site into that state by default, with no error logged on either side, which is why the failover page is where it needs saying.
There is exactly one signal. gitlab-rake gitlab:geo:check reports Container Registry Geo events ... last event at <timestamp>. It never turns red, but the timestamp stops advancing while the endpoint is absent and resumes when it returns. The added text in the metadata database section tells the reader to push an image after promotion and confirm that timestamp moves, which is the only check that distinguishes the two states.
Recovery is complete once the block exists: the next sync caught up the tag pushed during the silent window, because container repository sync operates on the whole repository rather than per tag.
Related issues
Related to !249940, which corrects the backend verification command this MR's preflight check links to. Worth reviewing together, though they touch different pages and different groups.
Overlaps !143468 (closed), which proposes deleting the rsync and backup/restore text from this same "Container registry" section. Review there settled on rewriting the section rather than removing it, which is what this MR does: the rsync stays, scoped to the local storage case it is still correct for, and the two backends it never covered get guidance of their own.
Evidence
Reproduced on a two-site Geo lab, GitLab 19.2.1, registry v4.40.2-gitlab, both sites on their own object storage and their own registry PostgreSQL, with container registry replication working end to end (matching catalogs, event-driven propagation observed).
The load-bearing claim, that the documented key is a silent no-op and a different key is the one that works, was executed on the Geo secondary, which carries no registry['notifications'] block. That is the node shape the documentation addresses for HA web nodes and for external registries.
| Step | gitlab.yml notification_secret: |
|---|---|
| Baseline, nothing set | empty |
After registry['notification_secret'] = 'PROBE1SECRET', reconfigure rc=0 |
still empty |
After gitlab_rails['registry_notification_secret'] = 'PROBE3SECRET', reconfigure rc=0 |
PROBE3SECRET |
The reconfigure succeeds in both cases. There is no warning, no error, and in the documented form no secret.
The geo_event name gate was confirmed by a controlled A/B on the primary of a two-site 19.2.1 pair, the node that carries the registry['notifications'] block: two reconfigures, with only the endpoint's name changed between them and the Authorization header untouched:
Endpoint name |
notification_secret in gitlab.yml |
|---|---|
registry_event |
empty, length 0 |
geo_event |
present, length 34 |
A second lab pair cross-checked the mechanism's edges. In the empty case that registry's own rendered config.yml read "name":"registry_event", so the notification endpoint was configured and working and it is specifically the secret derivation that did not fire. In the present case the value is byte-for-byte the endpoint's Authorization header, which confirms the whole expression rather than only the name condition:
Gitlab['gitlab_rails']['registry_notification_secret'] ||= endpoint['headers']['Authorization'].last if endpoint['name'] == 'geo_event'On that same pair, registry_notification_secret was absent from gitlab-secrets.json in both stages, so the value is derived at reconfigure time straight into gitlab.yml and never persisted. There is no cached copy, which means the empty reading is genuinely empty rather than a stale-value artifact.
So under any endpoint name other than geo_event, Rails holds no secret and every registry notification is rejected, and registry replication silently never becomes event-driven.
The failover statements were then observed on a 19.2.1 pair with both sites migrated to the metadata database: a real gitlab-ctl geo promote, followed by pushes to the promoted site with and without a notifications block, with the demoted site reattached as a secondary afterwards.
Three statements are read from source rather than executed, and each is named as such in the text:
- The
registry['notification_secret']key's absence, established by searchingomnibus-gitlab, which is a search result rather than an execution. The measured half of that claim is in the table above. - That the derived secret is never written to
gitlab-secrets.json, fromgather_gitlab_secretsinfiles/gitlab-cookbooks/package/libraries/helpers/secrets_helper.rb, which has noregistry_notification_secretkey. - The catalog pagination limits, read from
registry/handlers/catalog.goat the registry version the lab ran, not exercised against a registry holding more than one page.
No {{< history >}} block is included. Per the availability details guidance, history blocks track feature-level tier, status, and availability changes rather than content additions, and this MR documents existing behavior.
Author's checklist
- Follows the documentation style guide
-
markdownlint-cli2clean (0 errors, using the version CI pins) -
vale --minAlertLevel warningreports no warnings on any added line, checked with no rule exclusions - Editing existing pages only, so no new product availability details block is required
- Request review from the Geo technical writer
AI-Generated Content Disclosure: This MR was prepared with assistance from Claude Code. The output has been reviewed for correctness, verified on a lab against GitLab 19.2.1 except where the text names a claim as read from source, and validated against the documentation style guide.