Prioritize paid namespaces in Zoekt rollout selection
What does this MR do and why?
Zoekt rollout picks the oldest-enabled namespaces first. Most enabled namespaces are trials, so paid namespaces wait behind them for initial indexing. This MR makes rollout fill each batch cheapest-first: paid namespaces without replicas (GitLab.com only), then any namespace without replicas, then the existing full scan only if room is left. It also removes health-check deferrals that don't apply to the jobs involved, and makes the number of indices initial indexing starts per node per run an admin setting, so it can be raised without a deploy.
Four changes:
-
Selection order, behind the
zoekt_prioritized_rollout_selectionfeature flag (gitlab_com_derisk, default off, rollout issue #631182).Search::Zoekt::SelectionService#fetch_enabled_namespace_for_indexingnow fills a batch in three steps. With the flag off, only step 3 runs, which is the current behavior:Step Scope What it selects 1 GitLab.com only ( ::Gitlab::Saas.feature_available?(:exact_code_search))Rollout-allowed namespaces with an active paid subscription and no replicas, ordered by id, up tomax_batch_size.2 All instances Rollout-allowed namespaces with no replicas, excluding ones already selected, ordered by id(oldest first, as today), filling the rest of the batch.3 All instances The existing each_batch_with_mismatched_replicasscan, unchanged, only if the batch still has room. Handles other mismatches (too many replicas, or fewer than desired but not zero), skipping ids already selected.New
EnabledNamespacescopes:with_paid_subscription(EXISTSonGitlabSubscription.with_a_paid_hosted_plan.not_expired, joined onroot_namespace_id),without_replicas(after_id:)(NOT EXISTSonzoekt_replicas, bounded byid > after_idon both sides),ordered_by_id.Steps 1 and 2 each keep a cursor in
Rails.cacheso they skip the old namespaces that already have replicas. The cursor is set to just before the first namespace selected, or to the max id when nothing is selected, and the cache key changes every 10 minutes (CURSOR_CACHE_TTL), so each query rescans from the start once per window.Namespaces with no replica come first on purpose, so every namespace becomes searchable before any gets an extra replica. This also holds if
default_number_of_replicasgoes above 1: namespaces short of a second replica wait until no zero-replica namespaces are left to fill the batch. -
RolloutWorkerhealth check narrowed to tables it writes:zoekt_enabled_namespaces,zoekt_replicas,zoekt_indices. It previously also deferred onzoekt_nodes,zoekt_repositories,zoekt_tasks, which it never writes. -
SchedulingWorkerhealth check narrowed to the tables each task writes. NewWRITE_TASK_TABLESmaps the four tasks that write to the tables they touch, including rows changed by foreign key cascades:Task Tables auto_index_self_managedzoekt_enabled_namespacesevictionzoekt_enabled_namespaces,zoekt_replicas,zoekt_indicesremove_expired_subscriptionszoekt_enabled_namespaces,zoekt_replicas,zoekt_indicesupdate_replica_stateszoekt_replicasThe
initiatejob (no task) and every other task get no tables (block form ofdefer_on_database_health_signal;WRITE_TASK_TABLES.fetch(task, [])), since they only read, log, or dispatch events, and the 16 event workers they dispatch to run their own health checks on the tables they write. Database-wide indicators still apply to every job regardless of task:autovacuum_active_on_tableis the only table-specific one;patroni_apdex,prometheus_alert_indicator,wal_rate,write_ahead_logare global and unaffected by this change. -
New admin application setting
zoekt_initial_indexing_limit. Replaces theSchedulingService::INITIAL_INDEXING_LIMITconstant with an application setting (integer, 1 to 100, default 1), stored in thezoekt_settingsjsonb column. It's hidden from the admin UI (admin_ui: falseinSearch::Zoekt::Settings, same aszoekt_max_restarts_15m), but can be set through the application settings API and shows up ingitlab:zoekt:info. Default 1 means no behaviour change at deploy. The limit applies per node, so one run dispatches up to the limit times the number of nodes with pending indices. Theinitial_indexingtask gets no table health check inSchedulingWorker. Instead, each dispatched event'sSearch::Zoekt::InitialIndexingEventWorkerjob defers onzoekt_repositorieshealth.
zoekt_rollout_batch_size (currently 1 on GitLab.com) is not changed here. The empty-replica loop fix (!257616 (merged)) reached gprd on 2026-09-25. On 2026-09-28, gprd still logged about 15,000 RolloutWorker "Batch is completed with success" lines in 24 hours, so the loop fix does not explain those runs. If most of them provision a new namespace, batch size 1 already keeps up with demand (~6,400 enabled per day). That is not confirmed yet. Raise it only if measured throughput falls short.
Why (gprd, read-only console, 2026-09-25)
EnabledNamespace.with_missing_indices.count= 122,311, mostly trials. 2,000–10,000 trial namespaces are enabled per day and removed roughly 30 days later when the trial expires. Rollout picks oldest-enabled first, so paid namespaces wait behind them. 663 paid namespaces were waiting with mismatched replicas.- Correction (2026-09-28): gprd logs show
RolloutWorkerlogged "Batch is completed with success" about 15,000 times over 24 hours, roughly one run every 6 s, with a matching count of Sidekiquntil_executeddedup lines for the same worker. So runs are not paced by database-health deferrals; the worker loops almost back to back, each run re-enqueuing itself viaperform_asyncwhile its own lock is still held.- These runs happened after !257616 (merged) reached gprd (2026-09-25 at 10:59 UTC), so the loop fix does not explain them. Not yet confirmed: how many of them provision a new namespace. Count the distinct
namespace_idvalues in the success log lines.
- These runs happened after !257616 (merged) reached gprd (2026-09-25 at 10:59 UTC), so the loop fix does not explain them. Not yet confirmed: how many of them provision a new namespace. Count the distinct
- Initial indexing starts one index per node per run, and all new indices land on the currently freest node. Observed roughly one index every 10–18 minutes on a single node.
- Query cost comparison: on gprd a paid-first pass using the existing per-range scan took 10.45 s (scans the whole table every run); a single
with_mismatched_replicasquery withLIMIT 32took 2.51 s with ~619k buffers (122,950 subscription lookups, because the subscription check runs for every row before grouping); theNOT EXISTS+LIMITquery used here took 306 ms with ~69k buffers (merge anti-join drops namespaces that already have a replica, so only 279 subscription lookups; plan). Both returned the same paid namespaces. Today's full selection takes 5.92 s. - Production uses one replica per namespace, so on gprd
without_replicascovers nearly every namespace waiting for rollout.
References
- Related: #630717 (Zoekt rollout backlog; analysis in an internal note on that issue).
- Related: !257616 (merged) (fixes the empty-replica rollout loop; on gprd since 2026-09-25; independent of this MR).
Stability impact
| Area | Impact |
|---|---|
| Selection query | While namespaces without replicas exist (nearly always today), steps 1–2 fill the batch and the ~5.9 s full scan is skipped, so selection gets cheaper overall. With the cursor, steps 1 and 2 read tens to hundreds of buffers; the unbounded scan (0.3–3.4 s cold, ~69k buffers) runs once per query per 10 minutes. SelectionService runs twice per rollout run (RolloutService#execute and #result). |
| Starvation | Namespaces needing step 3 (too many replicas, overrides, or a second replica if the default is raised) wait while zero-replica namespaces keep arriving. This is intended, and the rollback row still applies. |
RolloutWorker deferrals |
Runs during autovacuum on zoekt_tasks/zoekt_repositories, which it never writes, so this is safe. It still defers on its own tables and on all global indicators. |
SchedulingWorker tasks |
The initiate job and dispatch-only tasks no longer defer during autovacuum on any Zoekt table, so they run on their normal schedule during autovacuum. The four writing tasks defer only on the tables they write, instead of all six. Jobs are deduplicated (until_executed), so queued jobs don't pile up, but more jobs cycle through Redis as deferred-then-retried. A new task that writes inside the job must be added to WRITE_TASK_TABLES, or it gets no table check. |
zoekt_initial_indexing_limit |
No change at deploy (default 1). Biggest risk once raised: raising it to N (at most 100) gives up to Nx more InitialIndexingEvents per node per run. Each creates zoekt_repositories rows (capped at 10,000 per event before re-emitting) and later zoekt_tasks for that node. Trial namespaces are typically tiny, but paid namespaces can be large, and new indices concentrate on one node (node concurrency is 20). Can be lowered again through the API without a deploy. |
| Safeguards retained | InitialIndexingEventWorker still defers on zoekt_repositories health; per-event insert cap of 10,000 stays; a node's task queue still drains at its own concurrency limit. |
| Node storage | Unchanged risk profile. Planning still checks unclaimed storage, and the one-lock-per-rollout-run model is unchanged, so there's no concurrent planning. |
| Rollback | Disable zoekt_prioritized_rollout_selection to return selection to the current behavior without a deploy. All four parts are independent and revertible. The rollout batch size setting can also be lowered at any time without a deploy. |
What to watch after deploy
RolloutWorkerduration and job rate.SchedulingWorkerjobs deferred per hour.Search::Zoekt::Index.pending.group(:zoekt_node_id).count, to check indexing tasks aren't overloading one node.- After raising
zoekt_initial_indexing_limit, watch the same per-node pending index count and task queue on the receiving node. - Zoekt task queue depth and latency on the node receiving new indices.
EnabledNamespace.with_missing_indicestrend over time.- Paid namespaces still waiting:
EnabledNamespace.with_rollout_allowed.with_paid_subscription.without_replicas.count. - Step 2 query time on gprd (
EnabledNamespace.with_rollout_allowed.without_replicas.ordered_by_id.with_limit(32)).
Database queries
This MR adds two queries to Search::Zoekt::SelectionService and makes the limit of one existing query configurable. No schema changes, no new indexes, no migrations. Selection runs in Search::Zoekt::RolloutWorker (cron Sidekiq worker, not a web request); RolloutService calls SelectionService twice per run (rollout_service.rb:35 and :74).
Query 1 — paid namespaces without replicas (GitLab.com only)
New. SelectionService#fetch_paid_namespaces_without_replicas, step 1 of selection.
SELECT "zoekt_enabled_namespaces".* FROM "zoekt_enabled_namespaces"
WHERE ("zoekt_enabled_namespaces"."last_rollout_failed_at" IS NULL
OR "zoekt_enabled_namespaces"."last_rollout_failed_at" < '2026-09-27 12:48:30')
AND (EXISTS (SELECT 1 FROM "gitlab_subscriptions"
WHERE "gitlab_subscriptions"."hosted_plan_name_uid" IN (3, 4, 5, 6, 7, 8, 9, 10, 11)
AND ("gitlab_subscriptions"."trial" = FALSE OR "gitlab_subscriptions"."trial" IS NULL)
AND (end_date IS NULL OR end_date >= '2026-09-28')
AND "gitlab_subscriptions"."namespace_id" = "zoekt_enabled_namespaces"."root_namespace_id"))
AND (NOT EXISTS (SELECT 1 FROM "zoekt_replicas"
WHERE "zoekt_replicas"."zoekt_enabled_namespace_id" = "zoekt_enabled_namespaces"."id"
AND "zoekt_replicas"."zoekt_enabled_namespace_id" > 111111))
AND "zoekt_enabled_namespaces"."id" > 111111
ORDER BY "zoekt_enabled_namespaces"."id" ASC LIMIT 1LIMIT is zoekt_rollout_batch_size (1 on GitLab.com, default 32 elsewhere). 111111 is the cursor (0 on the first run in each 10-minute window).
Indexes and gprd measurement
Indexes used: zoekt_enabled_namespaces_pkey for the ORDER BY id + LIMIT, index_zoekt_replicas_on_enabled_namespace_id for the NOT EXISTS, and the unique index_gitlab_subscriptions_on_namespace_id for the EXISTS.
- With the cursor, nothing paid waiting (worst case, checks every row above the cursor): 69 ms, ~450 buffers. Both index scans start at the cursor, and only the 100 rows above it get a subscription lookup. Plan: https://console.postgres.ai/gitlab/projects/gitlab-production-main/sessions/58964/commands/164403
- Without a cursor (first run in each window): 306 ms, ~69k buffers, 279 subscription lookups. Plan: https://console.postgres.ai/shared/e38a55ce-b82d-4f09-aa85-1f702f7d5e2e
Query 2 — any namespace without replicas
New. SelectionService#fetch_namespaces_without_replicas, step 2 of selection.
SELECT "zoekt_enabled_namespaces".* FROM "zoekt_enabled_namespaces"
WHERE ("zoekt_enabled_namespaces"."last_rollout_failed_at" IS NULL
OR "zoekt_enabled_namespaces"."last_rollout_failed_at" < '2026-09-27 12:48:30')
AND (NOT EXISTS (SELECT 1 FROM "zoekt_replicas"
WHERE "zoekt_replicas"."zoekt_enabled_namespace_id" = "zoekt_enabled_namespaces"."id"
AND "zoekt_replicas"."zoekt_enabled_namespace_id" > 111111))
AND "zoekt_enabled_namespaces"."id" > 111111
AND "zoekt_enabled_namespaces"."id" NOT IN (1, 2)
ORDER BY "zoekt_enabled_namespaces"."id" ASC LIMIT 30The NOT IN list is step 1's picks (at most the batch size); when step 1 returns nothing, Rails emits 1=1 instead. LIMIT is the batch size minus step 1's count. 111111 is the cursor, as in query 1. Same indexes as query 1.
Plans
- With the cursor at the oldest namespace without a replica (backlog): the main query reads ~6 buffers and returns its row right away. The plan computes the cursor in two subqueries (~53k buffers each) because Joe can't return a value to paste; the app passes a literal. Plan: https://console.postgres.ai/gitlab/projects/gitlab-production-main/sessions/58964/commands/164406
- With the cursor near the max id (no backlog): 1.3 ms, 17 buffers. Plan: https://console.postgres.ai/gitlab/projects/gitlab-production-main/sessions/58964/commands/164405
- Without a cursor (first run in each window): 3.36 s cold, ~68k buffers. Plan: https://console.postgres.ai/shared/1e64b041-f618-405e-a722-0454a3b593f9
Query 3 — initial indexing, limit now configurable
Changed. SchedulingService#initial_indexing, once per online node with pending indices.
SELECT "zoekt_indices".* FROM "zoekt_indices"
WHERE "zoekt_indices"."zoekt_node_id" = 2000234 AND "zoekt_indices"."state" = 0
ORDER BY "zoekt_indices"."id" ASC LIMIT 20Only the LIMIT changed: it's now the admin setting zoekt_initial_indexing_limit (default 1, 1 to 100) instead of the constant 1.
Indexes and plan
Uses index_zoekt_indices_on_zoekt_node_id_and_id (zoekt_node_id, id), which matches the filter and the order, then filters state, with no sort. Measured on gprd on a node with pending indices and LIMIT 20: 1.5 ms, 44 buffers.
Plan: https://console.postgres.ai/shared/22a5f916-c368-4d79-aa8f-751892b4b070
Unchanged
- The step 3
each_batch_with_mismatched_replicasscan is the same SQL as master, now skipped when steps 1–2 fill the batch. Master's full selection was measured at ~5.92 s on gprd. RolloutWorker/SchedulingWorkerchanges only touch health-check table lists, no SQL.- The new setting is read via the cached
ApplicationSetting.current.
Known trade-offs
- The first run of queries 1 and 2 in each 10-minute window has no cursor and exceeds the 100 ms guideline (up to 3.4 s cold). It runs in a cron worker, once per window, and replaces the ~5.9 s selection while a backlog exists.
- No index covers "has no replica"; the
NOT EXISTSis an anti-join, so cost grows withzoekt_enabled_namespacessize. - The
ORinwith_rollout_allowed(existing scope) likely prevents use ofindex_zens_on_last_rollout_failed_at; plans walk the primary key in id order and filter. - A namespace below the cursor that becomes eligible (a trial converting to paid, a replica removed, or a rollout retry) waits up to 10 minutes for the next rescan, or is picked up by step 3.
How to set up and validate locally
- On GDK with Zoekt enabled, simulate SaaS (
GITLAB_SIMULATE_SAAS=1) and enable the exact code search SaaS feature as usual for Zoekt development. Enable the flag withFeature.enable(:zoekt_prioritized_rollout_selection). - In the Rails console, create two trial groups and one paid group (with a
gitlab_subscription). Enable all three withSearch::Zoekt::EnabledNamespace.create!(root_namespace_id: ...), creating the trials first so they're older. - Run
Search::Zoekt::SelectionService.execute(max_batch_size: 1).enabled_namespacesand confirm the paid namespace is returned instead of the oldest trial. - Destroy the three enabled namespaces from step 2, since they still have no replicas and would fill the batch. Then create an enabled namespace with 2 replicas plus two with none. Run
Search::Zoekt::SelectionService.execute(max_batch_size: 2).enabled_namespacesand confirm only the two without replicas are returned; withmax_batch_size: 3the one with 2 replicas is also returned. - Check
Search::Zoekt::RolloutWorker.database_health_check_attrs[:tables]reflects the narrowed table list, and callSearch::Zoekt::SchedulingWorker.database_health_check_attrs[:block]with['eviction'],['initial_indexing']and[], confirming it returns the three eviction tables, then empty, then empty.
MR acceptance checklist
Evaluate this MR against the MR acceptance checklist.