Prioritize paid namespaces in Zoekt rollout selection

What does this MR do and why?

Zoekt rollout picks the oldest-enabled namespaces first. Most enabled namespaces are trials, so paid namespaces wait behind them for initial indexing. This MR makes rollout fill each batch cheapest-first: paid namespaces without replicas (GitLab.com only), then any namespace without replicas, then the existing full scan only if room is left. It also removes health-check deferrals that don't apply to the jobs involved, and makes the number of indices initial indexing starts per node per run an admin setting, so it can be raised without a deploy.

Four changes:

  1. Selection order, behind the zoekt_prioritized_rollout_selection feature flag (gitlab_com_derisk, default off, rollout issue #631182). Search::Zoekt::SelectionService#fetch_enabled_namespace_for_indexing now fills a batch in three steps. With the flag off, only step 3 runs, which is the current behavior:

    Step Scope What it selects
    1 GitLab.com only (::Gitlab::Saas.feature_available?(:exact_code_search)) Rollout-allowed namespaces with an active paid subscription and no replicas, ordered by id, up to max_batch_size.
    2 All instances Rollout-allowed namespaces with no replicas, excluding ones already selected, ordered by id (oldest first, as today), filling the rest of the batch.
    3 All instances The existing each_batch_with_mismatched_replicas scan, unchanged, only if the batch still has room. Handles other mismatches (too many replicas, or fewer than desired but not zero), skipping ids already selected.

    New EnabledNamespace scopes: with_paid_subscription (EXISTS on GitlabSubscription.with_a_paid_hosted_plan.not_expired, joined on root_namespace_id), without_replicas(after_id:) (NOT EXISTS on zoekt_replicas, bounded by id > after_id on both sides), ordered_by_id.

    Steps 1 and 2 each keep a cursor in Rails.cache so they skip the old namespaces that already have replicas. The cursor is set to just before the first namespace selected, or to the max id when nothing is selected, and the cache key changes every 10 minutes (CURSOR_CACHE_TTL), so each query rescans from the start once per window.

    Namespaces with no replica come first on purpose, so every namespace becomes searchable before any gets an extra replica. This also holds if default_number_of_replicas goes above 1: namespaces short of a second replica wait until no zero-replica namespaces are left to fill the batch.

  2. RolloutWorker health check narrowed to tables it writes: zoekt_enabled_namespaces, zoekt_replicas, zoekt_indices. It previously also deferred on zoekt_nodes, zoekt_repositories, zoekt_tasks, which it never writes.

  3. SchedulingWorker health check narrowed to the tables each task writes. New WRITE_TASK_TABLES maps the four tasks that write to the tables they touch, including rows changed by foreign key cascades:

    Task Tables
    auto_index_self_managed zoekt_enabled_namespaces
    eviction zoekt_enabled_namespaces, zoekt_replicas, zoekt_indices
    remove_expired_subscriptions zoekt_enabled_namespaces, zoekt_replicas, zoekt_indices
    update_replica_states zoekt_replicas

    The initiate job (no task) and every other task get no tables (block form of defer_on_database_health_signal; WRITE_TASK_TABLES.fetch(task, [])), since they only read, log, or dispatch events, and the 16 event workers they dispatch to run their own health checks on the tables they write. Database-wide indicators still apply to every job regardless of task: autovacuum_active_on_table is the only table-specific one; patroni_apdex, prometheus_alert_indicator, wal_rate, write_ahead_log are global and unaffected by this change.

  4. New admin application setting zoekt_initial_indexing_limit. Replaces the SchedulingService::INITIAL_INDEXING_LIMIT constant with an application setting (integer, 1 to 100, default 1), stored in the zoekt_settings jsonb column. It's hidden from the admin UI (admin_ui: false in Search::Zoekt::Settings, same as zoekt_max_restarts_15m), but can be set through the application settings API and shows up in gitlab:zoekt:info. Default 1 means no behaviour change at deploy. The limit applies per node, so one run dispatches up to the limit times the number of nodes with pending indices. The initial_indexing task gets no table health check in SchedulingWorker. Instead, each dispatched event's Search::Zoekt::InitialIndexingEventWorker job defers on zoekt_repositories health.

zoekt_rollout_batch_size (currently 1 on GitLab.com) is not changed here. The empty-replica loop fix (!257616 (merged)) reached gprd on 2026-09-25. On 2026-09-28, gprd still logged about 15,000 RolloutWorker "Batch is completed with success" lines in 24 hours, so the loop fix does not explain those runs. If most of them provision a new namespace, batch size 1 already keeps up with demand (~6,400 enabled per day). That is not confirmed yet. Raise it only if measured throughput falls short.

Why (gprd, read-only console, 2026-09-25)
  • EnabledNamespace.with_missing_indices.count = 122,311, mostly trials. 2,000–10,000 trial namespaces are enabled per day and removed roughly 30 days later when the trial expires. Rollout picks oldest-enabled first, so paid namespaces wait behind them. 663 paid namespaces were waiting with mismatched replicas.
  • Correction (2026-09-28): gprd logs show RolloutWorker logged "Batch is completed with success" about 15,000 times over 24 hours, roughly one run every 6 s, with a matching count of Sidekiq until_executed dedup lines for the same worker. So runs are not paced by database-health deferrals; the worker loops almost back to back, each run re-enqueuing itself via perform_async while its own lock is still held.
    • These runs happened after !257616 (merged) reached gprd (2026-09-25 at 10:59 UTC), so the loop fix does not explain them. Not yet confirmed: how many of them provision a new namespace. Count the distinct namespace_id values in the success log lines.
  • Initial indexing starts one index per node per run, and all new indices land on the currently freest node. Observed roughly one index every 10–18 minutes on a single node.
  • Query cost comparison: on gprd a paid-first pass using the existing per-range scan took 10.45 s (scans the whole table every run); a single with_mismatched_replicas query with LIMIT 32 took 2.51 s with ~619k buffers (122,950 subscription lookups, because the subscription check runs for every row before grouping); the NOT EXISTS + LIMIT query used here took 306 ms with ~69k buffers (merge anti-join drops namespaces that already have a replica, so only 279 subscription lookups; plan). Both returned the same paid namespaces. Today's full selection takes 5.92 s.
  • Production uses one replica per namespace, so on gprd without_replicas covers nearly every namespace waiting for rollout.

References

  • Related: #630717 (Zoekt rollout backlog; analysis in an internal note on that issue).
  • Related: !257616 (merged) (fixes the empty-replica rollout loop; on gprd since 2026-09-25; independent of this MR).

Stability impact

Area Impact
Selection query While namespaces without replicas exist (nearly always today), steps 1–2 fill the batch and the ~5.9 s full scan is skipped, so selection gets cheaper overall. With the cursor, steps 1 and 2 read tens to hundreds of buffers; the unbounded scan (0.3–3.4 s cold, ~69k buffers) runs once per query per 10 minutes. SelectionService runs twice per rollout run (RolloutService#execute and #result).
Starvation Namespaces needing step 3 (too many replicas, overrides, or a second replica if the default is raised) wait while zero-replica namespaces keep arriving. This is intended, and the rollback row still applies.
RolloutWorker deferrals Runs during autovacuum on zoekt_tasks/zoekt_repositories, which it never writes, so this is safe. It still defers on its own tables and on all global indicators.
SchedulingWorker tasks The initiate job and dispatch-only tasks no longer defer during autovacuum on any Zoekt table, so they run on their normal schedule during autovacuum. The four writing tasks defer only on the tables they write, instead of all six. Jobs are deduplicated (until_executed), so queued jobs don't pile up, but more jobs cycle through Redis as deferred-then-retried. A new task that writes inside the job must be added to WRITE_TASK_TABLES, or it gets no table check.
zoekt_initial_indexing_limit No change at deploy (default 1). Biggest risk once raised: raising it to N (at most 100) gives up to Nx more InitialIndexingEvents per node per run. Each creates zoekt_repositories rows (capped at 10,000 per event before re-emitting) and later zoekt_tasks for that node. Trial namespaces are typically tiny, but paid namespaces can be large, and new indices concentrate on one node (node concurrency is 20). Can be lowered again through the API without a deploy.
Safeguards retained InitialIndexingEventWorker still defers on zoekt_repositories health; per-event insert cap of 10,000 stays; a node's task queue still drains at its own concurrency limit.
Node storage Unchanged risk profile. Planning still checks unclaimed storage, and the one-lock-per-rollout-run model is unchanged, so there's no concurrent planning.
Rollback Disable zoekt_prioritized_rollout_selection to return selection to the current behavior without a deploy. All four parts are independent and revertible. The rollout batch size setting can also be lowered at any time without a deploy.
What to watch after deploy
  • RolloutWorker duration and job rate.
  • SchedulingWorker jobs deferred per hour.
  • Search::Zoekt::Index.pending.group(:zoekt_node_id).count, to check indexing tasks aren't overloading one node.
  • After raising zoekt_initial_indexing_limit, watch the same per-node pending index count and task queue on the receiving node.
  • Zoekt task queue depth and latency on the node receiving new indices.
  • EnabledNamespace.with_missing_indices trend over time.
  • Paid namespaces still waiting: EnabledNamespace.with_rollout_allowed.with_paid_subscription.without_replicas.count.
  • Step 2 query time on gprd (EnabledNamespace.with_rollout_allowed.without_replicas.ordered_by_id.with_limit(32)).

Database queries

This MR adds two queries to Search::Zoekt::SelectionService and makes the limit of one existing query configurable. No schema changes, no new indexes, no migrations. Selection runs in Search::Zoekt::RolloutWorker (cron Sidekiq worker, not a web request); RolloutService calls SelectionService twice per run (rollout_service.rb:35 and :74).

Query 1 — paid namespaces without replicas (GitLab.com only)

New. SelectionService#fetch_paid_namespaces_without_replicas, step 1 of selection.

SELECT "zoekt_enabled_namespaces".* FROM "zoekt_enabled_namespaces"
WHERE ("zoekt_enabled_namespaces"."last_rollout_failed_at" IS NULL
       OR "zoekt_enabled_namespaces"."last_rollout_failed_at" < '2026-09-27 12:48:30')
  AND (EXISTS (SELECT 1 FROM "gitlab_subscriptions"
               WHERE "gitlab_subscriptions"."hosted_plan_name_uid" IN (3, 4, 5, 6, 7, 8, 9, 10, 11)
                 AND ("gitlab_subscriptions"."trial" = FALSE OR "gitlab_subscriptions"."trial" IS NULL)
                 AND (end_date IS NULL OR end_date >= '2026-09-28')
                 AND "gitlab_subscriptions"."namespace_id" = "zoekt_enabled_namespaces"."root_namespace_id"))
  AND (NOT EXISTS (SELECT 1 FROM "zoekt_replicas"
                   WHERE "zoekt_replicas"."zoekt_enabled_namespace_id" = "zoekt_enabled_namespaces"."id"
                     AND "zoekt_replicas"."zoekt_enabled_namespace_id" > 111111))
  AND "zoekt_enabled_namespaces"."id" > 111111
ORDER BY "zoekt_enabled_namespaces"."id" ASC LIMIT 1

LIMIT is zoekt_rollout_batch_size (1 on GitLab.com, default 32 elsewhere). 111111 is the cursor (0 on the first run in each 10-minute window).

Indexes and gprd measurement

Indexes used: zoekt_enabled_namespaces_pkey for the ORDER BY id + LIMIT, index_zoekt_replicas_on_enabled_namespace_id for the NOT EXISTS, and the unique index_gitlab_subscriptions_on_namespace_id for the EXISTS.

Query 2 — any namespace without replicas

New. SelectionService#fetch_namespaces_without_replicas, step 2 of selection.

SELECT "zoekt_enabled_namespaces".* FROM "zoekt_enabled_namespaces"
WHERE ("zoekt_enabled_namespaces"."last_rollout_failed_at" IS NULL
       OR "zoekt_enabled_namespaces"."last_rollout_failed_at" < '2026-09-27 12:48:30')
  AND (NOT EXISTS (SELECT 1 FROM "zoekt_replicas"
                   WHERE "zoekt_replicas"."zoekt_enabled_namespace_id" = "zoekt_enabled_namespaces"."id"
                     AND "zoekt_replicas"."zoekt_enabled_namespace_id" > 111111))
  AND "zoekt_enabled_namespaces"."id" > 111111
  AND "zoekt_enabled_namespaces"."id" NOT IN (1, 2)
ORDER BY "zoekt_enabled_namespaces"."id" ASC LIMIT 30

The NOT IN list is step 1's picks (at most the batch size); when step 1 returns nothing, Rails emits 1=1 instead. LIMIT is the batch size minus step 1's count. 111111 is the cursor, as in query 1. Same indexes as query 1.

Plans

Query 3 — initial indexing, limit now configurable

Changed. SchedulingService#initial_indexing, once per online node with pending indices.

SELECT "zoekt_indices".* FROM "zoekt_indices"
WHERE "zoekt_indices"."zoekt_node_id" = 2000234 AND "zoekt_indices"."state" = 0
ORDER BY "zoekt_indices"."id" ASC LIMIT 20

Only the LIMIT changed: it's now the admin setting zoekt_initial_indexing_limit (default 1, 1 to 100) instead of the constant 1.

Indexes and plan

Uses index_zoekt_indices_on_zoekt_node_id_and_id (zoekt_node_id, id), which matches the filter and the order, then filters state, with no sort. Measured on gprd on a node with pending indices and LIMIT 20: 1.5 ms, 44 buffers.

Plan: https://console.postgres.ai/shared/22a5f916-c368-4d79-aa8f-751892b4b070

Unchanged

  • The step 3 each_batch_with_mismatched_replicas scan is the same SQL as master, now skipped when steps 1–2 fill the batch. Master's full selection was measured at ~5.92 s on gprd.
  • RolloutWorker/SchedulingWorker changes only touch health-check table lists, no SQL.
  • The new setting is read via the cached ApplicationSetting.current.
Known trade-offs
  • The first run of queries 1 and 2 in each 10-minute window has no cursor and exceeds the 100 ms guideline (up to 3.4 s cold). It runs in a cron worker, once per window, and replaces the ~5.9 s selection while a backlog exists.
  • No index covers "has no replica"; the NOT EXISTS is an anti-join, so cost grows with zoekt_enabled_namespaces size.
  • The OR in with_rollout_allowed (existing scope) likely prevents use of index_zens_on_last_rollout_failed_at; plans walk the primary key in id order and filter.
  • A namespace below the cursor that becomes eligible (a trial converting to paid, a replica removed, or a rollout retry) waits up to 10 minutes for the next rescan, or is picked up by step 3.

How to set up and validate locally

  1. On GDK with Zoekt enabled, simulate SaaS (GITLAB_SIMULATE_SAAS=1) and enable the exact code search SaaS feature as usual for Zoekt development. Enable the flag with Feature.enable(:zoekt_prioritized_rollout_selection).
  2. In the Rails console, create two trial groups and one paid group (with a gitlab_subscription). Enable all three with Search::Zoekt::EnabledNamespace.create!(root_namespace_id: ...), creating the trials first so they're older.
  3. Run Search::Zoekt::SelectionService.execute(max_batch_size: 1).enabled_namespaces and confirm the paid namespace is returned instead of the oldest trial.
  4. Destroy the three enabled namespaces from step 2, since they still have no replicas and would fill the batch. Then create an enabled namespace with 2 replicas plus two with none. Run Search::Zoekt::SelectionService.execute(max_batch_size: 2).enabled_namespaces and confirm only the two without replicas are returned; with max_batch_size: 3 the one with 2 replicas is also returned.
  5. Check Search::Zoekt::RolloutWorker.database_health_check_attrs[:tables] reflects the narrowed table list, and call Search::Zoekt::SchedulingWorker.database_health_check_attrs[:block] with ['eviction'], ['initial_indexing'] and [], confirming it returns the three eviction tables, then empty, then empty.

MR acceptance checklist

Evaluate this MR against the MR acceptance checklist.

🤖 Generated with Claude Code

Edited by Ravi Kumar

Merge request reports

Loading
Loading