Geo: Deduplicate registry_count query in GeoNodeStatus

What does this MR do and why?

Geo: Deduplicate registry_count query in GeoNodeStatus

Relates to #606494 (closed)

On a Geo secondary, GeoNodeStatus collects both a _count and a _registry_count field for every replicable.

Since we can't query the primary's total from the secondary, both fields are backed by the exact same value:- replicator.registry_count and we've been running that count twice per replicator, per metrics collection, since the fields were introduced years ago.

This MR extends collect_metric to accept multiple field names for a single computed value, so the registry count runs once and is assigned to both fields. No behaviour change: same fields, same values, same Prometheus gauges (geo_<replicable> and geo_<replicable>_registry are still emitted separately, from the same number as before).

Why it matters

registry_count is a batch_count over the whole registry table. On a locally benchmarked job_artifact_registry with ~100M rows, a single full count took ~40–90s (43s and 94s across the two identical runs), issued as ~1000 batched COUNT statements plus MAX(id). The duplicate therefore adds roughly 40–90s and ~1000 redundant queries for every replicator with a table of this size, on every collection — and Geo::MetricsUpdateWorker runs every minute, so this is constant unneeded load on the tracking database and contributes to metric count lag on large sites.

With 19 replication-enabled replicators, this removes one full batch_count pass per replicator per cycle.

How to set up and validate locally

Requires a Geo primary + secondary GDK pair with some replicated data.

  1. On the secondary, confirm both fields still match and are populated:
    # rails console
    status = GeoNodeStatus.current_node_status
    status.lfs_objects_count == status.lfs_objects_registry_count  # => true
  2. Confirm the count only runs once per replicator — record SQL while collecting status and check the whole-table registry count appears once (previously twice):
    statements = []
    ActiveSupport::Notifications.subscribe('sql.active_record') do |*, payload|
      statements << payload[:sql] if payload[:sql].include?('lfs_object_registry')
    end
    GeoNodeStatus.current_node_status
    statements.grep(/SELECT MAX\("lfs_object_registry"."id"\) FROM "lfs_object_registry" \z/).count
    # before: 2, after: 1
    Verified on a real secondary: 19 → 17 statements for lfs_objects (the unfiltered MAX(id) + COUNT pair now runs once). On a ~100M-row registry that removed pass is ~1000 batched statements.
  3. Confirm Prometheus output is unchanged: compare GeoNodeStatus::PROMETHEUS_METRICS.map { |c, _| status[c] } before and after this change — verified on a secondary with a ~100M-row job_artifact_registry: 152 gauges, 149 status fields, identical apart from time-varying lag/timestamp fields.
  4. Specs:
    bin/rspec ee/spec/models/geo_node_status_spec.rb
    # 703 examples, 0 failures, 1 pending

MR acceptance checklist

Evaluate this MR against the MR acceptance checklist. It helps you analyze changes to reduce risks in quality, performance, reliability, security, and maintainability.

Related to #606494 (closed)

Edited by Scott Murray

Merge request reports

Loading
Loading