Geo: Deduplicate registry_count query in GeoNodeStatus
What does this MR do and why?
Geo: Deduplicate registry_count query in GeoNodeStatus
Relates to #606494 (closed)
On a Geo secondary, GeoNodeStatus collects both a _count and a _registry_count field for every replicable.
Since we can't query the primary's total from the secondary, both fields are backed by the exact same value:- replicator.registry_count and we've been running that count twice per replicator, per metrics collection, since the fields were introduced years ago.
This MR extends collect_metric to accept multiple field names for a single computed value, so the registry count runs once and is assigned to both fields. No behaviour change: same fields, same values, same Prometheus gauges (geo_<replicable> and geo_<replicable>_registry are still emitted separately, from the same number as before).
Why it matters
registry_count is a batch_count over the whole registry table. On a locally benchmarked job_artifact_registry with ~100M rows, a single full count took ~40–90s (43s and 94s across the two identical runs), issued as ~1000 batched COUNT statements plus MAX(id). The duplicate therefore adds roughly 40–90s and ~1000 redundant queries for every replicator with a table of this size, on every collection — and Geo::MetricsUpdateWorker runs every minute, so this is constant unneeded load on the tracking database and contributes to metric count lag on large sites.
With 19 replication-enabled replicators, this removes one full batch_count pass per replicator per cycle.
How to set up and validate locally
Requires a Geo primary + secondary GDK pair with some replicated data.
- On the secondary, confirm both fields still match and are populated:
# rails console status = GeoNodeStatus.current_node_status status.lfs_objects_count == status.lfs_objects_registry_count # => true - Confirm the count only runs once per replicator — record SQL while collecting status and check the whole-table registry count appears once (previously twice):
Verified on a real secondary: 19 → 17 statements for
statements = [] ActiveSupport::Notifications.subscribe('sql.active_record') do |*, payload| statements << payload[:sql] if payload[:sql].include?('lfs_object_registry') end GeoNodeStatus.current_node_status statements.grep(/SELECT MAX\("lfs_object_registry"."id"\) FROM "lfs_object_registry" \z/).count # before: 2, after: 1lfs_objects(the unfilteredMAX(id)+COUNTpair now runs once). On a ~100M-row registry that removed pass is ~1000 batched statements. - Confirm Prometheus output is unchanged: compare
GeoNodeStatus::PROMETHEUS_METRICS.map { |c, _| status[c] }before and after this change — verified on a secondary with a ~100M-rowjob_artifact_registry: 152 gauges, 149 status fields, identical apart from time-varying lag/timestamp fields. - Specs:
bin/rspec ee/spec/models/geo_node_status_spec.rb # 703 examples, 0 failures, 1 pending
MR acceptance checklist
Evaluate this MR against the MR acceptance checklist. It helps you analyze changes to reduce risks in quality, performance, reliability, security, and maintainability.
Related to #606494 (closed)