Persist billable usage locally on air-gapped instances

What does this MR do and why?

!248498 (merged) added the billable_usage_daily_aggregates table in %19.4, but nothing writes to it. This adds the in-process write path: on an instance licensed offline, a billable event folds into that table instead of being emitted to a collector the instance cannot reach. Feature teams keep calling track_billing_event and change nothing.

Behind the local_billing_persistence feature flag, default off.

References

  • Issue: gitlab-org/customers-gitlab-com#18551
  • Predecessor: !248498 (merged), which added the table this MR writes to
  • Epic: &23047, under &22758

How to set up and validate locally

Console recipe

Two flags, not one. record_usage returns early on billing_event_tracking before it reaches the local-persistence branch, so enabling only local_billing_persistence records nothing and says nothing.

Feature.enable(:billing_event_tracking)
Feature.enable(:local_billing_persistence)

# License.current is not guaranteed to return the same object each call,
# so pin one instance rather than stubbing the method on it.
lic = License.current
lic.define_singleton_method(:offline_cloud_license?) { true }
License.define_singleton_method(:current) { lic }

Gitlab::BillingEvents::Client.new.local_persistence_enabled? # => true

ns = Group.first

# Counter: three events fold into one row.
3.times do
  Gitlab::BillingEvents::Client.track_billing_event(
    event_type: 'secrets_read', category: 'ConsoleTest', unit_of_measure: 'request',
    quantity: 1, namespace: ns, metadata: { feature_qualified_name: 'secrets_read' })
end

# Snapshot: the second reading supersedes the first.
[1150, 1162].each do |q|
  Gitlab::BillingEvents::Client.track_billing_snapshot(
    event_type: 'secrets_stored', category: 'ConsoleTest', unit_of_measure: 'secret',
    quantity: q, namespace: ns, metadata: { feature_qualified_name: 'secrets_stored' })
end

Utilization::BillableUsage::DailyAggregate.order(:id).pluck(
  :event_type, :feature_qualified_name, :root_namespace_id,
  :operation_type, :quantity, :events_count)

Gives secrets_read at quantity: 3, events_count: 3 and secrets_stored at quantity: 1162, events_count: 2 — summing would have given 2312, so that row is where counter and snapshot visibly diverge.

track_billing_event swallows exceptions, so a mistake here is quiet. Check log/development.log for BillingEvents: billing event tracked to be absent (emission was replaced) and BillingEvents: internal event tracked to be present (the internal event still fires).

Clean up with Utilization::BillableUsage::DailyAggregate.where(event_type: %w[secrets_read secrets_stored]).delete_all.

Verified locally against the migrated schema:

case result
counter ×3 quantity=3, events_count=3
snapshot ×2, same scope quantity=1162, events_count=2
two root namespaces separate rows, neither overwritten
two operation types separate rows, distinct event_aggregate_uuid
event_aggregate_uuid stability unchanged across repeated writes to one row
missing feature_qualified_name logged, no row, no raise
single event over the quantity ceiling rejected, no row
22:30Z from UTC+02 usage_date derived in UTC

The service spec runs 14 examples, the FOSS client spec 29, and the EE client spec 12 — all green. The FOSS spec is the regression guard for the refactor, and now also covers track_billing_snapshot and the default of local_persistence_enabled?, which the refactor had introduced untested.

Design decisions

Emission is replaced, not supplemented. An air-gapped instance cannot reach the billing collector, and EventEligibilityChecker discards the billing context before any HTTP attempt is made when Snowplow and product usage data are both off, which is the normal configuration there. The client still logs success, so nothing signals the loss. Only emission is swapped out: the internal event works without connectivity and is still tracked on both paths.

Counter versus snapshot. track_billing_snapshot joins track_billing_event for point-in-time measures. Which method you call decides whether a repeated write in the same day accumulates or supersedes, so the distinction cannot be got wrong at the call site, and it needs no column on the table.

The aggregate key. Rows are keyed on (usage_date, event_type, feature_qualified_name, root_namespace_id, operation_type), and the aggregate's uuid is derived from that tuple. root_namespace_id matters because secrets_stored emits one snapshot per root namespace and a snapshot replaces rather than adds; without it those writes would overwrite each other. operation_type is recorded as reported, so CustomersDot can apply NON_BILLABLE_OPERATION_TYPES to separate rows.

No metadata is stored. Product confirmed on &22758 that air-gapped Duo Agent Platform pricing is flat and token-independent, and asked that token counts not be collected at all, which left nothing both populated here and read downstream. Emitters still send metadata; the aggregate reads feature_qualified_name and operation_type out of it and keeps none of it.

Emitter change. secrets_stored moves to the snapshot path, since it reports how many secrets exist rather than tallying reads. That is the only change to the Secrets Manager emitters here — feature_qualified_name and the removal of the realm gate both landed on master independently while this was in review.

Documentation. The developer page gains a section on counters versus snapshots, so the choice of method is deliberate at the call site, and a short note that an air-gapped instance stores usage locally. The feature_qualified_name requirement sits on the metadata parameter itself, since it holds on every path rather than only this one.

Known limitations

Two, both deliberate for this iteration: a snapshot cannot report zero, and accumulated quantity is unbounded. Neither is reachable by the only producer on this path today.

A snapshot cannot report zero. record_usage rejects quantity <= 0, so a scope that empties during the day leaves the earlier, higher reading standing as that day's figure rather than superseding it. Correct for a counter, imprecise for a snapshot. Allowing quantity == 0 on the snapshot path is the fix when a producer needs it.

Accumulated quantity is not checked against the Iglu ceiling. The model validates the incoming event, but the upsert adds to the stored value, bypasses validations, and meets no constraint beyond >= 0 on a numeric(14,4). A high-volume day could therefore build a row that exceeds 2,147,483,647 and cannot be exported. Secrets Manager, the only producer that reaches this path, will not approach it; a quantity <= 2147483647 check constraint would move the failure to write time if that changes.

Database

Low risk: one single-row upsert per billable event, on air-gapped instances only, behind a flag that is off by default. No migration, no backfill, no new index — !248498 (merged) added the table and both of its indexes.

The conflict clause sets updated_at explicitly because a custom on_duplicate string replaces Rails' generated SET clause wholesale, so without it updated_at would never move on conflict. The timestamps themselves come from the model's own record_timestamps, which resolves to CURRENT_TIMESTAMP — so EXCLUDED.updated_at and now() are equivalent here.

INSERT INTO "billable_usage_daily_aggregates"
  ("usage_date","event_type","feature_qualified_name","root_namespace_id","operation_type",
   "unit_of_measure","quantity","events_count","event_aggregate_uuid","created_at","updated_at")
VALUES ('2026-08-10', 'secrets_read', 'secrets_read', 1, NULL, 'request', 1.0, 1,
        'd93ecd1b-fe67-5d1f-bea6-aac5e4fdd2bb', CURRENT_TIMESTAMP, CURRENT_TIMESTAMP)
ON CONFLICT ("usage_date","event_type","feature_qualified_name","root_namespace_id","operation_type")
DO UPDATE SET
  quantity = billable_usage_daily_aggregates.quantity + EXCLUDED.quantity,
  events_count = billable_usage_daily_aggregates.events_count + EXCLUDED.events_count,
  updated_at = EXCLUDED.updated_at
RETURNING "id"
Insert on billable_usage_daily_aggregates  (cost=0.00..0.02 rows=1 width=208)
  Conflict Resolution: UPDATE
  Conflict Arbiter Indexes: index_billable_usage_daily_aggs_on_unique_tuple
  ->  Result  (cost=0.00..0.02 rows=1 width=208)

The arbiter resolves to the five-column tuple index with operation_type NULL, which is what that index's NULLS NOT DISTINCT is there for: without it, every NULL-scoped write would insert a new row instead of folding into the day's.

Expected row counts. One row per key per day, on air-gapped instances only. Secrets Manager contributes at most one secrets_stored row per root namespace per day, plus one secrets_read row per root namespace per day. Writes are synchronous, so every secrets_read on an instance contends on a single row per namespace; that suits the current producer, and a higher-volume one would batch behind a worker first.

Changelog: other EE: true

MR acceptance checklist

Evaluate this MR against the MR acceptance checklist.

Edited by Vijay Hawoldar

Merge request reports

Loading
Loading