Wire JobPolicyPreloader into the pipeline jobs resolver (FF)

What does this MR do and why?

The GraphQL query behind the pipeline Jobs tab (project.pipeline.jobs, 20 jobs per page) issues several main-database queries per job through Ci::DeployablePolicy#has_outdated_deployment, reached from detailedStatus, userPermissions, canPlayJob, playable. Ci::Preloaders::JobPolicyPreloader (!256453 (merged)) loads all at once; this MR wires it into Resolvers::Ci::JobsResolver via before_connection_authorization, scoped by lookahead-recorded preload groups.

Selected Preloads
id, name, status nothing
userPermissions { readBuild } nothing
userPermissions { cancelBuild } deployments
playable deployments
detailedStatus, canPlayJob deployments and downstream projects

Types::Ci::JobType is reached from more entry points than the pipeline jobs page, so this page's field shape shouldn't impose a fixed preload cost on every query that touches a job. before_connection_authorization blocks get the context as a third argument; 20 existing two-arg blocks are unaffected. Feature flag batch_pipeline_job_policy_checks (gitlab_com_derisk, project actor, default off) gates it; off, behavior is unchanged, no changelog. The recorded preload groups are keyed by project id in the context, so a query spanning two projects with the flag on for only one preloads correctly for each.

Why a connection-level hook instead of a per-field batch loader

A per-field BatchLoader fires on any of the four fields and leaves two code paths to a job: batch-loaded or plain object. The preloader memoizes has_outdated_deployment? forever, so two fields could disagree. A connection hook preloads once up front, keeping one job object and avoiding changes to job_type.rb/expose_permissions.rb.

How the selection maps to preload groups

Only write abilities are gated by Ci::DeployablePolicy, so read-only permissions need no deployment data; downstream projects imply deployments, since Ci::Bridge#playable? reads deployment approvals (downstream_projects: true forces deployments: true). A bridge's detailed status needs the downstream project preloaded because it resolves :play_job for its action and :read_pipeline on the downstream pipeline for its details path. userPermissions needs no mapping table: ability_field names each subfield after its ability, intersected with Ci::BuildPolicy.all_job_write_abilities.

No new query shapes: this wires in an existing preloader and narrows when it runs, and the shapes were reviewed in !256453 (merged).

Measurement

Measured on a local GDK against a real pipeline, built entirely through the API and a .gitlab-ci.yml — nothing created from the rails console. The steps below are reproducible.

Two projects, created through the API; the second is the trigger target for the cross-project bridges. A .gitlab-ci.yml committed through the commits API produces exactly the 20 jobs the tab requests — 17 jobs plus 3 bridges.

The .gitlab-ci.yml used
stages: [build, test, deploy, cleanup]

variables:
  GIT_STRATEGY: none

compile:
  stage: build
  script: echo compiled

package:
  stage: build
  script: echo packaged

rspec:
  stage: test
  parallel: 5
  script: echo rspec

jest:
  stage: test
  script: echo jest

lint:
  stage: test
  script: echo lint

sast:
  stage: test
  script: echo sast

dependency-scanning:
  stage: test
  script: echo deps

smoke:
  stage: test
  when: manual
  script: echo smoke

child-pipeline:
  stage: test
  trigger:
    include: .gitlab-ci-child.yml

deploy-staging:
  stage: deploy
  script: echo staging
  environment:
    name: staging
    url: https://staging.example.com

deploy-review:
  stage: deploy
  script: echo review
  environment:
    name: review/main
    url: https://review.example.com
    on_stop: stop-review

deploy-canary:
  stage: deploy
  script: echo canary
  environment:
    name: canary

deploy-production:
  stage: deploy
  when: manual
  script: echo production
  environment:
    name: production
    url: https://example.com

trigger-downstream:
  stage: deploy
  trigger:
    project: GROUP/DOWNSTREAM_PROJECT

trigger-downstream-manual:
  stage: deploy
  when: manual
  trigger:
    project: GROUP/DOWNSTREAM_PROJECT

stop-review:
  stage: cleanup
  when: manual
  script: echo stopping
  environment:
    name: review/main
    action: stop

Getting the measured pipeline into the right state matters and is easy to get wrong. The page is measured on an older pipeline whose manual production deploy has been superseded. Reaching that state takes three things, all of which has_outdated_deployment? requires: the newer pipeline must run on a different commit (older_than_last_successful_deployment? returns false when the SHA matches); the job must be incomplete (when ci_forward_deployment_rollback_allowed? is set, a deploy job that already succeeded can never be outdated, so it's the manual deploy-production that qualifies); and the environment needs a successful deployment to be outdated against, which means playing deploy-production on the newer pipeline. With all three in place, the older pipeline's deploy-production comes back with updateBuild: false and cancelBuild: false while playable stays true — Ci::DeployablePolicy withholding write abilities on a manual production deploy that a newer release has already superseded. That's the case the page pays to load deployments for.

The query is the page's own, app/assets/javascripts/ci/pipeline_details/jobs/graphql/queries/get_pipeline_jobs.query.graphql, with its PageInfo fragment inlined — not a query written for the measurement. Worth stating plainly: it selects detailedStatus, playable and userPermissions { readBuild readJobArtifacts updateBuild cancelBuild }, and it does not select canPlayJob. Under the mapping in this MR the page therefore resolves to both preload groups, by way of detailedStatus.

Counts are read by toggling the flag with the admin Features API (POST /api/v4/features/batch_pipeline_job_policy_checks with value=true and project=<path> to enable, DELETE to clear), pausing after each toggle since flag state is cached in the web process. Counts come from log/development_json.log, which records db_count, db_main_count, db_ci_count, db_cached_count and duration_s for every request — no performance bar, no console. One warm-up request per flag state, then three measured requests.

20 jobs, one page, as an admin. Medians of three runs:

flag off flag on saved
queries 146 105 41 (−28%)
main database 88 66 22
CI database 58 38 20
additionally cached 34 19 15

Runs were consistent: 146/148/146 off, 103/105/105 on.

A few points worth calling out:

  1. The CI-database saving is easy to miss: the preloader also batches job_definition for every job and ci_stage plus downstream_pipeline for the bridges, so this isn't purely a main-database change.
  2. No latency claim. Wall time was 0.83–0.90s with the flag off and 0.81–1.07s with it on; database time 0.22–0.24s against 0.17–0.37s. On a warm local GDK that's noise. The result is the query count.
  3. The same query returns a byte-identical response with the flag on and off across all 20 jobs, including the withheld write abilities on the outdated production deploy. That's the assertion that matters most.
  4. On GDK, main-database reads are served by the replica, so they appear under db_main_replica_count rather than db_main_count, which reads 0. Someone reproducing this will otherwise think the main-database figure is missing.

Testing

spec/requests/api/graphql/ci/jobs_spec.rb covers N+1 counts, flag-on/off equivalence, preload-scope assertions, and three new regression examples: downstream projects are not preloaded when not needed, no N+1 on the detailed status of manual bridges, and only the flag-enabled project's preloader runs in a two-project query.

Testing details

The outdated-deploy job has a newer successful deployment to its environment, so has_outdated_deployment? withholds its write abilities and the fixture holds false values rather than being uniformly true. The current user is a member of the manual bridge's downstream project, so canPlayJob is true there and would flip to false if the preloader failed to resolve that project. The preloader specs and their EE counterpart assert that each job resolves to the environment it deploys to, and that one instance is shared within an environment but not across environments, which matters because EE's Ci::DeployablePolicy reads persisted_environment for protected_from? and protected_by?. The three regression examples were each run red against the unfixed code before the fix.

References

  • Builds on !256453 (merged) (preloader) and !256450 (merged), both merged
  • !256452 also adds include LooksAhead to this resolver for a different purpose (CI-database association preloads), gated by its own separate feature flag, so the two changes need to be sequenced or reconciled; !256451 (merged) is the other sibling
  • Related: #629310
  • Rollout: #629928

Screenshots or screen recordings

No UI change. What changes is the query shape, not the response.

How to set up and validate locally

bin/rspec spec/requests/api/graphql/ci/jobs_spec.rb spec/models/ci/preloaders/job_policy_preloader_spec.rb ee/spec/models/ee/ci/preloaders/job_policy_preloader_spec.rb
  • GraphQL type and resolver classes are held by the schema built at boot, so after checking out the branch restart the Rails server (gdk restart rails-web), or the old code keeps running even with the flag on. Feature flag state is also cached in the web process for about a minute after Feature.enable/Feature.disable.
  • To compare: open a pipeline's Jobs tab with performance bar enabled, pick getPipelineJobs, open pg details, toggle the flag, and reload.
  • The measurement in the Measurement section above is reproduced with the API steps described there.

MR acceptance checklist

Evaluate this MR against the MR acceptance checklist. It helps you analyze changes to reduce risks in quality, performance, reliability, security, and maintainability.

🤖 Generated with Claude Code

Edited by Hordur Freyr Yngvason

Merge request reports

Loading
Loading