Wire JobPolicyPreloader into the pipeline jobs resolver (FF)
What does this MR do and why?
The GraphQL query behind the pipeline Jobs tab (project.pipeline.jobs, 20 jobs per page) issues several main-database queries per job through Ci::DeployablePolicy#has_outdated_deployment, reached from detailedStatus, userPermissions, canPlayJob, playable. Ci::Preloaders::JobPolicyPreloader (!256453 (merged)) loads all at once; this MR wires it into Resolvers::Ci::JobsResolver via before_connection_authorization, scoped by lookahead-recorded preload groups.
| Selected | Preloads |
|---|---|
id, name, status |
nothing |
userPermissions { readBuild } |
nothing |
userPermissions { cancelBuild } |
deployments |
playable |
deployments |
detailedStatus, canPlayJob |
deployments and downstream projects |
Types::Ci::JobType is reached from more entry points than the pipeline jobs page, so this page's field shape shouldn't impose a fixed preload cost on every query that touches a job. before_connection_authorization blocks get the context as a third argument; 20 existing two-arg blocks are unaffected. Feature flag batch_pipeline_job_policy_checks (gitlab_com_derisk, project actor, default off) gates it; off, behavior is unchanged, no changelog. The recorded preload groups are keyed by project id in the context, so a query spanning two projects with the flag on for only one preloads correctly for each.
Why a connection-level hook instead of a per-field batch loader
A per-field BatchLoader fires on any of the four fields and leaves two code paths to a job: batch-loaded or plain object. The preloader memoizes has_outdated_deployment? forever, so two fields could disagree. A connection hook preloads once up front, keeping one job object and avoiding changes to job_type.rb/expose_permissions.rb.
How the selection maps to preload groups
Only write abilities are gated by Ci::DeployablePolicy, so read-only permissions need no deployment data; downstream projects imply deployments, since Ci::Bridge#playable? reads deployment approvals (downstream_projects: true forces deployments: true). A bridge's detailed status needs the downstream project preloaded because it resolves :play_job for its action and :read_pipeline on the downstream pipeline for its details path. userPermissions needs no mapping table: ability_field names each subfield after its ability, intersected with Ci::BuildPolicy.all_job_write_abilities.
No new query shapes: this wires in an existing preloader and narrows when it runs, and the shapes were reviewed in !256453 (merged).
Measurement
Measured on a local GDK against a real pipeline, built entirely through the API and a .gitlab-ci.yml — nothing created from the rails console. The steps below are reproducible.
Two projects, created through the API; the second is the trigger target for the cross-project bridges. A .gitlab-ci.yml committed through the commits API produces exactly the 20 jobs the tab requests — 17 jobs plus 3 bridges.
The .gitlab-ci.yml used
stages: [build, test, deploy, cleanup]
variables:
GIT_STRATEGY: none
compile:
stage: build
script: echo compiled
package:
stage: build
script: echo packaged
rspec:
stage: test
parallel: 5
script: echo rspec
jest:
stage: test
script: echo jest
lint:
stage: test
script: echo lint
sast:
stage: test
script: echo sast
dependency-scanning:
stage: test
script: echo deps
smoke:
stage: test
when: manual
script: echo smoke
child-pipeline:
stage: test
trigger:
include: .gitlab-ci-child.yml
deploy-staging:
stage: deploy
script: echo staging
environment:
name: staging
url: https://staging.example.com
deploy-review:
stage: deploy
script: echo review
environment:
name: review/main
url: https://review.example.com
on_stop: stop-review
deploy-canary:
stage: deploy
script: echo canary
environment:
name: canary
deploy-production:
stage: deploy
when: manual
script: echo production
environment:
name: production
url: https://example.com
trigger-downstream:
stage: deploy
trigger:
project: GROUP/DOWNSTREAM_PROJECT
trigger-downstream-manual:
stage: deploy
when: manual
trigger:
project: GROUP/DOWNSTREAM_PROJECT
stop-review:
stage: cleanup
when: manual
script: echo stopping
environment:
name: review/main
action: stopGetting the measured pipeline into the right state matters and is easy to get wrong. The page is measured on an older pipeline whose manual production deploy has been superseded. Reaching that state takes three things, all of which has_outdated_deployment? requires: the newer pipeline must run on a different commit (older_than_last_successful_deployment? returns false when the SHA matches); the job must be incomplete (when ci_forward_deployment_rollback_allowed? is set, a deploy job that already succeeded can never be outdated, so it's the manual deploy-production that qualifies); and the environment needs a successful deployment to be outdated against, which means playing deploy-production on the newer pipeline. With all three in place, the older pipeline's deploy-production comes back with updateBuild: false and cancelBuild: false while playable stays true — Ci::DeployablePolicy withholding write abilities on a manual production deploy that a newer release has already superseded. That's the case the page pays to load deployments for.
The query is the page's own, app/assets/javascripts/ci/pipeline_details/jobs/graphql/queries/get_pipeline_jobs.query.graphql, with its PageInfo fragment inlined — not a query written for the measurement. Worth stating plainly: it selects detailedStatus, playable and userPermissions { readBuild readJobArtifacts updateBuild cancelBuild }, and it does not select canPlayJob. Under the mapping in this MR the page therefore resolves to both preload groups, by way of detailedStatus.
Counts are read by toggling the flag with the admin Features API (POST /api/v4/features/batch_pipeline_job_policy_checks with value=true and project=<path> to enable, DELETE to clear), pausing after each toggle since flag state is cached in the web process. Counts come from log/development_json.log, which records db_count, db_main_count, db_ci_count, db_cached_count and duration_s for every request — no performance bar, no console. One warm-up request per flag state, then three measured requests.
20 jobs, one page, as an admin. Medians of three runs:
| flag off | flag on | saved | |
|---|---|---|---|
| queries | 146 | 105 | 41 (−28%) |
| main database | 88 | 66 | 22 |
| CI database | 58 | 38 | 20 |
| additionally cached | 34 | 19 | 15 |
Runs were consistent: 146/148/146 off, 103/105/105 on.
A few points worth calling out:
- The CI-database saving is easy to miss: the preloader also batches
job_definitionfor every job andci_stageplusdownstream_pipelinefor the bridges, so this isn't purely a main-database change. - No latency claim. Wall time was 0.83–0.90s with the flag off and 0.81–1.07s with it on; database time 0.22–0.24s against 0.17–0.37s. On a warm local GDK that's noise. The result is the query count.
- The same query returns a byte-identical response with the flag on and off across all 20 jobs, including the withheld write abilities on the outdated production deploy. That's the assertion that matters most.
- On GDK, main-database reads are served by the replica, so they appear under
db_main_replica_countrather thandb_main_count, which reads 0. Someone reproducing this will otherwise think the main-database figure is missing.
Testing
spec/requests/api/graphql/ci/jobs_spec.rb covers N+1 counts, flag-on/off equivalence, preload-scope assertions, and three new regression examples: downstream projects are not preloaded when not needed, no N+1 on the detailed status of manual bridges, and only the flag-enabled project's preloader runs in a two-project query.
Testing details
The outdated-deploy job has a newer successful deployment to its environment, so has_outdated_deployment? withholds its write abilities and the fixture holds false values rather than being uniformly true. The current user is a member of the manual bridge's downstream project, so canPlayJob is true there and would flip to false if the preloader failed to resolve that project. The preloader specs and their EE counterpart assert that each job resolves to the environment it deploys to, and that one instance is shared within an environment but not across environments, which matters because EE's Ci::DeployablePolicy reads persisted_environment for protected_from? and protected_by?. The three regression examples were each run red against the unfixed code before the fix.
References
- Builds on !256453 (merged) (preloader) and !256450 (merged), both merged
- !256452 also adds
include LooksAheadto this resolver for a different purpose (CI-database association preloads), gated by its own separate feature flag, so the two changes need to be sequenced or reconciled; !256451 (merged) is the other sibling - Related: #629310
- Rollout: #629928
Screenshots or screen recordings
No UI change. What changes is the query shape, not the response.
How to set up and validate locally
bin/rspec spec/requests/api/graphql/ci/jobs_spec.rb spec/models/ci/preloaders/job_policy_preloader_spec.rb ee/spec/models/ee/ci/preloaders/job_policy_preloader_spec.rb- GraphQL type and resolver classes are held by the schema built at boot, so after checking out the branch restart the Rails server (
gdk restart rails-web), or the old code keeps running even with the flag on. Feature flag state is also cached in the web process for about a minute afterFeature.enable/Feature.disable. - To compare: open a pipeline's Jobs tab with performance bar enabled, pick
getPipelineJobs, openpgdetails, toggle the flag, and reload. - The measurement in the Measurement section above is reproduced with the API steps described there.
MR acceptance checklist
Evaluate this MR against the MR acceptance checklist. It helps you analyze changes to reduce risks in quality, performance, reliability, security, and maintainability.