Draft: Preload pipeline jobs page CI associations with LooksAhead (FF)

What does this MR do and why?

Pipeline jobs page fields such as tags, retryable, artifacts, stage, refPath and stuck each loaded job_definition, job_artifacts, downstream_pipeline or project per job. Resolvers::Ci::JobsResolver now uses LooksAhead to preload these once per page, only for the selected fields.

The relation mixes builds, bridges and generic statuses, and Rails' preloader raises when a class lacks an association. Ci::Preloaders::CommitStatusRelationExtension overrides Rails' preload_associations hook to route the preloads through CommitStatusPreloader. The hook is :nodoc:, so a unit spec catches a rename. It also passes the project already loaded through the pipeline as available_records:, so jobs, their artifacts and downstream pipelines in the same project reuse it instead of loading it again. Strict-loading relations skip this, because that project is not strict loading.

The feature flag batch_preload_pipeline_job_associations (gitlab_com_derisk, project actor, default off) gates this. The policy-backed fields were batched separately in !255978 (merged). A new request spec example covers the full page and fails with either flag off.

The preloads replace per-row LIMIT 1 lookups with these queries, on the same indexed columns:

  • p_ci_job_artifacts WHERE job_id IN (...) AND partition_id = ? (all artifacts, and the archive file_type)
  • p_ci_job_definition_instances JOIN p_ci_job_definitions WHERE job_id IN (...)
  • ci_sources_pipelines WHERE source_job_id IN (...)
  • projects WHERE id IN (...) for projects not already loaded, with their namespaces and namespace routes

Measurement

The pipeline Jobs tab's own query (getPipelineJobs, 20 jobs, as an admin) on a local GDK, in all four flag states. Medians of three requests after one warm-up:

batch_preload_pipeline_job_associations batch_pipeline_job_policy_checks Queries Main DB CI DB Cached
off off 146 88 58 34
on off 100 77 23 25
off on 104 66 38 19
on on 83 60 23 18
  • With the policy flag off, this MR saves 46 queries; with it on, 21. Both flags together take the page from 146 to 83 (−43%).
  • Most of the saving is on the CI database (58 → 23), and the main database drops too. The response is byte-identical in all four states.
Where the savings come from
  • CI database: per-job lookups of job_artifacts, job_definition and downstream_pipeline become one query each. With the policy flag on, which already batches job_definition and the bridges' downstream pipelines, the 17 per-job job_artifacts lookups become one query, and one new job_messages preload makes the net saving 15.
  • Main database: the jobs, their artifacts and the child pipeline all point at the measured project, and now reuse the instance the page already loaded. With the policy flag alone, the child pipeline's project is a second copy, and checks on it load its namespace, group, namespace settings, organization and route again.
How the numbers were produced

The script below builds the scenario through the API only and then measures. It creates two projects. The first has a .gitlab-ci.yml that gives exactly the 20 jobs the tab requests: 17 jobs, a child pipeline trigger and two cross-project triggers into the second project.

The measured pipeline is an older one. A newer pipeline on a different commit has deployed to production, so the older pipeline's manual deploy-production job is outdated: it returns playable: true with updateBuild: false and cancelBuild: false. This makes the page load deployments, which is the case the policy preloads exist for.

For each flag state, the script sets both flags for the project through the admin Features API and waits 75 seconds, because Puma caches flag state for about a minute. It then sends the page's query once to warm up and three times to measure, reading db_count, db_main_count, db_main_replica_count, db_ci_count and db_cached_count from log/development_json.log by request ID.

GDK serves main-database reads from the replica, so the Main DB column adds db_main_replica_count to db_main_count.

References

Screenshots or screen recordings

No UI change.

How to set up and validate locally

bin/rspec spec/requests/api/graphql/ci/jobs_spec.rb spec/models/ci/preloaders/commit_status_relation_extension_spec.rb

To reproduce the measurement:

  1. Check out this branch and restart Rails (gdk restart rails-web). The GraphQL schema is built at boot, so a running server keeps the old resolver.

  2. Make sure the GDK runner is up. The scenario runs real pipelines.

  3. Create a personal access token for an admin with the api and admin_mode scopes.

  4. Save the script below as jobs_page_bench.sh and run it from the gitlab directory. It needs curl and jq and takes about 15 minutes. It prints the table above, plus each state's response hash and the outdated deploy's permissions.

    GITLAB_TOKEN=<token> ./jobs_page_bench.sh
jobs_page_bench.sh
#!/usr/bin/env bash
# Measures the pipeline Jobs tab query (getPipelineJobs) in all four states of
# batch_preload_pipeline_job_associations x batch_pipeline_job_policy_checks.
#
# Run from the gitlab directory of a GDK with a runner. Needs curl and jq.
#   GITLAB_TOKEN=<admin PAT, scopes api + admin_mode> ./jobs_page_bench.sh
set -euo pipefail

GITLAB_URL=${GITLAB_URL:-https://gdk.test:3000}
: "${GITLAB_TOKEN:?set GITLAB_TOKEN to an admin personal access token}"
LOG=${LOG:-log/development_json.log}
FLAG_WAIT=${FLAG_WAIT:-75} # Puma caches feature flag state for about a minute
RUNS=3

api() { curl -fsSk -H "PRIVATE-TOKEN: $GITLAB_TOKEN" "$@"; }

wait_for() { # wait_for <description> <command...>: retry until the command succeeds
  local what=$1; shift
  for _ in $(seq 1 120); do "$@" && return 0; sleep 5; done
  echo "timed out waiting for $what" >&2; exit 1
}

pipeline_for_sha() {
  api "$GITLAB_URL/api/v4/projects/$1/pipelines?sha=$2" | jq -er '.[0].id'
}

pipeline_done() {
  local status
  status=$(api "$GITLAB_URL/api/v4/projects/$1/pipelines/$2" | jq -r .status)
  [[ $status =~ ^(success|failed|canceled|skipped|manual)$ ]]
}

job_id() {
  api "$GITLAB_URL/api/v4/projects/$1/pipelines/$2/jobs?per_page=100&include_retried=false" |
    jq -er --arg name "$3" --arg status "$4" '.[] | select(.name == $name and .status == $status) | .id'
}

commit() { # commit <project id> <message> <actions json>
  api -X POST -H 'Content-Type: application/json' "$GITLAB_URL/api/v4/projects/$1/repository/commits" \
    -d "$(jq -n --arg msg "$2" --argjson actions "$3" '{branch: "main", commit_message: $msg, actions: $actions}')" |
    jq -r .id
}

create_project() {
  api -X POST "$GITLAB_URL/api/v4/projects" -d "name=$1" -d initialize_with_readme=true -d visibility=private |
    jq -r '"\(.id) \(.path_with_namespace)"'
}

set_flag() { # set_flag <name> <true|false> <project path>
  api -X POST "$GITLAB_URL/api/v4/features/$1" -d "value=$2" --data-urlencode "project=$3" > /dev/null
}

# 1. Scenario: a pipeline with the job mix of the Jobs tab, superseded by a newer one
suffix=$(date +%s)
read -r downstream_id downstream_path < <(create_project "jobs-bench-downstream-$suffix")
read -r project_id project_path < <(create_project "jobs-bench-$suffix")
echo "projects: $project_path, $downstream_path"

commit "$downstream_id" "Add downstream pipeline" "$(jq -n '[{action: "create", file_path: ".gitlab-ci.yml",
  content: "downstream-job:\n  stage: build\n  script: echo downstream\n"}]')" > /dev/null

ci=$(cat <<YAML
stages: [build, test, deploy, cleanup]

variables:
  GIT_STRATEGY: none

compile:
  stage: build
  script: echo compiled

package:
  stage: build
  script: echo packaged

rspec:
  stage: test
  parallel: 5
  script: echo rspec

jest:
  stage: test
  script: echo jest

lint:
  stage: test
  script: echo lint

sast:
  stage: test
  script: echo sast

dependency-scanning:
  stage: test
  script: echo deps

smoke:
  stage: test
  when: manual
  script: echo smoke

child-pipeline:
  stage: test
  trigger:
    include: .gitlab-ci-child.yml

deploy-staging:
  stage: deploy
  script: echo staging
  environment:
    name: staging
    url: https://staging.example.com

deploy-review:
  stage: deploy
  script: echo review
  environment:
    name: review/main
    url: https://review.example.com
    on_stop: stop-review

deploy-canary:
  stage: deploy
  script: echo canary
  environment:
    name: canary

deploy-production:
  stage: deploy
  when: manual
  script: echo production
  environment:
    name: production
    url: https://example.com

trigger-downstream:
  stage: deploy
  trigger:
    project: $downstream_path

trigger-downstream-manual:
  stage: deploy
  when: manual
  trigger:
    project: $downstream_path

stop-review:
  stage: cleanup
  when: manual
  script: echo stopping
  environment:
    name: review/main
    action: stop
YAML
)
child=$'child-build:\n  stage: build\n  script: echo child build\n\nchild-test:\n  stage: test\n  script: echo child test\n'

old_sha=$(commit "$project_id" "Add pipeline with the job mix of the Jobs tab" "$(jq -n --arg ci "$ci" --arg child "$child" \
  '[{action: "create", file_path: ".gitlab-ci.yml", content: $ci},
    {action: "create", file_path: ".gitlab-ci-child.yml", content: $child}]')")
wait_for "the first pipeline" pipeline_for_sha "$project_id" "$old_sha" > /dev/null
old_pipeline=$(pipeline_for_sha "$project_id" "$old_sha")
wait_for "pipeline $old_pipeline to finish" pipeline_done "$project_id" "$old_pipeline"

# A newer pipeline on a different commit deploys to production, which makes the
# older pipeline's manual production deploy outdated.
new_sha=$(commit "$project_id" "Second revision" "$(jq -n '[{action: "update", file_path: "README.md", content: "Second revision\n"}]')")
wait_for "the second pipeline" pipeline_for_sha "$project_id" "$new_sha" > /dev/null
new_pipeline=$(pipeline_for_sha "$project_id" "$new_sha")
wait_for "deploy-production to be playable" job_id "$project_id" "$new_pipeline" deploy-production manual > /dev/null
api -X POST "$GITLAB_URL/api/v4/projects/$project_id/jobs/$(job_id "$project_id" "$new_pipeline" deploy-production manual)/play" > /dev/null
wait_for "deploy-production to succeed" job_id "$project_id" "$new_pipeline" deploy-production success > /dev/null

old_iid=$(api "$GITLAB_URL/api/v4/projects/$project_id/pipelines/$old_pipeline" | jq -r .iid)
echo "measuring pipeline $old_iid (superseded by $(api "$GITLAB_URL/api/v4/projects/$project_id/pipelines/$new_pipeline" | jq -r .iid))"

# 2. Measurement: the page's own query, counts from the request log
query=$(sed '/^#import/d' app/assets/javascripts/ci/pipeline_details/jobs/graphql/queries/get_pipeline_jobs.query.graphql
        cat app/assets/javascripts/graphql_shared/fragments/page_info.fragment.graphql)
body=$(jq -n --arg q "$query" --arg path "$project_path" --arg iid "$old_iid" \
  '{query: $q, variables: {fullPath: $path, iid: $iid}}')

response=$(mktemp)
measure() { # prints one JSON line of counts for one request
  local headers rid line
  headers=$(mktemp)
  api -D "$headers" -H 'Content-Type: application/json' "$GITLAB_URL/api/graphql" -d "$body" > "$response"
  jq -e '.errors == null' "$response" > /dev/null
  rid=$(awk 'tolower($1) == "x-request-id:" { print $2 }' "$headers" | tr -d '\r')
  for _ in $(seq 1 25); do
    line=$(grep -F "\"correlation_id\":\"$rid\"" "$LOG" | grep -F '"path":"/api/graphql"' || true)
    [[ -n $line ]] && break
    sleep 0.2
  done
  jq -c --arg hash "$(jq -S .data "$response" | shasum | cut -c1-12)" \
    --argjson jobs "$(jq '.data.project.pipeline.jobs.nodes | length' "$response")" \
    --argjson prod "$(jq -c '.data.project.pipeline.jobs.nodes[] | select(.name == "deploy-production") | {playable} + .userPermissions' "$response")" \
    '{total: .db_count, main: ((.db_main_count // 0) + (.db_main_replica_count // 0)),
      ci: ((.db_ci_count // 0) + (.db_ci_replica_count // 0)), cached: .db_cached_count, $jobs, $hash, $prod}' <<< "$line"
}

echo
echo "| \`batch_preload_pipeline_job_associations\` | \`batch_pipeline_job_policy_checks\` | Queries | Main DB | CI DB | Cached | Raw totals |"
echo "|---|---|---|---|---|---|---|"
for state in "false false" "true false" "false true" "true true"; do
  read -r preload policy <<< "$state"
  set_flag batch_preload_pipeline_job_associations "$preload" "$project_path"
  set_flag batch_pipeline_job_policy_checks "$policy" "$project_path"
  sleep "$FLAG_WAIT"

  measure > /dev/null # warm-up
  runs=$(for _ in $(seq 1 $RUNS); do measure; done | jq -s .)
  jq -r --arg preload "$preload" --arg policy "$policy" '
    def median(f): map(f) | sort | .[length / 2 | floor];
    def onoff: if . == "true" then "**on**" else "off" end;
    "| \($preload | onoff) | \($policy | onoff) | \(median(.total)) | \(median(.main)) | \(median(.ci)) | \(median(.cached)) | \(map(.total) | join("/")) |"
  ' <<< "$runs"
  jq -r '"jobs: \(map(.jobs) | unique | join(",")), response hashes: \(map(.hash) | unique | join(",")), deploy-production: \(.[0].prod)"' <<< "$runs" >&2
done

set_flag batch_preload_pipeline_job_associations false "$project_path"
set_flag batch_pipeline_job_policy_checks false "$project_path"
rm -f "$response"

MR acceptance checklist

Evaluate this MR against the MR acceptance checklist. It helps you analyze changes to reduce risks in quality, performance, reliability, security, and maintainability.

Edited by Hordur Freyr Yngvason

Merge request reports

Loading
Loading