Raise the GLQL dashboard request concurrency to 8
What does this MR do and why?
This MR raises QUEUES[EXECUTION_QUEUE_DASHBOARD].concurrency in app/assets/javascripts/glql/core/executor.js from 6 to 8. Nothing else in that file changes. In spec/frontend/glql/core/executor_spec.js, the assertion on concurrencyLimit moves from 6 to 8. Two result cache specs block every dashboard slot with requests that never finish. They now start eight blocking requests instead of six. Their expected request counts and release indexes move up by two to match. The whole diff is 2 files, 15 lines added and 15 removed.
The default queue glql-queue-default stays at 1. GLQL blocks embedded in issues, epics and merge requests use it. The value of 1 is a readiness commitment from the GLQL beta review in issue 517546.
The queue is TaskQueue in app/assets/javascripts/glql/utils/task_queue.js. It is a static per-page queue. Every GLQL panel on one analytics dashboard shares it. The value caps how many POST /api/glql requests one viewer's browser keeps in flight at once. Requests beyond the cap wait in the browser. A task that has started keeps its slot until the server answers. Waiting tasks run lowest priority first, then in arrival order.
app/assets/javascripts/analytics/analytics_dashboards/components/visualizations/glql.vue binds :queue="$options.EXECUTION_QUEUE_DASHBOARD" on every request it makes. That covers the main query, the previous-period comparison query and load-more pages. Pages within one chart stay sequential, because each page needs the previous page's cursor. This MR only changes parallelism across panels.
History: !256168 (merged) gave dashboards their own queue at 4 on 2026-09-23. !257850 (merged) raised it to 6 on 2026-09-25. Browser reruns on gitlab.com showed a maximum in flight of exactly 4 and then exactly 6 on every tab. So the cap is what limits parallelism on the DAP Impact dashboard.
The 2026-09-30 browser rerun at 6 shows why a higher cap helps. The DAP Impact Adoption tab sends 15 requests on tab switch, 11 panel queries and 4 previous-period comparison queries. Nothing on Adoption is deferred. The tab is about 2900 px tall and the near-view margin is one viewport height, so all 15 requests queue at once. With 6 slots, 15 requests take three rounds. With 8 they take two.
Adoption was the one tab missing its goal from the parent work item.
| Tab at 6 in flight | Requests | Time to all results | Goal |
|---|---|---|---|
| Overview | 12 | 2.7 to 3.2 s | p50 under 3 s |
| Adoption | 15 | 3.7 to 4.3 s | under 3 s |
| Work, first view | 17, plus 16 more when scrolled to the bottom | 2.3 to 3.1 s | under 3 s |
| Spend | 7 | 0.8 to 1.3 s | under 1.5 s |
The median request took 0.5 to 0.7 s on Adoption. The slowest request was the previous-period comparison of the Group and project comparison table. It waited 1.5 s for a free slot, then ran 2.4 s, and finished at 3.9 s.
Why 8 is safe
The request count per dashboard visit does not change. Only how many run at the same time changes. Exposure to gitlab.com per-user rate limits is therefore unchanged, because those limits count requests each minute and each hour, not concurrency. The published burst limit for authenticated traffic is 100 requests each minute on Free, 1,250 on Premium and 2,000 on Ultimate. A visit to all four DAP Impact tabs sent 67 requests in the 2026-09-30 rerun, including the Work scroll-in batch.
The GLQL rate limiter in Analytics::Glql::QueryService keys on the SHA of the query text. The rule limit_glql_queries_by_query_sha allows 1 hit per 15 minutes per query text, and a hit is counted only when a query is aborted by the database statement timeout, ActiveRecord::QueryAborted. It does not count concurrency or request volume. This MR does not touch it.
gitlab.com serves over HTTP/2, confirmed with a HEAD request, so the browser does not cap connections per host at 6. On an HTTP/1.1 deployment the browser would hold the extra two requests itself. There 8 would behave like 6. That does no harm and gives no gain.
The per-viewer cost is up to two more concurrent /api/glql requests. Each holds one Puma thread for the request duration, typically 0.5 to 1.5 s. Each runs one or two ClickHouse queries one after the other: the data page and, on a page that has a next page, a COUNT. So the peak is 8 ClickHouse queries at once per viewer, up from 6. The DAP Impact dashboard is an experiment behind the dap_impact_v1 flag, so viewers are few.
Specs spec/frontend/glql/core/executor_spec.js, spec/frontend/glql/utils/task_queue_spec.js and spec/frontend/glql/utils/result_cache_spec.js pass on the MR commit 66723248. That is 62 tests and 0 failures across the three files.
Related MRs change how deep the queue gets. !258139 (closed) preloads offscreen panels after the visible ones, so a whole tab's panels queue at once. !258689 (merged) orders waiting requests by grid position so the visible panels go first. Both make the queue deeper, so the concurrency cap matters more, not less. Both are still open. !257812 (merged) merged into master on 2026-09-30. It caches dashboard results behind one shared Apollo client, so switching back to a tab reuses results. It replaced the CONCURRENCY_LIMITS map with a QUEUES map. That map holds a concurrency and a cache flag per queue. It kept the dashboard concurrency at 6. This MR is rebased on top of !257812 (merged) and keeps its QUEUES shape. Only the value changes.
Expected effect
In the API measurement, 8 in flight cut time to all results on the Adoption requests by about 1.0 to 1.1 s on the median run. That is from about 4.4 s to about 3.3 s. Per-request medians did not move outside run-to-run noise. So two more requests in flight from one viewer caused no visible server-side contention.
On the dashboard I expect Adoption time to all results to go from 3.7 to 4.3 s to roughly 2.7 to 3.3 s. This is an estimate from the API runs, not a browser measurement. Adoption would then sit at the edge of its 3 s goal rather than clearly missing it.
| Tab | Requests | Rounds at 6 | Rounds at 8 | Expected change |
|---|---|---|---|---|
| Overview | 12 | 2 | 2 | Little, the last few requests start sooner |
| Adoption | 15 | 3 | 2 | About 1 s faster, estimate |
| Work, first view | 17 | 3 | 3 | Modest, 1 request in the last round instead of 5 |
| Spend | 7 | 2 | 1 | Small, the seventh request no longer waits for a slot |
The round counts assume roughly equal request durations. The requests are not equal, so treat them as rough.
Measurement method and per-request detail
I rebuilt the 15 Adoption requests as GraphQL queries against group(fullPath: "gitlab-org") { analytics { duoWorkflows(...) { aggregated(...) } } }. The default range was 2026-08-31 to 2026-09-30. The previous period was 2026-07-31 to 2026-08-30.
I fired them from a laptop with glab api graphql through xargs -P N, in dashboard order. This has the same semantics as TaskQueue. At most N requests are in flight, and the next one starts when one finishes. I ran three runs each at 6 and 8, and one run each at 4 and at 15. The run at 15 is unbounded.
Differences from the browser:
- Requests go to
/api/graphqlinstead of/api/glql. - Each request pays a small process start cost.
- The Group and project comparison queries select
usersCountonly. ThecreditsUsedSummetric is not authorized for the measuring user at group level. - The order of arrival is fixed. In the browser it is driven by panel mounting.
The 6-slot runs came out at 3.7 to 4.7 s. The browser measured 3.7 to 4.3 s at 6. So the method reproduces the dashboard well enough to compare.
| In flight | Runs | Time to all results | Median request | Slowest request | Errors |
|---|---|---|---|---|---|
| 4 | 1 | 4.74 s | 0.92 s | 2.12 s | 0 |
| 6 | 3 | 4.66 s, 4.41 s, 3.71 s | 0.92 s, 0.94 s, 0.92 s | 2.77 s, 2.64 s, 2.06 s | 0 |
| 8 | 3 | 3.59 s, 3.29 s, 2.98 s | 1.15 s, 0.92 s, 0.85 s | 2.18 s, 2.14 s, 1.86 s | 0 |
| 15 | 1 | 10.65 s | 0.93 s | 10.59 s | 1 |
The slowest request was the Group and project comparison previous-period query. It went from 2.1 to 2.8 s at 6 to 1.9 to 2.2 s at 8. It started a round earlier and no longer waited behind the queue.
Per-request medians across the three runs at 6 and 8 were within about 0.1 to 0.4 s of each other. The difference had no consistent sign.
The run at 15 is a reference only. One request, the Session intensity by tier table, took 10.6 s and came back with an error, "Siphon replication is not enabled on this instance". The other 14 finished normally in under 2 s. That is one occurrence in about 120 requests, at the unbounded setting only. It is not evidence about 8. It is a reason to keep a cap.
Server load
The total work per dashboard visit is the same, compressed into fewer rounds. At gitlab-org the Adoption tab's 15 requests finish in about 3.3 s instead of about 4.4 s. So instantaneous load per viewer is up to a third higher for about a quarter less time. Multi-viewer contention could not be measured from one client.
- Requests per visit: unchanged.
- Peak in flight per viewer: 8, up from 6.
- Peak ClickHouse queries at once per viewer: 8, up from 6. The one or two queries inside a request run sequentially.
- Rate limiter in
Analytics::Glql::QueryService: untouched.
References
- Parent work item: https://gitlab.com/gitlab-org/gitlab/-/work_items/630507
- Measurements thread, 2026-09-30 browser rerun at 6 in flight: https://gitlab.com/gitlab-org/gitlab/-/work_items/630507#note_3928829381
- Dedicated dashboard queue at concurrency 4: !256168 (merged)
- Previous bump to 6: !257850 (merged)
- GLQL beta readiness review that fixed the embedded queue at 1: https://gitlab.com/gitlab-org/gitlab/-/issues/517546
Screenshots or screen recordings
Not applicable, no visual change.
How to set up and validate locally
- Open an analytics dashboard with GLQL panels. For example, open the DAP Impact dashboard at
/explore/analytics_dashboards/dap_impact?scope=<top-level group>with ClickHouse data in GDK. - Open DevTools, go to the Network tab, filter on
/api/glql, and reload. - Up to 8 requests are in flight at once. On master the cap is 6.
MR acceptance checklist
Evaluate this MR against the MR acceptance checklist. It helps you analyze changes to reduce risks in quality, performance, reliability, security, and maintainability.