Draft: Enable MCP server access for existing root groups
Note
This MR stays in Draft until the temporary audit-event index has been built on GitLab.com. That index is created asynchronously on weekends.
What does this MR do and why?
In GitLab 19.5, when the MCP server becomes generally available, existing top-level groups get Allow connection to GitLab (MCP client access) turned on. The only groups left off are those where someone deliberately turned it off.
The parent MR (already merged) turns the setting on for all new groups. This MR catches existing groups up to that default.
Why existing groups are off today:
namespace_settings.mcp_server_enabledwas added in 19.0 with no default, andGroup#mcp_server_enabled?treats a blank value as off.- Background migrations then wrote
experiment_features_enabled IS TRUE AND duo_features_enabled IS TRUEinto existing top-level groups. On GitLab.com,BackfillMcpServerEnabled(19.0) ran from 2026-04-27, andResyncMcpServerEnabled(19.2) re-applied the rule from 2026-06-25. The backfill was later turned into a no-op, so self-managed instances only run the resync. - From 2026-06-23 to 2026-09-04, a
before_savecallback recalculated the same rule whenever group settings were saved.
So for nearly every group, "off" was a side effect of beta and Duo settings or of a missing value, not a customer choice. On GitLab.com that's most top-level groups (see Rows affected).
What the migration changes
The batched background migration McpServerDefaultTrueForGroups sets namespace_settings.mcp_server_enabled = true for top-level groups where it is currently false or blank, except groups whose most recent mcp_server_enabled_updated audit event records a change to off. Only the latest event counts: if an Owner turned MCP off and later back on, any off value after that came from a migration.
- Duo availability plays no part. The backfills tied MCP to Duo on purpose, but today they are entirely separate settings, so groups with Duo set to Always off are turned on too.
- Beta and experiment features play no part either.
- It only ever turns the setting from off or blank to on, never the other way.
- Subgroups and personal namespaces are left alone, because only top-level groups are checked when an MCP request comes in.
Known limits
- There's no way to tell "never decided" from "saw it was off and was happy with that".
- The group setting has only been audited since 2026-06-18, so any change made before then can't be seen. The migration only checks events from that date onward.
- Between 2026-06-23 and 2026-09-04, a callback wrote MCP audit events whenever someone saved Duo settings. Some groups will therefore be skipped even though no one turned MCP off. This errs on the safe side.
- Audit events expire, so retention needs to be confirmed.
Not covered here
- The instance-level setting for self-managed and Dedicated (
application_settings.mcp_server_settings). That's a separate follow-up. - The migration is only queued on GitLab.com (
Gitlab.com_except_jh?). Group-level MCP control is only enforced there (mcp_server_saas_only), so running it on self-managed, Dedicated, or JiHu would change nothing users see. Those installs use the instance setting.
Query
It batches over namespaces.id (batch size 1000, sub-batch size 100) and runs two UPDATEs per sub-batch. The first sets blank values to on. The second turns on false values unless the group's latest MCP audit event turned it off. That lookup uses the (group_id, created_at, id) index and is limited to partitions from 2026-06-18 onward. This is the same shape as the earlier MCP backfills.
References
- Parent MR: !256761 (merged)
- Work item: https://gitlab.com/gitlab-org/gitlab/-/work_items/630175
- Customer report: #600594 (closed)
How to set up and validate locally
-
Run GDK in SaaS mode, so the group MCP setting shows in the UI: add
export GITLAB_SIMULATE_SAAS=1toenv.runitand rungdk restart. -
Check out this branch and run
GITLAB_SIMULATE_SAAS=1 bundle exec rails db:migrate. The migration is only queued on GitLab.com, so without the variable nothing is queued. -
In the Rails console, prepare two top-level groups:
off = Group.find_by_full_path('group-a').namespace_settings off.update!(mcp_server_enabled: false) duo_off = Group.find_by_full_path('group-b').namespace_settings duo_off.update!(mcp_server_enabled: false, duo_features_enabled: false, lock_duo_features_enabled: true)For
group-c, go to Settings > General > Permissions and group features in the UI and clear Allow connection to GitLab. Doing it through the UI records the audit event. -
Run the migration:
bundle exec rake "gitlab:background_migrations:finalize[McpServerDefaultTrueForGroups,namespaces,id,[]]" -
Check the results:
group-aandgroup-bare nowtrue, andgroup-cstaysfalse.
Database review
Queries
Each sub-batch of 100 namespaces runs these queries. Batches are 1000 rows. The first query is the standard cursor query from the batching framework.
1. Pick the top-level groups in the sub-batch
SELECT "namespaces"."id"
FROM "namespaces"
WHERE ("namespaces"."id") <= (:end_id) AND ("namespaces"."id") >= (:start_id)
AND "namespaces"."type" = 'Group'
AND "namespaces"."parent_id" IS NULL
ORDER BY "namespaces"."id" ASC
LIMIT 100Query plan: https://console.postgres.ai/gitlab/projects/gitlab-production-main/sessions/58371/commands/163080
Batch starting at gitlab-org (id 9970 to 11000): an index-only scan on index_groups_on_parent_id_id returns 100 rows in 11 ms, most of it cold-cache reads.
Plan
Limit (cost=6.78..6.88 rows=41 width=4) (actual time=11.068..11.085 rows=100 loops=1)
Buffers: shared hit=31 read=10
I/O Timings: read=10.936 write=0.000
-> Sort (cost=6.78..6.88 rows=41 width=4) (actual time=11.066..11.073 rows=100 loops=1)
Sort Key: namespaces.id
Sort Method: quicksort Memory: 25kB
Buffers: shared hit=31 read=10
I/O Timings: read=10.936 write=0.000
-> Index Only Scan using index_groups_on_parent_id_id on public.namespaces (cost=0.56..5.68 rows=41 width=4) (actual time=6.590..11.039 rows=123 loops=1)
Index Cond: ((namespaces.parent_id IS NULL) AND (namespaces.id <= 11000) AND (namespaces.id >= 9970))
Heap Fetches: 0
Buffers: shared hit=28 read=10
I/O Timings: read=10.936 write=0.000
Settings: work_mem = '230MB', jit = 'off', random_page_cost = '1.5', seq_page_cost = '4', effective_cache_size = '472585MB'2. Turn MCP on for those groups
Blank values first:
UPDATE "namespace_settings"
SET "mcp_server_enabled" = TRUE
WHERE "namespace_settings"."namespace_id" IN (:root_group_ids)
AND "namespace_settings"."mcp_server_enabled" IS NULLThen false values, unless an Owner turned MCP off:
UPDATE "namespace_settings"
SET "mcp_server_enabled" = TRUE
WHERE "namespace_settings"."namespace_id" IN (:root_group_ids)
AND "namespace_settings"."mcp_server_enabled" = FALSE
AND ((
SELECT group_audit_events.details LIKE '%:to: false%'
FROM group_audit_events
WHERE group_audit_events.group_id = namespace_settings.namespace_id
AND group_audit_events.created_at >= '2026-06-18'
AND group_audit_events.event_name = 'mcp_server_enabled_updated'
ORDER BY group_audit_events.created_at DESC, group_audit_events.id DESC
LIMIT 1
) IS NOT TRUE)The audit subquery runs once per candidate row and uses idx_group_audit_events_on_project_created_at_id (group_id, created_at, id). Only partitions from 2026-06-18 onward are scanned.
Query plan: https://console.postgres.ai/gitlab/projects/gitlab-production-main/sessions/58371/commands/163081
This plan was measured on both updates combined into one (mcp_server_enabled IS NOT TRUE). The two separate updates touch the same rows and run the same audit lookup. The same batch updates 99 of 100 rows in 215 ms on a cold cache, which is within the 1 s limit for background migrations. Most of that time is spent reading from disk. About 78 ms comes from the IN (SELECT ...) used here in place of the literal ID list, and the real migration doesn't run that part. The audit lookup is a backward index scan per partition, at 0.49 ms per row. The batch writes about 800 KB of WAL.
Plan
Update on public.namespace_settings (cost=77.68..10462.41 rows=0 width=0) (actual time=215.541..215.546 rows=0 loops=1)
Buffers: shared hit=4468 read=197 dirtied=101
WAL: records=346 fpi=97 bytes=802275
I/O Timings: read=210.944 write=0.000
-> Nested Loop (cost=77.68..10462.41 rows=50 width=35) (actual time=133.284..195.182 rows=99 loops=1)
Buffers: shared hit=3239 read=137 dirtied=3
WAL: records=3 fpi=3 bytes=23715
I/O Timings: read=192.731 write=0.000
-> HashAggregate (cost=77.12..78.12 rows=100 width=32) (actual time=78.528..78.566 rows=100 loops=1)
Group Key: "ANY_subquery".id
Batches: 1 Memory Usage: 24kB
Buffers: shared hit=29 read=70
I/O Timings: read=77.939 write=0.000
-> Subquery Scan on "ANY_subquery" (cost=0.57..76.87 rows=100 width=32) (actual time=6.765..78.435 rows=100 loops=1)
Buffers: shared hit=29 read=70
I/O Timings: read=77.939 write=0.000
-> Limit (cost=0.57..75.87 rows=100 width=4) (actual time=6.752..78.356 rows=100 loops=1)
Buffers: shared hit=29 read=70
I/O Timings: read=77.939 write=0.000
-> Index Scan using index_namespaces_on_type_and_id on public.namespaces (cost=0.57..3203542.43 rows=[redacted] width=4) (actual time=6.751..78.331 rows=100 loops=1)
Index Cond: (((namespaces.type)::text = 'Group'::text) AND (namespaces.id >= 9970))
Filter: (namespaces.parent_id IS NULL)
Buffers: shared hit=29 read=70
I/O Timings: read=77.939 write=0.000
-> Index Scan using namespace_settings_pkey on public.namespace_settings (cost=0.56..103.84 rows=1 width=10) (actual time=1.165..1.166 rows=1 loops=100)
Index Cond: (namespace_settings.namespace_id = "ANY_subquery".id)
Filter: ((namespace_settings.mcp_server_enabled IS NOT TRUE) AND ((SubPlan 1) IS NOT TRUE))
Buffers: shared hit=3210 read=67 dirtied=3
WAL: records=3 fpi=3 bytes=23715
I/O Timings: read=114.792 write=0.000
SubPlan 1
-> Limit (cost=2.99..100.26 rows=1 width=17) (actual time=0.489..0.489 rows=0 loops=99)
Buffers: shared hit=2759 read=16
I/O Timings: read=47.168 write=0.000
-> Append (cost=2.99..975.69 rows=10 width=17) (actual time=0.489..0.489 rows=0 loops=99)
Buffers: shared hit=2759 read=16
I/O Timings: read=47.168 write=0.000
-> Index Scan Backward using group_audit_events_202703_group_id_created_at_id_idx on gitlab_partitions_dynamic.group_audit_events_202703 group_audit_events_10 (actual time=0.001..0.001 rows=0 loops=99)
... (one backward index scan per monthly partition, 202606 to 202703; 0.001 to 0.199 ms each) ...
-> Index Scan Backward using group_audit_events_202606_group_id_created_at_id_idx on gitlab_partitions_dynamic.group_audit_events_202606 group_audit_events_1 (cost=0.56..80.96 rows=1 width=17) (actual time=0.137..0.137 rows=0 loops=99)
Index Cond: ((group_audit_events_1.group_id = namespace_settings.namespace_id) AND (group_audit_events_1.created_at >= '2026-06-18 00:00:00+00'::timestamp with time zone))
Filter: (group_audit_events_1.event_name = 'mcp_server_enabled_updated'::text)
Buffers: shared hit=392 read=4
I/O Timings: read=13.377 write=0.000
Settings: effective_cache_size = '472585MB', work_mem = '230MB', jit = 'off', random_page_cost = '1.5', seq_page_cost = '4'Rows affected on GitLab.com: nearly every top-level group. The counts are in an internal comment, because they're SAFE data.
Rollback and what to do if something goes wrong
downremoves the queued migration. The data change can't be undone automatically, because the migration doesn't record which rows it changed.- In an incident, pause the migration with ChatOps (
/chatops run batched_background_migrations pause <id>). Owners can turn MCP off for their own group in Settings > General > Permissions and group features.
What users would see if it's wrong
- A group turned on that shouldn't be: members can connect AI tools to GitLab through MCP. They only get access to what they can already see, but the Owner didn't intend to allow it.
- A group skipped that should be on: members get "MCP server not enabled for any of your groups" until an Owner turns the setting on. That's how things work today.
MR acceptance checklist
Evaluate this MR against the MR acceptance checklist. It helps you analyze changes to reduce risks in quality, performance, reliability, security, and maintainability.
- Database review required (batched background migration).
- The specs pass locally (20 examples).