Run the slack_assistant flow server-side instead of in CI

The problem

Two merge requests set up "GitLab Duo for Slack" as an API-only flow:

  • !254799 (merged) (merged) registers slack_assistant/v1 as a foundational flow behind the slack_duo_api_flow feature flag. It declares coding_environment: none.
  • !254801 (merged) (open) routes a Slack mention to that flow when the flag is on.

After both, a Slack mention still starts a CI job. coding_environment: none only makes that job skip the git clone and setup. The flow's tools reach GitLab over the API — no repository, no shell, no filesystem — so the job's only real achievement is skipping its own setup, while the user waits for a runner to pick it up.

What this MR does

Runs the flow through Workhorse's server-side execution endpoint (POST /api/v4/ai/duo_workflows/workflows/:workflow_id/execute) instead of starting a CI job. Workhorse holds the gRPC stream to the Duo Workflow Service; no pipeline is created. The flow runs as the requesting user.

Same feature flag, slack_duo_api_flow. With it off, nothing changes: the Duo Developer flow runs in CI exactly as today.

Pieces

  1. Ai::Messaging::ExecuteServerSideFlowService — the server-side counterpart to Ai::Catalog::ExecuteWorkflowService. Creates the workflow, then enqueues the turn worker. Agent privileges, pre-approved privileges, allow_agent_to_request_user and environment all come from the flow's registry entry, taken the same way Ai::DuoWorkflows::CreateAndStartWorkflowService takes them. Running in Workhorse rather than CI must not change what the flow is allowed to do.

  2. Ai::Messaging::ServerSideTurnWorker — runs one turn. Unlike most workers it blocks for the turn's duration, because holding the HTTP connection to Workhorse is the job. It is a worker of its own, rather than inline in the Slack event worker, so the run is not bounded by that worker's retry semantics and so fleet concurrency can be capped. Capped at 50 concurrent jobs, provisional until real traffic informs a better number. Deduplicated until_executed.

  3. Ai::Messaging::ResultDelivery — the terminal-outcome delivery logic (the atomic delivery claim, extracting the final agent message, the retry-on-failed-delivery behaviour) extracted from Ai::Messaging::CallbackWorker so both workers share it. A pure code move: the only differences are two comments generalised to cover both callers. The internal DeliveryRetryError constant is now Ai::Messaging::ResultDelivery::DeliveryRetryError, which is the name that appears in logs.

    Delivery is inline in the turn worker rather than event-driven because a turn ends at the :input_required status, which fires no event. Unlike a CI-executed flow there is no WorkflowFinishedEvent or WorkloadFinishedEvent for CallbackWorker to ride. Live progress still streams through ProgressDeliveryWorker, driven by the checkpoints the flow writes as it runs.

  4. ProgressDeliveryWorker now also stops on the delivery claim, not only on a terminal status, because a turn's end state (:input_required) is not terminal.

  5. The service account is resolved as before but is neither granted project membership nor used, since the flow runs as the requesting user.

How the server-side path is selected

The adapter matches on the resolved flow reference, using the generated registry accessor Ai::Catalog::FoundationalFlow.slack_assistant_v1, rather than reading slack_duo_api_flow a second time. The flag is evaluated once, where routing is decided; a second read against a separately derived actor could disagree with the flow the goal was built for. (The same argument the routing MR makes for threading the reference through goal assembly.)

Note on the commit stack

This branch temporarily carries a cherry-pick of the routing commit from !254801 (merged) so the change can be built and reviewed on top of it. That commit disappears on rebase once the routing MR merges.

It also targets !254712 (merged), which adds the Ai::DuoWorkflows::ServerSideExecutionService HTTP client this MR calls. Review that first.

What is not in this MR

Multi-turn conversations. Each mention starts a fresh workflow. The Redis thread-to-workflow mapping that lets a Slack thread resume an existing session is follow-up work.

Approval pauses. A turn can stop to ask the user to approve a tool call or a plan, and Slack has no way to answer that yet, so such a run reports the existing :no_response error. Handling it means Approve and Deny buttons in the thread, which need the same thread-to-workflow mapping as resume, so it is left for that follow-up. Note the registry entry pre-approves everything it grants, so this is not reachable through ordinary tool use.

Test coverage

Specs added for the new service and worker; extended for the adapter base class, the callback and progress workers, and the Slack adapter. 392 examples in the touched messaging specs pass, along with the feature-flag definition and Sidekiq worker meta specs.

How to set up and validate locally

  1. You need the GitLab for Slack app wired to your GDK, and the Duo Workflow Service running with the slack_assistant flow available.
  2. Enable the flag: Feature.enable(:slack_duo_api_flow).
  3. Seed the foundational flow entry and enable the flow for your group (steps 2-4 of !254799 (merged)).
  4. Mention the bot in a Slack thread.
  5. Confirm the answer arrives in the thread.
  6. Confirm no CI pipeline was created for the session — this is the difference from the routing MR, where a job runs and skips its clone.
  7. Confirm the session detail shows the session ran slack_assistant/v1.
  8. Disable the flag and mention again: Duo Developer runs in CI as before.
Edited by Igor Drozdov

Merge request reports

Loading
Loading