Tell the messaging adapter when a run fails or pauses

What does this MR do and why?

The messaging adapter turns Duo session events into Slack messages and @GitLabDuo note replies. Today it hears two events from the session itself — started, finished — and infers everything else from the CI job: job ended → run is over; job failed → run failed.

Two things break that:

  • A run paused for a tool approval ends its CI job cleanly. The adapter read that as "done", posted whatever the agent said last as the answer, and claimed the delivery — so when the session later really finished, the real answer was never posted.
  • A run that fails without a CI job (the headless runtime) left the thread spinning forever.

The session tells the adapter directly

Three events, all scoped to sessions with a messaging_callback_context exactly like the existing WorkflowStartedEvent/WorkflowFinishedEvent, so ordinary CI sessions emit nothing:

  • WorkflowFailedEvent on drop/stop → CallbackWorker calls on_flow_failed (:flow_failed / :flow_stopped). Slack and the note adapter get a distinct "session was stopped" message.
  • WorkflowApprovalRequiredEvent on require_tool_call_approval → new Adapters::Base#on_approval_required hook.
  • WorkflowInputRequiredEvent on require_input → new Adapters::Base#on_input_required hook. This is how a conversational turn ends, and it is the event the server-side runtime needs before it can stop delivering a turn inline (!254733 discussion).

Both hooks default to a no-op. Nothing implements them yet, so neither changes behaviour in this MR.

One owner for the outcome

The delivered_at claim decides who reports a terminal outcome, and three callers were bypassing it. Each of them can now report the same failure twice, because drop/stop publish an event:

  • ServerSideTurnWorker called the adapter directly. It now goes through deliver_failure, so a Workhorse failure is reported once rather than once by the worker and again by CallbackWorker.
  • Adapters::Base#with_lifecycle_hooks claims before reporting a synchronous failure, so a later drop of the same session stays quiet.
  • A drop from :created now publishes instead of being skipped. On CI a start fails inside the request that created the session, which the claim above covers. A server-side session instead waits in :created across a Sidekiq boundary, so a drop from there is a real failure that nothing else reports — filtering it out left the surface waiting forever.

CI backstop

A WorkloadFinishedEvent for a session that is paused or awaiting input is a pause — nothing delivered, nothing claimed. The real answer is delivered when the session finishes later.

awaiting_user? states that as the complement of terminal and active rather than as a list of waiting statuses, so a status added later counts as waiting. Missing one costs the user their answer; an extra one costs nothing.

Existing Slack and @GitLabDuo behaviour for sessions that simply finish or fail is unchanged.

References

How to set up and validate locally

Verified on the GDK against a real Slack workspace, with slack_duo_api_flow enabled so the session runs on the server-side runtime — no CI runner needed.

The scenario is the MR's headline promise: a run that dies no longer leaves the thread spinning.

  1. Mention the app in Slack. The thread gets 👀 and a progress message.

  2. While the session runs, stop it from the Rails console:

    wf = Ai::DuoWorkflows::Workflow.where(workflow_definition: 'slack_assistant/v1').order(:id).last
    wf.status_name  # => :running
    wf.stop!
  3. Within a second or two the thread shows ❌ instead of 👀, and the progress message is edited in place to The session was stopped before it finished. (error_text(:flow_stopped), reachable only through WorkflowFailedEvent → CallbackWorker → deliver_failure). On master nothing is published on stop, so the thread stays silent.

  4. Confirm the failure was reported once. ServerSideTurnWorker is still holding the Workhorse connection when the session is stopped, so it reaches its own delivery path afterwards and must stand down:

    grep "Duo Messaging" log/application_json.log | tail -5
    Duo Messaging: Terminal outcome already delivered | workflow_id: 3095

    In log/sidekiq.log the same run shows CallbackWorker finishing with external_http_count: 4 (the Slack calls) and ServerSideTurnWorker finishing with external_http_count: 1 — its Workhorse call only, no Slack calls. On master the worker calls the adapter directly and posts a second error.

The approval and input hooks are no-ops until !255537 (merged) implements them, so there is nothing to observe for those yet.

MR acceptance checklist

Evaluate this MR against the MR acceptance checklist. It helps you analyze changes to reduce risks in quality, performance, reliability, security, and maintainability.

Edited by Thomas Schmidt

Merge request reports

Loading
Loading