Draft: Add a simulated deploy driver for local development

What does this MR do and why?

Let's Rails engineers merge a simulated deploy driver into Rails to enable exploratory testing of the CD feature.

This MR is not intended to be merged, instead, ask your AI agent to temporarily merge the changes into your local GDK's gitlab repository.

GDK setup

Untested — these steps sketch the general idea, not a verified path.

  • Check out Relay/KAS locally: https://gitlab.com/gitlab-org/cluster-integration/gitlab-agent
  • In your GDK folder, tell your AI agent that KAS needs to be turned on with AutoFlow, and give it your local gitlab-agent directory so it can fetch the latest master and build it.
  • Enable the registry with notifications_enabled: true.
  • Turn on these feature flags: ai_native_deploy, organization_switching, ui_for_organizations.
  • Log in with a user who has access to the default organization.
  • Governance and human approval gates stay off: Cd::Rollouts::WorkflowKwargs sends no feature_flags kwarg, so a com.gitlab.cd.steps.approval step (new in v0.7.0) is refused with com.gitlab.cd.reason.approval_unavailable rather than parking for a person to approve it.

Seed script

  1. Copy the script from the collapsed seed_simulated_cd.rb block below.
  2. Save it as seed_simulated_cd.rb in the GDK's gitlab directory.
  3. Run bundle exec rails runner seed_simulated_cd.rb.
  4. Optionally, attribute the release to a user other than root by setting CD_SEED_USER, for example: CD_SEED_USER=my-user bundle exec rails runner seed_simulated_cd.rb.

It creates two environments, gtlb_staging and gtlb_production, the My_Site application, with its webserver service and nginx artifact source. A release containing the nginx 1.31.3 version.

seed_simulated_cd.rb
# frozen_string_literal: true

# Seeds a simulated CD setup in the default organization: the gtlb_staging and
# gtlb_production environments, the My_Site application with one nginx-backed
# webserver service, and an nginx 1.31.3 release. Paste a flow definition into
# the flow editor in the UI to deploy it.
#
# Safe to re-run: anything that already exists is reused.
#
#   bundle exec rails runner seed_simulated_cd.rb

unwrap = lambda do |response, key|
  raise "#{key}: #{Array(response.message).join(', ')}" if response.error?

  response.payload.fetch(key)
end

organization = Organizations::Organization.default_organization
raise 'No default organization found' unless organization

current_user = User.find_by(username: ENV.fetch('CD_SEED_USER', 'root'))

# Environments belong to the organization, not the application, so they are
# created first and shared by every application in it.
{ 'gtlb_staging' => :staging, 'gtlb_production' => :production }.each do |name, tier|
  next if Cd::Environment.in_organization(organization).with_name(name).exists?

  unwrap.call(
    Cd::Environments::CreateService.new(
      parent: organization,
      current_user: current_user,
      params: { name: name, tier: tier, driver_ref: 'simulated', driver_config: {} }
    ).execute,
    :environment
  )

  puts "Created environment #{name}"
end

application = Cd::Application.in_organization(organization).find_by(name: 'My_Site')

unless application
  application = unwrap.call(
    Cd::Applications::CreateService.new(
      parent: organization,
      current_user: current_user,
      params: {
        name: 'My_Site',
        services: [
          {
            name: 'webserver',
            # source_config is what a rollout reads to name the artifact and its type,
            # and the GraphQL mutation cannot set it, so the source is seeded here.
            artifact_sources: [
              {
                name: 'nginx',
                source_ref: 'docker.io/library/nginx',
                source_config: { 'type' => 'oci_image', 'name' => 'nginx' }
              }
            ]
          }
        ]
      }
    ).execute,
    :application
  )

  puts 'Created application My_Site'
end

artifact_source = application.services.find_by!(name: 'webserver').artifact_sources.find_by!(name: 'nginx')

version = artifact_source.versions.find_or_create_by!(name: '1.31.3') do |record|
  record.reference = 'docker.io/library/nginx:1.31.3'
end

release_name = 'nginx-1.31.3'

unless application.version_sets.exists?(name: release_name)
  unwrap.call(
    Cd::VersionSets::CreateService.new(
      parent: application,
      current_user: current_user,
      params: { name: release_name, version_ids: [version.id] }
    ).execute,
    :version_set
  )

  puts "Created release #{release_name}"
end

puts "Done. Add a flow definition to #{application.name} in the UI, then roll out #{release_name}."

Flow definition 1 - happy path, deploy to staging and prod

This flow rolls webserver out to gtlb_staging with a canary.deploy at weight 100, waits 5 seconds, then canaries gtlb_production at 25%, waits 7 seconds, and promotes to 100%. Each service takes 2 seconds, so the rollout runs about 18 seconds.

The staging step used to be a com.gitlab.cd.simulated.rolling.deploy. The Argo Rollouts driver dropped its own rolling step type in orchestration engine v0.7.0, and the simulated driver dropped its matching type to stay consistent with it. A single canary.deploy at weight 100 is a complete weight ladder in one step, so it passes the ladder check and deploys the service to completion — that's what the staging step uses now.

Porting a flow definition from the Argo Rollouts driver to this one is now a find-and-replace on the type strings, plus emptying each service's per-environment config to {}.

This flow emits 21 events:

  1. stage_started for the staging stage.
  2. step_started, service_started, service_succeeded, step_succeeded for the staging canary.deploy.
  3. stage_succeeded for the staging stage.
  4. step_started and step_succeeded for the top-level wait step. It sits outside any stage, so its payload carries no stage_name and no environment.
  5. stage_started for the production stage.
  6. step_started, service_started, service_succeeded, step_succeeded for the production canary.deploy at weight 25.
  7. step_started and step_succeeded for the wait nested inside the production stage — this one does carry stage_name.
  8. step_started, service_started, service_succeeded, step_succeeded for the canary.promote at weight 100.
  9. stage_succeeded for the production stage.
  10. rollout_succeeded, with an empty payload — the only event that carries no position.
flow.yaml
environments:
  gtlb_staging:
    services:
      webserver: {}
  gtlb_production:
    services:
      webserver: {}
steps:
  - type: com.gitlab.cd.steps.stage
    name: staging
    steps:
      - type: com.gitlab.cd.simulated.canary.deploy
        environment: gtlb_staging
        services:
          - name: webserver
            weight: 100
  - type: com.gitlab.cd.steps.wait
    seconds: 5
  - type: com.gitlab.cd.steps.stage
    name: production
    steps:
      - type: com.gitlab.cd.simulated.canary.deploy
        environment: gtlb_production
        services:
          - name: webserver
            weight: 25
      - type: com.gitlab.cd.steps.wait
        seconds: 7
      - type: com.gitlab.cd.simulated.canary.promote
        environment: gtlb_production
        services:
          - name: webserver
            weight: 100

Flow definition 2 - failure path, a canary that never reaches 100%

The workflow starts, and validates that the canary values don't end at 100%. A step_failed event with an error/reason is emitted.

Exactly 1 event: step_failed, with reason com.gitlab.cd.simulated.reason.canary_ladder_incomplete, blamed on position [0, 0] (the canary.deploy). The driver's validator runs before the engine walks anything, so there are no stage events and no service events at all, and no rollout_succeeded. Nothing is read, committed, or synced.

flow.yaml
environments:
  gtlb_production:
    services:
      webserver: {}
steps:
  - type: com.gitlab.cd.steps.stage
    name: production
    steps:
      - type: com.gitlab.cd.simulated.canary.deploy
        environment: gtlb_production
        services:
          - name: webserver
            weight: 25
      - type: com.gitlab.cd.steps.wait
        seconds: 5
      - type: com.gitlab.cd.simulated.canary.promote
        environment: gtlb_production
        services:
          - name: webserver
            weight: 50

Flow definition 3 - failure path, a service that fails to deploy

The previous example shows a flow rejected before it starts. This one shows a flow that passes validation, starts, and fails mid-step. The error is returned whenever deploying canary with weight 66%.

6 events: stage_started → step_started → service_started → (about 2 seconds) → service_failed → step_failed → stage_failed.

Reason com.gitlab.cd.simulated.reason.deploy_failed appears on both service_failed and step_failed. stage_failed carries no reason. The canary.promote at position [0, 1] never runs — a failing step ends the walk. There is no rollout_succeeded.

The trigger is the reserved weight 66: any canary step naming a service at weight 66 returns a canned failure after doing its simulated work.

A flow whose only canary step is at weight 66 or 67 never reaches the canned failure or the abort: the ladder check rejects it up front for not climbing to 100%. This flow pairs weight 66 with a canary.promote to 100 for exactly that reason.

flow.yaml
environments:
  gtlb_production:
    services:
      webserver: {}
steps:
  - type: com.gitlab.cd.steps.stage
    name: production
    steps:
      - type: com.gitlab.cd.simulated.canary.deploy
        environment: gtlb_production
        services:
          - name: webserver
            weight: 66
      - type: com.gitlab.cd.simulated.canary.promote
        environment: gtlb_production
        services:
          - name: webserver
            weight: 100

Flow definition 4 - failure path, workflow is abruptly halted

The previous example shows a flow that fails with an emitted event. This one shows a flow that stops the workflow entirely. No events are emitted after the abort.

3 events: stage_started → step_started → service_started, then nothing.

The trigger is the reserved weight 67: the driver calls Starlark's fail(), which is uncatchable, so the program ends where it stands and the engine never reaches its own reporting. Rails is told nothing more, and the rollout is left in_progress. The message survives only in the kas log.

Same ladder trap as flow definition 3: weight 67 alone would be rejected before it could abort, so this flow pairs it with a canary.promote to 100.

flow.yaml
environments:
  gtlb_production:
    services:
      webserver: {}
steps:
  - type: com.gitlab.cd.steps.stage
    name: production
    steps:
      - type: com.gitlab.cd.simulated.canary.deploy
        environment: gtlb_production
        services:
          - name: webserver
            weight: 67
      - type: com.gitlab.cd.simulated.canary.promote
        environment: gtlb_production
        services:
          - name: webserver
            weight: 100

Reading the events

Two things worth knowing before you look at the raw payloads:

  1. A failure(reason, message) from a driver arrives on the wire with its message under the key error, not message.
  2. There is deliberately no rollout_failed event. A failed deploy is a step_failed with no rollout_succeeded behind it.

Every event is POSTed to POST /api/v4/rollouts/<rollout_id> with topic com.gitlab.cd.deployment. The driver itself emits nothing; the engine emits all of it.

Edited by Cameron Swords

Merge request reports

Loading
Loading