Track project import lifecycle internal events
What
Fires internal events for the project import lifecycle, so every import transition can be tracked in real time regardless of importer, using the event names and properties agreed with the Data team in #617884.
Five events are emitted across the lifecycle:
start_project_importfinish_project_importfail_project_importcancel_project_importtimeout_project_import
Two import paths are covered:
Standard importers — app/models/project_import_state.rb, via
after_transition hooks. Covers GitHub, Bitbucket, Bitbucket Server, Gitea,
Fogbugz, file-based (gitlab_project), Manifest, and Git/Repository-by-URL.
project.mirror?is skipped:ProjectImportStateis backed by theproject_mirror_datatable and its state machine is reused by recurring pull-mirror updates, so without this guard every mirror sync would be counted as an import.- Only
Gitlab::ImportSources.importable_project_typesare tracked, which excludes template-based creation (gitlab_built_in_project_template). mark_as_timed_outsets a transienttimed_outflag and transitions tofailedvia the existingfail_opevent, same asmark_as_failed— no new state.ProjectImportStateis shared with recurring pull-mirror updates, and the mirror recovery logic (CE'sscheduleevent, EE's retry/capacity bookkeeping) only knows aboutfailed, so a genuinely newtimeoutstate would break mirror recovery. Theafter_transition :failedhandler picks the tracked event name (timeout_project_importvsfail_project_import) from the flag instead of the target state.Gitlab::Import::StuckProjectImportJobsWorkernow flags an import as timed out instead of just failed; otherStuckImportJobincluders (Jira) are unaffected.
Direct Transfer — Bulk Imports:
lib/bulk_imports/projects/pipelines/project_pipeline.rb—startedwhen the project record is created.lib/bulk_imports/common/pipelines/entity_finisher.rb—finishedwhen aproject_entityreaches the finished state.app/models/bulk_imports/entity.rb—failed,cancelled, andtimeout(cleanup_stale).
BulkImports::Entity and EntityFinisher are shared with Offline Transfer,
which imports projects through its own, nearly identical pipeline
(lib/import/offline/projects/pipelines/project_pipeline.rb). Rather than
excluding Offline Transfer, that pipeline now emits start_project_import
too, so both transfer mechanisms are covered by the same five events. The
two are distinguished by label, which is read from project.import_type
(gitlab_project_migration for Direct Transfer, offline_transfer for
Offline Transfer) instead of a hardcoded string. The shared
project/user/namespace/additional_properties hash is extracted into
BulkImports::Entity#project_import_event_attributes, reused by
ProjectPipeline (both DT and Offline Transfer), EntityFinisher, and
Entity itself.
Every event carries project / user / namespace identifiers, a label
of the importer name, and a hashed import_source property
(Gitlab::Import::SourceIdentifier, wrapping Gitlab::CryptoHelper.sha256)
so repeated imports of the same source can be detected without sending the
raw host/path to Snowplow (import_source counts as Red data under the
Data Classification Standard).
Scope notes
- This MR covers #627868 (closed) (the
import_sourcehashing helper) and #627872 (closed) (the standard-importer + Direct Transfer project lifecycle events); it is the first of several planned slices of #617884.request_channel,source_hosting,imported_objects_count, andfidelity_rateare deferred to follow-up MRs (tracked as separate child tasks under #617884). - Builds on the closed !236324 (closed), whose author intentionally deferred it to consolidate under #617884.
References
How to set up and validate locally
-
Enable Snowplow Micro so events are actually collected:
gdk config set snowplow_micro.enabled true gdk reconfigure gdk restart -
Import a Project
- New project
- Import project
-
Import a Project using Direct Transfer
- New group
-
Observe events in Snowplow Micro (http://gdk.test:9091/micro/good)