CI Health Incidents — Operations, Feedback & Evolution
**Pillars:** [Development Health Signal Platform](https://handbook.gitlab.com/handbook/engineering/infrastructure-platforms/developer-experience/development-analytics/#development-health-signal-platform), [Tooling Stewardship](https://handbook.gitlab.com/handbook/engineering/infrastructure-platforms/developer-experience/development-analytics/#tooling-stewardship) **Handbook:** [FY27 Q3/Q4 Roadmap](https://handbook.gitlab.com/handbook/engineering/infrastructure-platforms/developer-experience/development-analytics/#fy27-q3q4-roadmap) ### Problem Statement CI Health Incidents detection is live and being opened to engineers. The detection system (built in [Q1 -- Epic #24](https://gitlab.com/groups/gitlab-org/quality/analytics/-/epics/24)) surfaces actionable master-broken situations with ~91% noise reduction vs. the old system. This epic covers the next phase: operating the system in production, learning from incidents, evolving detection based on engineer feedback, and deciding the future of the old master-broken incidents process. Pipeline stability is a key focus across engineering: master-broken incidents are a primary driver. We want to help engineers and EMs quantify impact, improve time to resolution, and add guardrails to prevent recurrence. DA is also actively shaping the in-product flaky test detection initiative ([work item 606069](https://gitlab.com/gitlab-org/gitlab/-/work_items/606069)), contributing our statistical flakiness definition, failure signatures, and blast radius data as the signal layer the product needs. ### Participants - (DRI) - (EM) ### Exit Criteria - Engineers across the organization adopt CI Health Incidents as their primary signal for master-broken situations - The master-broken incidents process has been replaced by CI Health Incidents - The system improves iteratively based on real-world feedback and signal quality gaps - We begin to understand how to learn from incidents to prevent recurrence ### Topics Topics can be worked in parallel. #### Signal quality & coverage | Item | Tracking | Status | |------|----------|--------| | Add Slack reactions to CI health incident lifecycle | [ci-alerts #411](https://gitlab.com/gitlab-org/quality/analytics/ci-alerts/-/work_items/411) | ✅ Done | | Traceless job/pipeline failure detection | [team #547](https://gitlab.com/gitlab-org/quality/analytics/team/-/work_items/547) | 📋 Planned | #### Master-broken incidents: transition & deprecation | Item | Tracking | Status | |------|----------|--------| | Ownership & response process | [team #592](https://gitlab.com/gitlab-org/quality/analytics/team/-/work_items/592) | 📋 Planned | | Audit master-broken incidents: assess migration to CI Health, decide on naming, and update documentation | [team #541](https://gitlab.com/gitlab-org/quality/analytics/team/-/work_items/541) | 📋 Planned | | Decide on fate of existing master-broken incidents: deprecate, redirect, or merge; update handbook and docs | [team #618](https://gitlab.com/gitlab-org/quality/analytics/team/-/work_items/618) | 📋 Planned | #### Learning from incidents to prevent recurrence | Item | Tracking | Status | |------|----------|--------| | Investigate how to learn from CI health incidents and prevent recurrence | [team #619](https://gitlab.com/gitlab-org/quality/analytics/team/-/work_items/619) | 📋 Planned | #### Advanced automation | Item | Tracking | Status | |------|----------|--------| | Improve AI-assisted MR identification skill: make it reliable and shareable (move to canonical) | [team #620](https://gitlab.com/gitlab-org/quality/analytics/team/-/work_items/620) | 📋 Planned | | Auto-create revert MRs once causing MR is identified | [team #545](https://gitlab.com/gitlab-org/quality/analytics/team/-/work_items/545) | 📋 Planned | ### Development Log <!-- STATUS NOTE START --> <!-- STATUS NOTE END -->
epic