CI Health Incidents — Operations, Feedback & Evolution
**Pillars:** [Development Health Signal Platform](https://handbook.gitlab.com/handbook/engineering/infrastructure-platforms/developer-experience/development-analytics/#development-health-signal-platform), [Tooling Stewardship](https://handbook.gitlab.com/handbook/engineering/infrastructure-platforms/developer-experience/development-analytics/#tooling-stewardship)
**Handbook:** [FY27 Q3/Q4 Roadmap](https://handbook.gitlab.com/handbook/engineering/infrastructure-platforms/developer-experience/development-analytics/#fy27-q3q4-roadmap)
### Problem Statement
CI Health Incidents detection is live and being opened to engineers. The detection system (built in [Q1 -- Epic #24](https://gitlab.com/groups/gitlab-org/quality/analytics/-/epics/24)) surfaces actionable master-broken situations with ~91% noise reduction vs. the old system. This epic covers the next phase: operating the system in production, learning from incidents, evolving detection based on engineer feedback, and deciding the future of the old master-broken incidents process.
Pipeline stability is a key focus across engineering: master-broken incidents are a primary driver. We want to help engineers and EMs quantify impact, improve time to resolution, and add guardrails to prevent recurrence. DA is also actively shaping the in-product flaky test detection initiative ([work item 606069](https://gitlab.com/gitlab-org/gitlab/-/work_items/606069)), contributing our statistical flakiness definition, failure signatures, and blast radius data as the signal layer the product needs.
### Participants
- (DRI)
- (EM)
### Exit Criteria
- Engineers across the organization adopt CI Health Incidents as their primary signal for master-broken situations
- The master-broken incidents process has been replaced by CI Health Incidents
- The system improves iteratively based on real-world feedback and signal quality gaps
- We begin to understand how to learn from incidents to prevent recurrence
### Topics
Topics can be worked in parallel.
#### Signal quality & coverage
| Item | Tracking | Status |
|------|----------|--------|
| Add Slack reactions to CI health incident lifecycle | [ci-alerts #411](https://gitlab.com/gitlab-org/quality/analytics/ci-alerts/-/work_items/411) | ✅ Done |
| Traceless job/pipeline failure detection | [team #547](https://gitlab.com/gitlab-org/quality/analytics/team/-/work_items/547) | 📋 Planned |
#### Master-broken incidents: transition & deprecation
| Item | Tracking | Status |
|------|----------|--------|
| Ownership & response process | [team #592](https://gitlab.com/gitlab-org/quality/analytics/team/-/work_items/592) | 📋 Planned |
| Audit master-broken incidents: assess migration to CI Health, decide on naming, and update documentation | [team #541](https://gitlab.com/gitlab-org/quality/analytics/team/-/work_items/541) | 📋 Planned |
| Decide on fate of existing master-broken incidents: deprecate, redirect, or merge; update handbook and docs | [team #618](https://gitlab.com/gitlab-org/quality/analytics/team/-/work_items/618) | 📋 Planned |
#### Learning from incidents to prevent recurrence
| Item | Tracking | Status |
|------|----------|--------|
| Investigate how to learn from CI health incidents and prevent recurrence | [team #619](https://gitlab.com/gitlab-org/quality/analytics/team/-/work_items/619) | 📋 Planned |
#### Advanced automation
| Item | Tracking | Status |
|------|----------|--------|
| Improve AI-assisted MR identification skill: make it reliable and shareable (move to canonical) | [team #620](https://gitlab.com/gitlab-org/quality/analytics/team/-/work_items/620) | 📋 Planned |
| Auto-create revert MRs once causing MR is identified | [team #545](https://gitlab.com/gitlab-org/quality/analytics/team/-/work_items/545) | 📋 Planned |
### Development Log
<!-- STATUS NOTE START -->
<!-- STATUS NOTE END -->
epic