2021-06-21: GKE unable to scale due to lack of SSD availability
Timeline
All times UTC.
2021-06-20
10:30- First attempt at cluster-autoscaler failing to expand nodes in our production cluster (sidekiq related)
2021-06-21
05:51- Signs of SSD exhaustion from our CI runners begins08:05- Apdex and rail queuing being an upward trend08:19- Google Support notifies us that they are low on SSD stock08:43- ahmad declares incident in Slack.09:01- Discovered that the API deployment is not healthy, unable to scale nodes to add to Pod capacity09:55- We add another node pool usingpd-standarddisks10:37- The new node pool is not scaling upward, to remediate pressure on the API, we manually scaled that node pool across all three zones13:42- We decide to bring online new node pools withpd-standarddisks as deployments are blocked as other workloads are impacted in a similar fashion18:16- During testing in our canary stage, the cluster-autoscaler does not work with our new node pools, we broadcast that we are not allowing production deployments to occur - support case for addressing the cluster-autoscaler is opened with Google18:26- Infrastructure changes to node pools and nodeSelectors for remaining workloads begin
2021-06-22
16:28- First deploy into canary completes19:21- First deploy to production completes since the beginning of this incident
Corrective Actions
- Transition Kubernetes services to
pd-standarddisks: delivery#1837 (moved) - Capturing resource exhaustion errors as a metric: https://ops.gitlab.net/gitlab-com/gitlab-com-infrastructure/-/merge_requests/2666
Incident Review
Summary
This is an incident that includes multiple facets with a few other issues found along the way. We started off in #4937 (closed) where our CI Runners were unable to spin up new machines. The root cause there was a shortage of SSD availability in our chosen region to operate out of us-east1. This didn't start impacting the API service until the next day when our API service needed to scale upward to take care of heavier load that grows with user traffic. During that period of time, new nodes were unable to be brought online, which impacted the API as it was starting to become stressed. This was resolved manually with great collaboration to bring online non ssd backed nodes and manually scaling the node pools for which the API services runs on. For a period of approximately 146 minutes, the API service was considered degraded. The user experience during this time would have been slow response times. There's no indication that the error rates increased.
After resolving the API, we were still in hot water as Auto-Deploy was unable to deploy as Kubernetes required us to spin up new nodes in order to schedule the new Pods. This was unable to happen across our entire fleet in both staging, canary, and production. Due to this, Auto-Deploy was completely blocked. We proceeded throughout the course of the day to bring online new node pools using pd-standard disks. This uncovered a new issue. The cluster-autoscaler was unable to scale due to a bug where is one node pool that matches the selector cannot scale, the cluster-autoscaler does not attempt to scale another matching node pool where labels may match. This is a known issue with Google with no ETA on a fix.
We decided to proceed forward with moving all of our Kubernetes resources onto pd-standard disks. We left ourselves running primarily on these node pools for one day. We did not notice any poor behavior, thus we proceed to clean up and remove the SSD backed node pools.
- Service(s) affected: ServiceAPI ServiceSidekiq ServiceGit ServiceGCP AutoDeploy (at risk and above)
- Team attribution: ~"team::Core-Infra"
- Time to detection: 22 Hours
- Minutes downtime or degradation: 146 minutes for the ServiceAPI - 1.5 days for AutoDeploy
Metrics
API Apdex Degredation
Kubernetes Cluster Autoscaler Errors
Customer Impact
- Who was impacted by this incident? (i.e. external customers, internal customers)
- All customers for 146 minutes - AutoDeploy for 1.5 days
- What was the customer experience during the incident? (i.e. preventing them from doing X, incorrect display of Y, ...)
- Sluggish Response times
- How many customers were affected?
- n/a
- If a precise customer impact number is unknown, what is the estimated impact (number and ratio of failed requests, amount of traffic drop, ...)?
- All requests during the degradation period
What were the root causes?
- GCP ran out of SSD's
Incident Response Analysis
- How was the incident detected?
- A lot of log digging. Initial errors indicated a problem with the cluster-autoscaler unable to scale because of some nodes are in a backoff state, but the logging does not indicate why node pools are in backoff. This was found via other means and resulted in us seeing errors at the instance group level. But the only error there was a
ZONE_RESOURCE_POOL_EXHAUSTED, nothing told us WHAT resource we had exhausted. Utlimately only GCP support was able to convey this information.
- A lot of log digging. Initial errors indicated a problem with the cluster-autoscaler unable to scale because of some nodes are in a backoff state, but the logging does not indicate why node pools are in backoff. This was found via other means and resulted in us seeing errors at the instance group level. But the only error there was a
- How could detection time be improved?
- Alerting on resource availability
- How was the root cause diagnosed?
- See above
- How could time to diagnosis be improved?
- Alerting on resource availability
- How did we reach the point where we knew how to mitigate the impact?
- GCP Support indicated the resource exhaustion was not impacting
pd-standarddisk types
- GCP Support indicated the resource exhaustion was not impacting
- How could time to mitigation be improved?
- n/a - while we resolve the API issue as quickly as we could, we remained at risk for other workloads
- What went well?
- This was an incident that spun around the world 3 times. Kudos for all team members involved in passing notes to the oncoming team members to ensure that we worked through this around the clock.
Post Incident Analysis
- Did we have other events in the past with the same root cause?
- No
- Do we have existing backlog items that would've prevented or greatly reduced the impact of this incident?
- No
- Was this incident triggered by a change (deployment of code or change to infrastructure)? If yes, link the issue.
- Neither - ServiceGCP
Lessons Learned
- We uncovered other bugs during remediation that slowed down our ability to manage this incident as a whole
- We lack the ability to monitor resource availability
- We lack the ability to reserve resources
- We learned that our workloads appear to work okay on
pd-standarddisks vspd-ssd
Guidelines
Resources
- If the Situation Zoom room was utilised, recording will be automatically uploaded to Incident room Google Drive folder (private)

