Websockets service (ActionCable) is not resilient to redis failovers
Summary
The Websockets service which is built on ActionCable does not handle redis failovers very well. When a failover occurs, the puma processes crash. This results in high error rates, paging the SRE on-call.
The process crashes with this exception:
/srv/gitlab/vendor/bundle/ruby/2.7.0/gems/redis-4.4.0/lib/redis/connection/ruby.rb:63:in `block in _read_from_socket': Connection reset by peer (Errno::ECONNRESET)Impact
- The websockets service crashes which creates unavailability of "live update" functionality.
- This creates operational load for SRE on-call (e.g. gitlab-com/gl-infra/production#6310 (closed)).
- End-users may experience lack of live updates but will not see any errors, as the JS client reconnects gracefully.
Recommendation
There is an upstream issue to address this in rails.
Short-term: We should find a workaround to add redis re-connection to our actioncable puma processes.
Long-term: Work with upstream on getting this fixed in rails.
Verification
The issue should be reproducible by running redis in a HA sentinel configuration and triggering a failover (redis-cli -p 26379 sentinel failover gitlab-redis).