Provision KAS with Redis connection
With the advent of gitlab-org/charts/gitlab!1773 (merged), we should deploy KAS with Redis support enabled.
This will enable the KAS Rate-Limiting feature, which is a dependency of gitlab-org/gitlab#297682 (closed)
Overall information
Each kas instance holds up to 5 connections to Redis. This is configurable. You can see other parameters there too - dial, read, write, idle timeouts. If there is no or limited activity, we may only have e.g. 1 open connection (depending on the idle timeout, which is 50 seconds by default). Number of requests per seconds and the overall amount of data depend on the number of connected agentk Pods. All data is transient with TTL set and is explicitly removed as soon as it's no longer needed.
Rough estimation of data and request per second requirements for Redis
Redis requests due to rate limiting
Request/second estimation
It's 2 requests to Redis per 1 gRPC request to kas from each agentk Pod. Types of requests:
-
Configuration: Requests are long running. If there are no configuration updates (the usual situation - agent configuration should be pretty static), then the connection is closed and re-opened every 30 minutes. If there is an update then after it's sent to the agent, the connection is re-established (i.e. a request is made) and it sits there until the next update (or 30 minutes).
-
GitOps: Requests are long running. Works exactly the same as above with 1 gRPC request per configured manifest repository. If GitOps is not configured (i.e. 0 repos), no requests are made. Our expectation is that most users will use 1 manifest repository i.e. 1 gRPC request/30 minutes if no changes or 1 gRPC request per commit to that repository.
-
Reverse gRPC tunnel: Same model as above. Each
agentkPodestablishes 10 gRPC tunnel connections tokas. They stay open for 30 minutes and are then re-established. Nothing uses them at the moment. This is for gitlab-org&5528 and other features in the future, in development that's why unused. -
Cillium: 1 gRPC request per Cilium alert. How often - depends on the user's cluster and configuration of Cillium rules.
We expect a typical load of 11+ gRPC requests i.e. 22+ Redis requests per 30 minute interval + GitOps + Cillium requests per connected agentk Pod. That is ~0.03+ requests per second per connected agentk Pod (lower bound).
Rate limiter sets TTL for data to 59 seconds. "30 minutes" above, means 30 minutes with 5% jitter.
Size estimations
Rate limiter uses ~25 bytes per agent token (we store half the token in Redis, and a int count).
Redis requests due to tracking of connected agents
Request/second estimation
kas tracks connected agentk Pods in Redis for various purposes:
-
connectionsByAgentId- to look up information about connectedagentkPods by agent id. -
connectionsByProjectId- same as above but by configuration project id. -
tunnelsByAgentId- to look up information about connected reverse tunnels by agent id. EachagentkPodestablishes 10 tunnels.
All 3 use the same code for storage. Data is stored in a hash:
- Key of the hash is constructed using the agent/project id.
- key of the value inside of the hash is just a random 64 bit integer as a string
- Value in the hash is a generic wrapper to track the expiration time (because Redis does not support TTL on key-value pairs inside of a hash) and the wrapped value holds the actual information.
- Hash has a TTL on the whole hash set to 5 minutes. This is re-set every time a value is written to the hash and periodically (see below). If
kasstops data will be eventually gone and that's the desired behavior. - Data inside of the hash is refreshed and GCed every 4 and 10 minutes. GC is needed so that instances of
kastake care of stale data that another instance ofkaswrote that then e.g. crashed is cleaned up. Refresh is needed so that data fromkasinstance A that is still needed is not GCed bykasinstance B. Amount of data read and written and the number of requests to Redis depend on the number of things tracked i.e. number of connectedagentkPods. Reads useHSCANfor most efficient iteration and writes use pipelining for best network/connection utilization.
Per connected agentk Pod we expect a typical load of:
- 3+ Redis requests per 4 minute interval for refresh of 3 hashes above. This scales sublinearly because there are likely more than 1
agentkPodper project/agent id so they are batched together on both the read and write paths. - 3+ Redis requests per 10 minute interval for GC of hashes above. Same as above re. scaling.
- We actually don't have any features that perform hash lookups at the moment. We are planning and working on: agent list page, agent info page, Kubernetes CI tunnel, etc. They will all access data in Redis.
That is ~0.012+ + ~0.005+ = ~0.017+ requests per second per connected agentk Pod (lower bound).
Size estimations
From gitlab-org/cluster-integration/gitlab-agent!331 (merged):
-
connectionsByAgentIduseConnectedAgentInfoas the value. Typical size including the wrapper is ~160 bytes. -
connectionsByProjectId- same as above - ~160 bytes. -
tunnelsByAgentIduseTunnelInfoas the value. Typical size including the wrapper is ~130 bytes.
Both sizes will be growing a little bit over time, but nothing crazy.
Summary
Lower bound is ~0.03+ rps for rate limiting and ~0.017+ rps for hashes is ~0.047 rps per connected agentk Pod.
~25 + ~160 + ~160 + ~130 bytes = ~475 bytes of storage per connected agentk Pod.