Pin Artemis production worker log level to info
The production Artemis worker pool deadlocked again on 2026-09-22, the
same failure assessed on 2026-08-04. All five pods sat Running 1/1 Ready while processing nothing: py-spy dumps taken 90 seconds apart came
back byte-identical on every thread, default held 1685 ready messages at
consumer_utilisation 3.9e-7, and 239 of 261 guest requests had not moved
in over an hour.
Debug records exceed PIPE_BUF (4096), so the worker threads in each
dramatiq child tear frames on its length-framed log pipe. A writer then
blocks in send_bytes while it holds the logging handler lock, and the
remaining threads queue behind it. Every wedged pod stranded open Postgres
transactions, the oldest pinning row locks on guest_requests for three
days.
ARTEMIS_LOG_LEVEL=debug reached the live deployment through a kubectl
edit on 2025-07-31 and appears nowhere in git. TFT-5005 assumed the next
terragrunt apply would reconcile it away. It does not: env merges by
name, Helm keeps entries it never owned, and the value outlived every
upgrade including revision 403 on 2026-09-17.
Rendering the name from worker_extra_env hands ownership back to Helm,
so info wins on the next apply. The artemis-core chart never renders
ARTEMIS_LOG_LEVEL, which is why logging.level in values.yaml.tftpl
has no effect and cannot carry this instead.
Production only. Staging and dev keep debug logging for troubleshooting.
The .debug image tag stays everywhere, per the artemis!1552 TODO.
Verified by rendering artemis-core 0.0.5 against the deployed values plus
this change: the worker container receives ARTEMIS_LOG_LEVEL: info.
Assisted-by: Claude Code