Pin Artemis production worker log level to info

The production Artemis worker pool deadlocked again on 2026-09-22, the same failure assessed on 2026-08-04. All five pods sat Running 1/1 Ready while processing nothing: py-spy dumps taken 90 seconds apart came back byte-identical on every thread, default held 1685 ready messages at consumer_utilisation 3.9e-7, and 239 of 261 guest requests had not moved in over an hour.

Debug records exceed PIPE_BUF (4096), so the worker threads in each dramatiq child tear frames on its length-framed log pipe. A writer then blocks in send_bytes while it holds the logging handler lock, and the remaining threads queue behind it. Every wedged pod stranded open Postgres transactions, the oldest pinning row locks on guest_requests for three days.

ARTEMIS_LOG_LEVEL=debug reached the live deployment through a kubectl edit on 2025-07-31 and appears nowhere in git. TFT-5005 assumed the next terragrunt apply would reconcile it away. It does not: env merges by name, Helm keeps entries it never owned, and the value outlived every upgrade including revision 403 on 2026-09-17.

Rendering the name from worker_extra_env hands ownership back to Helm, so info wins on the next apply. The artemis-core chart never renders ARTEMIS_LOG_LEVEL, which is why logging.level in values.yaml.tftpl has no effect and cannot carry this instead.

Production only. Staging and dev keep debug logging for troubleshooting. The .debug image tag stays everywhere, per the artemis!1552 TODO.

Verified by rendering artemis-core 0.0.5 against the deployed values plus this change: the worker container receives ARTEMIS_LOG_LEVEL: info.

TFT-5005

Assisted-by: Claude Code

Edited by Miroslav Vadkerti

Merge request reports

Loading
Loading