Bump updated only on net changes and in bulk request updates

Found while investigating the internal API 504s on 2026-10-06. The updated column fails in two opposite ways.

Stage entries rewritten on every heartbeat

Each worker heartbeat PUT /v0.1/requests/<id> rewrites every pipeline_stage_entries row of the request, including entries with nothing to change. A request with 244 plans gets 244 row UPDATEs in one SERIALIZABLE transaction. On production, all 132 entries of one stage got a new updated within 9 ms, 103 of them pending and unchanged. crdb_internal.node_statement_statistics counted 1,339,511 executions of UPDATE pipeline_stage_entries SET updated = _ WHERE id = _ against 15,481 of UPDATE requests, about 86 stage-entry writes per PUT.

_sync_pipeline_entries assigns five attributes on every existing entry. SQLAlchemy marks the entry dirty on assignment and fires before_update for it even when no value changed. The sqlalchemy_utils.Timestamp listener then sets updated, and the flush emits a real UPDATE for each entry.

This MR replaces sqlalchemy_utils.Timestamp with a local mixin that defines the same columns. Its listener sets updated only when Session.is_modified() reports a net change. Alembic autogenerate detects no schema changes.

requests.updated frozen since insert

update_test_request writes the request row with a bulk Query.update(), which skips ORM events, so requests.updated has never moved since 88bc512e (api: add update test request endpoint). On production, 820 of 986 running requests had updated older than 30 minutes. The cancel path in delete_test_request has the same gap. Both now set updated explicitly, in the statement they already run.

Typing fallout

With a typed mixin, mypy now checks the model classes. The MR drops their # type: ignore comments and fixes what mypy reports: a shadowed loop variable, a missing Dict[str, Any] annotation, a redundant cast(), and uselist=True on Stages.pipeline_entries, which matches the runtime default.

Behaviour changes

  • updated on runs, stages, tokens and users no longer moves on saves without a net change. _upsert_run_db reassigns artifacts on every PUT, so runs.updated stops tracking heartbeats. requests.updated takes over that role.
  • requests.updated moves on every PUT and cancel. The API returns it, and the updated_before and updated_after filters read it.

Verification

I reproduced the bug and the fix on CockroachDB v20.1.3, the production version, with the schema built by alembic upgrade head. An unchanged 6-entry pipeline went from 6 UPDATEs to 0, and a single status change produces exactly 1. ChoiceType values and timezone-aware datetimes with a +02:00 offset compare equal after the round trip.

The four new tests in tests/test_crud_test_request.py fail on main and pass here. psycopg2 batches same-shaped UPDATEs into one executemany call, so the counter adds up parameter sets instead of statement executions.

After deploy, the ratio of UPDATE pipeline_stage_entries to UPDATE requests executions should drop from about 86 to 1-2. On 20.1, crdb_internal.node_statement_statistics is per node and resets every 2 hours, so compare both counts from the same node within one window.

Generated-by: Claude Code

Edited by Miroslav Vadkerti

Merge request reports

Loading
Loading