Run risk classification flow end to end
Part of Customer Zero of MR Risk Assessment.
Why this exists
This issue is the only thing standing between the current state and a feature that works. Today, every assessment stops at "New risk assessment queued": the flow's claims get stored, but nothing ever turns them into a score, so no merge request ever shows a risk level. The scoring library is built and tested, but nothing calls it.
What you should see
Open a merge request on a project with the risk classification flow enabled. Within the run's lifetime, the widget moves from "New risk assessment queued" to a full result: a risk level, a confidence level, a rationale, and a signal breakdown. A run that never reports back is marked failed after 30 minutes, instead of sitting queued forever.
How the run reaches a result
Starting a run now also schedules its timeout: the flow's before_start step finds or creates the assessment row and moves it to queued. A worker, Ai::RiskClassification::CalculateScoreWorker, does the scoring: it runs the signal extractor, scores the result, and writes the score, confidence, signal breakdown, missing signals, rationale, and assessed_at in one update, then marks the row complete. A second worker, TimeoutWorker, marks a run failed after 30 minutes if it never reports back.
Two guards protect that result. First, the scoring worker writes only while the row is still queued. This stops a score that arrives late from landing on a row that has moved on, which would otherwise leave a row showing a full result while its status said something else. Second, the timeout is checked against the row's own idle time, not the age of the timeout job, and a timeout that fires too early defers itself instead of failing the row. This stops a timer left over from an earlier run from failing a classification that started after it. On top of both, a row that already reached complete can never be marked failed, so a late timeout can never erase a result that arrived in time.
domain_tags is a non-null column, exposed over GraphQL, that nothing writes today. Without it, a reviewer cannot see which domains a change touches. The scoring worker now fills it from the signal breakdown: any domain whose contribution scored above zero is added to the list.
Implementation in review: !253762 (merged).