perf(api,docs): the published ceiling overstates the problem — the 16,000-task breach is a proven measurement artifact and the 60 s cliff does not reproduce

Summary

Re-measuring the published envelope at HEAD 9dca6c52 shows the performance story is materially better than what is documented, and that at least one of its headline numbers is a measurement artifact rather than a product limit.

Claim Published / believed Measured 2026-09-15
Whole-project load @ 2,000 tasks 60,159 ms (the "cliff") 4,098 ms
Task list first page @ 8,000 8,253 ms — breach, ceiling 681 ms — passes
Task list first page @ 16,000 96.67% error rate — ceiling 30/30 HTTP 200, 5,214 ms
Task list first page @ 10,000 never measured 397 ms

The 16,000-task breach is an artifact — this one is settled

capacity-tasks.json recorded err=0.9667 at 16,000 and the sweep declared a ceiling. A targeted re-run at 10,000 / 12,000 / 16,000 returned 200 on every single request, with the API container at OOM=false, restarts=0, and memory flat at ~144-148 MiB of a 7.7 GiB limit. Nothing was exhausted.

The cause is JWT expiry inside a long sweep:

  • Client._login() caches the access token in a module-level _TOKEN_CACHE keyed by base URL, and there is no refresh path. Every step of a sweep reuses the token minted at the first step.
  • ACCESS_TOKEN_LIFETIME is 15 minutes (settings/base.py:850).
  • The tasks sweep ran 17 min 51 s (20:38:31 → capacity-tasks.json mtime 20:56:22). 16,000 is its final step, so it executed after the token expired.
  • full_load ran in a separate process with a fresh token and took 13 min 55 s — just under the lifetime — which is why it reported err=0.0 at every step, including its genuine latency breach.
  • The diagnostic ran 6 min 28 s on a fresh token: no errors at any size.
  • p50 at the failing step was 4.8 ms — the speed of an auth rejection, not of database work.
  • The harness's control sample stayed green at 6.7 ms throughout, because it polls /api/v1/health/, which is unauthenticated (urls.py:91). The control is structurally incapable of detecting this.

Consequence: any capacity sweep that runs longer than 15 minutes fabricates a ceiling at whatever step it happens to be on. The published tasks ceiling of 16,000 should be withdrawn, not merely re-measured. Tracked against the harness in #3826.

The 60 s cliff does not reproduce

Same harness, same step, current code: 4,098 ms. Recomputing #3385 (closed)'s exponent table (excluding the unreliable 500-task point by bracketing 250 → 1,000): k ≈ 0.87, 1.52, 1.30, against the baseline's k ≈ 5.02 on the final doubling. The step shape that issue is built around is absent.

This does not establish which of "something fixed it" or "it was always an artifact" is true — those have different consequences, and #3385 (closed) already lists the artifact hypothesis as the thing to exclude first. Evidence is posted there; this issue does not close it.

Where runner contention does and does not apply

The contention argument is sound, but it has to be pointed at the right instrument:

  • It applies to the k6 nightly. sizing.md:172 records task_list p95 across four nightlies on identical data as 5,961 / 12,815 / 16,584 / 33,043 ms — a 5.5x spread, which that page already attributes to shared-runner contention. Supporting evidence: no job in .gitlab-ci.yml carries a tags: key, so GitLab free-schedules across three self-hosted runners (nuc-runner-2, runner-04-NUC, gitlab-runner-3-new). .gitlab-ci.yml:334 records the same commit with the same split running 326 s on one and 783 s on another — a 2.4x gap. #3671 measures a 3.8x semgrep swing cleanly correlated with which runner a job lands on, and flags nuc-runner-2 as the noisy member.
  • It does not explain the 60 s figure. That number came from the capacity harness, which runs off-CI by design on a developer workstation — the committed result records macOS-26.5.1-arm64-arm-64bit and "Shared developer workstation, not an idle host", at load ~16-18. Host contention is plausible there, but it is a different machine from the CI runners, and attributing it to runner hardware would be wrong.

A new anomaly, deliberately not explained here

The diagnostic found a step between 10,000 and 12,000 that I cannot account for:

Tasks Warm-up requests Measured p50
10,000 390-418 ms 396.7 ms
12,000 660-714 ms 5000.0 ms
16,000 1511-1592 ms 5214.2 ms

At 12,000 and 16,000 the measured samples are far slower than that step's own warm-up requests, which is backwards, and 12,000's p50 lands on exactly 5000.0 ms — too round for a latency measurement. One hypothesis worth testing: the Celery worker recomputing CPM for the freshly seeded project, starving the single uvicorn process the capacity stack deliberately runs. That is untested. It could also be another artifact of the same family as the token bug.

This matters for #3825: it is the first real signal about behavior above 8,000 tasks, and it is not yet trustworthy enough to plan against.

What should change

  • Withdraw the 16,000-task ceiling — it is an artifact, not a limit
  • Update administration/sizing.md's task-list rows and drop the "cliff, not a slope" framing (tracked in #3827)
  • Fix the harness so this class cannot recur: refresh the token per step, retain per-sample status codes, and make the control sample authenticated so it can see auth failures (#3826)
  • Explain or exclude the 10k → 12k step before any number above 8,000 is published
  • Leave #3385 (closed) open — this issue supplies evidence toward its first acceptance criterion but does not satisfy it

What this issue deliberately does not claim

That the ceiling is "fixed". #3384 (closed) records that the last three attempts in this area closed with a fix that did not match the title (#2277 (closed), #2807 (closed), #2767 (closed)), and warns against publishing "a forecast at a precision we have not earned". These are single, unrepeated runs on a contended workstation, and two of the four sizes carry known defects. The honest claim is narrower: the documented numbers overstate the problem, and at least one of them measures the instrument rather than the product.

Note on milestone

Filed to 0.4 as requested. No release:: value applied — that decision is the maintainer's.

Provenance

Measured 2026-09-15 at HEAD 9dca6c52 on the capacity harness (with the #3826 workaround) plus a targeted diagnostic with per-request status capture. Related: #3385 (closed), #3826, #3827, #3825, #3671, #2826 (closed), #2814 (closed), #2815 (closed), #3384 (closed).