ci(runners): P0 & broken-target preemption; P1-P9 yield without bumping in-progress #268

Merged
JMR-dev merged 2 commits from ci-p0-only-preemption into main 2026-07-03 22:10:43 +00:00
JMR-dev commented 2026-07-03 22:08:54 +00:00 (Migrated from github.com)

What

Refines the traffic-control runner-priority job (added in #265) so that only P0 preempts in-progress work, P1–P9 yield without bumping anyone, and a new broken label deprioritises a stuck PR to the bottom (and lets any higher-priority PR reclaim its wasted runner). Workflow-only change (.github/workflows/ci.yml); no app/build changes.

New semantics

Effective priority of a PR = broken ⇒ 10 (bottom, below P9, overriding any P0–P9); else the lowest-numbered P0–P9 label present (P0 = highest); else default P5.

Preemption (cancellation) rule — a strictly-lower-priority OTHER PR's active/queued CI run is cancelled iff (THIS PR is P0) OR (that PR is broken):

  • P0 — emergency only (app broken in production / emergency security update). P0 keeps the original preemption: it cancels all strictly-lower-priority in-progress/queued runs to grab runners immediately.
  • broken targets — because a broken PR's run can't merge, its run is wasted, so any higher-priority PR (not just P0) may cancel it to reclaim the runner.
  • P1–P9 vs non-broken targets — no cancellation. A higher-priority PR never evicts a lower-priority run that is already going; it takes the next free slot instead.

P1–P9 yield-without-bumping (bounded hold-back). Instead of cancelling, a P1–P9 PR's traffic-control defers its own heavy jobs (they all needs: traffic-control) while any strictly-higher-priority OTHER open PR still has an active/queued CI run — so the higher-priority PR's heavy jobs reach the runner queue first. It polls every HOLD_BACK_POLL_SECONDS (15s) up to HOLD_BACK_BUDGET_SECONDS (180s), then proceeds regardless. The budget is kept well under the job's timeout-minutes (bumped 5 → 6) so the loop always exit 0s before the hard timeout (a timed-out job would skip the heavy jobs and fail ci-passed).

The broken label

broken is a manually-applied signal — only the maintainer / repo owner sets it — meaning "this PR is stuck/failing; deprioritise it to the bottom so others aren't blocked behind it while it's fixed." A broken PR:

  • never preempts anyone — even if it is also labelled P0, broken wins (a stuck PR can't be an emergency merge);
  • always yields — every other PR, even lower P-levels, advances ahead of it;
  • may have its in-progress run cancelled by any higher-priority PR (its run is wasted).

Removing the label restores its normal P-priority on the next evaluation. Example this addresses: a P3 PR with failing CI was making lower-priority PRs wait behind it — marking it broken lets them proceed and reclaim its runner.

Preserved safety (unchanged)

  • Never touches main / push runs (--event pull_request + headBranch != "main" filters).
  • Never cancels this PR's own run (skips self by PR number and by GITHUB_RUN_ID).
  • Never cancels an equal-or-higher-priority PR (only strictly-lower, prio > self).
  • PR-controlled strings (branch names, labels) are only ever read via env / gh JSON into shell vars — never interpolated as code. The broken check is a fixed-string comparison inside jq.
  • continue-on-error + set +e + always exit 0; fail-open on any API hiccup.
  • traffic-control stays a non-required, best-effort job (absent from ci-passed.needs); ci-passed behaviour is unchanged.

Honest limitations / tradeoffs

  • GitHub Actions has no native priority queue — runner assignment is roughly FIFO. The hold-back is therefore a best-effort head-start, not a hard guarantee: under sustained contention the bounded wait can expire before the higher-priority PR drains, and the lower-priority PR then proceeds anyway.
  • A hold-back occupies a runner while it waits. traffic-control is a tiny gh-only job, and the wait is bounded (≤180s, hard-capped by timeout-minutes: 6), but it is a real, if small, cost — which is exactly why the wait is bounded rather than open-ended.
  • Same-priority PRs do not yield to each other (only strictly higher).

Validation

  • actionlint clean.
  • The embedded step script was extracted and exercised against a mocked gh CLI + fake clock (7 scenarios): P0 preempts strictly-lower & spares equal P0; P5 holds back then proceeds on drain; immediate-proceed with only equal/lower others; P1 waits out the budget against a never-draining P0; safety filters (main / push / own-run / completed / equal-or-higher all skipped, only the legit in-progress run cancelled); a normal P5 preempts a broken PR (whose P3 label would otherwise outrank it) while still yielding to a real P2; and a P0+broken PR preempts nobody and yields to everyone. All passed.
  • Builds / emulators intentionally not run (workflow-only change).

Not arming auto-merge — leaving this for owner review.

🤖 Generated with Claude Code

## What Refines the `traffic-control` runner-priority job (added in #265) so that **only P0 preempts in-progress work**, P1–P9 **yield without bumping** anyone, and a new **`broken`** label deprioritises a stuck PR to the bottom (and lets any higher-priority PR reclaim its wasted runner). Workflow-only change (`.github/workflows/ci.yml`); no app/build changes. ## New semantics Effective priority of a PR = **`broken` ⇒ 10** (bottom, below P9, overriding any P0–P9); else the lowest-numbered `P0`–`P9` label present (P0 = highest); else default **P5**. **Preemption (cancellation) rule** — a *strictly-lower-priority* OTHER PR's active/queued CI run is cancelled **iff** `(THIS PR is P0)` **OR** `(that PR is broken)`: - **P0 — emergency only** (app broken in production / emergency security update). P0 keeps the original preemption: it cancels **all** strictly-lower-priority in-progress/queued runs to grab runners immediately. - **`broken` targets** — because a broken PR's run can't merge, its run is wasted, so **any** higher-priority PR (not just P0) may cancel it to reclaim the runner. - **P1–P9 vs non-broken targets** — **no cancellation.** A higher-priority PR never evicts a lower-priority run that is already going; it takes the next free slot instead. **P1–P9 yield-without-bumping (bounded hold-back).** Instead of cancelling, a P1–P9 PR's `traffic-control` **defers its own heavy jobs** (they all `needs: traffic-control`) while any strictly-higher-priority OTHER open PR still has an active/queued CI run — so the higher-priority PR's heavy jobs reach the runner queue first. It polls every `HOLD_BACK_POLL_SECONDS` (15s) up to `HOLD_BACK_BUDGET_SECONDS` (180s), then proceeds regardless. The budget is kept well under the job's `timeout-minutes` (bumped 5 → 6) so the loop always `exit 0`s before the hard timeout (a timed-out job would skip the heavy jobs and fail `ci-passed`). ### The `broken` label `broken` is a **manually-applied** signal — only the maintainer / repo owner sets it — meaning *"this PR is stuck/failing; deprioritise it to the bottom so others aren't blocked behind it while it's fixed."* A broken PR: - **never preempts** anyone — even if it is also labelled `P0`, `broken` wins (a stuck PR can't be an emergency merge); - **always yields** — every other PR, even lower P-levels, advances ahead of it; - **may have its in-progress run cancelled** by any higher-priority PR (its run is wasted). Removing the label restores its normal P-priority on the next evaluation. Example this addresses: a `P3` PR with failing CI was making lower-priority PRs wait behind it — marking it `broken` lets them proceed *and* reclaim its runner. ## Preserved safety (unchanged) - Never touches **main / push** runs (`--event pull_request` + `headBranch != "main"` filters). - Never cancels **this PR's own** run (skips self by PR number and by `GITHUB_RUN_ID`). - Never cancels an **equal-or-higher-priority** PR (only strictly-lower, `prio > self`). - PR-controlled strings (branch names, labels) are only ever read via env / `gh` JSON into shell vars — never interpolated as code. The `broken` check is a fixed-string comparison inside `jq`. - `continue-on-error` + `set +e` + always `exit 0`; fail-open on any API hiccup. - `traffic-control` stays a **non-required, best-effort** job (absent from `ci-passed.needs`); `ci-passed` behaviour is unchanged. ## Honest limitations / tradeoffs - **GitHub Actions has no native priority queue** — runner assignment is roughly FIFO. The hold-back is therefore a **best-effort head-start, not a hard guarantee**: under sustained contention the bounded wait can expire before the higher-priority PR drains, and the lower-priority PR then proceeds anyway. - **A hold-back occupies a runner while it waits.** `traffic-control` is a tiny gh-only job, and the wait is bounded (≤180s, hard-capped by `timeout-minutes: 6`), but it is a real, if small, cost — which is exactly why the wait is bounded rather than open-ended. - Same-priority PRs do not yield to each other (only *strictly* higher). ## Validation - `actionlint` clean. - The embedded step script was extracted and exercised against a mocked `gh` CLI + fake clock (7 scenarios): P0 preempts strictly-lower & spares equal P0; P5 holds back then proceeds on drain; immediate-proceed with only equal/lower others; P1 waits out the budget against a never-draining P0; safety filters (main / push / own-run / completed / equal-or-higher all skipped, only the legit in-progress run cancelled); a normal P5 preempts a `broken` PR (whose P3 label would otherwise outrank it) while still yielding to a real P2; and a `P0+broken` PR preempts nobody and yields to everyone. All passed. - Builds / emulators intentionally not run (workflow-only change). Not arming auto-merge — leaving this for owner review. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
Sign in to join this conversation.