ci(runners): priority-based runner orchestration via P0–P9 labels #265

Merged
JMR-dev merged 3 commits from ci-priority-runner-orchestration into main 2026-07-03 21:16:34 +00:00
JMR-dev commented 2026-07-03 21:04:05 +00:00 (Migrated from github.com)

What & why

The E2E matrix (~9 emulator jobs) plus build/unit/static jobs make each PR's CI runner-heavy. When many PRs run at once — especially when autoupdate.yml rebases them all as main advances — they starve the shared GitHub Actions runner pool with no notion of importance. This adds priority-based runner orchestration: when runners are contended, a PR's P0–P9 label decides precedence (P0 = highest, P9 = lowest).

Mechanism

A new lightweight traffic-control job runs first — the five heavy jobs (Debug build, Unit tests, Static analysis, E2E, E2E (API 37 preview)) now needs: traffic-control, so they only start consuming runners after it. It:

  1. Reads the triggering PR's effective priority = the highest-priority (lowest-numbered) P0–P9 label it carries, defaulting to P5 when the PR has no P-label.
  2. Snapshots all open PRs (gh pr list --json number,headRefName,labels).
  3. For every other open PR that is strictly lower priority (numerically greater), finds its active CI runs (gh run list --workflow ci.yml --branch <head> --event pull_request, status ≠ completed) and cancels them (gh run cancel), freeing those runners for this higher-priority PR.

The freed runners aren't reserved (GitHub has no such primitive), but removing lower-priority competitors from the in-progress/queue set materially improves the odds this PR's jobs get scheduled. A preempted PR simply re-runs on its next push / autoupdate rebase.

Attacker-influenced values (branch names, labels) are only ever passed via env: / read from gh JSON into shell variables — never interpolated into the script as code — so there's no script-injection surface.

Hard safety rules (all enforced in the script)

  • Never cancels a run on main / a push event — filters --event pull_request and re-checks headBranch != "main" in jq.
  • Never cancels this PR's own run — skips self by PR number, and skips the current $GITHUB_RUN_ID.
  • Never cancels an equal-or-higher-priority PR — only other_prio > self_prio.
  • Only strictly-lower-priority other open PRs' active runs are touched, and only the ci.yml workflow (autoupdate/release runs are left alone).

Does not break the merge gate

  • traffic-control is not a required check and is not in ci-passed's needs, so it can never block a merge.
  • Least privilege: it gets job-scoped actions: write (cancel runs) + pull-requests: read (read labels); all other jobs keep the top-level contents: read.
  • It's best-effort: every gh call is guarded, the script disables errexit and always exit 0, and the step is continue-on-error — an API error, a missing permission, or a fork PR can never fail the job.
  • Belt-and-suspenders: because the heavy jobs needs: it, a (should-never-happen) traffic-control failure would mark them skipped; ci-passed now treats skipped as a gate failure too, so the fail-safe direction is block the merge, never pass it untested.

The E2E matrix and Static analysis are otherwise untouched.

How to label PRs

Add exactly one P0–P9 label to a PR to set its CI priority (P0 = highest). No label ⇒ P5. If a PR carries several P-labels, the highest-priority (lowest number) wins.

Note: the on: pull_request trigger uses the default activity types (opened, synchronize, reopened) — not labeled — so changing a label does not re-trigger CI and won't affect an already-scheduled run. To apply a new priority immediately, push a commit (or re-run the workflow). Adding labeled to the trigger was deliberately avoided: it would launch a full CI run on every label edit.

Trade-offs & limitations (GitHub Actions constraints)

  • Cooperative, not true preemption. GitHub offers no priority runner queue. This cancels lower-priority competitors; it doesn't reserve capacity, and there's a few-seconds API-latency window before a cancel lands.
  • Cancel, don't hold-back. A low-priority PR still starts; it's cancelled only once a higher-priority PR's traffic-control runs. If nothing higher-priority is active, low-priority runs proceed (idle capacity isn't wasted).
  • Possible re-run churn / starvation. Under sustained high-priority load a low-priority PR can be repeatedly cancelled. Accepted by design; bounded because autoupdate only re-triggers when main advances.
  • Fork PRs. GITHUB_TOKEN is read-only for fork PRs regardless of declared permissions, so cancels there no-op with a warning. This repo uses same-repo feat-…/fix-… branches, so this isn't a practical issue.
  • Requires actions: write for GITHUB_TOKEN. If an org policy caps the token below write, cancels fail (warning only, no harm).
  • Ordering cost. Every PR pays the short traffic-control job (~a handful of seconds) before heavy jobs start — negligible versus multi-minute E2E, and it's the point: free runners before requesting them.

Validation

  • actionlint clean on all workflows.
  • Priority algorithm (label→priority, defaults, ties, anchored P0–P9 match, decision matrix) unit-checked (16/16) and the bash control flow (TSV parse, numeric gating, self-run-id skip, cancel counting) exercised with a stubbed harness. No emulators/builds were run, per scope.

🤖 Generated with Claude Code

## What & why The E2E matrix (~9 emulator jobs) plus build/unit/static jobs make each PR's CI runner-heavy. When many PRs run at once — especially when `autoupdate.yml` rebases them all as `main` advances — they starve the shared GitHub Actions runner pool with no notion of importance. This adds **priority-based runner orchestration**: when runners are contended, a PR's `P0`–`P9` label decides precedence (P0 = highest, P9 = lowest). ## Mechanism A new lightweight **`traffic-control`** job runs **first** — the five heavy jobs (`Debug build`, `Unit tests`, `Static analysis`, `E2E`, `E2E (API 37 preview)`) now `needs: traffic-control`, so they only start consuming runners after it. It: 1. Reads the triggering PR's effective priority = the **highest-priority (lowest-numbered)** `P0`–`P9` label it carries, defaulting to **P5** when the PR has no P-label. 2. Snapshots all open PRs (`gh pr list --json number,headRefName,labels`). 3. For every **other** open PR that is **strictly lower priority** (numerically greater), finds its active CI runs (`gh run list --workflow ci.yml --branch <head> --event pull_request`, status ≠ completed) and **cancels** them (`gh run cancel`), freeing those runners for this higher-priority PR. The freed runners aren't *reserved* (GitHub has no such primitive), but removing lower-priority competitors from the in-progress/queue set materially improves the odds this PR's jobs get scheduled. A preempted PR simply re-runs on its next push / autoupdate rebase. Attacker-influenced values (branch names, labels) are only ever passed via `env:` / read from `gh` JSON into shell variables — never interpolated into the script as code — so there's no script-injection surface. ### Hard safety rules (all enforced in the script) - **Never** cancels a run on `main` / a `push` event — filters `--event pull_request` and re-checks `headBranch != "main"` in `jq`. - **Never** cancels this PR's **own** run — skips self by PR number, and skips the current `$GITHUB_RUN_ID`. - **Never** cancels an **equal-or-higher**-priority PR — only `other_prio > self_prio`. - Only strictly-lower-priority **other** open PRs' active runs are touched, and only the `ci.yml` workflow (autoupdate/release runs are left alone). ### Does not break the merge gate - `traffic-control` is **not** a required check and is **not** in `ci-passed`'s `needs`, so it can never block a merge. - Least privilege: it gets job-scoped `actions: write` (cancel runs) + `pull-requests: read` (read labels); all other jobs keep the top-level `contents: read`. - It's **best-effort**: every `gh` call is guarded, the script disables `errexit` and always `exit 0`, and the step is `continue-on-error` — an API error, a missing permission, or a fork PR can never fail the job. - Belt-and-suspenders: because the heavy jobs `needs:` it, a (should-never-happen) traffic-control failure would mark them `skipped`; `ci-passed` now treats `skipped` as a gate failure too, so the fail-safe direction is **block the merge**, never pass it untested. The E2E matrix and Static analysis are otherwise untouched. ## How to label PRs Add exactly one `P0`–`P9` label to a PR to set its CI priority (P0 = highest). No label ⇒ **P5**. If a PR carries several P-labels, the highest-priority (lowest number) wins. Note: the `on: pull_request` trigger uses the default activity types (`opened`, `synchronize`, `reopened`) — **not** `labeled` — so changing a label does **not** re-trigger CI and won't affect an already-scheduled run. To apply a new priority immediately, push a commit (or re-run the workflow). Adding `labeled` to the trigger was deliberately avoided: it would launch a full CI run on every label edit. ## Trade-offs & limitations (GitHub Actions constraints) - **Cooperative, not true preemption.** GitHub offers no priority runner queue. This cancels lower-priority *competitors*; it doesn't reserve capacity, and there's a few-seconds API-latency window before a cancel lands. - **Cancel, don't hold-back.** A low-priority PR still *starts*; it's cancelled only once a higher-priority PR's `traffic-control` runs. If nothing higher-priority is active, low-priority runs proceed (idle capacity isn't wasted). - **Possible re-run churn / starvation.** Under sustained high-priority load a low-priority PR can be repeatedly cancelled. Accepted by design; bounded because autoupdate only re-triggers when `main` advances. - **Fork PRs.** `GITHUB_TOKEN` is read-only for fork PRs regardless of declared permissions, so cancels there no-op with a warning. This repo uses same-repo `feat-…`/`fix-…` branches, so this isn't a practical issue. - **Requires `actions: write` for `GITHUB_TOKEN`.** If an org policy caps the token below write, cancels fail (warning only, no harm). - **Ordering cost.** Every PR pays the short `traffic-control` job (~a handful of seconds) before heavy jobs start — negligible versus multi-minute E2E, and it's the point: free runners *before* requesting them. ## Validation - `actionlint` clean on all workflows. - Priority algorithm (label→priority, defaults, ties, anchored `P0`–`P9` match, decision matrix) unit-checked (16/16) and the bash control flow (TSV parse, numeric gating, self-run-id skip, cancel counting) exercised with a stubbed harness. No emulators/builds were run, per scope. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
Sign in to join this conversation.