Ends the merge cascade and moves ownership of CI triggering to the traffic-controller.
autoupdate.yml updates branches with the built-in GITHUB_TOKEN (was a PAT). A GITHUB_TOKEN push does not start workflow runs (GitHub's anti-recursion rule), so updating every behind PR no longer re-triggers every PR's CI → the cascade (autoupdate rebases all PRs × ci.yml cancel-in-progress) is gone. Branches still go up to date, satisfying "require branches up to date".
New scheduler ci-trigger.yml → traffic_control.py --mode trigger deliberately (re-)triggers CI for the highest-priority PR(s) whose head SHA has absent/stale required checks — a few at a time (an inflight cap, default 2), in the existing P0–P9 / broken-draft⇒P10 / oldest-first priority order. A poor-man's merge queue built on the priority core from #342. It runs:
workflow_run after "Auto-update PR branches" completes — the race-free moment (autoupdate has finished moving branches to their new, checkless head SHAs), so the pass sees exactly the PRs that now need a run;
schedule (cron, 15 min) — a fail-open backstop;
workflow_dispatch — manual kick.
Triggering uses workflow_dispatch, which is EXEMPT from the anti-recursion rule — so even the built-in GITHUB_TOKEN (granted actions: write, exactly as the existing traffic-control job already is) starts the run. The dispatch runs on the PR's head branch, so its checks land on the PR head SHA and satisfy branch protection's required checks. No PAT, no secret, no repo/scope change is required (see "Maintainer note").
ci.yml gains a workflow_dispatch trigger (pr/head_sha/reason inputs) and a per-PR concurrency group (ci-pr-<number|input|ref>) that unifies pull_request and dispatch runs for the same PR (either supersedes a stale run of the other). Its on: pull_request trigger is kept.
Split: pure decision core + thin gh shell (reuses #342)
select_triggers() (pure, unit-tested) reuses effective_priority, _ordering_key, runs_to_cancel, ACTIVE, PullRequest. classify_sha_runs() (pure) decides "needy" from a head SHA's runs. The shell (run_trigger, gather_trigger_snapshot, _dispatch_ci) is the only part touching gh and reuses _gh_json / _cancel_run / _positive_int. --mode orchestrate (the in-run job) is unchanged.
How each gotcha is handled
Stale/absent required checks on the new head SHA after a GITHUB_TOKEN update — the scheduler detects them (classify_sha_runs: a head SHA with no active run and no verdict run = needy) and reliably (re-)triggers via gh workflow run ci.yml --ref <head-branch>. Verified empirically that pull_request runs report the PR head SHA (not the merge SHA), so head-SHA matching is correct.
Starvation — every eligible PR eventually runs. Triggering a PR gives its head SHA a run, so it leaves the needy set; between merges the needy set only shrinks, and the scheduler re-runs on every autoupdate + a cron backstop, so the set drains in a bounded number of passes. Ordering is by priority (oldest-first within a level), so higher-priority PRs are served first, never exclusively forever (a served PR stops being needy until its next push/update). Proved by TriggerStarvationTests.test_every_needy_pr_is_triggered_within_bounded_passes (mixed priorities, cap 2 — all drain; lower-priority PRs are not starved).
A brand-new PR's first run — handled entirely by ci.yml's retained on: pull_request (opened): a new PR gets CI immediately, independent of the scheduler. If that first run is later cancelled (e.g. preemption), the SHA becomes needy and the scheduler picks it up; the cron backstop is a second safety net.
Fork PRs (no PAT/secret access) — the scheduler detects isCrossRepository and skips them (can't dispatch --ref on a branch outside the base repo), logging that they're left to on: pull_request. Fork PRs still get CI on open and on each fork push via on: pull_request, so they are never wedged. (test_fork_pr_skipped.)
Fail-open (core safety property) — two layers: (a) structural — ci.yml keeps on: pull_request, so any human push always triggers CI regardless of the scheduler; CI can never become permanently un-triggerable even if the scheduler is broken/absent. (b) script — every gh call is guarded and run_trigger always returns 0 (and the top-level except in __main__ swallows anything), so a misfire never wedges anything; the next autoupdate/cron pass retries.
P0 emergency (reuses the priority core)
A needy P0 bypasses the inflight cap and preempts its strictly-lower active runs (reusing runs_to_cancel + _cancel_run) so a runner frees immediately; the preempted PRs re-enter the needy set and are re-triggered next pass in priority order. (test_p0_bypasses_cap_and_preempts_lower, test_p0_does_not_preempt_equal_priority.)
Deviation from the suggested title / PAT (flagged)
The ticket assumed a PAT would be needed to trigger. It is not: workflow_dispatch is exempt from anti-recursion, so the built-in token triggers CI — a strictly better outcome (no maintainer secret/scope change). The title's parenthetical was adjusted from "PAT triggers by priority" to "priority-ordered dispatch" for accuracy. AUTOUPDATE_TOKEN is now unused by any workflow and may be deleted (optional cleanup; autoupdate.yml's environment: CI_CD is likewise now vestigial). Rename the PR if you prefer the original wording.
Maintainer note — no config change required
ci-trigger.yml requests permissions: actions: write for the built-in GITHUB_TOKEN; the existing traffic-control job in ci.ymlalready relies on the same grant, so no repo/org "workflow permissions" change is needed. If, unexpectedly, dispatch is rejected in this repo's settings, the fix is a one-line permissions allowance — CI stays triggerable via on: pull_request meanwhile.
import yaml; yaml.safe_load(...) → ci.yml, autoupdate.yml, ci-trigger.yml all parse.
py_compile both scripts → OK.
--mode trigger --dry-run scenarios: post-merge (2 running + 3 needy + 1 fork → triggers the single P1 in the free slot, skips the fork); P0 emergency (bypasses a full cap, preempts both lower runs); fork-only-needy (triggers nothing, degrades gracefully).
Like #342/#346, this is CI infrastructure (Python + YAML) with no app-runtime surface, so it ships with the pure-Python unit tests run by the traffic-control-tests gate rather than emulator E2E.
Closes #349
## Design
Ends the merge cascade and moves ownership of CI **triggering** to the traffic-controller.
1. **`autoupdate.yml` updates branches with the built-in `GITHUB_TOKEN`** (was a PAT). A `GITHUB_TOKEN` push does not start workflow runs (GitHub's anti-recursion rule), so updating every behind PR no longer re-triggers every PR's CI → the cascade (`autoupdate rebases all PRs` × `ci.yml cancel-in-progress`) is gone. Branches still go up to date, satisfying "require branches up to date".
2. **New scheduler `ci-trigger.yml` → `traffic_control.py --mode trigger`** deliberately (re-)triggers CI for the highest-priority PR(s) whose head SHA has absent/stale required checks — **a few at a time** (an inflight cap, default 2), in the **existing** P0–P9 / broken-draft⇒P10 / oldest-first priority order. A poor-man's merge queue built on the priority core from #342. It runs:
- **`workflow_run` after "Auto-update PR branches" completes** — the race-free moment (autoupdate has finished moving branches to their new, checkless head SHAs), so the pass sees exactly the PRs that now need a run;
- **`schedule` (cron, 15 min)** — a fail-open backstop;
- **`workflow_dispatch`** — manual kick.
3. **Triggering uses `workflow_dispatch`, which is EXEMPT from the anti-recursion rule** — so even the built-in `GITHUB_TOKEN` (granted `actions: write`, exactly as the existing `traffic-control` job already is) starts the run. The dispatch runs on the PR's head **branch**, so its checks land on the PR **head SHA** and satisfy branch protection's required checks. **No PAT, no secret, no repo/scope change is required** (see "Maintainer note").
`ci.yml` gains a `workflow_dispatch` trigger (`pr`/`head_sha`/`reason` inputs) and a **per-PR concurrency group** (`ci-pr-<number|input|ref>`) that unifies `pull_request` and dispatch runs for the same PR (either supersedes a stale run of the other). Its `on: pull_request` trigger is **kept**.
### Split: pure decision core + thin gh shell (reuses #342)
`select_triggers()` (pure, unit-tested) reuses `effective_priority`, `_ordering_key`, `runs_to_cancel`, `ACTIVE`, `PullRequest`. `classify_sha_runs()` (pure) decides "needy" from a head SHA's runs. The shell (`run_trigger`, `gather_trigger_snapshot`, `_dispatch_ci`) is the only part touching `gh` and reuses `_gh_json` / `_cancel_run` / `_positive_int`. `--mode orchestrate` (the in-run job) is **unchanged**.
## How each gotcha is handled
- **Stale/absent required checks on the new head SHA after a `GITHUB_TOKEN` update** — the scheduler detects them (`classify_sha_runs`: a head SHA with no active run and no *verdict* run = needy) and reliably (re-)triggers via `gh workflow run ci.yml --ref <head-branch>`. Verified empirically that `pull_request` runs report the PR **head** SHA (not the merge SHA), so head-SHA matching is correct.
- **Starvation — every eligible PR eventually runs.** Triggering a PR gives its head SHA a run, so it **leaves the needy set**; between merges the needy set only shrinks, and the scheduler re-runs on every autoupdate + a cron backstop, so the set drains in a bounded number of passes. Ordering is by priority (oldest-first within a level), so higher-priority PRs are served *first*, never *exclusively forever* (a served PR stops being needy until its next push/update). Proved by `TriggerStarvationTests.test_every_needy_pr_is_triggered_within_bounded_passes` (mixed priorities, cap 2 — all drain; lower-priority PRs are not starved).
- **A brand-new PR's first run** — handled entirely by `ci.yml`'s retained `on: pull_request` (`opened`): a new PR gets CI immediately, independent of the scheduler. If that first run is later cancelled (e.g. preemption), the SHA becomes needy and the scheduler picks it up; the cron backstop is a second safety net.
- **Fork PRs (no PAT/secret access)** — the scheduler detects `isCrossRepository` and **skips** them (can't dispatch `--ref` on a branch outside the base repo), logging that they're left to `on: pull_request`. Fork PRs still get CI on open and on each fork push via `on: pull_request`, so they are **never wedged**. (`test_fork_pr_skipped`.)
- **Fail-open (core safety property)** — two layers: (a) *structural* — `ci.yml` keeps `on: pull_request`, so any human push always triggers CI regardless of the scheduler; CI can **never** become permanently un-triggerable even if the scheduler is broken/absent. (b) *script* — every `gh` call is guarded and `run_trigger` always returns 0 (and the top-level `except` in `__main__` swallows anything), so a misfire never wedges anything; the next autoupdate/cron pass retries.
## P0 emergency (reuses the priority core)
A needy **P0** bypasses the inflight cap and **preempts** its strictly-lower active runs (reusing `runs_to_cancel` + `_cancel_run`) so a runner frees immediately; the preempted PRs re-enter the needy set and are re-triggered next pass in priority order. (`test_p0_bypasses_cap_and_preempts_lower`, `test_p0_does_not_preempt_equal_priority`.)
## Deviation from the suggested title / PAT (flagged)
The ticket assumed a PAT would be needed to trigger. It is **not**: `workflow_dispatch` is exempt from anti-recursion, so the built-in token triggers CI — a strictly better outcome (no maintainer secret/scope change). The title's parenthetical was adjusted from "PAT triggers by priority" to "priority-ordered dispatch" for accuracy. **`AUTOUPDATE_TOKEN` is now unused** by any workflow and may be deleted (optional cleanup; `autoupdate.yml`'s `environment: CI_CD` is likewise now vestigial). Rename the PR if you prefer the original wording.
### Maintainer note — no config change required
`ci-trigger.yml` requests `permissions: actions: write` for the built-in `GITHUB_TOKEN`; the existing `traffic-control` job in `ci.yml` **already** relies on the same grant, so no repo/org "workflow permissions" change is needed. If, unexpectedly, dispatch is rejected in this repo's settings, the fix is a one-line `permissions` allowance — CI stays triggerable via `on: pull_request` meanwhile.
## Tests & validation (no emulator/Gradle)
- `python -m unittest discover -s .github/scripts` → **59 pass** (24 new: `ClassifyShaRunsTests`, `SelectTriggersTests`, `TriggerStarvationTests`, `TriggerSnapshotParsingTests`; existing `--mode orchestrate` tests unchanged).
- `import yaml; yaml.safe_load(...)` → `ci.yml`, `autoupdate.yml`, `ci-trigger.yml` all parse.
- `py_compile` both scripts → OK.
- `--mode trigger --dry-run` scenarios: post-merge (2 running + 3 needy + 1 fork → triggers the single P1 in the free slot, skips the fork); P0 emergency (bypasses a full cap, preempts both lower runs); fork-only-needy (triggers nothing, degrades gracefully).
Like #342/#346, this is CI infrastructure (Python + YAML) with no app-runtime surface, so it ships with the pure-Python unit tests run by the `traffic-control-tests` gate rather than emulator E2E.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Closes #349
Design
Ends the merge cascade and moves ownership of CI triggering to the traffic-controller.
autoupdate.ymlupdates branches with the built-inGITHUB_TOKEN(was a PAT). AGITHUB_TOKENpush does not start workflow runs (GitHub's anti-recursion rule), so updating every behind PR no longer re-triggers every PR's CI → the cascade (autoupdate rebases all PRs×ci.yml cancel-in-progress) is gone. Branches still go up to date, satisfying "require branches up to date".New scheduler
ci-trigger.yml→traffic_control.py --mode triggerdeliberately (re-)triggers CI for the highest-priority PR(s) whose head SHA has absent/stale required checks — a few at a time (an inflight cap, default 2), in the existing P0–P9 / broken-draft⇒P10 / oldest-first priority order. A poor-man's merge queue built on the priority core from #342. It runs:workflow_runafter "Auto-update PR branches" completes — the race-free moment (autoupdate has finished moving branches to their new, checkless head SHAs), so the pass sees exactly the PRs that now need a run;schedule(cron, 15 min) — a fail-open backstop;workflow_dispatch— manual kick.Triggering uses
workflow_dispatch, which is EXEMPT from the anti-recursion rule — so even the built-inGITHUB_TOKEN(grantedactions: write, exactly as the existingtraffic-controljob already is) starts the run. The dispatch runs on the PR's head branch, so its checks land on the PR head SHA and satisfy branch protection's required checks. No PAT, no secret, no repo/scope change is required (see "Maintainer note").ci.ymlgains aworkflow_dispatchtrigger (pr/head_sha/reasoninputs) and a per-PR concurrency group (ci-pr-<number|input|ref>) that unifiespull_requestand dispatch runs for the same PR (either supersedes a stale run of the other). Itson: pull_requesttrigger is kept.Split: pure decision core + thin gh shell (reuses #342)
select_triggers()(pure, unit-tested) reuseseffective_priority,_ordering_key,runs_to_cancel,ACTIVE,PullRequest.classify_sha_runs()(pure) decides "needy" from a head SHA's runs. The shell (run_trigger,gather_trigger_snapshot,_dispatch_ci) is the only part touchingghand reuses_gh_json/_cancel_run/_positive_int.--mode orchestrate(the in-run job) is unchanged.How each gotcha is handled
Stale/absent required checks on the new head SHA after a
GITHUB_TOKENupdate — the scheduler detects them (classify_sha_runs: a head SHA with no active run and no verdict run = needy) and reliably (re-)triggers viagh workflow run ci.yml --ref <head-branch>. Verified empirically thatpull_requestruns report the PR head SHA (not the merge SHA), so head-SHA matching is correct.Starvation — every eligible PR eventually runs. Triggering a PR gives its head SHA a run, so it leaves the needy set; between merges the needy set only shrinks, and the scheduler re-runs on every autoupdate + a cron backstop, so the set drains in a bounded number of passes. Ordering is by priority (oldest-first within a level), so higher-priority PRs are served first, never exclusively forever (a served PR stops being needy until its next push/update). Proved by
TriggerStarvationTests.test_every_needy_pr_is_triggered_within_bounded_passes(mixed priorities, cap 2 — all drain; lower-priority PRs are not starved).A brand-new PR's first run — handled entirely by
ci.yml's retainedon: pull_request(opened): a new PR gets CI immediately, independent of the scheduler. If that first run is later cancelled (e.g. preemption), the SHA becomes needy and the scheduler picks it up; the cron backstop is a second safety net.Fork PRs (no PAT/secret access) — the scheduler detects
isCrossRepositoryand skips them (can't dispatch--refon a branch outside the base repo), logging that they're left toon: pull_request. Fork PRs still get CI on open and on each fork push viaon: pull_request, so they are never wedged. (test_fork_pr_skipped.)Fail-open (core safety property) — two layers: (a) structural —
ci.ymlkeepson: pull_request, so any human push always triggers CI regardless of the scheduler; CI can never become permanently un-triggerable even if the scheduler is broken/absent. (b) script — everyghcall is guarded andrun_triggeralways returns 0 (and the top-levelexceptin__main__swallows anything), so a misfire never wedges anything; the next autoupdate/cron pass retries.P0 emergency (reuses the priority core)
A needy P0 bypasses the inflight cap and preempts its strictly-lower active runs (reusing
runs_to_cancel+_cancel_run) so a runner frees immediately; the preempted PRs re-enter the needy set and are re-triggered next pass in priority order. (test_p0_bypasses_cap_and_preempts_lower,test_p0_does_not_preempt_equal_priority.)Deviation from the suggested title / PAT (flagged)
The ticket assumed a PAT would be needed to trigger. It is not:
workflow_dispatchis exempt from anti-recursion, so the built-in token triggers CI — a strictly better outcome (no maintainer secret/scope change). The title's parenthetical was adjusted from "PAT triggers by priority" to "priority-ordered dispatch" for accuracy.AUTOUPDATE_TOKENis now unused by any workflow and may be deleted (optional cleanup;autoupdate.yml'senvironment: CI_CDis likewise now vestigial). Rename the PR if you prefer the original wording.Maintainer note — no config change required
ci-trigger.ymlrequestspermissions: actions: writefor the built-inGITHUB_TOKEN; the existingtraffic-controljob inci.ymlalready relies on the same grant, so no repo/org "workflow permissions" change is needed. If, unexpectedly, dispatch is rejected in this repo's settings, the fix is a one-linepermissionsallowance — CI stays triggerable viaon: pull_requestmeanwhile.Tests & validation (no emulator/Gradle)
python -m unittest discover -s .github/scripts→ 59 pass (24 new:ClassifyShaRunsTests,SelectTriggersTests,TriggerStarvationTests,TriggerSnapshotParsingTests; existing--mode orchestratetests unchanged).import yaml; yaml.safe_load(...)→ci.yml,autoupdate.yml,ci-trigger.ymlall parse.py_compileboth scripts → OK.--mode trigger --dry-runscenarios: post-merge (2 running + 3 needy + 1 fork → triggers the single P1 in the free slot, skips the fork); P0 emergency (bypasses a full cap, preempts both lower runs); fork-only-needy (triggers nothing, degrades gracefully).Like #342/#346, this is CI infrastructure (Python + YAML) with no app-runtime surface, so it ships with the pure-Python unit tests run by the
traffic-control-testsgate rather than emulator E2E.🤖 Generated with Claude Code