ci: traffic-controller owns CI triggering (GITHUB_TOKEN updates, priority-ordered dispatch)

End the merge cascade and give the traffic-controller ownership of CI *triggering*.

- autoupdate.yml updates PR branches with the built-in GITHUB_TOKEN instead of a PAT,
  so an update push no longer auto-retriggers CI (GitHub's anti-recursion rule) — the
  cascade (every merge re-runs every PR, cancel-in-progress thrashing them) is gone.
- New scheduler ci-trigger.yml -> traffic_control.py --mode trigger (re-)triggers CI
  for the highest-priority PR(s) whose head SHA has absent/stale checks, a few at a
  time (inflight cap), in the existing P0-P9 / broken-draft priority order — a
  poor-man's merge queue reusing the priority core. It runs after autoupdate finishes
  (workflow_run, race-free) plus a cron backstop plus manual dispatch.
- Triggering uses workflow_dispatch, which is EXEMPT from anti-recursion, so the
  built-in GITHUB_TOKEN (actions: write) starts the run — NO PAT / secret change needed.
- ci.yml gains a workflow_dispatch trigger (pr/head_sha/reason inputs) and a per-PR
  concurrency group unifying pull_request and dispatch runs; its on: pull_request path
  is kept so brand-new PRs, human pushes, and fork PRs always get CI (fail-open).

Pure select_triggers / classify_sha_runs decision core added to traffic_control.py with
24 new unit tests (priority order, oldest-first fairness, inflight cap, fork skip, P0
bypass+preempt, head-SHA needy classification, and a liveness/anti-starvation simulation).

Closes #349

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
2026-07-05 01:27:40 -05:00
co-authored by Claude Opus 4.8
parent 833dfc030a
commit 05d06eb45b
5 changed files with 653 additions and 14 deletions
+14 -8
View File
@@ -6,13 +6,17 @@ name: Auto-update PR branches
# touched (PR_FILTER: all) — this is no longer limited to PRs with GitHub auto-merge
# enabled.
#
# IMPORTANT: for the branch update to RE-TRIGGER the PR's CI (so it can pass and merge),
# this must run with a PAT, not the default GITHUB_TOKEN — pushes made by GITHUB_TOKEN do
# not start new workflow runs (GitHub's anti-recursion rule), so the updated PR would sit
# with stale checks. Create a fine-grained PAT scoped to this repo with
# contents:read/write + pull-requests:read/write and add it as the AUTOUPDATE_TOKEN secret.
# Without it this falls back to GITHUB_TOKEN, which updates the branch but will NOT re-run
# the PR's checks.
# IMPORTANT (issue #349): the branch update runs with the default GITHUB_TOKEN — ON PURPOSE.
# A GITHUB_TOKEN push does NOT start new workflow runs (GitHub's anti-recursion rule), so
# updating every behind PR here NO LONGER re-triggers every PR's CI. That deliberately breaks
# the old merge-cascade (every merge -> autoupdate rebases all PRs with a PAT -> all re-run ->
# ci.yml's cancel-in-progress kills each in-flight run -> PRs thrash and can't converge).
# Branches still go up to date (satisfying "require branches up to date"); they just don't
# auto-run CI on the new head SHA. Re-triggering that SHA's CI is now OWNED by the traffic-
# controller scheduler (`.github/workflows/ci-trigger.yml` -> `traffic_control.py --mode
# trigger`), which triggers the updated PRs deliberately, in priority order, a few at a time.
# So this workflow must NOT use the PAT for the update push (that would re-introduce the
# cascade). The AUTOUPDATE_TOKEN secret is no longer needed by this workflow.
on:
push:
@@ -38,7 +42,9 @@ jobs:
- name: Update all behind PRs
uses: chinthakagodawita/autoupdate@0707656cd062a3b0cf8fa9b2cda1d1404d74437e # v1.7.0
env:
GITHUB_TOKEN: ${{ secrets.AUTOUPDATE_TOKEN || secrets.GITHUB_TOKEN }}
# Default GITHUB_TOKEN — NOT a PAT — so this update push does not auto-retrigger CI
# (anti-recursion). See the header: re-triggering is owned by ci-trigger.yml.
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
PR_FILTER: "all"
PR_READY_STATE: "all"
MERGE_CONFLICT_ACTION: "ignore"
+85
View File
@@ -0,0 +1,85 @@
# SPDX-License-Identifier: GPL-3.0-or-later
name: CI trigger (traffic-controller)
# The traffic-controller SCHEDULER (issue #349). It OWNS CI *triggering*. After main advances,
# autoupdate.yml updates every behind PR's branch with the built-in GITHUB_TOKEN which, by
# GitHub's anti-recursion rule, does NOT start CI — so those PRs sit with absent/stale required
# checks on their new head SHA and cannot merge. This workflow then (re-)triggers CI for the
# highest-priority such PR(s), a few at a time (an inflight cap), in the existing P0–P9 /
# broken-draft priority order — a poor-man's merge queue that replaces the old "every merge
# re-runs every PR" thundering herd (the cascade; see the ci-merge-cascade note + issue #349).
#
# HOW IT TRIGGERS: `traffic_control.py --mode trigger` runs `gh workflow run ci.yml --ref
# <pr-head-branch>`. A workflow_dispatch is EXEMPT from the anti-recursion rule, so even the
# built-in GITHUB_TOKEN's dispatch DOES start the run — no PAT is required (this job grants its
# token `actions: write`). The dispatched run executes on the PR's head branch, so its checks
# land on the PR head SHA and satisfy branch protection's required checks.
#
# WHEN IT RUNS:
# • workflow_run, after "Auto-update PR branches" completes — the race-free moment: autoupdate
# has finished moving branches to their new (checkless) head SHAs, so this pass sees exactly
# the PRs that now need a run. (A bare `push: main` trigger would race autoupdate and often
# read the pre-update SHAs, missing them until the next pass.)
# • schedule (cron) — a backstop so no PR is ever permanently un-triggered even if a
# workflow_run is missed/skipped (part of the fail-open guarantee), and so a brand-new PR
# whose first `on: pull_request` run got cancelled is still picked up.
# • workflow_dispatch — manual kick.
#
# FAIL-OPEN: the script guards every gh call and always exits 0; and structurally, ci.yml keeps
# its `on: pull_request` trigger, so a human push (and a brand-new PR) always triggers CI
# regardless of this scheduler — CI can never become permanently un-triggerable. Fork PRs (no
# token/secret access) are skipped here and left to `on: pull_request`, so they are never wedged.
on:
workflow_run:
workflows: ["Auto-update PR branches"]
types: [completed]
schedule:
# Backstop cadence (UTC). GitHub may delay scheduled runs under load; that is fine — this
# is only a safety net behind the immediate workflow_run trigger above.
- cron: "*/15 * * * *"
workflow_dispatch:
# Trigger-only; this workflow never gates a merge. `actions: write` lets the built-in
# GITHUB_TOKEN dispatch ci.yml (workflow_dispatch) and cancel strictly-lower runs when a P0
# emergency preempts. `pull-requests: read` + `contents: read` cover the PR/label enumeration.
permissions:
contents: read
pull-requests: read
actions: write
# One trigger pass at a time. Do NOT cancel an in-flight pass (cancel-in-progress: false):
# a half-finished pass could leave some needy PRs un-triggered until the next pass.
concurrency:
group: ci-trigger
cancel-in-progress: false
jobs:
trigger:
name: Trigger CI by priority
runs-on: ubuntu-latest
timeout-minutes: 10 # generous backstop; the script only enumerates + dispatches, no waits
steps:
- name: Check out source
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
- name: Set up Python
uses: actions/setup-python@ece7cb06caefa5fff74198d8649806c4678c61a1 # v6.3.0
with:
python-version: "3.x"
# gh is auto-configured from GH_TOKEN / GH_REPO. The script guards every gh call and
# always exits 0, so a hiccup (API error, missing permission, fork PR) can never wedge CI
# — and even a total failure here leaves ci.yml's `on: pull_request` path intact.
- name: Trigger CI for the highest-priority PR(s) needing a run
env:
GH_TOKEN: ${{ github.token }}
GH_REPO: ${{ github.repository }}
# Poor-man's merge-queue width: at most this many PRs run CI concurrently under the
# scheduler (a P0 emergency bypasses this cap). Kept conservative because each PR
# fans out to the whole E2E matrix (~8 API levels + preview); this is the main knob
# to raise for throughput vs runner budget. The coordinator drives runner allocation.
MAX_INFLIGHT_RUNS: "2"
# The workflow file the scheduler enumerates runs for and dispatches.
CI_WORKFLOW_FILE: "ci.yml"
run: python3 .github/scripts/traffic_control.py --mode trigger
+32 -3
View File
@@ -4,10 +4,33 @@ name: CI
on:
pull_request:
branches: [main]
# The traffic-controller SCHEDULER (ci-trigger.yml, issue #349) (re-)triggers CI for a
# specific PR via this workflow_dispatch after a GITHUB_TOKEN auto-update has left the PR's
# head SHA with absent/stale checks. Dispatched on the PR's head BRANCH, so the run's checks
# land on the PR head SHA and satisfy branch protection. `on: pull_request` above is KEPT so
# brand-new PRs, human pushes, and fork PRs still get CI directly — this is the fail-open
# guarantee: CI is always triggerable even if the scheduler is broken or absent.
workflow_dispatch:
inputs:
pr:
description: "PR number this run is for (set by the traffic-controller scheduler)."
required: false
type: string
head_sha:
description: "Expected head SHA (informational, for traceability in the run log)."
required: false
type: string
reason:
description: "Why this run was dispatched (informational)."
required: false
type: string
# A new push to a PR cancels any in-flight run for that PR.
# A new trigger for a PR cancels that PR's own in-flight run (a newer head SHA supersedes).
# The group is keyed to the PR NUMBER so a `pull_request` run and a scheduler
# `workflow_dispatch` run for the SAME PR share one concurrency group (either supersedes a
# stale run of the other); it falls back to the ref when no PR number is in context.
concurrency:
group: ci-${{ github.ref }}
group: ci-pr-${{ github.event.pull_request.number || inputs.pr || github.ref }}
cancel-in-progress: true
permissions:
@@ -86,7 +109,13 @@ jobs:
env:
GH_TOKEN: ${{ github.token }}
GH_REPO: ${{ github.repository }}
SELF_PR: ${{ github.event.pull_request.number }}
# On a `pull_request` run this is the PR number and the script does its full in-run
# runner-priority orchestration. On a scheduler `workflow_dispatch` run (issue #349)
# the event is not `pull_request`, so the script no-ops here (`--mode orchestrate`
# only acts on pull_request events) — priority was ALREADY applied at trigger time by
# ci-trigger.yml, so re-doing the in-run hold-back would just waste runner time. The
# `|| inputs.pr` keeps the number in the log for a dispatched run.
SELF_PR: ${{ github.event.pull_request.number || inputs.pr }}
# P1–P9 bounded hold-back knobs, read by traffic_control.py. BUDGET must stay
# comfortably below timeout-minutes so the poll loop always exits 0 before the
# hard job timeout fires — a timed-out job would skip the heavy jobs and fail