ci: traffic-controller owns CI triggering (GITHUB_TOKEN updates, priority-ordered dispatch)

End the merge cascade and give the traffic-controller ownership of CI *triggering*.

- autoupdate.yml updates PR branches with the built-in GITHUB_TOKEN instead of a PAT,
  so an update push no longer auto-retriggers CI (GitHub's anti-recursion rule) — the
  cascade (every merge re-runs every PR, cancel-in-progress thrashing them) is gone.
- New scheduler ci-trigger.yml -> traffic_control.py --mode trigger (re-)triggers CI
  for the highest-priority PR(s) whose head SHA has absent/stale checks, a few at a
  time (inflight cap), in the existing P0-P9 / broken-draft priority order — a
  poor-man's merge queue reusing the priority core. It runs after autoupdate finishes
  (workflow_run, race-free) plus a cron backstop plus manual dispatch.
- Triggering uses workflow_dispatch, which is EXEMPT from anti-recursion, so the
  built-in GITHUB_TOKEN (actions: write) starts the run — NO PAT / secret change needed.
- ci.yml gains a workflow_dispatch trigger (pr/head_sha/reason inputs) and a per-PR
  concurrency group unifying pull_request and dispatch runs; its on: pull_request path
  is kept so brand-new PRs, human pushes, and fork PRs always get CI (fail-open).

Pure select_triggers / classify_sha_runs decision core added to traffic_control.py with
24 new unit tests (priority order, oldest-first fairness, inflight cap, fork skip, P0
bypass+preempt, head-SHA needy classification, and a liveness/anti-starvation simulation).

Closes #349

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
2026-07-05 01:27:40 -05:00
co-authored by Claude Opus 4.8
parent 833dfc030a
commit 05d06eb45b
5 changed files with 653 additions and 14 deletions
+32 -3
View File
@@ -4,10 +4,33 @@ name: CI
on:
pull_request:
branches: [main]
# The traffic-controller SCHEDULER (ci-trigger.yml, issue #349) (re-)triggers CI for a
# specific PR via this workflow_dispatch after a GITHUB_TOKEN auto-update has left the PR's
# head SHA with absent/stale checks. Dispatched on the PR's head BRANCH, so the run's checks
# land on the PR head SHA and satisfy branch protection. `on: pull_request` above is KEPT so
# brand-new PRs, human pushes, and fork PRs still get CI directly — this is the fail-open
# guarantee: CI is always triggerable even if the scheduler is broken or absent.
workflow_dispatch:
inputs:
pr:
description: "PR number this run is for (set by the traffic-controller scheduler)."
required: false
type: string
head_sha:
description: "Expected head SHA (informational, for traceability in the run log)."
required: false
type: string
reason:
description: "Why this run was dispatched (informational)."
required: false
type: string
# A new push to a PR cancels any in-flight run for that PR.
# A new trigger for a PR cancels that PR's own in-flight run (a newer head SHA supersedes).
# The group is keyed to the PR NUMBER so a `pull_request` run and a scheduler
# `workflow_dispatch` run for the SAME PR share one concurrency group (either supersedes a
# stale run of the other); it falls back to the ref when no PR number is in context.
concurrency:
group: ci-${{ github.ref }}
group: ci-pr-${{ github.event.pull_request.number || inputs.pr || github.ref }}
cancel-in-progress: true
permissions:
@@ -86,7 +109,13 @@ jobs:
env:
GH_TOKEN: ${{ github.token }}
GH_REPO: ${{ github.repository }}
SELF_PR: ${{ github.event.pull_request.number }}
# On a `pull_request` run this is the PR number and the script does its full in-run
# runner-priority orchestration. On a scheduler `workflow_dispatch` run (issue #349)
# the event is not `pull_request`, so the script no-ops here (`--mode orchestrate`
# only acts on pull_request events) — priority was ALREADY applied at trigger time by
# ci-trigger.yml, so re-doing the in-run hold-back would just waste runner time. The
# `|| inputs.pr` keeps the number in the log for a dispatched run.
SELF_PR: ${{ github.event.pull_request.number || inputs.pr }}
# P1–P9 bounded hold-back knobs, read by traffic_control.py. BUDGET must stay
# comfortably below timeout-minutes so the poll loop always exits 0 before the
# hard job timeout fires — a timed-out job would skip the heavy jobs and fail