ci: traffic-controller owns CI triggering (GITHUB_TOKEN updates, PAT triggers by priority) — end the cascade #349

Closed
opened 2026-07-05 04:18:41 +00:00 by JMR-dev · 1 comment
JMR-dev commented 2026-07-05 04:18:41 +00:00 (Migrated from github.com)

Problem

autoupdate.yml updates every behind PR using a PAT (AUTOUPDATE_TOKEN), which re-triggers each PR's CI on every merge (a GITHUB_TOKEN push would not). With ci.yml's concurrency: cancel-in-progress, every merge cancels + restarts all other open PRs' runs → a cascade that thrashes CI and stops PRs converging (worse with the slow E2E matrix + flaky API-37). Merge queue would solve this but is organization-only and JMR-dev is a user account → unavailable. (See the ci-merge-cascade note.)

Fix — change who owns the CI trigger

  1. autoupdate.yml updates branches with GITHUB_TOKEN (drop the PAT for the update push). Branches stay current (satisfy "require up to date") but the update does NOT auto-retrigger CI → no cascade.
  2. The traffic-controller owns CI triggering, via the PAT. Since only a PAT / App / human push starts a run, give the PAT to the traffic-controller so it deliberately triggers CI for PRs in priority order (the P0–P9 + broken/draft⇒P10 + same-level running-first/oldest logic already in traffic_control.py), a few at a time instead of the herd. A poor-man's merge queue on the priority logic we already have.

Design notes

  • After a GITHUB_TOKEN auto-update, the PR's required checks are stale/absent for the new head SHA, so it cannot merge until CI runs on that SHA — the scheduler MUST reliably trigger the chosen PR's run via the PAT (gh workflow run / re-dispatch / marker commit) and must avoid starvation (every eligible PR eventually runs).
  • Today traffic-control is an in-run job that orders runner access + cancels lower runs. This shifts it (or adds a companion workflow) toward a scheduler that triggers runs — likely on: push→main + schedule (or PR events), PAT-authenticated, invoking traffic_control.py to choose the next PR(s). Reuse the existing priority core; add the "which to trigger now" decision + the PAT trigger I/O.
  • Higher-risk than a typical change — it is the core CI-trigger flow; a bug means PRs do not get CI / cannot merge. Wants careful design, a fail-open fallback (PRs still manually triggerable), and coordinator review before arming.
  • Sequence: land AFTER the current logging epic + its PRs settle, so changing the CI-trigger flow does not compound the in-flight churn.

Builds on #342 (traffic-controller module) + #346 (its CI tests).

## Problem `autoupdate.yml` updates every behind PR using a **PAT (`AUTOUPDATE_TOKEN`)**, which re-triggers each PR's CI on every merge (a `GITHUB_TOKEN` push would not). With `ci.yml`'s `concurrency: cancel-in-progress`, every merge cancels + restarts all other open PRs' runs → a cascade that thrashes CI and stops PRs converging (worse with the slow E2E matrix + flaky API-37). Merge queue would solve this but is **organization-only** and JMR-dev is a user account → unavailable. (See the `ci-merge-cascade` note.) ## Fix — change who owns the CI trigger 1. **`autoupdate.yml` updates branches with `GITHUB_TOKEN`** (drop the PAT for the *update* push). Branches stay current (satisfy "require up to date") but the update does NOT auto-retrigger CI → **no cascade**. 2. **The traffic-controller owns CI *triggering*, via the PAT.** Since only a PAT / App / human push starts a run, give the PAT to the traffic-controller so it deliberately triggers CI for PRs **in priority order** (the P0–P9 + broken/draft⇒P10 + same-level running-first/oldest logic already in `traffic_control.py`), a few at a time instead of the herd. A poor-man's merge queue on the priority logic we already have. ## Design notes - After a `GITHUB_TOKEN` auto-update, the PR's required checks are stale/absent for the new head SHA, so it cannot merge until CI runs on that SHA — the scheduler MUST reliably trigger the chosen PR's run via the PAT (`gh workflow run` / re-dispatch / marker commit) and must avoid **starvation** (every eligible PR eventually runs). - Today `traffic-control` is an in-run job that orders runner access + cancels lower runs. This shifts it (or adds a companion workflow) toward a **scheduler** that *triggers* runs — likely `on: push→main` + `schedule` (or PR events), PAT-authenticated, invoking `traffic_control.py` to choose the next PR(s). Reuse the existing priority core; add the "which to trigger now" decision + the PAT trigger I/O. - **Higher-risk than a typical change** — it is the core CI-trigger flow; a bug means PRs do not get CI / cannot merge. Wants careful design, a fail-open fallback (PRs still manually triggerable), and coordinator review before arming. - **Sequence:** land AFTER the current logging epic + its PRs settle, so changing the CI-trigger flow does not compound the in-flight churn. Builds on #342 (traffic-controller module) + #346 (its CI tests).
JMR-dev commented 2026-07-05 04:23:30 +00:00 (Migrated from github.com)

Dispatch plan (coordinator): Opus agent, after the logging epic (#324 areas + #331 guard #348) and its docs/CI PRs (#339/#344) have all merged — so changing the CI-trigger flow does not compound in-flight churn. Held in Ready until then. Watch for the gotchas: stale/absent required checks on the new head SHA after a GITHUB_TOKEN update, starvation, a brand-new PR's first run, and fork PRs (no PAT access).

Dispatch plan (coordinator): **Opus** agent, **after** the logging epic (#324 areas + #331 guard #348) and its docs/CI PRs (#339/#344) have all merged — so changing the CI-trigger flow does not compound in-flight churn. Held in Ready until then. Watch for the gotchas: stale/absent required checks on the new head SHA after a GITHUB_TOKEN update, starvation, a brand-new PR's first run, and fork PRs (no PAT access).
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: JMR-dev/LibreMail#349