The dominant merge-blocking flake was the "Set up Android SDK" step
(android-actions/setup-android v4.0.1) dying BEFORE the emulator starts:
Wrong version in preinstalled sdkmanager
Warning: ... preparing SDK package Android Emulator: Error reading Zip
content from a SeekableByteChannel.
Error: The process '.../sdkmanager' failed with exit code 1
Root cause: the action's default cmdline-tools version (20.0) rarely matches the
runner image's preinstalled one, so it logs "Wrong version in preinstalled
sdkmanager" and re-fetches cmdline-tools with NO checksum; it then runs its
default `sdkmanager tools platform-tools` install. Any of those downloads can be
a corrupt/truncated zip, which sdkmanager turns into an un-retried exit 1. v4.0.1
is the latest release, so this is fixed by configuration + hardening, not a bump.
Harden with verify -> reject -> retry, never trusting sdkmanager's exit code
alone, via a new stdlib-only helper .github/scripts/setup_android_sdk.py:
- bootstrap: download the pinned cmdline-tools zip, verify size + SHA-256
(authoritative pin, cross-checked against Google's published SHA-1), and
install it to $ANDROID_SDK_ROOT/cmdline-tools/20.0 -- the exact path
setup-android probes first, so the action reuses the verified tree and never
does its own unverified "Wrong version" re-download. A mismatch (corrupt OR
wrong version) deletes the bad zip + any half-extracted dir and re-downloads.
- install: sdkmanager --install with retry + backoff; on a corrupt package zip it
purges the partial/corrupt package dir (and sdkmanager's temp dirs) before
retrying, forcing a fresh download instead of a re-read.
- setup-android now runs with packages: "" (no flaky tools/platform-tools
install) and cmdline-tools-version: "14742923" (reuse the verified bootstrap).
- actions/cache restore + success-gated save so only a verified SDK is ever
cached (integrity gates the cache); shrinks the re-download/corruption surface.
Applied to every SDK-setup job (debug-build, unit-tests, static-analysis, e2e
matrix, e2e-preview). Emulator BOOT logic, #372 API-37 sharding, and #388
diagnostics are untouched. Pure-logic helpers are unit-tested
(test_setup_android_sdk.py, run by the traffic-control-tests job).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
#350 made ci-trigger.yml dispatch ci.yml with the built-in GITHUB_TOKEN, on the
claim that a workflow_dispatch is anti-recursion-exempt so no PAT is needed. In
practice a GITHUB_TOKEN-triggered run is held in `action_required` awaiting manual
approval and never runs un-attended, so auto-updated PRs' CI never ran (stalled
#285). The original #349 design was right: dispatch with a PAT so the run executes
as the authorized owner with no approval gate.
- ci-trigger.yml: the trigger step's GH_TOKEN is now
`${{ secrets.AUTOUPDATE_TOKEN || github.token }}` (was `${{ github.token }}`).
AUTOUPDATE_TOKEN (the PAT) is REQUIRED for the scheduler; the `|| github.token`
fallback stays fail-open but only starts CI if repo settings don't gate
GITHUB_TOKEN-triggered runs.
- autoupdate.yml: branch update stays on GITHUB_TOKEN (must NOT retrigger CI --
that would re-introduce the cascade). Clarified that AUTOUPDATE_TOKEN is still
required by the repo (by ci-trigger.yml) so the secret isn't deleted.
- Corrected the now-wrong "no PAT needed / workflow_dispatch anti-recursion-exempt"
comments in ci-trigger.yml and the traffic_control.py docstrings.
updates = GITHUB_TOKEN, triggering = PAT.
Validation: all three workflow YAMLs parse clean; traffic-control unit tests still
pass (59 tests) -- the change is workflow-env only, script logic unchanged.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
End the merge cascade and give the traffic-controller ownership of CI *triggering*.
- autoupdate.yml updates PR branches with the built-in GITHUB_TOKEN instead of a PAT,
so an update push no longer auto-retriggers CI (GitHub's anti-recursion rule) — the
cascade (every merge re-runs every PR, cancel-in-progress thrashing them) is gone.
- New scheduler ci-trigger.yml -> traffic_control.py --mode trigger (re-)triggers CI
for the highest-priority PR(s) whose head SHA has absent/stale checks, a few at a
time (inflight cap), in the existing P0-P9 / broken-draft priority order — a
poor-man's merge queue reusing the priority core. It runs after autoupdate finishes
(workflow_run, race-free) plus a cron backstop plus manual dispatch.
- Triggering uses workflow_dispatch, which is EXEMPT from anti-recursion, so the
built-in GITHUB_TOKEN (actions: write) starts the run — NO PAT / secret change needed.
- ci.yml gains a workflow_dispatch trigger (pr/head_sha/reason inputs) and a per-PR
concurrency group unifying pull_request and dispatch runs; its on: pull_request path
is kept so brand-new PRs, human pushes, and fork PRs always get CI (fail-open).
Pure select_triggers / classify_sha_runs decision core added to traffic_control.py with
24 new unit tests (priority order, oldest-first fairness, inflight cap, fork skip, P0
bypass+preempt, head-SHA needy classification, and a liveness/anti-starvation simulation).
Closes#349
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Restore (and extend to drafts) the old bash's broken-reclaim behaviour that the
initial Python refactor had dropped. runs_to_cancel now cancels an OTHER PR's
active/queued runs when EITHER:
(a) THIS PR is P0 and that PR is strictly-lower (reclaim every lower runner); OR
(b) that PR is broken/draft (effective priority 10) and THIS PR is strictly-higher
(effective priority < 10) — a wasted run any ready PR may reclaim.
P1-P9 still never bump a *normal* (non-broken/draft) lower run; a broken/draft PR
(P10) preempts nothing (nothing is strictly-lower than the bottom, and the
equal-or-higher invariant means a P10 never cancels another P10). Self / main-push /
equal-or-higher invariants unchanged.
Updates the module docstring + ci.yml comments (the "only P0 preempts" wording
becomes: P0 preempts everything strictly-lower; additionally, any strictly-higher PR
preempts a broken/draft run) and the job step/permission/needs comments. Adds unit
tests: P3 reclaims a broken P10 run and a draft P10 run; P3 does not bump a normal P5
run; a P10 self preempts nothing; plus an end-to-end P5-reclaims-draft-then-waits
scenario. 37 unit tests pass; ci.yml parses clean; --dry-run shows a P3 cancelling a
draft (and broken) run while still yielding to a higher P1.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The priority-based runner orchestration ("traffic-control") lived as a large
inline-bash step in ci.yml — a two-pass preemption + hold-back script that was
effectively untestable in YAML. Move it into .github/scripts/traffic_control.py,
structured as a pure decision CORE + a thin gh-I/O SHELL:
* Pure functions (no network/clock/subprocess), unit-testable in isolation:
- effective_priority(pr): lowest-numbered P0-P9, default P5; broken OR draft => 10.
- runs_to_cancel(this_pr, all_prs, self_run_id): PASS 1 — run ids to cancel,
empty unless THIS PR is P0; only strictly-lower running/queued runs; never self
(by number or run id), never equal-or-higher.
- wait_blockers(this_pr, all_prs): PASS 2 — yield to any strictly-higher PR with
an active/queued run, and to same-level peers ordered ahead (running-first,
then oldest createdAt). Empty => proceed.
* Shell (run_live): gathers the snapshot via gh, applies cancels, runs the bounded
hold-back poll loop; always exits 0. --dry-run feeds the core a snapshot JSON and
prints decisions with zero network.
* No jq/bash dependency (cross-platform, per the repo's Python-stdlib convention).
ci.yml's traffic-control job now checks out the repo and runs the module. Job
permissions gain `contents: read` (for checkout) alongside the existing
`actions: write` / `pull-requests: read`; env and downstream `needs:` wiring
unchanged; step stays `continue-on-error`.
33 stdlib unittest cases cover priority resolution, P0-only preemption, the
self/main/equal-or-higher invariants, and the same-level running-first/oldest
ordering.
Behaviour is preserved except the ticket's refinements: (1) drafts now count as
P10 (bottom); (2) an explicit same-level running-first-then-oldest tiebreaker; and
(3) per the ticket's order-of-operations, ONLY P0 preempts — the old bash also let
any higher-priority PR cancel a `broken` target's run, which no longer happens
(a broken/draft run is only cancelled by a P0, via the same strictly-lower rule).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>