Commit Graph
7 Commits
Author SHA1 Message Date
JMR-dev 58447d7d12 ci(e2e): gate the emulator on window focus to fix the RootViewPicker flake (#468)
Root cause: intermittently the launched activity window has has-window-focus=false
for the WHOLE instrumented run, so Espresso's RootViewPicker (onView().check(),
Intents.intended(), pressBack(), focus-dependent clipboard) times out after 10s and
fails EVERY focus-dependent test at once while the ~280 pure-Compose semantics tests
(which don't need window focus) pass. A failing E2E (35) leg's logcat (PR #470, run
28985259521) shows has-window-focus=true ZERO times across the whole session and both
the first attempt and the once-retry fail identically -- a persistent environmental
state, not a per-test transient. The prior mitigation, a single fire-and-forget
`adb shell input keyevent 82` (MENU) right after boot, is too weak: MENU no longer
dismisses the modern (API 30+) keyguard and, delivered before SystemUI/keyguard comes
up, is simply dropped -- so the insecure keyguard / non-interactive display persists
and no app window ever takes focus.

Fix: a single shared helper, .github/scripts/emulator_focus_gate.py, invoked
identically by BOTH E2E jobs (the e2e API 29-36 matrix AND e2e-preview API 37) and by
the local preflight runners (local_instrumented.py / api37_e2e.py), so it cannot drift.
It wakes the display (KEYCODE_WAKEUP), dismisses + disables the keyguard
(wm dismiss-keyguard, locksettings set-disabled true), keeps the screen on
(svc power stayon true + max screen_off_timeout), zeroes the animation scales, then
polls dumpsys power/window until the device is interactive AND a real window holds
input focus (mCurrentFocus non-null) -- re-nudging each iteration -- before the suite
runs. Applied uniformly, this also gives e2e-preview the animation-disable the matrix
already had. The gate is soft (bounded wait, then proceeds with a ::warning:: and the
final device state) and non-fatal (`|| true`), preserving #454's guarantee that the
unlock never aborts the boot; it leaves #454's manual boot, #460's path-filter and
#464's wedge-capture untouched.

The pure readiness parser is unit-tested by test_emulator_focus_gate.py (run by the
traffic-control-tests job). Determinism is validated by this PR's own matrix run.

Closes #468
2026-07-08 20:32:16 -05:00
JMR-dev 5c00a8da70 ci(e2e): route matrix emulator + system-image install through the #389-hardened installer
The E2E (33) matrix leg deadlocked the merge queue for ~2h when
android-emulator-runner's un-guarded "Create AVD and generate snapshot" step
died with "Error on ZipFile unknown archive" installing a corrupt Android
Emulator SDK zip. #389 hardened the platform/build-tools install (SHA-verify ->
reject-corrupt -> purge -> re-download) but left the emulator + system-image
install to the action, un-guarded.

Pre-install "emulator" + "system-images;android-<api>;google_apis;x86_64"
through setup_android_sdk.py before the emulator-runner steps, so a corrupt zip
is self-healed here and the action then finds both packages already installed
and skips its fragile fetch. Runs on both AVD-cache hit and miss (the emulator
binary + image live under the SDK root, not the ~/.android AVD-snapshot cache,
so they must be present for even a cached AVD to boot). Not added to the
android-sdk-v1 cache (kept small); re-install is a fast sdkmanager no-op when
already present.

The e2e-preview (API 37) job already routes emulator + system image through the
hardened installer, so no change there. setup_android_sdk.py already handles
these package ids generically; add a unit assertion pinning the matrix's
google_apis/x86_64 id to the correct purge path.

Closes #443
2026-07-08 10:39:38 -05:00
JMR-devandClaude Opus 4.8 aa628831be ci: retry + cache Android SDK/emulator setup to survive corrupt-zip sdkmanager failures (#389)
The dominant merge-blocking flake was the "Set up Android SDK" step
(android-actions/setup-android v4.0.1) dying BEFORE the emulator starts:

    Wrong version in preinstalled sdkmanager
    Warning: ... preparing SDK package Android Emulator: Error reading Zip
    content from a SeekableByteChannel.
    Error: The process '.../sdkmanager' failed with exit code 1

Root cause: the action's default cmdline-tools version (20.0) rarely matches the
runner image's preinstalled one, so it logs "Wrong version in preinstalled
sdkmanager" and re-fetches cmdline-tools with NO checksum; it then runs its
default `sdkmanager tools platform-tools` install. Any of those downloads can be
a corrupt/truncated zip, which sdkmanager turns into an un-retried exit 1. v4.0.1
is the latest release, so this is fixed by configuration + hardening, not a bump.

Harden with verify -> reject -> retry, never trusting sdkmanager's exit code
alone, via a new stdlib-only helper .github/scripts/setup_android_sdk.py:

- bootstrap: download the pinned cmdline-tools zip, verify size + SHA-256
  (authoritative pin, cross-checked against Google's published SHA-1), and
  install it to $ANDROID_SDK_ROOT/cmdline-tools/20.0 -- the exact path
  setup-android probes first, so the action reuses the verified tree and never
  does its own unverified "Wrong version" re-download. A mismatch (corrupt OR
  wrong version) deletes the bad zip + any half-extracted dir and re-downloads.
- install: sdkmanager --install with retry + backoff; on a corrupt package zip it
  purges the partial/corrupt package dir (and sdkmanager's temp dirs) before
  retrying, forcing a fresh download instead of a re-read.
- setup-android now runs with packages: "" (no flaky tools/platform-tools
  install) and cmdline-tools-version: "14742923" (reuse the verified bootstrap).
- actions/cache restore + success-gated save so only a verified SDK is ever
  cached (integrity gates the cache); shrinks the re-download/corruption surface.

Applied to every SDK-setup job (debug-build, unit-tests, static-analysis, e2e
matrix, e2e-preview). Emulator BOOT logic, #372 API-37 sharding, and #388
diagnostics are untouched. Pure-logic helpers are unit-tested
(test_setup_android_sdk.py, run by the traffic-control-tests job).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-06 18:59:05 -05:00
JMR-devandClaude Opus 4.8 2fcee291ce ci: trigger CI with the PAT so dispatches don't need manual approval (#351)
#350 made ci-trigger.yml dispatch ci.yml with the built-in GITHUB_TOKEN, on the
claim that a workflow_dispatch is anti-recursion-exempt so no PAT is needed. In
practice a GITHUB_TOKEN-triggered run is held in `action_required` awaiting manual
approval and never runs un-attended, so auto-updated PRs' CI never ran (stalled
#285). The original #349 design was right: dispatch with a PAT so the run executes
as the authorized owner with no approval gate.

- ci-trigger.yml: the trigger step's GH_TOKEN is now
  `${{ secrets.AUTOUPDATE_TOKEN || github.token }}` (was `${{ github.token }}`).
  AUTOUPDATE_TOKEN (the PAT) is REQUIRED for the scheduler; the `|| github.token`
  fallback stays fail-open but only starts CI if repo settings don't gate
  GITHUB_TOKEN-triggered runs.
- autoupdate.yml: branch update stays on GITHUB_TOKEN (must NOT retrigger CI --
  that would re-introduce the cascade). Clarified that AUTOUPDATE_TOKEN is still
  required by the repo (by ci-trigger.yml) so the secret isn't deleted.
- Corrected the now-wrong "no PAT needed / workflow_dispatch anti-recursion-exempt"
  comments in ci-trigger.yml and the traffic_control.py docstrings.

updates = GITHUB_TOKEN, triggering = PAT.

Validation: all three workflow YAMLs parse clean; traffic-control unit tests still
pass (59 tests) -- the change is workflow-env only, script logic unchanged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-05 13:28:10 -05:00
JMR-devandClaude Opus 4.8 05d06eb45b ci: traffic-controller owns CI triggering (GITHUB_TOKEN updates, priority-ordered dispatch)
End the merge cascade and give the traffic-controller ownership of CI *triggering*.

- autoupdate.yml updates PR branches with the built-in GITHUB_TOKEN instead of a PAT,
  so an update push no longer auto-retriggers CI (GitHub's anti-recursion rule) — the
  cascade (every merge re-runs every PR, cancel-in-progress thrashing them) is gone.
- New scheduler ci-trigger.yml -> traffic_control.py --mode trigger (re-)triggers CI
  for the highest-priority PR(s) whose head SHA has absent/stale checks, a few at a
  time (inflight cap), in the existing P0-P9 / broken-draft priority order — a
  poor-man's merge queue reusing the priority core. It runs after autoupdate finishes
  (workflow_run, race-free) plus a cron backstop plus manual dispatch.
- Triggering uses workflow_dispatch, which is EXEMPT from anti-recursion, so the
  built-in GITHUB_TOKEN (actions: write) starts the run — NO PAT / secret change needed.
- ci.yml gains a workflow_dispatch trigger (pr/head_sha/reason inputs) and a per-PR
  concurrency group unifying pull_request and dispatch runs; its on: pull_request path
  is kept so brand-new PRs, human pushes, and fork PRs always get CI (fail-open).

Pure select_triggers / classify_sha_runs decision core added to traffic_control.py with
24 new unit tests (priority order, oldest-first fairness, inflight cap, fork skip, P0
bypass+preempt, head-SHA needy classification, and a liveness/anti-starvation simulation).

Closes #349

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-05 01:27:40 -05:00
JMR-devandClaude Opus 4.8 a34db54b97 fix(ci): let any strictly-higher PR reclaim a broken/draft run (#342)
Restore (and extend to drafts) the old bash's broken-reclaim behaviour that the
initial Python refactor had dropped. runs_to_cancel now cancels an OTHER PR's
active/queued runs when EITHER:
  (a) THIS PR is P0 and that PR is strictly-lower (reclaim every lower runner); OR
  (b) that PR is broken/draft (effective priority 10) and THIS PR is strictly-higher
      (effective priority < 10) — a wasted run any ready PR may reclaim.

P1-P9 still never bump a *normal* (non-broken/draft) lower run; a broken/draft PR
(P10) preempts nothing (nothing is strictly-lower than the bottom, and the
equal-or-higher invariant means a P10 never cancels another P10). Self / main-push /
equal-or-higher invariants unchanged.

Updates the module docstring + ci.yml comments (the "only P0 preempts" wording
becomes: P0 preempts everything strictly-lower; additionally, any strictly-higher PR
preempts a broken/draft run) and the job step/permission/needs comments. Adds unit
tests: P3 reclaims a broken P10 run and a draft P10 run; P3 does not bump a normal P5
run; a P10 self preempts nothing; plus an end-to-end P5-reclaims-draft-then-waits
scenario. 37 unit tests pass; ci.yml parses clean; --dry-run shows a P3 cancelling a
draft (and broken) run while still yielding to a higher P1.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-04 22:19:01 -05:00
JMR-devandClaude Opus 4.8 8b7f895d6a ci: extract traffic-controller into a testable Python module (#342)
The priority-based runner orchestration ("traffic-control") lived as a large
inline-bash step in ci.yml — a two-pass preemption + hold-back script that was
effectively untestable in YAML. Move it into .github/scripts/traffic_control.py,
structured as a pure decision CORE + a thin gh-I/O SHELL:

* Pure functions (no network/clock/subprocess), unit-testable in isolation:
  - effective_priority(pr): lowest-numbered P0-P9, default P5; broken OR draft => 10.
  - runs_to_cancel(this_pr, all_prs, self_run_id): PASS 1 — run ids to cancel,
    empty unless THIS PR is P0; only strictly-lower running/queued runs; never self
    (by number or run id), never equal-or-higher.
  - wait_blockers(this_pr, all_prs): PASS 2 — yield to any strictly-higher PR with
    an active/queued run, and to same-level peers ordered ahead (running-first,
    then oldest createdAt). Empty => proceed.
* Shell (run_live): gathers the snapshot via gh, applies cancels, runs the bounded
  hold-back poll loop; always exits 0. --dry-run feeds the core a snapshot JSON and
  prints decisions with zero network.
* No jq/bash dependency (cross-platform, per the repo's Python-stdlib convention).

ci.yml's traffic-control job now checks out the repo and runs the module. Job
permissions gain `contents: read` (for checkout) alongside the existing
`actions: write` / `pull-requests: read`; env and downstream `needs:` wiring
unchanged; step stays `continue-on-error`.

33 stdlib unittest cases cover priority resolution, P0-only preemption, the
self/main/equal-or-higher invariants, and the same-level running-first/oldest
ordering.

Behaviour is preserved except the ticket's refinements: (1) drafts now count as
P10 (bottom); (2) an explicit same-level running-first-then-oldest tiebreaker; and
(3) per the ticket's order-of-operations, ONLY P0 preempts — the old bash also let
any higher-priority PR cancel a `broken` target's run, which no longer happens
(a broken/draft run is only cancelled by a P0, via the same strictly-lower rule).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-04 22:08:13 -05:00