ci(e2e): converge matrix emulator boot on e2e-preview manual boot #454

Merged
JMR-dev merged 1 commits from ci-448-emulator-boot-race into main 2026-07-08 18:41:18 +00:00
JMR-dev commented 2026-07-08 17:50:07 +00:00 (Migrated from github.com)

What

Closes #448.

The matrix e2e job (API 29-36) booted its emulator through reactivecircus/android-emulator-runner, whose un-guarded, fatal adb shell input keyevent 82 races system_server's input binder republish on snapshot resume -> No service published for: input. That is the intermittent boot race that flaked the merge queue and -- once #446 finally let runs reach the boot phase on API 33 -- failed both "Run E2E tests" and its "retry after emulator boot race" step.

This converges the matrix job on the e2e-preview (API 37) job's proven hand-provisioned manual boot -- the fix the old retry-step comment already named as "the definitive fix ... tracked separately".

Chosen approach: Option A (converge on e2e-preview)

Of the issue's two options I took Option A (fuller: replace android-emulator-runner's boot), not Option B (harden the action in place), because:

  • The maintainer's own in-code comment diagnosed the race as the action's internal, un-guarded keyevent 82 firing before any user script: -- so Option B (a gate inside the action's script) cannot defuse it; only owning the boot can.
  • e2e-preview cold-boots and is proven stable + REQUIRED; the race is documented as snapshot-resume-specific, so cold boot removes the root cause.
  • It is self-validating: this PR's own matrix run exercises the new boot on every API level 29-36.

What changed (only .github/workflows/ci.yml, only the e2e job)

Removed (the android-emulator-runner boot path):

  • Cache AVD snapshot + Create AVD and generate snapshot for caching
  • Run E2E tests + Run E2E tests (retry after emulator boot race)

Added (mirrors e2e-preview):

  • Create AVD -- avdmanager create avd -n test -k system-images;android-<api>;google_apis;x86_64 -d pixel_2, with ANDROID_AVD_HOME pinned + exported (same "Unknown AVD name" guard e2e-preview uses).
  • Boot emulator and run E2E:
    • adb start-server up front (adb-bind race guard),
    • a 2-attempt COLD boot loop (-no-snapshot), each with one bounded timeout 300 adb wait-for-device shell 'until sys.boot_completed' (a stuck emulator fails fast instead of hanging to the 50-min cap),
    • a non-fatal adb shell input keyevent 82 || true unlock (kills the race) plus animation-disable settings (parity with the old disable-animations: true),
    • a readiness gate before connectedDebugAndroidTest,
    • gradle retries once on a test failure (parity with the old retry).
    • Emulator flags mirror e2e-preview: -no-window -no-audio -no-boot-anim -no-snapshot -accel on -gpu swiftshader_indirect -camera-back/front none -verbose -debug init,avd_config,kernel.

Kept untouched: #446's hardened pre-install emulator + system image step, the shared android-sdk-v1 cache, Enable KVM, the failure-diagnostics dump, and all report/artifact uploads (the diagnostics artifact now also carries emulator-api<api>.log + boot-diagnostics-api<api>.txt). Device/target/arch (pixel_2 / google_apis / x86_64) and the API matrix (29-36) stay in lockstep with testOptions.managedDevices in app/build.gradle.kts.

Tradeoff (flagged)

Cold boot drops the per-leg AVD snapshot cache -- deliberately, because snapshot resume is the documented root cause of the race and mirroring e2e-preview means -no-snapshot. Net boot-time impact is ~neutral: cache-miss legs previously cold-booted twice (generate + resume) and now boot once; only warm cache-hit legs pay a small cold-boot delta, well within the 50-min cap. No wedge-capture wrapper is added here -- it was reverted from this matrix job in 90dfb18 for hanging all 8 legs; only e2e-preview keeps it.

Validation

  • .github/workflows/ci.yml parses clean under PyYAML; the e2e job step list and API matrix (29-36) verified.
  • The bash mirrors e2e-preview's proven idioms nearly line-for-line (set -euo pipefail, the for attempt in 1 2 boot loop, the single-quoted boot-wait, the run || { ...; run; } retry); GitHub's default run: shell is bash -eo pipefail. (No local bash -n -- this dev box has no usable bash: WSL2 cannot start without virtualization, and there is no Git Bash. The matrix run on this PR is the real gate.)
  • No app code touched -> no unit/instrumented tests or AppLog changes needed (CI-infra change).
  • Self-validating: the matrix E2E on this PR boots every API level 29-36 with the new pattern, so a broken boot fails its own E2E and cannot merge.
## What Closes #448. The matrix `e2e` job (API 29-36) booted its emulator through `reactivecircus/android-emulator-runner`, whose **un-guarded, fatal** `adb shell input keyevent 82` races `system_server`'s `input` binder republish on snapshot resume -> `No service published for: input`. That is the intermittent **boot race** that flaked the merge queue and -- once #446 finally let runs reach the boot phase on API 33 -- failed **both** "Run E2E tests" and its "retry after emulator boot race" step. This converges the matrix job on the **`e2e-preview` (API 37) job's proven hand-provisioned manual boot** -- the fix the old retry-step comment already named as "the definitive fix ... tracked separately". ## Chosen approach: Option A (converge on e2e-preview) Of the issue's two options I took **Option A** (fuller: replace android-emulator-runner's boot), not Option B (harden the action in place), because: - The maintainer's own in-code comment diagnosed the race as the action's internal, un-guarded `keyevent 82` firing **before** any user `script:` -- so Option B (a gate *inside* the action's script) cannot defuse it; only owning the boot can. - `e2e-preview` cold-boots and is **proven stable + REQUIRED**; the race is documented as snapshot-**resume**-specific, so cold boot removes the root cause. - It is **self-validating**: this PR's own matrix run exercises the new boot on every API level 29-36. ## What changed (only `.github/workflows/ci.yml`, only the `e2e` job) Removed (the android-emulator-runner boot path): - `Cache AVD snapshot` + `Create AVD and generate snapshot for caching` - `Run E2E tests` + `Run E2E tests (retry after emulator boot race)` Added (mirrors `e2e-preview`): - **`Create AVD`** -- `avdmanager create avd -n test -k system-images;android-<api>;google_apis;x86_64 -d pixel_2`, with `ANDROID_AVD_HOME` pinned + exported (same "Unknown AVD name" guard e2e-preview uses). - **`Boot emulator and run E2E`**: - `adb start-server` up front (adb-bind race guard), - a **2-attempt COLD boot loop** (`-no-snapshot`), each with **one bounded** `timeout 300 adb wait-for-device shell 'until sys.boot_completed'` (a stuck emulator fails fast instead of hanging to the 50-min cap), - a **non-fatal** `adb shell input keyevent 82 || true` unlock (kills the race) plus animation-disable settings (parity with the old `disable-animations: true`), - a **readiness gate** before `connectedDebugAndroidTest`, - gradle **retries once** on a test failure (parity with the old retry). - Emulator flags mirror e2e-preview: `-no-window -no-audio -no-boot-anim -no-snapshot -accel on -gpu swiftshader_indirect -camera-back/front none -verbose -debug init,avd_config,kernel`. Kept untouched: #446's hardened **pre-install emulator + system image** step, the shared **`android-sdk-v1`** cache, `Enable KVM`, the failure-diagnostics dump, and all report/artifact uploads (the diagnostics artifact now also carries `emulator-api<api>.log` + `boot-diagnostics-api<api>.txt`). Device/target/arch (`pixel_2` / `google_apis` / `x86_64`) and the API matrix (29-36) stay in lockstep with `testOptions.managedDevices` in `app/build.gradle.kts`. ## Tradeoff (flagged) Cold boot **drops the per-leg AVD snapshot cache** -- deliberately, because snapshot resume is the documented root cause of the race and mirroring e2e-preview means `-no-snapshot`. Net boot-time impact is ~neutral: cache-miss legs previously cold-booted **twice** (generate + resume) and now boot once; only warm cache-hit legs pay a small cold-boot delta, well within the 50-min cap. No wedge-capture wrapper is added here -- it was reverted from this matrix job in `90dfb18` for hanging all 8 legs; only e2e-preview keeps it. ## Validation - `.github/workflows/ci.yml` parses clean under **PyYAML**; the `e2e` job step list and API matrix (29-36) verified. - The bash mirrors e2e-preview's proven idioms nearly line-for-line (`set -euo pipefail`, the `for attempt in 1 2` boot loop, the single-quoted boot-wait, the `run || { ...; run; }` retry); GitHub's default `run:` shell is `bash -eo pipefail`. (No local `bash -n` -- this dev box has no usable bash: WSL2 cannot start without virtualization, and there is no Git Bash. The matrix run on this PR is the real gate.) - No app code touched -> no unit/instrumented tests or AppLog changes needed (CI-infra change). - **Self-validating**: the matrix E2E on this PR boots every API level 29-36 with the new pattern, so a broken boot fails its own E2E and cannot merge.
mergify[bot] commented 2026-07-08 18:36:53 +00:00 (Migrated from github.com)

Merge Queue Status

  • ✅ Entered queue — 2026-07-08 18:36 UTC · Rule: default · triggered by merge protections
  • 🚫 Left the queue — 2026-07-08 18:41 UTC · at 2b958404bb3be6169b7b5f7d82f50168af9fb565

This pull request spent 4 minutes 41 seconds in the queue, with no time running CI.

Reason

Pull request #454 has been merged manually at 2c215c696a

Hint

You were too fast!

Tick the box to put this pull request back in the merge queue (same as @mergifyio queue).

  • Requeue this pull request
<!--- DO NOT EDIT -*- Mergify Payload -*- {"version": 1, "state": "dequeued", "queue_rule_name": "default", "queued_at": "2026-07-08T18:36:52.307653+00:00", "estimated_time_of_merge": null, "speculative_check_pr": null, "required_conditions": []} -*- Mergify Payload End -*- --> # Merge Queue Status - ✅ **Entered queue** — `2026-07-08 18:36 UTC` · Rule: `default` · triggered by merge protections - 🚫 **Left the queue** — `2026-07-08 18:41 UTC` · at `2b958404bb3be6169b7b5f7d82f50168af9fb565` This pull request spent **4 minutes 41 seconds** in the queue, with no time running CI. ## Reason Pull request #454 has been merged manually at *2c215c696a53f02d8f111cf89a5474ca97b6a277* ## Hint You were too fast! Tick the box to put this pull request back in the merge queue (same as `@mergifyio queue`). - [ ] Requeue this pull request <!-- mergify:queue-control:requeue -->
Sign in to join this conversation.