ci(e2e): extend the #334 API-37 boot/E2E diagnostics to all emulator E2E jobs (API 29-36 matrix) #387

Closed
opened 2026-07-06 20:40:49 +00:00 by JMR-dev · 0 comments
JMR-dev commented 2026-07-06 20:40:49 +00:00 (Migrated from github.com)

Problem

Emulator E2E flakes happen across all the E2E jobs, but only e2e-preview (API 37) can explain itself. The #334 boot-diagnostics work instrumented e2e-preview only (ci.yml ~L367-466): it captures the emulator boot log (emulator -verbose -debug init,avd_config,kernel), streams logcat to a file, dumps full system state on a boot timeout (emulator -accel-check, /dev/kvm, GPU mode, free -h, df -h, AVD config.ini, emulator.log tail), dumps logcat -d + emulator.log on failure, and uploads all of it as an artifact if: always().

The e2e matrix (API 29-36) uses reactivecircus/android-emulator-runner and has none of this — it only uploads the test report. When it flakes there is nothing to say why.

Concrete example: run 28820763671, job 85471668352 = E2E (31) → sdkmanager failed with exit code 1, zero diagnostic trail (and not even emulator-specific).

Goal

Every CI emulator E2E job captures enough diagnostics that a failure/flake is never a mystery — "there is never a question of why a failure/flake happened."

Bring the #334 diagnostic pattern to the e2e matrix:

  • Stream logcat -v time to a file during the test run (per API level).
  • On failure: dump logcat -d + emulator/system state (accel-check, /dev/kvm, free -h, df -h) to the step log.
  • if: always() upload a per-API e2e-diagnostics-api<level> artifact (logcat + captured state).
  • Make transient SDK-install / emulator-boot failures diagnosable too (capture the sdkmanager/action output; the E2E (31) sdkmanager exit-1 above should leave a trail; a bounded retry of a transient SDK install is in scope if clean).

Constraints

  • Do not modify the e2e-preview job — PR #372 is restructuring it (sharding) and it already has these diagnostics. Base on main, touch the e2e matrix only, to avoid colliding with #372.
  • Keep reactivecircus/android-emulator-runner and the existing boot-race retry; ADD diagnostics via the step script: + if: failure()/if: always() steps around it (do not rip it out / convert to manual boot).
  • A shared composite action to DRY the diagnostics is fine ONLY if it does not force edits to e2e-preview (else it collides with #372) — otherwise keep it inline in the matrix.

Acceptance

Each e2e matrix leg uploads a per-API diagnostics artifact on failure/always, and a forced failure shows the state dump in the step log. Validated by the PR's own CI run (the matrix must still pass green).

Relates to #334 (the api37 diagnostics template), #372 (sharding — do not collide).

## Problem Emulator E2E flakes happen across **all** the E2E jobs, but only `e2e-preview` (API 37) can explain itself. The #334 boot-diagnostics work instrumented `e2e-preview` only (ci.yml ~L367-466): it captures the emulator boot log (`emulator -verbose -debug init,avd_config,kernel`), streams `logcat` to a file, dumps full system state on a boot timeout (`emulator -accel-check`, `/dev/kvm`, GPU mode, `free -h`, `df -h`, AVD `config.ini`, `emulator.log` tail), dumps `logcat -d` + `emulator.log` on failure, and uploads all of it as an artifact `if: always()`. The **`e2e` matrix (API 29-36)** uses `reactivecircus/android-emulator-runner` and has **none** of this — it only uploads the test report. When it flakes there is nothing to say why. Concrete example: run 28820763671, job 85471668352 = `E2E (31)` → `sdkmanager failed with exit code 1`, zero diagnostic trail (and not even emulator-specific). ## Goal Every CI emulator E2E job captures enough diagnostics that a failure/flake is **never a mystery** — "there is never a question of why a failure/flake happened." Bring the #334 diagnostic pattern to the `e2e` matrix: - Stream `logcat -v time` to a file during the test run (per API level). - On failure: dump `logcat -d` + emulator/system state (accel-check, `/dev/kvm`, `free -h`, `df -h`) to the step log. - `if: always()` upload a per-API `e2e-diagnostics-api<level>` artifact (logcat + captured state). - Make transient SDK-install / emulator-boot failures diagnosable too (capture the `sdkmanager`/action output; the `E2E (31)` sdkmanager exit-1 above should leave a trail; a bounded retry of a transient SDK install is in scope if clean). ## Constraints - Do **not** modify the `e2e-preview` job — PR #372 is restructuring it (sharding) and it already has these diagnostics. Base on `main`, touch the `e2e` matrix only, to avoid colliding with #372. - Keep `reactivecircus/android-emulator-runner` and the existing boot-race retry; ADD diagnostics via the step `script:` + `if: failure()`/`if: always()` steps around it (do not rip it out / convert to manual boot). - A shared composite action to DRY the diagnostics is fine ONLY if it does not force edits to `e2e-preview` (else it collides with #372) — otherwise keep it inline in the matrix. ## Acceptance Each `e2e` matrix leg uploads a per-API diagnostics artifact on failure/always, and a forced failure shows the state dump in the step log. Validated by the PR's own CI run (the matrix must still pass green). Relates to #334 (the api37 diagnostics template), #372 (sharding — do not collide).
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: JMR-dev/LibreMail#387