E2E legs intermittently wedge (hang) with no fast-fail until the job force-kill (seen API-29 #374, API-34 #396). GitHub's post-force-kill step behavior is unreliable, so #388's if: failure()/if: always() diagnostics don't reliably capture a wedge — and they don't capture wedge-specific state anyway. This PR captures definitive, reliable wedge diagnostics to PROVE the root cause (leading hypothesis: snapshot-restore boot race — sys.boot_completed=1 before system_server republishes binder services).
EVIDENCE ONLY — no boot-readiness guard / fix (per the maintainer: prove the cause first). Closes#404.
Mechanism
Wrap the connectedDebugAndroidTest invocation in an explicit timeout -k 30s 1200 (20 min) — comfortably above a normal run (~13-15 min), well below the hard cap — in all run paths: the e2e matrix first-attempt and retry steps, and each e2e-preview shard (attempt 1 + its retry).
normal run (<20 min) → finishes before the wrapper; timeout passes gradle's real exit code through; unaffected.
wedge (>20 min) → timeout trips (exit 124) instead of the force-kill, so the capture is guaranteed to run while the emulator is still alive (in the e2e matrix the runner action tears the emulator down the instant its script returns, so a post-step can't probe it — the capture must be inline).
On exit 124, capture_wedge writes the smoking gun (each probe guarded || true, appended so a wedge on attempt 1 survives even if the retry passes):
running/last instrumented test (from the streamed logcat TestRunner tail)
SIGQUIT (kill -3) thread dumps of org.libremail.app + org.libremail.app.test (pidof -> kill -3 -> ART dumps threads to logcat + /data/anr/), then cat /data/anr/* + logcat -d
dumpsys activity, dumpsys window
service list + service check input/window/activity (the boot-race crux — are binder services published?)
whether "Create AVD" was a snapshot cache-hit (e2e matrix; N/A for the cold-booting preview)
emulator accel / kvm / mem / disk (parity with #388)
Then it exits with the real status, so #388's diagnostics and the existing retry still fire unchanged. The || status=$? idiom (not if timeout; then) makes the exit code survive regardless of set -e, verified for success / normal-fail / wedge both with and without set -e.
Artifacts
e2e matrix: wedge-diagnostics-api<level> (if-no-files-found: ignore — no wedge, no file, so healthy runs stay quiet)
Separate from #388's e2e-diagnostics / boot-diagnostics artifacts; same pinned upload-artifact SHA.
Validation
actionlint clean; YAML parses; all three script bodies pass bash -n; the exit-code idiom behaviorally verified for success / normal-fail / wedge (124), under and without set -e. A real wedge is a flake and can't be forced — validated by construction + syntax.
Coordination
Edits are confined to the E2E test-run step; #402's E2E path-filter edits are in the job if/needs + ci-passed gate — non-conflicting. Sequence the ci.yml merges (#402 then this); auto-merge intentionally not armed.
## What & why
E2E legs intermittently **wedge** (hang) with no fast-fail until the job force-kill (seen API-29 #374, API-34 #396). GitHub's post-force-kill step behavior is unreliable, so #388's `if: failure()`/`if: always()` diagnostics don't reliably capture a wedge — and they don't capture wedge-specific state anyway. This PR captures **definitive, reliable** wedge diagnostics to PROVE the root cause (leading hypothesis: snapshot-restore boot race — `sys.boot_completed=1` before `system_server` republishes binder services).
**EVIDENCE ONLY — no boot-readiness guard / fix** (per the maintainer: prove the cause first). Closes #404.
## Mechanism
Wrap the `connectedDebugAndroidTest` invocation in an explicit `timeout -k 30s 1200` (20 min) — comfortably above a normal run (~13-15 min), well below the hard cap — in **all** run paths: the `e2e` matrix first-attempt **and** retry steps, and each `e2e-preview` shard (attempt 1 + its retry).
- **normal run (<20 min)** → finishes before the wrapper; `timeout` passes gradle's real exit code through; unaffected.
- **wedge (>20 min)** → `timeout` trips (**exit 124**) instead of the force-kill, so the capture is **guaranteed to run while the emulator is still alive** (in the `e2e` matrix the runner action tears the emulator down the instant its `script` returns, so a post-step can't probe it — the capture must be inline).
On exit 124, `capture_wedge` writes the smoking gun (each probe guarded `|| true`, appended so a wedge on attempt 1 survives even if the retry passes):
- running/last instrumented test (from the streamed logcat `TestRunner` tail)
- **SIGQUIT (`kill -3`) thread dumps** of `org.libremail.app` + `org.libremail.app.test` (`pidof` -> `kill -3` -> ART dumps threads to logcat + `/data/anr/`), then `cat /data/anr/*` + `logcat -d`
- `dumpsys activity`, `dumpsys window`
- `service list` + `service check input`/`window`/`activity` (**the boot-race crux** — are binder services published?)
- `getprop sys.boot_completed` + `getprop | grep init.svc`
- whether "Create AVD" was a snapshot cache-hit (`e2e` matrix; N/A for the cold-booting preview)
- emulator accel / kvm / mem / disk (parity with #388)
Then it `exit`s with the **real** status, so **#388's diagnostics and the existing retry still fire** unchanged. The `|| status=$?` idiom (not `if timeout; then`) makes the exit code survive regardless of `set -e`, verified for success / normal-fail / wedge both with and without `set -e`.
## Artifacts
- `e2e` matrix: **`wedge-diagnostics-api<level>`** (`if-no-files-found: ignore` — no wedge, no file, so healthy runs stay quiet)
- `e2e-preview`: **`wedge-diagnostics-api37-preview-shard<n>`** (shard-unique)
Separate from #388's `e2e-diagnostics` / boot-diagnostics artifacts; same pinned `upload-artifact` SHA.
## Validation
`actionlint` clean; YAML parses; all three script bodies pass `bash -n`; the exit-code idiom behaviorally verified for success / normal-fail / wedge (124), under and without `set -e`. A real wedge is a flake and can't be forced — validated by construction + syntax.
## Coordination
Edits are confined to the E2E test-run **step**; #402's E2E path-filter edits are in the job `if`/`needs` + `ci-passed` gate — non-conflicting. Sequence the `ci.yml` merges (#402 then this); auto-merge intentionally **not** armed.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
⚠️ Overnight finding (2026-07-07) — do NOT merge as-is; needs rework. Auto-merge disarmed.
This PR's own CI run showed all 8 reactivecircus matrix E2E legs (API 29–36) hung in_progress at 37+ min, while the API-37 preview shards (hand-rolled cold-boot, not reactivecircus) PASSED. Earlier batch PRs (without this wrapper) ran the matrix legs fine. → Strong signal that this PR's timeout -k 30s 1200 wrapper around the E2E gradle step breaks the reactivecircus android-emulator-runner (which owns the emulator lifecycle + snapshot resume).
Two concrete problems:
The wrapper's 20-min timeout never fired (legs stuck at 37 > 20 min) → the hang is in the emulator boot/setup phase, before the wrapped gradle step — exactly where the snapshot-resume race lives (see the input keyevent 82 note already in ci.yml). A gradle-step timeout wrapper structurally can't catch a boot-phase wedge.
The matrix E2E job has no timeout-minutes (defaults to 6h), so a hung leg blocks 8 runners for hours. (The API-37 preview job does set timeout-minutes: 35 — that's why only the matrix hangs.)
Rework direction:
Add a job-level timeout-minutes to the matrix E2E job so a wedge auto-fails (and the if: failure()#388 diagnostics + a capture step actually run) instead of hanging 6h.
Capture via a background watchdog (spawn a step that sleeps N min then dumps kill -3/dumpsys/service list and getprops) rather than wrapping the gradle step in timeout — the watchdog runs alongside the emulator-runner instead of fighting it, and covers the boot phase.
Actions taken overnight: cancelled the hung run (was blocking runners; no if: failure() diagnostics are recoverable from a hung-then-cancelled step anyway), and disarmed auto-merge. Relates #404, #388, ci-e2e-infra-flakes.
## ⚠️ Overnight finding (2026-07-07) — do NOT merge as-is; needs rework. Auto-merge disarmed.
This PR's **own** CI run showed all **8 reactivecircus matrix E2E legs (API 29–36) hung `in_progress` at 37+ min**, while the **API-37 preview shards** (hand-rolled cold-boot, *not* reactivecircus) **PASSED**. Earlier batch PRs (without this wrapper) ran the matrix legs fine. → Strong signal that **this PR's `timeout -k 30s 1200` wrapper around the E2E gradle step breaks the reactivecircus `android-emulator-runner`** (which owns the emulator lifecycle + snapshot resume).
**Two concrete problems:**
1. **The wrapper's 20-min timeout never fired** (legs stuck at 37 > 20 min) → the hang is in the emulator **boot/setup** phase, *before* the wrapped gradle step — exactly where the snapshot-resume race lives (see the `input keyevent 82` note already in `ci.yml`). A gradle-step `timeout` wrapper structurally can't catch a boot-phase wedge.
2. The **matrix E2E job has no `timeout-minutes`** (defaults to 6h), so a hung leg blocks 8 runners for hours. (The API-37 preview job *does* set `timeout-minutes: 35` — that's why only the matrix hangs.)
**Rework direction:**
- Add a **job-level `timeout-minutes`** to the matrix E2E job so a wedge **auto-fails** (and the `if: failure()` #388 diagnostics + a capture step actually run) instead of hanging 6h.
- Capture via a **background watchdog** (spawn a step that sleeps N min then dumps `kill -3`/`dumpsys`/`service list` and getprops) rather than wrapping the gradle step in `timeout` — the watchdog runs alongside the emulator-runner instead of fighting it, and covers the boot phase.
**Actions taken overnight:** cancelled the hung run (was blocking runners; no `if: failure()` diagnostics are recoverable from a hung-then-cancelled step anyway), and disarmed auto-merge. Relates #404, #388, [[ci-e2e-infra-flakes]].
Tick the box to add this pull request to the merge queue (same as @mergifyio queue).
Queue this pull request
Tick the box to add this pull request to the merge queue (same as `@mergifyio queue`).
- [ ] Queue this pull request <!-- mergify:queue-control:queue -->
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
What & why
E2E legs intermittently wedge (hang) with no fast-fail until the job force-kill (seen API-29 #374, API-34 #396). GitHub's post-force-kill step behavior is unreliable, so #388's
if: failure()/if: always()diagnostics don't reliably capture a wedge — and they don't capture wedge-specific state anyway. This PR captures definitive, reliable wedge diagnostics to PROVE the root cause (leading hypothesis: snapshot-restore boot race —sys.boot_completed=1beforesystem_serverrepublishes binder services).EVIDENCE ONLY — no boot-readiness guard / fix (per the maintainer: prove the cause first). Closes #404.
Mechanism
Wrap the
connectedDebugAndroidTestinvocation in an explicittimeout -k 30s 1200(20 min) — comfortably above a normal run (~13-15 min), well below the hard cap — in all run paths: thee2ematrix first-attempt and retry steps, and eache2e-previewshard (attempt 1 + its retry).timeoutpasses gradle's real exit code through; unaffected.timeouttrips (exit 124) instead of the force-kill, so the capture is guaranteed to run while the emulator is still alive (in thee2ematrix the runner action tears the emulator down the instant itsscriptreturns, so a post-step can't probe it — the capture must be inline).On exit 124,
capture_wedgewrites the smoking gun (each probe guarded|| true, appended so a wedge on attempt 1 survives even if the retry passes):TestRunnertail)kill -3) thread dumps oforg.libremail.app+org.libremail.app.test(pidof->kill -3-> ART dumps threads to logcat +/data/anr/), thencat /data/anr/*+logcat -ddumpsys activity,dumpsys windowservice list+service check input/window/activity(the boot-race crux — are binder services published?)getprop sys.boot_completed+getprop | grep init.svce2ematrix; N/A for the cold-booting preview)Then it
exits with the real status, so #388's diagnostics and the existing retry still fire unchanged. The|| status=$?idiom (notif timeout; then) makes the exit code survive regardless ofset -e, verified for success / normal-fail / wedge both with and withoutset -e.Artifacts
e2ematrix:wedge-diagnostics-api<level>(if-no-files-found: ignore— no wedge, no file, so healthy runs stay quiet)e2e-preview:wedge-diagnostics-api37-preview-shard<n>(shard-unique)Separate from #388's
e2e-diagnostics/ boot-diagnostics artifacts; same pinnedupload-artifactSHA.Validation
actionlintclean; YAML parses; all three script bodies passbash -n; the exit-code idiom behaviorally verified for success / normal-fail / wedge (124), under and withoutset -e. A real wedge is a flake and can't be forced — validated by construction + syntax.Coordination
Edits are confined to the E2E test-run step; #402's E2E path-filter edits are in the job
if/needs+ci-passedgate — non-conflicting. Sequence theci.ymlmerges (#402 then this); auto-merge intentionally not armed.🤖 Generated with Claude Code
⚠️ Overnight finding (2026-07-07) — do NOT merge as-is; needs rework. Auto-merge disarmed.
This PR's own CI run showed all 8 reactivecircus matrix E2E legs (API 29–36) hung
in_progressat 37+ min, while the API-37 preview shards (hand-rolled cold-boot, not reactivecircus) PASSED. Earlier batch PRs (without this wrapper) ran the matrix legs fine. → Strong signal that this PR'stimeout -k 30s 1200wrapper around the E2E gradle step breaks the reactivecircusandroid-emulator-runner(which owns the emulator lifecycle + snapshot resume).Two concrete problems:
input keyevent 82note already inci.yml). A gradle-steptimeoutwrapper structurally can't catch a boot-phase wedge.timeout-minutes(defaults to 6h), so a hung leg blocks 8 runners for hours. (The API-37 preview job does settimeout-minutes: 35— that's why only the matrix hangs.)Rework direction:
timeout-minutesto the matrix E2E job so a wedge auto-fails (and theif: failure()#388 diagnostics + a capture step actually run) instead of hanging 6h.kill -3/dumpsys/service listand getprops) rather than wrapping the gradle step intimeout— the watchdog runs alongside the emulator-runner instead of fighting it, and covers the boot phase.Actions taken overnight: cancelled the hung run (was blocking runners; no
if: failure()diagnostics are recoverable from a hung-then-cancelled step anyway), and disarmed auto-merge. Relates #404, #388, ci-e2e-infra-flakes.Tick the box to add this pull request to the merge queue (same as
@mergifyio queue).