ci: capture thread-dump + service state on an E2E wedge to prove the root cause (#404) #406

Merged
JMR-dev merged 3 commits from ci-404-wedge-diagnostics into main 2026-07-07 16:55:11 +00:00
JMR-dev commented 2026-07-07 02:32:27 +00:00 (Migrated from github.com)

What & why

E2E legs intermittently wedge (hang) with no fast-fail until the job force-kill (seen API-29 #374, API-34 #396). GitHub's post-force-kill step behavior is unreliable, so #388's if: failure()/if: always() diagnostics don't reliably capture a wedge — and they don't capture wedge-specific state anyway. This PR captures definitive, reliable wedge diagnostics to PROVE the root cause (leading hypothesis: snapshot-restore boot race — sys.boot_completed=1 before system_server republishes binder services).

EVIDENCE ONLY — no boot-readiness guard / fix (per the maintainer: prove the cause first). Closes #404.

Mechanism

Wrap the connectedDebugAndroidTest invocation in an explicit timeout -k 30s 1200 (20 min) — comfortably above a normal run (~13-15 min), well below the hard cap — in all run paths: the e2e matrix first-attempt and retry steps, and each e2e-preview shard (attempt 1 + its retry).

  • normal run (<20 min) → finishes before the wrapper; timeout passes gradle's real exit code through; unaffected.
  • wedge (>20 min) → timeout trips (exit 124) instead of the force-kill, so the capture is guaranteed to run while the emulator is still alive (in the e2e matrix the runner action tears the emulator down the instant its script returns, so a post-step can't probe it — the capture must be inline).

On exit 124, capture_wedge writes the smoking gun (each probe guarded || true, appended so a wedge on attempt 1 survives even if the retry passes):

  • running/last instrumented test (from the streamed logcat TestRunner tail)
  • SIGQUIT (kill -3) thread dumps of org.libremail.app + org.libremail.app.test (pidof -> kill -3 -> ART dumps threads to logcat + /data/anr/), then cat /data/anr/* + logcat -d
  • dumpsys activity, dumpsys window
  • service list + service check input/window/activity (the boot-race crux — are binder services published?)
  • getprop sys.boot_completed + getprop | grep init.svc
  • whether "Create AVD" was a snapshot cache-hit (e2e matrix; N/A for the cold-booting preview)
  • emulator accel / kvm / mem / disk (parity with #388)

Then it exits with the real status, so #388's diagnostics and the existing retry still fire unchanged. The || status=$? idiom (not if timeout; then) makes the exit code survive regardless of set -e, verified for success / normal-fail / wedge both with and without set -e.

Artifacts

  • e2e matrix: wedge-diagnostics-api<level> (if-no-files-found: ignore — no wedge, no file, so healthy runs stay quiet)
  • e2e-preview: wedge-diagnostics-api37-preview-shard<n> (shard-unique)

Separate from #388's e2e-diagnostics / boot-diagnostics artifacts; same pinned upload-artifact SHA.

Validation

actionlint clean; YAML parses; all three script bodies pass bash -n; the exit-code idiom behaviorally verified for success / normal-fail / wedge (124), under and without set -e. A real wedge is a flake and can't be forced — validated by construction + syntax.

Coordination

Edits are confined to the E2E test-run step; #402's E2E path-filter edits are in the job if/needs + ci-passed gate — non-conflicting. Sequence the ci.yml merges (#402 then this); auto-merge intentionally not armed.

🤖 Generated with Claude Code

## What & why E2E legs intermittently **wedge** (hang) with no fast-fail until the job force-kill (seen API-29 #374, API-34 #396). GitHub's post-force-kill step behavior is unreliable, so #388's `if: failure()`/`if: always()` diagnostics don't reliably capture a wedge — and they don't capture wedge-specific state anyway. This PR captures **definitive, reliable** wedge diagnostics to PROVE the root cause (leading hypothesis: snapshot-restore boot race — `sys.boot_completed=1` before `system_server` republishes binder services). **EVIDENCE ONLY — no boot-readiness guard / fix** (per the maintainer: prove the cause first). Closes #404. ## Mechanism Wrap the `connectedDebugAndroidTest` invocation in an explicit `timeout -k 30s 1200` (20 min) — comfortably above a normal run (~13-15 min), well below the hard cap — in **all** run paths: the `e2e` matrix first-attempt **and** retry steps, and each `e2e-preview` shard (attempt 1 + its retry). - **normal run (<20 min)** → finishes before the wrapper; `timeout` passes gradle's real exit code through; unaffected. - **wedge (>20 min)** → `timeout` trips (**exit 124**) instead of the force-kill, so the capture is **guaranteed to run while the emulator is still alive** (in the `e2e` matrix the runner action tears the emulator down the instant its `script` returns, so a post-step can't probe it — the capture must be inline). On exit 124, `capture_wedge` writes the smoking gun (each probe guarded `|| true`, appended so a wedge on attempt 1 survives even if the retry passes): - running/last instrumented test (from the streamed logcat `TestRunner` tail) - **SIGQUIT (`kill -3`) thread dumps** of `org.libremail.app` + `org.libremail.app.test` (`pidof` -> `kill -3` -> ART dumps threads to logcat + `/data/anr/`), then `cat /data/anr/*` + `logcat -d` - `dumpsys activity`, `dumpsys window` - `service list` + `service check input`/`window`/`activity` (**the boot-race crux** — are binder services published?) - `getprop sys.boot_completed` + `getprop | grep init.svc` - whether "Create AVD" was a snapshot cache-hit (`e2e` matrix; N/A for the cold-booting preview) - emulator accel / kvm / mem / disk (parity with #388) Then it `exit`s with the **real** status, so **#388's diagnostics and the existing retry still fire** unchanged. The `|| status=$?` idiom (not `if timeout; then`) makes the exit code survive regardless of `set -e`, verified for success / normal-fail / wedge both with and without `set -e`. ## Artifacts - `e2e` matrix: **`wedge-diagnostics-api<level>`** (`if-no-files-found: ignore` — no wedge, no file, so healthy runs stay quiet) - `e2e-preview`: **`wedge-diagnostics-api37-preview-shard<n>`** (shard-unique) Separate from #388's `e2e-diagnostics` / boot-diagnostics artifacts; same pinned `upload-artifact` SHA. ## Validation `actionlint` clean; YAML parses; all three script bodies pass `bash -n`; the exit-code idiom behaviorally verified for success / normal-fail / wedge (124), under and without `set -e`. A real wedge is a flake and can't be forced — validated by construction + syntax. ## Coordination Edits are confined to the E2E test-run **step**; #402's E2E path-filter edits are in the job `if`/`needs` + `ci-passed` gate — non-conflicting. Sequence the `ci.yml` merges (#402 then this); auto-merge intentionally **not** armed. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
JMR-dev commented 2026-07-07 07:26:54 +00:00 (Migrated from github.com)

⚠️ Overnight finding (2026-07-07) — do NOT merge as-is; needs rework. Auto-merge disarmed.

This PR's own CI run showed all 8 reactivecircus matrix E2E legs (API 29–36) hung in_progress at 37+ min, while the API-37 preview shards (hand-rolled cold-boot, not reactivecircus) PASSED. Earlier batch PRs (without this wrapper) ran the matrix legs fine. → Strong signal that this PR's timeout -k 30s 1200 wrapper around the E2E gradle step breaks the reactivecircus android-emulator-runner (which owns the emulator lifecycle + snapshot resume).

Two concrete problems:

  1. The wrapper's 20-min timeout never fired (legs stuck at 37 > 20 min) → the hang is in the emulator boot/setup phase, before the wrapped gradle step — exactly where the snapshot-resume race lives (see the input keyevent 82 note already in ci.yml). A gradle-step timeout wrapper structurally can't catch a boot-phase wedge.
  2. The matrix E2E job has no timeout-minutes (defaults to 6h), so a hung leg blocks 8 runners for hours. (The API-37 preview job does set timeout-minutes: 35 — that's why only the matrix hangs.)

Rework direction:

  • Add a job-level timeout-minutes to the matrix E2E job so a wedge auto-fails (and the if: failure() #388 diagnostics + a capture step actually run) instead of hanging 6h.
  • Capture via a background watchdog (spawn a step that sleeps N min then dumps kill -3/dumpsys/service list and getprops) rather than wrapping the gradle step in timeout — the watchdog runs alongside the emulator-runner instead of fighting it, and covers the boot phase.

Actions taken overnight: cancelled the hung run (was blocking runners; no if: failure() diagnostics are recoverable from a hung-then-cancelled step anyway), and disarmed auto-merge. Relates #404, #388, ci-e2e-infra-flakes.

## ⚠️ Overnight finding (2026-07-07) — do NOT merge as-is; needs rework. Auto-merge disarmed. This PR's **own** CI run showed all **8 reactivecircus matrix E2E legs (API 29–36) hung `in_progress` at 37+ min**, while the **API-37 preview shards** (hand-rolled cold-boot, *not* reactivecircus) **PASSED**. Earlier batch PRs (without this wrapper) ran the matrix legs fine. → Strong signal that **this PR's `timeout -k 30s 1200` wrapper around the E2E gradle step breaks the reactivecircus `android-emulator-runner`** (which owns the emulator lifecycle + snapshot resume). **Two concrete problems:** 1. **The wrapper's 20-min timeout never fired** (legs stuck at 37 > 20 min) → the hang is in the emulator **boot/setup** phase, *before* the wrapped gradle step — exactly where the snapshot-resume race lives (see the `input keyevent 82` note already in `ci.yml`). A gradle-step `timeout` wrapper structurally can't catch a boot-phase wedge. 2. The **matrix E2E job has no `timeout-minutes`** (defaults to 6h), so a hung leg blocks 8 runners for hours. (The API-37 preview job *does* set `timeout-minutes: 35` — that's why only the matrix hangs.) **Rework direction:** - Add a **job-level `timeout-minutes`** to the matrix E2E job so a wedge **auto-fails** (and the `if: failure()` #388 diagnostics + a capture step actually run) instead of hanging 6h. - Capture via a **background watchdog** (spawn a step that sleeps N min then dumps `kill -3`/`dumpsys`/`service list` and getprops) rather than wrapping the gradle step in `timeout` — the watchdog runs alongside the emulator-runner instead of fighting it, and covers the boot phase. **Actions taken overnight:** cancelled the hung run (was blocking runners; no `if: failure()` diagnostics are recoverable from a hung-then-cancelled step anyway), and disarmed auto-merge. Relates #404, #388, [[ci-e2e-infra-flakes]].
mergify[bot] commented 2026-07-07 16:45:27 +00:00 (Migrated from github.com)

Tick the box to add this pull request to the merge queue (same as @mergifyio queue).

  • Queue this pull request
Tick the box to add this pull request to the merge queue (same as `@mergifyio queue`). - [ ] Queue this pull request <!-- mergify:queue-control:queue -->
Sign in to join this conversation.