ci: capture thread-dump + service state on an E2E wedge to prove the root cause (#404) #406

Merged
JMR-dev merged 3 commits from ci-404-wedge-diagnostics into main 2026-07-07 16:55:11 +00:00
3 Commits
Author SHA1 Message Date
JMR-dev 90dfb189e6 fix(ci): drop matrix E2E wedge-capture that hangs all 8 legs
The `timeout -k 30s` wrapper + `capture_wedge()` added in 2f32657 for the
matrix `e2e` job (API 29-36) reproducibly wedges every leg, while the
manually-provisioned API 37 preview shard running the identical capture
logic passes. Revert the two matrix "Run E2E tests" steps' `script:`
blocks to main's plain script (just the backgrounded logcat stream +
`./gradlew connectedDebugAndroidTest`) and drop the now-dead "Upload wedge
diagnostics" step from the matrix job.

Kept untouched: the job-level `timeout-minutes: 50` backstop added in
27ede55, and the entire `e2e-preview` job (its own capture_wedge/watchdog
and wedge-diagnostics-api37-preview-shard* upload are unaffected).
2026-07-07 11:29:41 -05:00
JMR-dev 27ede55dbd ci(e2e): add job-level timeout to matrix E2E job (#404)
The matrix E2E job (api-level 29-36) had no timeout-minutes, so a wedge
hangs until GitHub's 6-hour default instead of being force-killed. The
sibling e2e-preview job already sets timeout-minutes: 35. A normal
matrix run is ~15-20 min and a retry-inclusive run ~40 min, so set
timeout-minutes: 50 to give headroom above the in-step wedge-capture
timeout (1200s) while still bounding worst-case runtime.
2026-07-07 07:46:33 -05:00
JMR-devandClaude Opus 4.8 2f32657aff ci: capture thread-dump + service state on an E2E wedge to prove the root cause (#404)
E2E legs intermittently WEDGE (hang) with no fast-fail until the job force-kill,
and GitHub's post-force-kill step behavior is unreliable, so #388's diagnostics
don't reliably capture the wedge — and don't capture wedge-specific state anyway.

Wrap the `connectedDebugAndroidTest` run (both the `e2e` matrix first-attempt +
retry, and each `e2e-preview` shard) in an explicit `timeout -k 30s 1200`
(20 min) — comfortably above a normal run (~13-15 min), well below the hard cap —
so a wedge trips the wrapper (exit 124), NOT the force-kill, GUARANTEEING the
capture runs while the emulator is still alive. On 124, capture_wedge grabs the
smoking gun into a `wedge-diagnostics-api<level>` artifact: the running/last test
(logcat TestRunner), SIGQUIT (kill -3) thread dumps of the app + instrumentation
processes (ART -> logcat + /data/anr), dumpsys activity/window, `service list` +
`service check input/window/activity` (the boot-race crux), sys.boot_completed +
init.svc.* state, the snapshot cache-hit note, and accel/kvm/mem/disk. Then it
exits with the real status so #388's diagnostics + the existing retry still fire;
a normal run finishes before the wrapper and is unaffected.

EVIDENCE ONLY — no boot-readiness guard/fix (maintainer: prove the cause first).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-07 07:46:33 -05:00