ci: capture definitive diagnostics on an E2E emulator wedge (prove the root cause before fixing) #404

Closed
opened 2026-07-07 02:07:35 +00:00 by JMR-dev · 1 comment
JMR-dev commented 2026-07-07 02:07:35 +00:00 (Migrated from github.com)

Context

E2E matrix legs intermittently wedge — a leg hangs with no fast-fail until the 35-min job timeout, blocking the PR (seen API-29 #374, API-34 #396). Leading hypothesis: snapshot-restore boot race (sys.boot_completed=1 restored before system_server republishes binder services -> a test/activity blocks on an unavailable service -> connectedDebugAndroidTest hangs, no per-test timeout). It is UNPROVEN — in-progress logs aren't fetchable, and cancelling the wedge (to unstick the cascade) destroys the evidence. #388 added general E2E diagnostics but they aren't reliably captured on a wedge (post-timeout step behavior is flaky) and don't capture wedge-specific state.

Goal

Capture definitive, reliable diagnostics of a wedge to PROVE the root cause — before any fix. (Maintainer: prove first; the boot-readiness guard is on hold.)

Implement (extends #388; e2e matrix + e2e-preview)

  • Wrap connectedDebugAndroidTest in an explicit wrapper timeout (e.g. timeout 1200) shorter than the 35-min hard cap, so a wedge trips the wrapper (not the force-kill) and the diagnostic dump is GUARANTEED to run + upload.
  • On the wrapper-timeout (wedge), capture the smoking gun: the running/last test, a thread/ANR dump of the app + instrumentation processes (kill -3 / dumpsys), dumpsys activity + dumpsys window, service list + service check input window activity, getprop sys.boot_completed + getprop | grep init.svc, and whether "Create AVD" was a snapshot cache-hit. Upload as an artifact.
  • Then fail the step so the existing retry still runs. The point is EVIDENCE, not a fix.

Acceptance

The next wedge yields an artifact that shows WHY the leg hung — e.g. a thread blocked on an unpublished binder service (proves the boot race) vs. a genuinely hung test / KVM starvation (different cause). Do NOT build the boot-readiness guard.

Relates: #388 (diagnostics), #399/#402 (concurrent ci.yml — coordinate the merge), the wedges on #374/#396.

## Context E2E matrix legs intermittently **wedge** — a leg hangs with no fast-fail until the 35-min job timeout, blocking the PR (seen API-29 #374, API-34 #396). Leading **hypothesis**: snapshot-restore boot race (`sys.boot_completed=1` restored before `system_server` republishes binder services -> a test/activity blocks on an unavailable service -> `connectedDebugAndroidTest` hangs, no per-test timeout). It is **UNPROVEN** — in-progress logs aren't fetchable, and cancelling the wedge (to unstick the cascade) destroys the evidence. #388 added general E2E diagnostics but they aren't reliably captured on a wedge (post-timeout step behavior is flaky) and don't capture wedge-specific state. ## Goal Capture **definitive, reliable** diagnostics of a wedge to PROVE the root cause — before any fix. (Maintainer: prove first; the boot-readiness guard is on hold.) ## Implement (extends #388; `e2e` matrix + `e2e-preview`) - Wrap `connectedDebugAndroidTest` in an explicit **wrapper timeout** (e.g. `timeout 1200`) shorter than the 35-min hard cap, so a wedge trips the wrapper (not the force-kill) and the diagnostic dump is GUARANTEED to run + upload. - On the wrapper-timeout (wedge), capture the smoking gun: the running/last test, a **thread/ANR dump** of the app + instrumentation processes (`kill -3` / `dumpsys`), `dumpsys activity` + `dumpsys window`, `service list` + `service check input window activity`, `getprop sys.boot_completed` + `getprop | grep init.svc`, and whether "Create AVD" was a snapshot cache-hit. Upload as an artifact. - Then fail the step so the existing retry still runs. The point is EVIDENCE, not a fix. ## Acceptance The next wedge yields an artifact that shows WHY the leg hung — e.g. a thread blocked on an unpublished binder service (proves the boot race) vs. a genuinely hung test / KVM starvation (different cause). **Do NOT build the boot-readiness guard.** Relates: #388 (diagnostics), #399/#402 (concurrent ci.yml — coordinate the merge), the wedges on #374/#396.
JMR-dev commented 2026-07-07 17:04:39 +00:00 (Migrated from github.com)

Delivered in scoped form via #406 (merged): matrix job gets a job-level timeout-minutes: 50 backstop (bounds any wedge to 50 min vs 6h) + the API-37 preview keeps the full capture_wedge dump. The matrix legs' in-flight capture was dropped because the timeout-wrapper reproducibly hung all 8 reactivecircus legs (confirmed twice). Matrix-leg capture (watchdog or manual-boot convergence) tracked as a follow-up: #421.

Delivered in **scoped** form via #406 (merged): matrix job gets a **job-level `timeout-minutes: 50`** backstop (bounds any wedge to 50 min vs 6h) + the **API-37 preview** keeps the full `capture_wedge` dump. The matrix legs' in-flight capture was **dropped** because the `timeout`-wrapper reproducibly hung all 8 reactivecircus legs (confirmed twice). Matrix-leg capture (watchdog or manual-boot convergence) tracked as a follow-up: #421.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: JMR-dev/LibreMail#404