E2E matrix legs intermittently wedge — a leg hangs with no fast-fail until the 35-min job timeout, blocking the PR (seen API-29 #374, API-34 #396). Leading hypothesis: snapshot-restore boot race (sys.boot_completed=1 restored before system_server republishes binder services -> a test/activity blocks on an unavailable service -> connectedDebugAndroidTest hangs, no per-test timeout). It is UNPROVEN — in-progress logs aren't fetchable, and cancelling the wedge (to unstick the cascade) destroys the evidence. #388 added general E2E diagnostics but they aren't reliably captured on a wedge (post-timeout step behavior is flaky) and don't capture wedge-specific state.
Goal
Capture definitive, reliable diagnostics of a wedge to PROVE the root cause — before any fix. (Maintainer: prove first; the boot-readiness guard is on hold.)
Wrap connectedDebugAndroidTest in an explicit wrapper timeout (e.g. timeout 1200) shorter than the 35-min hard cap, so a wedge trips the wrapper (not the force-kill) and the diagnostic dump is GUARANTEED to run + upload.
On the wrapper-timeout (wedge), capture the smoking gun: the running/last test, a thread/ANR dump of the app + instrumentation processes (kill -3 / dumpsys), dumpsys activity + dumpsys window, service list + service check input window activity, getprop sys.boot_completed + getprop | grep init.svc, and whether "Create AVD" was a snapshot cache-hit. Upload as an artifact.
Then fail the step so the existing retry still runs. The point is EVIDENCE, not a fix.
Acceptance
The next wedge yields an artifact that shows WHY the leg hung — e.g. a thread blocked on an unpublished binder service (proves the boot race) vs. a genuinely hung test / KVM starvation (different cause). Do NOT build the boot-readiness guard.
Relates: #388 (diagnostics), #399/#402 (concurrent ci.yml — coordinate the merge), the wedges on #374/#396.
## Context
E2E matrix legs intermittently **wedge** — a leg hangs with no fast-fail until the 35-min job timeout, blocking the PR (seen API-29 #374, API-34 #396). Leading **hypothesis**: snapshot-restore boot race (`sys.boot_completed=1` restored before `system_server` republishes binder services -> a test/activity blocks on an unavailable service -> `connectedDebugAndroidTest` hangs, no per-test timeout). It is **UNPROVEN** — in-progress logs aren't fetchable, and cancelling the wedge (to unstick the cascade) destroys the evidence. #388 added general E2E diagnostics but they aren't reliably captured on a wedge (post-timeout step behavior is flaky) and don't capture wedge-specific state.
## Goal
Capture **definitive, reliable** diagnostics of a wedge to PROVE the root cause — before any fix. (Maintainer: prove first; the boot-readiness guard is on hold.)
## Implement (extends #388; `e2e` matrix + `e2e-preview`)
- Wrap `connectedDebugAndroidTest` in an explicit **wrapper timeout** (e.g. `timeout 1200`) shorter than the 35-min hard cap, so a wedge trips the wrapper (not the force-kill) and the diagnostic dump is GUARANTEED to run + upload.
- On the wrapper-timeout (wedge), capture the smoking gun: the running/last test, a **thread/ANR dump** of the app + instrumentation processes (`kill -3` / `dumpsys`), `dumpsys activity` + `dumpsys window`, `service list` + `service check input window activity`, `getprop sys.boot_completed` + `getprop | grep init.svc`, and whether "Create AVD" was a snapshot cache-hit. Upload as an artifact.
- Then fail the step so the existing retry still runs. The point is EVIDENCE, not a fix.
## Acceptance
The next wedge yields an artifact that shows WHY the leg hung — e.g. a thread blocked on an unpublished binder service (proves the boot race) vs. a genuinely hung test / KVM starvation (different cause). **Do NOT build the boot-readiness guard.**
Relates: #388 (diagnostics), #399/#402 (concurrent ci.yml — coordinate the merge), the wedges on #374/#396.
Delivered in scoped form via #406 (merged): matrix job gets a job-level timeout-minutes: 50 backstop (bounds any wedge to 50 min vs 6h) + the API-37 preview keeps the full capture_wedge dump. The matrix legs' in-flight capture was dropped because the timeout-wrapper reproducibly hung all 8 reactivecircus legs (confirmed twice). Matrix-leg capture (watchdog or manual-boot convergence) tracked as a follow-up: #421.
Delivered in **scoped** form via #406 (merged): matrix job gets a **job-level `timeout-minutes: 50`** backstop (bounds any wedge to 50 min vs 6h) + the **API-37 preview** keeps the full `capture_wedge` dump. The matrix legs' in-flight capture was **dropped** because the `timeout`-wrapper reproducibly hung all 8 reactivecircus legs (confirmed twice). Matrix-leg capture (watchdog or manual-boot convergence) tracked as a follow-up: #421.
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Context
E2E matrix legs intermittently wedge — a leg hangs with no fast-fail until the 35-min job timeout, blocking the PR (seen API-29 #374, API-34 #396). Leading hypothesis: snapshot-restore boot race (
sys.boot_completed=1restored beforesystem_serverrepublishes binder services -> a test/activity blocks on an unavailable service ->connectedDebugAndroidTesthangs, no per-test timeout). It is UNPROVEN — in-progress logs aren't fetchable, and cancelling the wedge (to unstick the cascade) destroys the evidence. #388 added general E2E diagnostics but they aren't reliably captured on a wedge (post-timeout step behavior is flaky) and don't capture wedge-specific state.Goal
Capture definitive, reliable diagnostics of a wedge to PROVE the root cause — before any fix. (Maintainer: prove first; the boot-readiness guard is on hold.)
Implement (extends #388;
e2ematrix +e2e-preview)connectedDebugAndroidTestin an explicit wrapper timeout (e.g.timeout 1200) shorter than the 35-min hard cap, so a wedge trips the wrapper (not the force-kill) and the diagnostic dump is GUARANTEED to run + upload.kill -3/dumpsys),dumpsys activity+dumpsys window,service list+service check input window activity,getprop sys.boot_completed+getprop | grep init.svc, and whether "Create AVD" was a snapshot cache-hit. Upload as an artifact.Acceptance
The next wedge yields an artifact that shows WHY the leg hung — e.g. a thread blocked on an unpublished binder service (proves the boot race) vs. a genuinely hung test / KVM starvation (different cause). Do NOT build the boot-readiness guard.
Relates: #388 (diagnostics), #399/#402 (concurrent ci.yml — coordinate the merge), the wedges on #374/#396.
Delivered in scoped form via #406 (merged): matrix job gets a job-level
timeout-minutes: 50backstop (bounds any wedge to 50 min vs 6h) + the API-37 preview keeps the fullcapture_wedgedump. The matrix legs' in-flight capture was dropped because thetimeout-wrapper reproducibly hung all 8 reactivecircus legs (confirmed twice). Matrix-leg capture (watchdog or manual-boot convergence) tracked as a follow-up: #421.