The commit below describes "a back press aimed at a picker that had closed five seconds
earlier". The timing is right and the agency is wrong, and the agency is the interesting
half. Re-read out of the same logcat:
20:59:35.686 UiDevice: Retrieving node ... [RES='android:id/button1'].
20:59:35.689 UiObject2: Clicking on (927, 2274).
20:59:36.033 MainActivity RESUMED
20:59:36.350 VRI[PickActivity]: visibilityChanged ... newVisibility=false
`aerr_wait` and `aerr_close` both missed on that iteration and `dismissASystemErrorDialog`
fell through to `android:id/button1` -- the framework's generic AlertDialog positive button,
which is on every AlertDialog on the device. It hit one inside DocumentsUI, and that is what
closed the picker. The picker did not close on its own; this class closed it.
So the destroyed Activity and #271 are one incident rather than two findings that happened
to share a trace, and the loop's shape is three iterations rather than two: iteration 2
closes the picker through `button1` and then presses back into an app that is already in
front, iteration 3 dismisses the launcher's ANR dialog and presses again, and that press
finishes MainActivity.
It also sharpens what the fix does. With the re-read, iteration 2 returns -- the app is
focused within a second of the `button1` click -- so neither of the two presses that
followed it happens at all. The previous message implied the fix caught only the last one.
`requireAReadableScreen` now drops a Boolean return value, which this codebase treats as a
smell. It is correct there -- the `device.wait` on the next line is the re-probe -- and the
call site says so rather than leaving a reader to work out whether it was an oversight.
No behaviour change beyond the comment: the fix itself is unchanged and the sweep is re-run
because `app/src` is touched.
Refs #102, #271
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`dismissThePicker` guarded its back presses on `Activity.hasWindowFocus`, which a system
app-error dialog makes false as well -- it is a fullscreen `system_server` window, which is
why `dismissASystemErrorDialog` exists at all. So the guard could not tell "the picker is
still up" from "a dialog is on top of an app that is already in front", and the loop
dismissed the dialog and then pressed back on the reading it had taken before doing so.
Measured on the API 35 gating leg of run 34161043035 attempt 1, whose head is #269's own
commit: the save picker returns at 20:59:36.033, the launcher's ANR dialog is dismissed at
20:59:41.169, a back press goes out at 20:59:41.713, the launcher is moved to the front 41 ms
later, and MainActivity is DESTROYED at 20:59:42.278. Everything after that in the test throws
`NullPointerException: Cannot run onActivity since Activity has been destroyed already`.
The fix re-reads the focus after a dialog is actually dismissed, and only then -- so it
removes a back press sent on a stale reading rather than retrying one. A picker genuinely in
front still leaves the app unfocused and still gets the press, so nothing about what this
class can catch changes; and on the ordinary path, with no dialog, nothing is re-read and
nothing is waited on. `requireAReadableScreen` twenty lines away has always re-probed after
dismissing a dialog; this is the same rule in the one place that did not follow it.
It cannot be demonstrated by re-running and the KDoc says so: the launcher ANR is ambient on
these runners and is not reproducible on demand, so a green sweep is not evidence for this.
The trace is.
`docs/ci-failure-modes.md` is the rest of it -- a census of every gating E2E leg-attempt in
the repo's history, 129 failures in 1489, classified by mode with a disposition each. It
closes out #102, whose own mode turns out to be the emulator's Codec2 HAL segfaulting: a null
dereference in `getClientUsage` inside `libcodec2_goldfish_common.so` kills
`c2.goldfish.h264.decoder`, and Media3's 25 s export watchdog then aborts the export. Six
occurrences, six carrying that crash in the same job's log, 0.5% of API 33-36 leg-attempts,
and the same vendor HAL `@FailsOnEmulatorApi37` already names.
Three counting rules are in the document and in CLAUDE.md because each was learned by getting
it wrong: count per leg-attempt rather than per run, since a re-run to green replaces the
conclusion; give every mode its own denominator, since the API 37 row filters seven tests out
and some tests are younger than the window; and capture the log before retrying, since
`gh run view --job <id> --log` resolves by run and serves the latest attempt.
Refs #102
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>