Closes out #102, which asked for either a cause with a measurement behind it or a documented
decision to treat the failures as environmental "with the same honesty @FailsOnEmulatorApi37
gets". It gets both, plus one fix that came out of the reading.
The census, because the ticket's modes had never been re-counted
Every gating E2E leg-attempt in the repo's history, status_check.yml, 2026-08-20..2026-09-07: 1489 leg-attempts, 129 failures, 8.7%. Per leg-attempt and not per run, and that is measured
rather than asserted: those 129 sit in 83 distinct runs, and 45 of the 83 ended green once
someone re-ran them — so a census that counts failed runs finds 38 events where there were 129,
and the ones it deletes are exactly the failures somebody already decided were noise. Cancelled
legs are excluded; those are cancel-in-progress cancellations, not runs.
mode
all history
since 2026-08-27
last 2 days
disposition
SAF picker (incl. rotation)
52
17
7
#108 on 37; zero on 33-36 in 881 leg-attempts since #96
five shapes; three fixed, one fixed here, two are #270
emulator never came up (adb 224)
3
3
1
infra, before the suite
#102's own mode: the emulator's Codec2 HAL segfaults
transcodesH264ToH265AndReportsProgress does not starve. The per-test logcat of run 34000816016 a1, 62 ms after the test starts:
D MediaCodec: MediaCodec::reclaim(...) c2.goldfish.h264.decoder
F DEBUG : Cmdline: /vendor/bin/hw/android.hardware.media.c2@1.0-service-goldfish
F DEBUG : signal 0 (SIGSEGV) ... Cause: null pointer dereference
#00 C2Block2D::handle() const+4 libcodec2_vndk.so
#01 getClientUsage(std::shared_ptr<C2BlockPool> const&) libcodec2_goldfish_common.so
#02 android::C2GoldfishAvcDec::process(...) libcodec2_goldfish_avcdec.so
E CCodec : Codec2 component "c2.goldfish.h264.decoder" died.
The HAL dies and respawns; Media3 is left with a dead codec and its own 25-second export watchdog
aborts the export, which is the Muxer error in the job log. Six for six — every occurrence
carries the crash in the same job's log. Six in 1210 API 33-36 leg-attempts, 0.5%, and it is
the same vendor library @FailsOnEmulatorApi37 already names. The crashing code is /vendor/lib64/*; there is nothing here to fix, and the ticket's "two severities of one weakness"
reading turns out to have been right.
That also answers #190's open question about widening the marker — 0.5% below API 37 against
100% on it, so the marker is the right width. Posted there with the numbers.
The one thing that was ours: a back press sent on a stale reading
dismissThePicker guarded its back presses on Activity.hasWindowFocus, which a system app-error
dialog makes false as well — it is a fullscreen system_server window, which is the entire reason dismissASystemErrorDialog exists. So the guard could not tell "the picker is still up" from
"a dialog is on top of an app that is already in front":
20:59:35.689 UiObject2: Clicking on (927, 2274) <- iteration 2's dismissal, on button1
20:59:36.033 MainActivity RESUMED <- the picker is gone, by the test's own hand
20:59:36.350 VRI[PickActivity]: visibilityChanged ... newVisibility=false
20:59:37.068 UiDevice: Pressing back button. <- iteration 2 presses anyway
20:59:41.169 Input channel 'Application Not Responding: ...nexuslauncher' was disposed
20:59:41.713 UiDevice: Pressing back button. <- iteration 3
20:59:41.754 TopTaskTracker: onTaskMovedToFront: ... NexusLauncherActivity
20:59:42.278 MainActivity DESTROYED
That is API 35 of run 34161043035 attempt 1, whose head is #269's own commit. Read the first
two lines carefully: the picker did not close on its own — this class closed it, when dismissASystemErrorDialog fell through to android:id/button1 and clicked DocumentsUI's own
positive button (#271). From 36.033 there was nothing to back out of. The loop pressed anyway on
iteration 2, dismissed the launcher's ANR dialog on iteration 3 — #93's occluder, still ambient,
and the only remaining reason the focus read false — and pressed again. That press finished MainActivity, and everything afterwards threw Cannot run onActivity since Activity has been destroyed already.
With the re-read, iteration 2 returns and neither press happens.
The fix re-reads the focus after a dialog is actually dismissed, and only then. It removes a back
press rather than retrying one — a picker genuinely in front still leaves the app unfocused and
still gets pressed, so nothing this class can catch changes, and on the ordinary path with no
dialog nothing is re-read and nothing is waited on. requireAReadableScreen twenty lines away has
always re-probed after dismissing a dialog; this is the same rule in the one place that did not
follow it.
It cannot be demonstrated by re-running, and the KDoc says so rather than implying a green
sweep is evidence: the launcher ANR is ambient and not reproducible on demand. The trace is the
evidence.
Not done here, deliberately
No retry wrapper, no re-run loop, no widened timeout.#102 rules those out and they are not
in this diff.
#270 — the two save-test failure shapes that still have no mechanism. Both predate #269, and
the first move is a re-count, not a fix.
#271 — ERROR_DIALOG_BUTTONS ends with android:id/button1, which is any AlertDialog's
positive button; it was measured clicking one inside DocumentsUI. Narrowing it is a decision, not
a cleanup.
#102 itself is left open for the maintainer to close. The evidence is on the ticket.
Two rates are "not yet measured", not "fixed". The wedge is 0 in 106 API 33/34 leg-attempts
against a prior 1.8%, which is P ≈ 0.15 under no change. And #269's head has exactly one green
gating leg-attempt on API 35 (34161043035 a2) plus the attempt-1 failure this PR fixes; main's merge commit has had no gating run at all.
Testing
assembleDebug, testDebugUnitTest, compileDebugAndroidTestKotlin, ktlintCheck, detekt, lintDebug all green, and the local gate's instrumented sweep ran on API 33, 34, 35 and 36.
API 37 is NOT COVERED LOCALLY — no device attached — so CI's gating leg answers for it.
The changed code is an androidTest helper with no JVM seam: it drives UiDevice and ActivityScenario against a real picker, so there is nothing here a unit test could hold. Named
rather than implied.
Closes out **#102**, which asked for either a cause with a measurement behind it or a documented
decision to treat the failures as environmental "with the same honesty `@FailsOnEmulatorApi37`
gets". It gets both, plus one fix that came out of the reading.
## The census, because the ticket's modes had never been re-counted
Every gating E2E leg-attempt in the repo's history, `status_check.yml`, 2026-08-20..2026-09-07:
**1489 leg-attempts, 129 failures, 8.7%.** Per leg-attempt and not per run, and that is measured
rather than asserted: those 129 sit in 83 distinct runs, and **45 of the 83 ended green** once
someone re-ran them — so a census that counts failed *runs* finds 38 events where there were 129,
and the ones it deletes are exactly the failures somebody already decided were noise. Cancelled
legs are excluded; those are `cancel-in-progress` cancellations, not runs.
| mode | all history | since 2026-08-27 | last 2 days | disposition |
|---|---|---|---|---|
| SAF picker (incl. rotation) | 52 | 17 | 7 | #108 on 37; **zero** on 33-36 in 881 leg-attempts since #96 |
| framework down (`Can't find service`) | 25 | 11 | 4 | #108 |
| other named-test failure | 16 | 2 | 1 | each fixed by its own PR |
| wedge | 9 | 6 | 1 | #122/#219 — 0 in 106 since, **not yet demonstrated** |
| **Media3 export watchdog (#102's own mode)** | **6** | **3** | **1** | **environmental, measured** |
| SAF save tests | 7 | 7 | 7 | five shapes; three fixed, one fixed here, two are #270 |
| emulator never came up (`adb` 224) | 3 | 3 | 1 | infra, before the suite |
## #102's own mode: the emulator's Codec2 HAL segfaults
`transcodesH264ToH265AndReportsProgress` does not starve. The per-test logcat of run
`34000816016` a1, 62 ms after the test starts:
```
D MediaCodec: MediaCodec::reclaim(...) c2.goldfish.h264.decoder
F DEBUG : Cmdline: /vendor/bin/hw/android.hardware.media.c2@1.0-service-goldfish
F DEBUG : signal 0 (SIGSEGV) ... Cause: null pointer dereference
#00 C2Block2D::handle() const+4 libcodec2_vndk.so
#01 getClientUsage(std::shared_ptr<C2BlockPool> const&) libcodec2_goldfish_common.so
#02 android::C2GoldfishAvcDec::process(...) libcodec2_goldfish_avcdec.so
E CCodec : Codec2 component "c2.goldfish.h264.decoder" died.
```
The HAL dies and respawns; Media3 is left with a dead codec and its own 25-second export watchdog
aborts the export, which is the `Muxer error` in the job log. **Six for six** — every occurrence
carries the crash in the same job's log. Six in 1210 API 33-36 leg-attempts, **0.5%**, and it is
the same vendor library `@FailsOnEmulatorApi37` already names. The crashing code is
`/vendor/lib64/*`; there is nothing here to fix, and the ticket's "two severities of one weakness"
reading turns out to have been right.
That also answers **#190**'s open question about widening the marker — 0.5% below API 37 against
100% on it, so the marker is the right width. Posted there with the numbers.
## The one thing that was ours: a back press sent on a stale reading
`dismissThePicker` guarded its back presses on `Activity.hasWindowFocus`, which a system app-error
dialog makes false as well — it is a fullscreen `system_server` window, which is the entire reason
`dismissASystemErrorDialog` exists. So the guard could not tell "the picker is still up" from
"a dialog is on top of an app that is already in front":
```
20:59:35.689 UiObject2: Clicking on (927, 2274) <- iteration 2's dismissal, on button1
20:59:36.033 MainActivity RESUMED <- the picker is gone, by the test's own hand
20:59:36.350 VRI[PickActivity]: visibilityChanged ... newVisibility=false
20:59:37.068 UiDevice: Pressing back button. <- iteration 2 presses anyway
20:59:41.169 Input channel 'Application Not Responding: ...nexuslauncher' was disposed
20:59:41.713 UiDevice: Pressing back button. <- iteration 3
20:59:41.754 TopTaskTracker: onTaskMovedToFront: ... NexusLauncherActivity
20:59:42.278 MainActivity DESTROYED
```
That is API 35 of run `34161043035` attempt 1, whose head is **#269's own commit**. Read the first
two lines carefully: **the picker did not close on its own — this class closed it**, when
`dismissASystemErrorDialog` fell through to `android:id/button1` and clicked DocumentsUI's own
positive button (#271). From `36.033` there was nothing to back out of. The loop pressed anyway on
iteration 2, dismissed the launcher's ANR dialog on iteration 3 — #93's occluder, still ambient,
and the only remaining reason the focus read false — and pressed again. That press finished
`MainActivity`, and everything afterwards threw `Cannot run onActivity since Activity has been
destroyed already`.
With the re-read, **iteration 2 returns** and neither press happens.
The fix re-reads the focus after a dialog is actually dismissed, and only then. **It removes a back
press rather than retrying one** — a picker genuinely in front still leaves the app unfocused and
still gets pressed, so nothing this class can catch changes, and on the ordinary path with no
dialog nothing is re-read and nothing is waited on. `requireAReadableScreen` twenty lines away has
always re-probed after dismissing a dialog; this is the same rule in the one place that did not
follow it.
**It cannot be demonstrated by re-running, and the KDoc says so** rather than implying a green
sweep is evidence: the launcher ANR is ambient and not reproducible on demand. The trace is the
evidence.
## Not done here, deliberately
- **No retry wrapper, no re-run loop, no widened timeout.** #102 rules those out and they are not
in this diff.
- **#270** — the two save-test failure shapes that still have no mechanism. Both predate #269, and
the first move is a re-count, not a fix.
- **#271** — `ERROR_DIALOG_BUTTONS` ends with `android:id/button1`, which is any `AlertDialog`'s
positive button; it was measured clicking one inside DocumentsUI. Narrowing it is a decision, not
a cleanup.
- **#102 itself is left open** for the maintainer to close. The evidence is on the ticket.
- **Two rates are "not yet measured", not "fixed".** The wedge is 0 in 106 API 33/34 leg-attempts
against a prior 1.8%, which is P ≈ 0.15 under no change. And #269's head has exactly one green
gating leg-attempt on API 35 (`34161043035` a2) plus the attempt-1 failure this PR fixes;
`main`'s merge commit has had no gating run at all.
## Testing
`assembleDebug`, `testDebugUnitTest`, `compileDebugAndroidTestKotlin`, `ktlintCheck`, `detekt`,
`lintDebug` all green, and the local gate's instrumented sweep ran on API 33, 34, 35 and 36.
API 37 is `NOT COVERED LOCALLY` — no device attached — so CI's gating leg answers for it.
The changed code is an `androidTest` helper with no JVM seam: it drives `UiDevice` and
`ActivityScenario` against a real picker, so there is nothing here a unit test could hold. Named
rather than implied.
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Closes out #102, which asked for either a cause with a measurement behind it or a documented
decision to treat the failures as environmental "with the same honesty
@FailsOnEmulatorApi37gets". It gets both, plus one fix that came out of the reading.
The census, because the ticket's modes had never been re-counted
Every gating E2E leg-attempt in the repo's history,
status_check.yml, 2026-08-20..2026-09-07:1489 leg-attempts, 129 failures, 8.7%. Per leg-attempt and not per run, and that is measured
rather than asserted: those 129 sit in 83 distinct runs, and 45 of the 83 ended green once
someone re-ran them — so a census that counts failed runs finds 38 events where there were 129,
and the ones it deletes are exactly the failures somebody already decided were noise. Cancelled
legs are excluded; those are
cancel-in-progresscancellations, not runs.Can't find service)adb224)#102's own mode: the emulator's Codec2 HAL segfaults
transcodesH264ToH265AndReportsProgressdoes not starve. The per-test logcat of run34000816016a1, 62 ms after the test starts:The HAL dies and respawns; Media3 is left with a dead codec and its own 25-second export watchdog
aborts the export, which is the
Muxer errorin the job log. Six for six — every occurrencecarries the crash in the same job's log. Six in 1210 API 33-36 leg-attempts, 0.5%, and it is
the same vendor library
@FailsOnEmulatorApi37already names. The crashing code is/vendor/lib64/*; there is nothing here to fix, and the ticket's "two severities of one weakness"reading turns out to have been right.
That also answers #190's open question about widening the marker — 0.5% below API 37 against
100% on it, so the marker is the right width. Posted there with the numbers.
The one thing that was ours: a back press sent on a stale reading
dismissThePickerguarded its back presses onActivity.hasWindowFocus, which a system app-errordialog makes false as well — it is a fullscreen
system_serverwindow, which is the entire reasondismissASystemErrorDialogexists. So the guard could not tell "the picker is still up" from"a dialog is on top of an app that is already in front":
That is API 35 of run
34161043035attempt 1, whose head is #269's own commit. Read the firsttwo lines carefully: the picker did not close on its own — this class closed it, when
dismissASystemErrorDialogfell through toandroid:id/button1and clicked DocumentsUI's ownpositive button (#271). From
36.033there was nothing to back out of. The loop pressed anyway oniteration 2, dismissed the launcher's ANR dialog on iteration 3 — #93's occluder, still ambient,
and the only remaining reason the focus read false — and pressed again. That press finished
MainActivity, and everything afterwards threwCannot run onActivity since Activity has been destroyed already.With the re-read, iteration 2 returns and neither press happens.
The fix re-reads the focus after a dialog is actually dismissed, and only then. It removes a back
press rather than retrying one — a picker genuinely in front still leaves the app unfocused and
still gets pressed, so nothing this class can catch changes, and on the ordinary path with no
dialog nothing is re-read and nothing is waited on.
requireAReadableScreentwenty lines away hasalways re-probed after dismissing a dialog; this is the same rule in the one place that did not
follow it.
It cannot be demonstrated by re-running, and the KDoc says so rather than implying a green
sweep is evidence: the launcher ANR is ambient and not reproducible on demand. The trace is the
evidence.
Not done here, deliberately
in this diff.
the first move is a re-count, not a fix.
ERROR_DIALOG_BUTTONSends withandroid:id/button1, which is anyAlertDialog'spositive button; it was measured clicking one inside DocumentsUI. Narrowing it is a decision, not
a cleanup.
against a prior 1.8%, which is P ≈ 0.15 under no change. And #269's head has exactly one green
gating leg-attempt on API 35 (
34161043035a2) plus the attempt-1 failure this PR fixes;main's merge commit has had no gating run at all.Testing
assembleDebug,testDebugUnitTest,compileDebugAndroidTestKotlin,ktlintCheck,detekt,lintDebugall green, and the local gate's instrumented sweep ran on API 33, 34, 35 and 36.API 37 is
NOT COVERED LOCALLY— no device attached — so CI's gating leg answers for it.The changed code is an
androidTesthelper with no JVM seam: it drivesUiDeviceandActivityScenarioagainst a real picker, so there is nothing here a unit test could hold. Namedrather than implied.