CI legs fail on system services, not assertions: three modes today that look like one cause #102

Closed
opened 2026-08-25 14:25:31 +00:00 by JMR-dev · 19 comments
JMR-dev commented 2026-08-25 14:25:31 +00:00 (Migrated from github.com)

Split out of #101, where I first mis-attributed it to RealMediaBenchmark. This is the test that actually failed.

The failure

Media3EngineTest.transcodesH264ToH265AndReportsProgress fails intermittently on the gating API 34 leg:

Media3EngineTest > transcodesH264ToH265AndReportsProgress[test(AVD) - 14]  FAILED
  java.lang.IllegalStateException: androidx.media3.transformer.ExportException: Muxer error

A sibling signature on PR #95's API 36 leg in the same window:

java.lang.IllegalStateException: Abort: no output sample written in the last 25000 milliseconds

Both are the hardware transcode failing to make progress, on levels where it is expected to work.

Why this is awkward, and interesting

This is one of the two tests that already carries @FailsOnEmulatorApi37. It is excluded from
the API 37 gating leg because it dies inside that image's c2.goldfish.h264.decoder, and it runs in
the advisory job there instead. On API 33-36 it is a normal gating test and has been reliable all
session — dozens of green runs.

So the same test is: known-broken on one emulator image, reliable on four, and now intermittently
failing on two of those four.

The hypothesis worth testing first

no output sample written in the last 25000 milliseconds is a starvation signature, not a
correctness one.

#93's root cause, established in PR #96, was that the launcher ANRs on a loaded runner — loaded
enough that system_server leaves an "Application Not Responding" dialog that never clears. A runner
under that much pressure would also starve a hardware transcode of its 25-second sample budget.

If that is the same illness, then #96 fixed the accessibility symptom of runner load and this is
the same cause surfacing in the codec path. That would also explain why both appeared in the same
few hours rather than gradually.

Check that before writing a Media3-specific theory. The measurement is whether these failures
correlate with the ANR signature in the same job's logcat.

Frequency, honestly

Two sightings in one window (#95 API 36, #97 API 34). That is not yet a rate, and I overstated a
different flake earlier today on four sightings, so this ticket deliberately does not claim it is a
blocker. It is filed so the next sighting has somewhere to land and so the connection to #93 is on
record rather than rediscovered.

Done means

Either a cause with a measurement behind it — runner load, a Media3 version behaviour, an emulator
configuration — or a documented decision to treat it as environmental with the same honesty
@FailsOnEmulatorApi37 gets. Not a retry wrapper added because the test is annoying: the point of
this test is that a hardware transcode completes, and retrying until it does would delete the
assertion.

_Split out of #101, where I first mis-attributed it to `RealMediaBenchmark`. This is the test that actually failed._ ### The failure `Media3EngineTest.transcodesH264ToH265AndReportsProgress` fails intermittently on the **gating** API 34 leg: ``` Media3EngineTest > transcodesH264ToH265AndReportsProgress[test(AVD) - 14] FAILED java.lang.IllegalStateException: androidx.media3.transformer.ExportException: Muxer error ``` A sibling signature on PR #95's API 36 leg in the same window: ``` java.lang.IllegalStateException: Abort: no output sample written in the last 25000 milliseconds ``` Both are the hardware transcode failing to make progress, on levels where it is expected to work. ### Why this is awkward, and interesting **This is one of the two tests that already carries `@FailsOnEmulatorApi37`.** It is excluded from the API 37 gating leg because it dies inside that image's `c2.goldfish.h264.decoder`, and it runs in the advisory job there instead. On API 33-36 it is a normal gating test and has been reliable all session — dozens of green runs. So the same test is: known-broken on one emulator image, reliable on four, and now intermittently failing on two of those four. ### The hypothesis worth testing first `no output sample written in the last 25000 milliseconds` is a **starvation** signature, not a correctness one. #93's root cause, established in PR #96, was that **the launcher ANRs on a loaded runner** — loaded enough that `system_server` leaves an "Application Not Responding" dialog that never clears. A runner under that much pressure would also starve a hardware transcode of its 25-second sample budget. If that is the same illness, then #96 fixed the *accessibility* symptom of runner load and this is the same cause surfacing in the codec path. That would also explain why both appeared in the same few hours rather than gradually. **Check that before writing a Media3-specific theory.** The measurement is whether these failures correlate with the ANR signature in the same job's logcat. ### Frequency, honestly **Two sightings in one window** (#95 API 36, #97 API 34). That is not yet a rate, and I overstated a different flake earlier today on four sightings, so this ticket deliberately does not claim it is a blocker. It is filed so the next sighting has somewhere to land and so the connection to #93 is on record rather than rediscovered. ### Done means Either a cause with a measurement behind it — runner load, a Media3 version behaviour, an emulator configuration — or a documented decision to treat it as environmental with the same honesty `@FailsOnEmulatorApi37` gets. **Not** a retry wrapper added because the test is annoying: the point of this test is that a hardware transcode completes, and retrying until it does would delete the assertion.
JMR-dev commented 2026-08-25 15:05:28 +00:00 (Migrated from github.com)

A third failure mode, and the reason to widen this ticket rather than file a fourth.

PR #99's gating E2E API 37 leg failed today with the tests all passing:

Tests 50/56 completed. (2 skipped) (0 failed)
...
Shell command failed (20): am get-current-user
Device emulator-5554 failed to uninstall test APK org.libremediaconverter.
> Task :app:connectedDebugAndroidTest FAILED
--- native crashes (tail 60) ---

Zero test failures. am get-current-user returning 20 means system_server is not answering — the framework went down during teardown, after every test had already passed. The notAnnotation filter was intact (56 tests, the right number), so this is not a filtering regression from that PR's workflow edit.

Three modes, one shape

# mode signature status
1 SAF picker launcher ANRs, AccessibilityWindowManager drops every app window #93 — fixed by #96
2 Media3 transcode no output sample written in the last 25000 milliseconds this ticket
3 API 37 teardown am get-current-user fails; framework down, 0 test failures new, here

All three are system services failing to answer under load, not assertions failing. #96 established mode 1's root cause as a runner loaded enough that the launcher ANRs and system_server leaves a dialog that never clears. Modes 2 and 3 are what that same pressure would look like in the codec path and in teardown.

They also arrived together. None of these three was seen before today; all three appeared within a few hours, on diffs that cannot cause them (KDoc comments, a lookup table, a workflow step).

What that changes about this ticket

It is probably not a Media3 ticket. Test whether the three correlate with runner load before writing a codec-specific theory — #96 already proved the load hypothesis once, and the measurement is whether these failures carry the ANR signature in the same job's logcat.

If they do, the useful fix is at the harness level — fewer things competing on the runner, a longer settle before the suite starts, or a documented acceptance that these legs are load-sensitive — rather than three separate per-test patches.

Frequency, still honest

Mode 2: two sightings. Mode 3: one. Not claiming a rate, and deliberately not escalating severity on three data points — the correction to #49 earlier today came from doing exactly that on four.

**A third failure mode, and the reason to widen this ticket rather than file a fourth.** PR #99's gating **E2E API 37** leg failed today with the tests all passing: ``` Tests 50/56 completed. (2 skipped) (0 failed) ... Shell command failed (20): am get-current-user Device emulator-5554 failed to uninstall test APK org.libremediaconverter. > Task :app:connectedDebugAndroidTest FAILED --- native crashes (tail 60) --- ``` **Zero test failures.** `am get-current-user` returning 20 means `system_server` is not answering — the framework went down during teardown, after every test had already passed. The `notAnnotation` filter was intact (56 tests, the right number), so this is not a filtering regression from that PR's workflow edit. ### Three modes, one shape | # | mode | signature | status | |---|---|---|---| | 1 | SAF picker | launcher ANRs, `AccessibilityWindowManager` drops every app window | #93 — **fixed** by #96 | | 2 | Media3 transcode | `no output sample written in the last 25000 milliseconds` | this ticket | | 3 | API 37 teardown | `am get-current-user` fails; framework down, 0 test failures | new, here | All three are **system services failing to answer under load**, not assertions failing. #96 established mode 1's root cause as a runner loaded enough that the launcher ANRs and `system_server` leaves a dialog that never clears. Modes 2 and 3 are what that same pressure would look like in the codec path and in teardown. They also arrived together. None of these three was seen before today; all three appeared within a few hours, on diffs that cannot cause them (KDoc comments, a lookup table, a workflow step). ### What that changes about this ticket It is probably not a Media3 ticket. **Test whether the three correlate with runner load before writing a codec-specific theory** — #96 already proved the load hypothesis once, and the measurement is whether these failures carry the ANR signature in the same job's logcat. If they do, the useful fix is at the harness level — fewer things competing on the runner, a longer settle before the suite starts, or a documented acceptance that these legs are load-sensitive — rather than three separate per-test patches. ### Frequency, still honest Mode 2: two sightings. Mode 3: one. **Not claiming a rate**, and deliberately not escalating severity on three data points — the correction to #49 earlier today came from doing exactly that on four.
JMR-dev commented 2026-08-26 01:51:31 +00:00 (Migrated from github.com)

The three modes, now counted — and one of them moved

Census of every gating E2E leg-attempt since 2026-08-24 (advisory job excluded), classified by
anchoring the test name to the FAILED marker in each failing job's own log. 45 failures across
400 leg-attempts.
Counting runs rather than leg-attempts sees only 19 of those 45, because a
re-run to green erases the evidence — every figure here is per leg-attempt.

cause count
SafPickerRoundTripTest — picker and/or rotation (#93) 29
no test named — abort or infra 6
deliberate @FailsOnEmulatorApi37 probe for #83 4
doesNotOverwriteAPickTheUserHasAlreadyMade (#49) 3
transcodesH264ToH265AndReportsProgress (#102 Media3 mode) 2
routesAFastMp4JobByDeviceCapability 1

The finding that matters for this ticket

#96 fixed the SAF picker on API 33–36 and did not fix it on API 37.

Since #96 merged (2026-08-25T13:42Z) there have been zero picker failures on 33/34/35/36 and
two on API 37 — runs 32865281555 and 32899061992. I checked with git merge-base --is-ancestor that both head commits actually contain #96's fix, so this is not a stale branch.

API 37 is also the only leg that aborts in system_server (#108). Cross-referencing the two:
5 of the 8 API 37 legs carrying the TaskSnapshotPersister abort trace also had a test fail, and
in 4 of those 5 the failing test was the picker.

That is the concrete shape of this ticket's "one cause" hypothesis: on API 37 the picker test is
not failing for #93's original reason (an ANR dialog hiding every window from UiAutomator — that
cause was fixed and stayed fixed on four other API levels). It is failing because system_server
stops answering, which is the same event that produces #108's abort. The picker test is simply the
most system-service-dependent test in the suite, so it is the first to notice.

Consequence for triage: #108 should not be treated as a cosmetic post-run crash, and the API 37
picker failures should not be filed as a #93 regression. They are one problem, and #108 is the
better handle on it.

Method note

Historical attempts cannot be read with gh run view --job <id> --log — it resolves by run and
serves the latest attempt, so a re-run hands you a green log for a red attempt. Use
gh api --allow-escape-sequences /repos/{owner}/{repo}/actions/jobs/{job_id}/logs, with the job id
from /actions/runs/{run}/attempts/{n}/jobs. Without --allow-escape-sequences, gh writes nothing
and exits 0.

## The three modes, now counted — and one of them moved Census of every gating E2E leg-attempt since 2026-08-24 (advisory job excluded), classified by anchoring the test name to the `FAILED` marker in each failing job's own log. **45 failures across 400 leg-attempts.** Counting runs rather than leg-attempts sees only 19 of those 45, because a re-run to green erases the evidence — every figure here is per leg-attempt. | cause | count | |---|---| | `SafPickerRoundTripTest` — picker and/or rotation (#93) | 29 | | no test named — abort or infra | 6 | | deliberate `@FailsOnEmulatorApi37` probe for #83 | 4 | | `doesNotOverwriteAPickTheUserHasAlreadyMade` (#49) | 3 | | `transcodesH264ToH265AndReportsProgress` (#102 Media3 mode) | 2 | | `routesAFastMp4JobByDeviceCapability` | 1 | ## The finding that matters for this ticket **#96 fixed the SAF picker on API 33–36 and did not fix it on API 37.** Since #96 merged (2026-08-25T13:42Z) there have been **zero** picker failures on 33/34/35/36 and **two** on API 37 — runs `32865281555` and `32899061992`. I checked with `git merge-base --is-ancestor` that both head commits actually contain #96's fix, so this is not a stale branch. API 37 is also the only leg that aborts in `system_server` (#108). Cross-referencing the two: **5 of the 8 API 37 legs carrying the `TaskSnapshotPersister` abort trace also had a test fail**, and in 4 of those 5 the failing test was the picker. That is the concrete shape of this ticket's "one cause" hypothesis: on API 37 the picker test is not failing for #93's original reason (an ANR dialog hiding every window from UiAutomator — that cause was fixed and stayed fixed on four other API levels). It is failing because `system_server` stops answering, which is the same event that produces #108's abort. The picker test is simply the most system-service-dependent test in the suite, so it is the first to notice. **Consequence for triage:** #108 should not be treated as a cosmetic post-run crash, and the API 37 picker failures should not be filed as a #93 regression. They are one problem, and #108 is the better handle on it. ## Method note Historical attempts cannot be read with `gh run view --job <id> --log` — it resolves by run and serves the *latest* attempt, so a re-run hands you a green log for a red attempt. Use `gh api --allow-escape-sequences /repos/{owner}/{repo}/actions/jobs/{job_id}/logs`, with the job id from `/actions/runs/{run}/attempts/{n}/jobs`. Without `--allow-escape-sequences`, gh writes nothing and exits 0.
JMR-dev commented 2026-08-26 01:53:25 +00:00 (Migrated from github.com)

Correcting the causal claim in my previous comment

Two sentences above are not supported by the data in the same comment, and I am retracting them:

It is failing because system_server stops answering, which is the same event that produces
#108's abort.

the API 37 picker failures should not be filed as a #93 regression

The discriminator is in my own table. The two post-#96 API 37 picker failures are:

run / attempt abort trace picker failed
32865281555 a2 no yes
32899061992 a1 yes yes

One of the two has no TaskSnapshotPersister trace at all. (32865281555 in fact failed on API 37
twice — a1 with neither an abort trace nor a named test failure, then a2 with the picker.) The
"4 of 5 abort legs also failed the picker" figure does not rescue the claim, because those
co-occurrences are almost all pre-#96, when the picker was failing on every API level for #93's
own cause. Only the post-#96 API 37 failures discriminate, and that is n=2 with one against.

The second sentence is the harmful one: it would tell a future triager to discard the possibility of
a #93 regression on 1-for-2 evidence. Withdrawn.

A third mechanism I failed to enumerate

SafPickerRoundTripTest's own KDoc already documents an API-37-specific failure mode, and it is
not #108's:

the android-37.x images are where that stops being free: surfaceflinger aborts inside the guest's
Gralloc5 mapper, init SIGKILLs zygote with it, and the framework restarts underneath the run.

So there are at least three candidates for the post-#96 API 37 picker failures — #93's cause
surviving on 37, #108's system_server abort, and this surfaceflinger/Gralloc5 restart — and I
picked one without ruling out the other two.

I also checked whether #96's fix might be version-sensitive: dismissASystemErrorDialog (the
aerr_wait path) is not API-gated and runs identically on every leg, so a gated remedy is not
the explanation.

What actually survives, and is worth keeping

#96 fixed the picker on API 33–36 and API 37 has failed twice since. Zero picker failures on
33/34/35/36 post-fix; two on 37, both on head commits verified with git merge-base --is-ancestor
to contain the fix.

Which of the three mechanisms is responsible is unresolved. The abort trace being absent from one
of the two failures is precisely why it is unresolved, rather than being the reason to pick #108.

## Correcting the causal claim in my previous comment Two sentences above are not supported by the data in the same comment, and I am retracting them: > It is failing because `system_server` stops answering, which is the same event that produces > #108's abort. > the API 37 picker failures should not be filed as a #93 regression **The discriminator is in my own table.** The two post-#96 API 37 picker failures are: | run / attempt | abort trace | picker failed | |---|---|---| | `32865281555 a2` | **no** | yes | | `32899061992 a1` | yes | yes | One of the two has no `TaskSnapshotPersister` trace at all. (`32865281555` in fact failed on API 37 twice — a1 with neither an abort trace nor a named test failure, then a2 with the picker.) The "4 of 5 abort legs also failed the picker" figure does not rescue the claim, because those co-occurrences are almost all **pre-#96**, when the picker was failing on every API level for #93's own cause. Only the post-#96 API 37 failures discriminate, and that is n=2 with one against. The second sentence is the harmful one: it would tell a future triager to discard the possibility of a #93 regression on 1-for-2 evidence. Withdrawn. ## A third mechanism I failed to enumerate `SafPickerRoundTripTest`'s own KDoc already documents an API-37-specific failure mode, and it is **not** #108's: > the android-37.x images are where that stops being free: surfaceflinger aborts inside the guest's > Gralloc5 mapper, init SIGKILLs zygote with it, and the framework restarts underneath the run. So there are at least three candidates for the post-#96 API 37 picker failures — #93's cause surviving on 37, #108's `system_server` abort, and this surfaceflinger/Gralloc5 restart — and I picked one without ruling out the other two. I also checked whether #96's fix might be version-sensitive: `dismissASystemErrorDialog` (the `aerr_wait` path) is **not** API-gated and runs identically on every leg, so a gated remedy is not the explanation. ## What actually survives, and is worth keeping **#96 fixed the picker on API 33–36 and API 37 has failed twice since.** Zero picker failures on 33/34/35/36 post-fix; two on 37, both on head commits verified with `git merge-base --is-ancestor` to contain the fix. Which of the three mechanisms is responsible is **unresolved**. The abort trace being absent from one of the two failures is precisely why it is unresolved, rather than being the reason to pick #108.
JMR-dev commented 2026-08-26 02:16:40 +00:00 (Migrated from github.com)

Partial answer to this ticket's question, from the API 37 side

Two of the API 37 aborts are provably one bug, not two — details and evidence on #108.

The crash the E2E_DISABLE_SYSTEM_UI mitigation already handles and the crash in this ticket's
third mode are the same assertion in the same emulator mapper:

Abort message: 'Assertion failed: !rcEnc->featureInfo()->hasReadColorBufferDma'
  ... GoldfishMapper::readFromHost

They differ only in who calls it — RegionSamplingThread (SystemUI's nav-bar luma sampler, which
the mitigation removes by deleting the package) versus TaskSnapshotPer (WindowManager, inside
system_server, where there is no package to remove). All 8 recorded #108 legs show
RegionSamplingThread=0, so they are entirely the second caller.

What this does and does not settle. It supports "one cause" for the API 37 aborts. It says
nothing about the picker failures — I retracted that claim above, and the reason still stands: one
of the two post-#96 API 37 picker failures had no abort trace at all. A shared root for the aborts
does not retroactively make the picker a symptom of them.

## Partial answer to this ticket's question, from the API 37 side Two of the API 37 aborts are provably **one** bug, not two — details and evidence on #108. The crash the `E2E_DISABLE_SYSTEM_UI` mitigation already handles and the crash in this ticket's third mode are the same assertion in the same emulator mapper: ``` Abort message: 'Assertion failed: !rcEnc->featureInfo()->hasReadColorBufferDma' ... GoldfishMapper::readFromHost ``` They differ only in who calls it — `RegionSamplingThread` (SystemUI's nav-bar luma sampler, which the mitigation removes by deleting the package) versus `TaskSnapshotPer` (WindowManager, inside `system_server`, where there is no package to remove). All 8 recorded #108 legs show `RegionSamplingThread=0`, so they are entirely the second caller. **What this does and does not settle.** It supports "one cause" for the API 37 *aborts*. It says nothing about the picker failures — I retracted that claim above, and the reason still stands: one of the two post-#96 API 37 picker failures had no abort trace at all. A shared root for the aborts does not retroactively make the picker a symptom of them.
JMR-dev commented 2026-08-26 02:35:22 +00:00 (Migrated from github.com)

A fourth mode, on a docs-only PR

2026-08-26T02:07-02:30Z, run 32922652540, E2E API 34, on #115 — a change to CLAUDE.md and
nothing else, so no test in the suite could have been affected by the diff.

Starting 59 tests on test(AVD) - 14
... 52/59 completed. (2 skipped) (0 failed)
  expected: 59   received: 59   failed: unknown   completed cleanly: yes
ERROR        | bad color buffer handle 375
##[warning]E2E api34 WEDGED (api34)

Not an abort, not an assertion. All 59 tests ran, and then the leg hung in teardown until
WEDGE_TIMEOUT killed it at ~22 minutes. Gradle never printed a summary line, which is why the count
is unknown.

bad color buffer handle 375 two seconds later is the emulator's own complaint, and it puts this in
the same family as the #108 aborts — the hasReadColorBufferDma assertion is also a colour-buffer
readback fault in the goldfish mapper. I am not claiming they are the same bug; I made that
mistake once on this ticket already. What is fair to say is that three of this ticket's now-four
modes name the emulator's colour-buffer path, and none of them names anything in this app.

So the running list for this ticket is:

mode where status
SAF picker failing 33-36 (fixed by #96), still on 37 cause on 37 unresolved
Media3 transcode timeout 34, 36 open
system_server abort 37 only #108 — same gralloc assertion as the mitigated surfaceflinger crash, different caller
post-test wedge 34 new; filed the reporting half as #118

The only reason this one was legible is the report from #111 — before it, a wedged leg was an
anonymous red with no counts. The table did overstate it (completed cleanly: yes on a leg the wedge
had just killed), which is #118 and is a reporting fix, not a pass/fail change.

## A fourth mode, on a docs-only PR `2026-08-26T02:07-02:30Z`, run `32922652540`, **E2E API 34**, on #115 — a change to `CLAUDE.md` and nothing else, so no test in the suite could have been affected by the diff. ``` Starting 59 tests on test(AVD) - 14 ... 52/59 completed. (2 skipped) (0 failed) expected: 59 received: 59 failed: unknown completed cleanly: yes ERROR | bad color buffer handle 375 ##[warning]E2E api34 WEDGED (api34) ``` Not an abort, not an assertion. **All 59 tests ran, and then the leg hung in teardown** until `WEDGE_TIMEOUT` killed it at ~22 minutes. Gradle never printed a summary line, which is why the count is `unknown`. `bad color buffer handle 375` two seconds later is the emulator's own complaint, and it puts this in the same family as the #108 aborts — the `hasReadColorBufferDma` assertion is also a colour-buffer readback fault in the goldfish mapper. **I am not claiming they are the same bug**; I made that mistake once on this ticket already. What is fair to say is that three of this ticket's now-four modes name the emulator's colour-buffer path, and none of them names anything in this app. So the running list for this ticket is: | mode | where | status | |---|---|---| | SAF picker failing | 33-36 (fixed by #96), still on 37 | cause on 37 unresolved | | Media3 transcode timeout | 34, 36 | open | | `system_server` abort | 37 only | #108 — same gralloc assertion as the mitigated surfaceflinger crash, different caller | | **post-test wedge** | **34** | **new; filed the reporting half as #118** | The only reason this one was legible is the report from #111 — before it, a wedged leg was an anonymous red with no counts. The table did overstate it (`completed cleanly: yes` on a leg the wedge had just killed), which is #118 and is a reporting fix, not a pass/fail change.
JMR-dev commented 2026-08-26 03:24:32 +00:00 (Migrated from github.com)

The wedge mode has recurred, on a different API level

Second sighting: 2026-08-26T03:23Z, PR #117, E2E API 33 (the first was API 34 on #115). Same
shape — expected: 60, received: 60, failed: unknown, then WEDGED. No abort trace at all this
time, so it is not #108 in disguise.

That makes the mode table for this ticket:

mode seen on notes
SAF picker failing 33-36 (fixed by #96), 37 still cause on 37 unresolved
Media3 transcode timeout 34, 36 open
system_server abort 37 only #108 — same gralloc assertion as the mitigated surfaceflinger crash, different caller
post-test wedge 34, 33 no abort trace; #118 covers the reporting half only

The wedge is now the second mode confirmed on more than one API level, and the second whose evidence
came from a PR that could not plausibly have caused it (#115 was docs-only; #117 touches validation
suggestions and nothing near teardown).

Still not claiming a single cause. Three of the four modes name the emulator's colour-buffer path
and this one does not name anything yet — the wedge diagnostics artifact is uploaded per occurrence
(e2e-wedge-api33, e2e-wedge-api34) and nobody has read one. That is the obvious next evidence,
and it is cheap: two artifacts already exist.

## The wedge mode has recurred, on a different API level Second sighting: `2026-08-26T03:23Z`, PR #117, **E2E API 33** (the first was API 34 on #115). Same shape — `expected: 60, received: 60, failed: unknown`, then `WEDGED`. **No abort trace at all** this time, so it is not #108 in disguise. That makes the mode table for this ticket: | mode | seen on | notes | |---|---|---| | SAF picker failing | 33-36 (fixed by #96), 37 still | cause on 37 unresolved | | Media3 transcode timeout | 34, 36 | open | | `system_server` abort | 37 only | #108 — same gralloc assertion as the mitigated surfaceflinger crash, different caller | | post-test wedge | **34, 33** | no abort trace; #118 covers the reporting half only | The wedge is now the second mode confirmed on more than one API level, and the second whose evidence came from a PR that could not plausibly have caused it (#115 was docs-only; #117 touches validation suggestions and nothing near teardown). **Still not claiming a single cause.** Three of the four modes name the emulator's colour-buffer path and this one does not name anything yet — the wedge diagnostics artifact is uploaded per occurrence (`e2e-wedge-api33`, `e2e-wedge-api34`) and nobody has read one. That is the obvious next evidence, and it is cheap: two artifacts already exist.
JMR-dev commented 2026-08-26 03:27:17 +00:00 (Migrated from github.com)

The fourth mode now has a name, and my description of it was wrong

I read the two wedge diagnostics artifacts. Correcting what I wrote above: I said "every test
ran, then the leg hung in teardown", trusting received: N from the run-shape table. Both wedges
actually end on a test that never returned:

API 34, run 32921547861:  started: thePickedInputSurvivesARealRotation(SafPickerRoundTripTest)   [no finished:]
API 33, run 32924766720:  started: thePickedInputSurvivesARealRotation(SafPickerRoundTripTest)   [no finished:]

Same test, both times, on two PRs that could not have caused it (#115 docs-only, #117 validation
suggestions). Filed as #122.

And it is not this ticket's family. In both artifacts every binder service the wedge probe checks
is found — input, window, activity, media.player. Nothing crashed. That separates it
cleanly from the three colour-buffer modes:

mode seen on emulator crash?
SAF picker failing 33-36 (fixed by #96), 37 still on 37, sometimes
Media3 transcode timeout 34, 36 no
system_server abort (#108) 37 only yes — hasReadColorBufferDma
rotation test hangs (#122) 33, 34 no — framework fully alive

So this ticket's premise — "three modes that look like one cause" — is now two answered questions
and two open ones. #108 is the gralloc assertion with a second caller. #122 is a hanging test, not an
emulator fault at all. What remains genuinely open is the Media3 transcode timeout and why the SAF
picker still fails on 37.

That is a narrowing, not a closure: I would not close this until those two have the same treatment.

## The fourth mode now has a name, and my description of it was wrong I read the two wedge diagnostics artifacts. **Correcting what I wrote above**: I said "every test ran, then the leg hung in teardown", trusting `received: N` from the run-shape table. Both wedges actually end on a test that never returned: ``` API 34, run 32921547861: started: thePickedInputSurvivesARealRotation(SafPickerRoundTripTest) [no finished:] API 33, run 32924766720: started: thePickedInputSurvivesARealRotation(SafPickerRoundTripTest) [no finished:] ``` Same test, both times, on two PRs that could not have caused it (#115 docs-only, #117 validation suggestions). Filed as **#122**. **And it is not this ticket's family.** In both artifacts every binder service the wedge probe checks is `found` — `input`, `window`, `activity`, `media.player`. Nothing crashed. That separates it cleanly from the three colour-buffer modes: | mode | seen on | emulator crash? | |---|---|---| | SAF picker failing | 33-36 (fixed by #96), 37 still | on 37, sometimes | | Media3 transcode timeout | 34, 36 | no | | `system_server` abort (#108) | 37 only | yes — `hasReadColorBufferDma` | | **rotation test hangs (#122)** | **33, 34** | **no — framework fully alive** | So this ticket's premise — "three modes that look like one cause" — is now **two** answered questions and two open ones. #108 is the gralloc assertion with a second caller. #122 is a hanging test, not an emulator fault at all. What remains genuinely open is the Media3 transcode timeout and why the SAF picker still fails on 37. That is a narrowing, not a closure: I would not close this until those two have the same treatment.
JMR-dev commented 2026-08-26 04:46:53 +00:00 (Migrated from github.com)

The API 37 picker question is answered — and my retraction above was wrong, for an instructive reason

Every E2E API 37 gating failure since #96 merged, classified by whether the picker test failed and
whether the gralloc assertion (hasReadColorBufferDma) appears — 14 leg-attempts:

abort present no abort
picker failed 6 0
picker did not fail 6 2

The picker failed / no abort cell is empty. Six for six.

Why I previously said one had no abort

I grepped for TaskSnapshotPer — the thread name from #108 — rather than for the assertion itself.
Re-checking the case I cited as the counterexample (run 32865281555 attempt 2, job 97952658829):

hasReadColorBufferDma : 1
TaskSnapshotPer       : 0
RegionSamplingThread  : 2

It aborted through the other caller — the SystemUI region-sampling path that
E2E_DISABLE_SYSTEM_UI exists to remove, and which evidently did not take on that run. So the abort
was there; my filter could not see it.

That is a real methodological lesson and I would rather write it down than bury it: grepping the
thread name instead of the fault split one bug into two.
#108's title names
TaskSnapshotPersister, and I let the title become the search term. The assertion is the invariant;
the thread is just which caller happened to trip it.

What this settles, and what it does not

Settles: the post-#96 API 37 picker failures are not a #93 regression. #93's cause was an ANR
dialog hiding every window from UiAutomator, fixed in #96 and holding on API 33-36 with zero picker
failures there since. On 37 the picker fails only in the presence of a gralloc abort, which takes the
framework down underneath it — UiAutomation.getWindows() empty, Can't find service: package.

Does not settle: it remains an association, six for six, not a proved mechanism. And it does not
resurrect the causal claim I withdrew earlier in the stronger form I first wrote it: the aborts are
not triggered by the picker test. The diagnostics artifact shows three aborts firing before the
first test starts, and one that the picker test survived and finished 1.1 s later. Ambient fault,
occasional victim — the picker is simply the largest target, being the slowest test in the suite.

Remaining open on this ticket

One mode, not two: the Media3 transcode timeout (transcodesH264ToH265AndReportsProgress, seen on
API 34 and 36, no abort trace). The other three now have names — #108 (gralloc, two callers), #122
(rotation hang, framework fully alive), and this picker association.

## The API 37 picker question is answered — and my retraction above was wrong, for an instructive reason Every `E2E API 37` gating failure since #96 merged, classified by whether the picker test failed and whether the **gralloc assertion** (`hasReadColorBufferDma`) appears — 14 leg-attempts: | | abort present | no abort | |---|---|---| | **picker failed** | **6** | **0** | | picker did not fail | 6 | 2 | **The `picker failed / no abort` cell is empty.** Six for six. ## Why I previously said one had no abort I grepped for `TaskSnapshotPer` — the *thread name* from #108 — rather than for the assertion itself. Re-checking the case I cited as the counterexample (run `32865281555` attempt 2, job `97952658829`): ``` hasReadColorBufferDma : 1 TaskSnapshotPer : 0 RegionSamplingThread : 2 ``` It aborted through the **other** caller — the SystemUI region-sampling path that `E2E_DISABLE_SYSTEM_UI` exists to remove, and which evidently did not take on that run. So the abort was there; my filter could not see it. That is a real methodological lesson and I would rather write it down than bury it: **grepping the thread name instead of the fault split one bug into two.** #108's title names `TaskSnapshotPersister`, and I let the title become the search term. The assertion is the invariant; the thread is just which caller happened to trip it. ## What this settles, and what it does not **Settles:** the post-#96 API 37 picker failures are **not a #93 regression**. #93's cause was an ANR dialog hiding every window from UiAutomator, fixed in #96 and holding on API 33-36 with zero picker failures there since. On 37 the picker fails only in the presence of a gralloc abort, which takes the framework down underneath it — `UiAutomation.getWindows()` empty, `Can't find service: package`. **Does not settle:** it remains an association, six for six, not a proved mechanism. And it does not resurrect the causal claim I withdrew earlier in the stronger form I first wrote it: the aborts are **not triggered by the picker test**. The diagnostics artifact shows three aborts firing before the first test starts, and one that the picker test survived and finished 1.1 s later. Ambient fault, occasional victim — the picker is simply the largest target, being the slowest test in the suite. ## Remaining open on this ticket One mode, not two: **the Media3 transcode timeout** (`transcodesH264ToH265AndReportsProgress`, seen on API 34 and 36, no abort trace). The other three now have names — #108 (gralloc, two callers), #122 (rotation hang, framework fully alive), and this picker association.
JMR-dev commented 2026-08-26 05:04:41 +00:00 (Migrated from github.com)

The last unnamed mode has a name too

The transcode failure is Media3's own export watchdog, not an assertion of ours:

Caused by: java.lang.IllegalStateException: Abort: no output sample written in the
  last 25000 milliseconds. DebugTrace: "Tracing disabled"
  at androidx.media3.transformer.Transformer.lambda$maybeInitializeExportWatchdogTimer$0(Transformer.java:1303)

androidx.media3.transformer.ExportException: Muxer error

Transformer arms a 25-second watchdog and aborts the export when the encoder produces no output
sample for that long. So the emulator's encode stalls; Media3 notices and gives up. Nothing in this
app decided anything.

And the test it hits is already known to be broken on one image

transcodesH264ToH265AndReportsProgress is one of the two tests carrying
@FailsOnEmulatorApi37
(Media3EngineTest.kt:72). CLAUDE.md records why: two Media3 hardware
transcodes fail inside the emulator's own c2.goldfish.h264.decoder on the android-37 images.

So the same test that always fails on API 37's emulator codec intermittently fails on 34 and
36 — twice in the census window — with a stall long enough to trip a 25-second watchdog. The
straightforward reading is that this is one emulator-codec weakness expressed at two severities:
deterministic on 37, load-dependent below it. That is an inference from two data points and the
shared marker, not a proof.

What would confirm it: the stall should be visible in logcat as the c2.goldfish encoder
starving, in the e2e-diagnostics-apiNN artifact for a run that hit it. Two such runs exist. That is
the same cheap artifact read that settled #122.

This ticket's question is now answered

All four modes have names, and none of them is an assertion failing about this app's behaviour:

mode what it actually is ticket
SAF picker failing on 37 collateral of the gralloc abort — 6 of 6 picker failures coincide with one #108
system_server abort hasReadColorBufferDma in the goldfish mapper, two callers, ambient #108
post-test wedge on 33/34 the rotation test hangs, framework fully alive #122
Media3 transcode timeout Media3's 25 s export watchdog on a stalled emulator encode this comment

The premise this was filed under — "three modes today that look like one cause" — turned out to be
half right in an unhelpful way. They are not one cause. But they are all the emulator failing
underneath the suite
, in four unrelated subsystems: gralloc, whatever hangs the rotation, and the
goldfish codec. Not one bug; one class of bug.

I would close this in favour of #108 and #122 plus a new ticket for the codec stall if anyone wants
it chased, rather than keep a four-way umbrella open. Leaving that call to the maintainer.

## The last unnamed mode has a name too The transcode failure is **Media3's own export watchdog**, not an assertion of ours: ``` Caused by: java.lang.IllegalStateException: Abort: no output sample written in the last 25000 milliseconds. DebugTrace: "Tracing disabled" at androidx.media3.transformer.Transformer.lambda$maybeInitializeExportWatchdogTimer$0(Transformer.java:1303) androidx.media3.transformer.ExportException: Muxer error ``` `Transformer` arms a 25-second watchdog and aborts the export when the encoder produces no output sample for that long. So the emulator's encode stalls; Media3 notices and gives up. Nothing in this app decided anything. ## And the test it hits is already known to be broken on one image `transcodesH264ToH265AndReportsProgress` is **one of the two tests carrying `@FailsOnEmulatorApi37`** (`Media3EngineTest.kt:72`). CLAUDE.md records why: two Media3 hardware transcodes fail inside the emulator's own `c2.goldfish.h264.decoder` on the android-37 images. So the same test that **always** fails on API 37's emulator codec **intermittently** fails on 34 and 36 — twice in the census window — with a stall long enough to trip a 25-second watchdog. The straightforward reading is that this is one emulator-codec weakness expressed at two severities: deterministic on 37, load-dependent below it. That is an inference from two data points and the shared marker, not a proof. **What would confirm it:** the stall should be visible in logcat as the `c2.goldfish` encoder starving, in the `e2e-diagnostics-apiNN` artifact for a run that hit it. Two such runs exist. That is the same cheap artifact read that settled #122. ## This ticket's question is now answered All four modes have names, and **none of them is an assertion failing about this app's behaviour**: | mode | what it actually is | ticket | |---|---|---| | SAF picker failing on 37 | collateral of the gralloc abort — 6 of 6 picker failures coincide with one | #108 | | `system_server` abort | `hasReadColorBufferDma` in the goldfish mapper, two callers, ambient | #108 | | post-test wedge on 33/34 | the rotation test hangs, framework fully alive | #122 | | **Media3 transcode timeout** | **Media3's 25 s export watchdog on a stalled emulator encode** | this comment | The premise this was filed under — *"three modes today that look like one cause"* — turned out to be half right in an unhelpful way. They are not one cause. But they are all **the emulator failing underneath the suite**, in four unrelated subsystems: gralloc, whatever hangs the rotation, and the goldfish codec. Not one bug; one class of bug. I would close this in favour of #108 and #122 plus a new ticket for the codec stall if anyone wants it chased, rather than keep a four-way umbrella open. Leaving that call to the maintainer.
JMR-dev commented 2026-08-27 03:59:39 +00:00 (Migrated from github.com)

Tally from a batch of eight PRs opened today (#131, #144–#151) — offered because this ticket asks whether the modes share a cause, and a single day's batch is a decent sample of one workload.

Roughly fifty leg-runs. Every failure was on a diff that was docs-only or JVM-test-only, so none of them can be attributed to the change under test.

mode count shape
Wedge (#122) 3 received: N then gradle never returns; killed at 1200s; failed: unknown
SAF picker will not close 1 documentsui in front after 4 back presses; UiAutomation.getWindows() empty
Emulator never came up 1 adb … failed with exit code 224 during setup, before any test ran
Advisory API 37 leg every run by design, per CLAUDE.md

Wedges split API 33 ×1, API 34 ×2 — so not one level, which also corrects something I wrote on #122 earlier from a smaller sample.

What separates them, and it is not subtle: the wedge happens after every test has reported (received: 60), the picker failure is a real assertion inside a running suite, and the adb-224 case happens before the suite starts at all. Those are three different points in the lifecycle, which argues against the single shared cause this ticket floats — unless the shared cause is simply "the emulator on this runner is unreliable in several independent ways".

The practically useful finding: every one of the five retried clean on a plain re-run, no change to the diff. And on a wedged leg the tests themselves completed — what is lost is the verdict, not the coverage.

One trap worth putting in this ticket, since this is where people land when a leg fails: status_check.yml sets concurrency: cancel-in-progress: true, so re-running a job on an older run for the same ref cancels whatever newer run is in flight. gh pr checks then reports every cancelled job as fail, which reads exactly like a build break — nine jobs "failing" in 42 seconds on a one-file docs diff. Check that nothing newer is queued for the ref before retrying, or re-run the newest run instead.

**Tally from a batch of eight PRs opened today** (#131, #144–#151) — offered because this ticket asks whether the modes share a cause, and a single day's batch is a decent sample of one workload. Roughly fifty leg-runs. Every failure was on a diff that was **docs-only or JVM-test-only**, so none of them can be attributed to the change under test. | mode | count | shape | |---|---|---| | Wedge (#122) | 3 | `received: N` then gradle never returns; killed at 1200s; `failed: unknown` | | SAF picker will not close | 1 | `documentsui` in front after 4 back presses; `UiAutomation.getWindows()` empty | | Emulator never came up | 1 | `adb … failed with exit code 224` during setup, before any test ran | | Advisory API 37 leg | every run | by design, per `CLAUDE.md` | Wedges split **API 33 ×1, API 34 ×2** — so not one level, which also corrects something I wrote on #122 earlier from a smaller sample. **What separates them, and it is not subtle:** the wedge happens *after* every test has reported (`received: 60`), the picker failure is a real assertion inside a running suite, and the adb-224 case happens before the suite starts at all. Those are three different points in the lifecycle, which argues against the single shared cause this ticket floats — unless the shared cause is simply "the emulator on this runner is unreliable in several independent ways". **The practically useful finding:** every one of the five retried clean on a plain re-run, no change to the diff. And on a wedged leg the tests themselves completed — what is lost is the verdict, not the coverage. **One trap worth putting in this ticket**, since this is where people land when a leg fails: `status_check.yml` sets `concurrency: cancel-in-progress: true`, so re-running a job on an **older** run for the same ref cancels whatever newer run is in flight. `gh pr checks` then reports every cancelled job as `fail`, which reads exactly like a build break — nine jobs "failing" in 42 seconds on a one-file docs diff. Check that nothing newer is queued for the ref before retrying, or re-run the newest run instead.
JMR-dev commented 2026-08-27 04:16:17 +00:00 (Migrated from github.com)

Correcting my own tally above. The line "every one of the five retried clean on a plain re-run" was wrong when I posted it, and it was the one practically useful sentence in the comment, so it is worth fixing rather than leaving.

#146 (test/container-capabilities-audio) needed two re-runs, not one. Run 33036437394 reached attempt 3, and the gating E2E API 37 leg failed on attempts 1 and 2:

attempt gating E2E API 37 mode
1 failure —
2 failure adb … failed with exit code 224, 03:45:52Z — emulator never came up
3 success —

So the "Emulator never came up" row in my table is 2, not 1, and that mode has now repeated on a re-run of the same ref — which is a mildly interesting datum for this ticket, because it is the one mode that happens before the suite starts and therefore cannot be blamed on anything the tests do.

What survives unchanged: every failure was still on a docs-only or JVM-test-only diff, the three modes still sit at three different points in the lifecycle, and all eight branches are green now. What does not survive is "one re-run clears it" — for one branch it took two, and the retry reproduced the original mode rather than a new one.

(#150's single retry did come back clean; attempt 2 succeeded. I checked both rather than assume the correction generalised.)

**Correcting my own tally above.** The line "every one of the five retried clean on a plain re-run" was wrong when I posted it, and it was the one practically useful sentence in the comment, so it is worth fixing rather than leaving. **#146 (`test/container-capabilities-audio`) needed two re-runs, not one.** Run `33036437394` reached **attempt 3**, and the gating `E2E API 37` leg failed on attempts 1 *and* 2: | attempt | gating `E2E API 37` | mode | |---|---|---| | 1 | failure | — | | 2 | failure | `adb … failed with exit code 224`, 03:45:52Z — emulator never came up | | 3 | success | — | So the "Emulator never came up" row in my table is **2, not 1**, and that mode has now repeated on a re-run of the same ref — which is a mildly interesting datum for this ticket, because it is the one mode that happens before the suite starts and therefore cannot be blamed on anything the tests do. What survives unchanged: every failure was still on a docs-only or JVM-test-only diff, the three modes still sit at three different points in the lifecycle, and all eight branches are green now. What does not survive is "one re-run clears it" — for one branch it took two, and the retry reproduced the original mode rather than a new one. (#150's single retry did come back clean; attempt 2 succeeded. I checked both rather than assume the correction generalised.)
JMR-dev commented 2026-08-27 04:32:50 +00:00 (Migrated from github.com)

One more data point, same batch. #148's re-push (a one-file, JVM-test-only diff) hit the SAF picker mode again on the gating E2E API 37 leg:

SafPickerRoundTripTest > pickingAFileThroughTheSystemPickerFillsInTheFileCard
  the system picker would not close: after 4 back presses the app still does not
  have the window focus, and com.google.android.documentsui is in front.
  What could be seen: no windows at all (UiAutomation.getWindows() is empty)

  expected: 57   received: 57   failed: 1   completed cleanly: yes

Attempt 2 came back clean, so the running tally for this batch is now:

mode count retried clean first try?
Wedge (#122) 3 yes
SAF picker will not close 2 yes, both
Emulator never came up (adb 224) 2 no — #146 needed two re-runs

Each of the three modes has now repeated within a single day's batch, which is the part worth recording. Whatever this ticket concludes about a shared cause, none of the three is a one-off.

The shape row is doing its job, incidentally: received: 57, failed: 1, completed cleanly: yes said in one line that all 57 gating tests ran and exactly one asserted false — which is what separates this mode from the wedge (where the tests complete but the verdict is lost) and from adb-224 (where the suite never starts). Reading the marker rather than the job name, per CLAUDE.md.

**One more data point, same batch.** #148's re-push (a one-file, JVM-test-only diff) hit the **SAF picker** mode again on the gating `E2E API 37` leg: ``` SafPickerRoundTripTest > pickingAFileThroughTheSystemPickerFillsInTheFileCard the system picker would not close: after 4 back presses the app still does not have the window focus, and com.google.android.documentsui is in front. What could be seen: no windows at all (UiAutomation.getWindows() is empty) expected: 57 received: 57 failed: 1 completed cleanly: yes ``` Attempt 2 came back clean, so the running tally for this batch is now: | mode | count | retried clean first try? | |---|---|---| | Wedge (#122) | 3 | yes | | SAF picker will not close | **2** | yes, both | | Emulator never came up (`adb` 224) | 2 | **no** — #146 needed two re-runs | **Each of the three modes has now repeated within a single day's batch**, which is the part worth recording. Whatever this ticket concludes about a shared cause, none of the three is a one-off. The shape row is doing its job, incidentally: `received: 57, failed: 1, completed cleanly: yes` said in one line that all 57 gating tests ran and exactly one asserted false — which is what separates this mode from the wedge (where the tests complete but the verdict is lost) and from adb-224 (where the suite never starts). Reading the marker rather than the job name, per `CLAUDE.md`.
JMR-dev commented 2026-08-29 16:07:12 +00:00 (Migrated from github.com)

A mode this tally has not recorded before: a Media3 hardware transcode failing on API 34.

From #165, run 33261618358 attempt 1:

E2E API 34 | Media3EngineTest > transcodesH264ToH265AndReportsProgress
E2E API 34 |   expected: 60   received: 60   failed: 1   completed cleanly: yes
E2E API 37 |   adb ... failed with exit code 224

Two legs, two different modes, one run.

Why the API 34 one is worth a line here. CLAUDE.md records Media3 hardware transcodes failing "inside the emulator's own c2.goldfish.h264.decoder" as an API 37 property — two of the three @FailsOnEmulatorApi37 tests are exactly that. This is the same class of failure on API 34, which is a gating leg with no marker and no allowance.

If it recurs, the interesting question is whether the marker's premise is level-specific or whether API 34 has simply been lucky. One occurrence is not that evidence, which is why this is a note rather than a claim.

Attribution, checked rather than assumed: #165's diff extracts an actions builder from two screen composables and adds two JVM test files. Neither screen references Media3Engine or Transformer, and Unit tests and Static analysis both passed on the same attempt. It was worth checking regardless — this is the first change in wave 2 to touch main source that the instrumented suite actually drives, so "test-only diff" was no longer available as a shortcut.

Running tally for the two batches: wedge ×4, SAF picker ×3, adb-224 ×3, Media3-on-34 ×1.

**A mode this tally has not recorded before: a Media3 hardware transcode failing on API 34.** From #165, run `33261618358` attempt 1: ``` E2E API 34 | Media3EngineTest > transcodesH264ToH265AndReportsProgress E2E API 34 | expected: 60 received: 60 failed: 1 completed cleanly: yes E2E API 37 | adb ... failed with exit code 224 ``` Two legs, two different modes, one run. **Why the API 34 one is worth a line here.** `CLAUDE.md` records Media3 hardware transcodes failing "inside the emulator's own `c2.goldfish.h264.decoder`" as an **API 37** property — two of the three `@FailsOnEmulatorApi37` tests are exactly that. This is the same class of failure on **API 34**, which is a *gating* leg with no marker and no allowance. If it recurs, the interesting question is whether the marker's premise is level-specific or whether API 34 has simply been lucky. One occurrence is not that evidence, which is why this is a note rather than a claim. **Attribution, checked rather than assumed:** #165's diff extracts an actions builder from two screen composables and adds two JVM test files. Neither screen references `Media3Engine` or `Transformer`, and `Unit tests` and `Static analysis` both passed on the same attempt. It was worth checking regardless — this is the first change in wave 2 to touch `main` source that the instrumented suite actually drives, so "test-only diff" was no longer available as a shortcut. Running tally for the two batches: wedge ×4, SAF picker ×3, adb-224 ×3, Media3-on-34 ×1.
JMR-dev commented 2026-09-06 00:29:09 +00:00 (Migrated from github.com)

Still live, on a PR whose production diff cannot cause it

PR #218, run 34000816016, API 34 gating leg, job 101399258547:

Media3EngineTest > transcodesH264ToH265AndReportsProgress[test(AVD) - 14] FAILED
  androidx.media3.transformer.ExportException: Muxer error
    at androidx.media3.transformer.Transformer.lambda$maybeInitializeExportWatchdogTimer$0(Transformer.java:1301)

Same test, same exception, same API 34 gating leg as this ticket's opening report.

  expected: 60   received: 60   failed: 1   completed cleanly: yes

and, later in the same log, ERROR | Failed to find ColorBuffer: 288 — the graphics-stack noise
that keeps showing up around these.

Why this occurrence is worth adding

The production diff on #218 is: making LibreMediaConverterApp open, adding a protected open val sweepScope that still resolves to Dispatchers.IO in production, and publishing the Job that
onCreate already started. Nothing there touches Media3, a codec, a muxer, or a surface. It is about
as clean a demonstration as this ticket is going to get that the failure is independent of the change
under test — which is the thing #190 says costs a human 40 minutes per occurrence to re-establish.

Also worth noting against the "either the marker is too narrow or the API 36 occurrence is rarer"
question in #190: this is API 34, so the spread is now 34, 36 and 37 for the same test.

## Still live, on a PR whose production diff cannot cause it PR #218, run `34000816016`, **API 34 gating leg**, job `101399258547`: ``` Media3EngineTest > transcodesH264ToH265AndReportsProgress[test(AVD) - 14] FAILED androidx.media3.transformer.ExportException: Muxer error at androidx.media3.transformer.Transformer.lambda$maybeInitializeExportWatchdogTimer$0(Transformer.java:1301) ``` Same test, same exception, same API 34 gating leg as this ticket's opening report. ``` expected: 60 received: 60 failed: 1 completed cleanly: yes ``` and, later in the same log, `ERROR | Failed to find ColorBuffer: 288` — the graphics-stack noise that keeps showing up around these. ### Why this occurrence is worth adding The production diff on #218 is: making `LibreMediaConverterApp` `open`, adding a `protected open val sweepScope` that still resolves to `Dispatchers.IO` in production, and publishing the `Job` that `onCreate` already started. Nothing there touches Media3, a codec, a muxer, or a surface. It is about as clean a demonstration as this ticket is going to get that the failure is independent of the change under test — which is the thing #190 says costs a human 40 minutes per occurrence to re-establish. Also worth noting against the "either the marker is too narrow or the API 36 occurrence is rarer" question in #190: this is **API 34**, so the spread is now 34, 36 and 37 for the same test.
JMR-dev commented 2026-09-07 21:20:55 +00:00 (Migrated from github.com)

Sighting 2026-09-07 — API 35, PR #269, NPE: Activity has been destroyed

Recording this because it cost a re-run and would otherwise be re-diagnosed from scratch. Evidence
re-read from the raw log rather than relayed.

Run 34161043035, job 101862812172, E2E API 35. Red once, green on re-run
(101864620664) — same commit, no code change. APIs 33, 34, 36 and 37 all passed on that same
commit, and a local run-e2e.sh 35 on the same tree was 72/72/0.

20:59:36.917  ERROR | Failed to find ColorBuffer: 512
20:59:43.172  java.lang.NullPointerException: Cannot run onActivity since Activity has been
                     destroyed already
              at androidx.test.internal.util.Checks.checkNotNull(Checks.java:50)
              at androidx.test.core.app.ActivityScenario.lambda$onActivity$2(ActivityScenario.java:792)
              at android.app.Instrumentation$SyncRunnable.run(Instrumentation.java:2612)
              at android.os.Looper.loopOnce(Looper.java:232)
              at android.os.Looper.loop(Looper.java:317)
20:59:46.803  ERROR | Failed to find ColorBuffer: 562

The stack contains no frame from any test file. It is the main looper's, and the Activity was
already destroyed when onActivity asked for it — the shape this issue's title describes as failing
on system services rather than assertions, bracketed seven seconds either side by the colour-buffer
path this issue already names in three of its four modes. The suite ran to completion:
expected 72, received 72, failed 1, completed cleanly: yes.

The honest caveat

The failing test was SafPickerRoundTripTest.aFailedSaveDeletesTheDocumentItCouldNotWrite, which
PR #269 modifies.
So this is not a clean "unrelated test" sighting, and it should not be counted as
one. What argues it is this issue's class rather than that diff: no test frame in the stack, the
colour-buffer brackets, green on re-run of the identical commit, green on the other four API levels,
and green locally at API 35. What argues for caution: a pattern this well known is exactly what a
real regression could hide behind. It was re-run for that reason rather than waved through.

One process note worth recording

A re-run destroys the evidence. Job 101862812172 now serves the re-run's log — the original
20:59 failure is not retrievable at that job id any more. Anyone diagnosing one of these should
capture the log before pressing re-run, or the sighting cannot be checked afterwards. The excerpt
above survives only because it was saved locally at the time.

## Sighting 2026-09-07 — API 35, PR #269, `NPE: Activity has been destroyed` Recording this because it cost a re-run and would otherwise be re-diagnosed from scratch. Evidence re-read from the raw log rather than relayed. **Run `34161043035`, job `101862812172`, E2E API 35.** Red once, green on re-run (`101864620664`) — **same commit, no code change**. APIs 33, 34, 36 and 37 all passed on that same commit, and a local `run-e2e.sh 35` on the same tree was 72/72/0. ``` 20:59:36.917 ERROR | Failed to find ColorBuffer: 512 20:59:43.172 java.lang.NullPointerException: Cannot run onActivity since Activity has been destroyed already at androidx.test.internal.util.Checks.checkNotNull(Checks.java:50) at androidx.test.core.app.ActivityScenario.lambda$onActivity$2(ActivityScenario.java:792) at android.app.Instrumentation$SyncRunnable.run(Instrumentation.java:2612) at android.os.Looper.loopOnce(Looper.java:232) at android.os.Looper.loop(Looper.java:317) 20:59:46.803 ERROR | Failed to find ColorBuffer: 562 ``` **The stack contains no frame from any test file.** It is the main looper's, and the Activity was already destroyed when `onActivity` asked for it — the shape this issue's title describes as failing on system services rather than assertions, bracketed seven seconds either side by the colour-buffer path this issue already names in three of its four modes. The suite ran to completion: `expected 72, received 72, failed 1, completed cleanly: yes`. ### The honest caveat **The failing test was `SafPickerRoundTripTest.aFailedSaveDeletesTheDocumentItCouldNotWrite`, which PR #269 modifies.** So this is not a clean "unrelated test" sighting, and it should not be counted as one. What argues it is this issue's class rather than that diff: no test frame in the stack, the colour-buffer brackets, green on re-run of the identical commit, green on the other four API levels, and green locally at API 35. What argues for caution: a pattern this well known is exactly what a real regression could hide behind. It was re-run for that reason rather than waved through. ### One process note worth recording **A re-run destroys the evidence.** Job `101862812172` now serves the *re-run's* log — the original 20:59 failure is not retrievable at that job id any more. Anyone diagnosing one of these should capture the log before pressing re-run, or the sighting cannot be checked afterwards. The excerpt above survives only because it was saved locally at the time.
JMR-dev commented 2026-09-07 21:59:56 +00:00 (Migrated from github.com)

This ticket's own mode has a measured cause, and it is not starvation

transcodesH264ToH265AndReportsProgress fails because the emulator's Codec2 HAL process
segfaults.
Read out of the per-test logcat in e2e-report-api34 for run 34000816016
attempt 1 — the API 34 gating leg of 2026-09-06, 62 ms after the test starts:

00:20:39.752 I TestRunner: started: transcodesH264ToH265AndReportsProgress
00:20:39.757 I TransformerInternal: Init [AndroidXMedia3/1.11.0] [emu64xa, sdk_gphone64_x86_64, 34]
00:20:39.814 D MediaCodec: MediaCodec::reclaim(0x7f01432ddb40) c2.goldfish.h264.decoder
00:20:39.822 F DEBUG : Cmdline: /vendor/bin/hw/android.hardware.media.c2@1.0-service-goldfish
00:20:39.822 F DEBUG : pid: 349, tid: 6079, name: sh.h264.decoder
00:20:39.822 F DEBUG : signal 0 (SIGSEGV), code 1 (SEGV_MAPERR)
00:20:39.822 F DEBUG : Cause: null pointer dereference
  #00 C2Block2D::handle() const+4                       /vendor/lib64/libcodec2_vndk.so
  #01 getClientUsage(std::shared_ptr<C2BlockPool> const&)  /vendor/lib64/libcodec2_goldfish_common.so
  #02 android::C2GoldfishAvcDec::process(...)           /vendor/lib64/libcodec2_goldfish_avcdec.so
  #03 android::SimpleC2Component::processQueue()+2888   /vendor/lib64/libcodec2_goldfish_common.so
00:20:39.830 E tombstoned: Tombstone written to: tombstone_00
00:20:39.839 E CCodec  : Codec2 component "c2.goldfish.h264.decoder" died.
00:20:39.846 E MediaCodec: Codec reported err 0xffffffe0/DEAD_OBJECT ... while in state 10/RELEASING
00:20:39.911 D android.hardware.media.c2@1.0-service-goldfish: Goldfish C2 Service starting...

The decoder HAL dies and respawns; Media3 is left with a DEAD_OBJECT codec, writes no output
sample, and its 25-second export watchdog aborts the export — which is the
ExportException: Muxer error this ticket was filed on. Nothing in this app decided anything,
and it is not the runner starving a working codec: it is a null-pointer dereference in the
emulator's own vendor codec library.

It is 6 for 6

Every gating leg-attempt that has ever failed this way carries the same HAL crash in the same
job's log — grepped for c2@1.0-service-goldfish in the --- native crashes (tail 60) --- dump
that e2e-run.sh writes on failure:

run / attempt leg date
32855014836 a1 API 36 2026-08-25
32857067112 a1 API 34 2026-08-25
32919928048 a1 API 36 2026-08-26
33261618358 a1 API 34 2026-08-29
33588264439 a1 API 36 2026-09-02
34000816016 a1 API 34 2026-09-06

Six occurrences in 1210 gating leg-attempts on API 33-36 across 2026-08-20..09-07 — 0.5%,
and 0/301 on API 33, 3/302 on API 34, 0/304 on API 35, 3/303 on API 36. (The API 37 row does not
run this test, so it is not in the denominator. Two other failures of this test are excluded and
said so below.)

What this settles, and what it does not

Settles: the comment above that called this "one emulator-codec weakness
expressed at two severities: deterministic on 37, load-dependent below it" was right about the
subsystem
, and the @FailsOnEmulatorApi37 marker's stated reason — "fail inside the emulator's
own c2.goldfish.h264.decoder" — is the same vendor HAL. It is one weakness, and this is its
33-36 expression.

Does not settle: why the HAL dereferences null. MediaCodec::reclaim is logged 8 ms before
the crash, and a reclaim is the resource manager taking a codec instance away from a client — so
"a reclaim races C2GoldfishAvcDec::process and the block pool is gone under it" is the obvious
hypothesis and is untested. I am recording it as a hypothesis and not acting on it, because
this ticket has twice published a causal claim its own data did not support.

Not a candidate for a fix here. The crashing code is /vendor/lib64/* inside the system image.
This is environmental, in the same sense @FailsOnEmulatorApi37 is, and now with the same kind of
evidence behind it.

Two failures of the same test that are not this mode

Anchoring on the test name alone would have counted eight. Excluded, with the reason:

  • 32545625459 a1 (API 37, 2026-08-22) — 37 tests failed together, framework down, gralloc
    assertion present. That is #108, and this test was collateral.
  • 32669190757 a1 (API 35, 2026-08-23) — predates #111's shape report, and the log carries the
    bare FAILED marker with no message at all. Unclassifiable, not classified.

Live census of every gating leg-attempt, 2026-08-20 .. 2026-09-07

Per leg-attempt, as this ticket requires; cancelled legs excluded (those are the
cancel-in-progress cancellations the tally above warns about, not runs). 1489 gating
leg-attempts, 129 failures, 8.7%.
Classified by anchoring the failing test name to the FAILED
marker in each failing job's own log, then by the message under it.

mode all history since 2026-08-27 last 2 days
SAF picker (will not close / never showed / rotation) 52 17 7
framework down, no test named or test as collateral 25 11 4
other named-test failure 16 2 1
wedge — gradle never returned 9 6 1
Media3 export watchdog (this ticket) 6 3 1
SAF save tests — Activity gone / composition unreachable 7 7 7
emulator never came up (adb 224) 3 3 1
unclassified (no test named, no signature) 9 2 0
denominator (gating leg-attempts) 1489 796 387

The mode column sums to 127, not 129: the two transcode failures excluded above are left out of it rather than filed under a mode they are not.

Per-mode denominators are not the same denominator, which is the counting lesson from doing
this: three of these modes cannot occur on all five rows. The save tests carry
@FailsOnEmulatorApi37 and only existed from 2026-09-06T15:12, so their rate is 7 in 88
leg-attempts on API 33-36 — not 7 in 1489. The wedge has only ever happened on API 33/34.

Per-mode disposition

mode disposition evidence
Media3 export watchdog environmental, measured — goldfish Codec2 HAL SIGSEGV, 6/6 above
API 37 picker + framework-down #108 — 24 of the 30 API 37 failures since 2026-08-27 carry hasReadColorBufferDma; of the 6 that do not, 3 are adb 224, 2 have no signature at all, and 1 has the framework-down signature without the assertion reaching the crash tail job logs
wedge consistent with #219 having fixed it, not demonstrated — 9 occurrences in 497 API 33/34 leg-attempts before 2026-09-06T02:00Z (1.8%), 0 in the 106 since. The last one (34001741668, 2026-09-06T00:36) is thePickedInputSurvivesARealRotation again, and git merge-base --is-ancestor 32ab54d <head> says that head does not contain #219's fix. At the prior rate, P(0 in 106) ≈ 0.15, so this is suggestive and not yet evidence e2e-wedge-api34 of 34001741668: started: the rotation test, no finished:, every binder service found
adb 224 infra, before the suite starts — 3 occurrences, all API 37, nothing to attribute job logs
SAF save tests at least four distinct shapes, three of them already fixed or pre-fix — split below below

The newest mode splits into four, and only one of the seven is on today's main

The @FailsOnEmulatorApi37-marked save tests (#226, #250) landed on 2026-09-06 and account for
every gating failure on API 33-36 since. Filed as one mode, they are not one:

run / attempt head message shape
34041593697 a1 (35) fa10d94 — the commit that added the test IllegalStateException: No compose hierarchies found, thrown directly pre-b23ff0f
34056545386 a1 (35) cbbaf74 (#250) ComposeTimeoutException ... after 120000 ms the CONVERSION_TIMEOUT_MS case 19e3539 fixed. Not a system-service failure at all — API 35's software encode measured 134.8 s against a 120 s bound
34057706195 a1 (34) 19e3539 No compose hierarchies found ×2, thrown directly contains b23ff0f, so that fix did not close this shape
34067653670 a1 (35), 34146936252 a1 (35) d45abe7, 73482520 waited 300000ms for a node tagged action.saveFile, no composition error awaitNode only appends the composition error when fetchSemanticsNodes threw, so the composition was readable throughout. Confirmed in 34067653670's per-test logcat: MainActivity is RESUMED at 23:49:52.981 and stays resumed until the rule tears it down 300 s later, and no conversion runs at all in that window. Open
34146936252 a2 (35) 73482520 waited 300000ms ...; last composition error: No compose hierarchies app PAUSED at 17:38:30.522 and never resumes; the back press at 17:38:32.479 follows Waiting 5000ms for ... com.google.android.permissioncontroller, i.e. the permission-dialog helper #269 replaced with pm grant
34161043035 a1 (35) 9fd96d0 — #269 itself NPE: Cannot run onActivity since Activity has been destroyed already traced to a mechanism; see below

So "mode 5 at 8% of leg-attempts" would be wrong — most of these are shapes their own
follow-up commit had already answered, on heads that predate it.

The one on current main has a mechanism, and it is #93's launcher ANR with a new victim

From the per-test logcat in e2e-report-api35 of 34161043035 attempt 1 (attempt 2 was
green; the attempt-1 artifact survives with its own id, which is how it can still be read):

20:59:36.033  MainActivity RESUMED          <- the save picker has returned; the app is in front
20:59:37.068  UiDevice: Pressing back button.
20:59:38.092  InteractionController: Timed out waiting 1000ms for command and events.
20:59:41.094  UiDevice: Retrieving node with selector: BySelector [RES='android:id/aerr_wait']
20:59:41.100  UiObject2: Clicking on (540, 1359).
20:59:41.169  Input channel object '9128b0b Application Not Responding:
                com.google.android.apps.nexuslauncher (client)' was disposed
20:59:41.713  UiDevice: Pressing back button.
20:59:41.754  TopTaskTracker: onTaskMovedToFront: ... cmp=com.google.android.apps.nexuslauncher
20:59:41.755  MainActivity PAUSED
20:59:42.224  MainActivity STOPPED
20:59:42.278  MainActivity DESTROYED
20:59:42.599  TestRunner: failed

dismissThePicker's loop is:

repeat(BACK_PRESSES) {
    if (awaitAppFocus()) return
    dismissASystemErrorDialog()
    device.pressBack()
}

awaitAppFocus() is composeRule.activity.hasWindowFocus(). The launcher's ANR dialog makes
that false too
— it is a fullscreen system_server window, which is the whole reason
dismissASystemErrorDialog exists. So the guard reads false, the dialog is removed at
41.169, and the back press at 41.713 is then sent on a reading taken before the occluder
was removed — into an app that is already in front with nothing to go back to. It finishes
MainActivity, the launcher comes to the front, and the rest of the test has no Activity.

The method's own KDoc names this hazard — "a third from Recent would finish MainActivity and
take the rest of the test with it"
— and the guard is what is supposed to prevent it. It cannot,
because it cannot tell "the picker is still up" from "a system dialog is on top".

requireAReadableScreen, twenty lines away, already does the right thing: it re-reads after
dismissASystemErrorDialog() before doing anything else. dismissThePicker does not. PR to
follow — the fix is to re-read the focus after removing a dialog and skip the back press if the
app already has it, which removes an action rather than retrying one.

One more finding, recorded rather than acted on

ERROR_DIALOG_BUTTONS ends with android:id/button1, the framework's generic AlertDialog
positive button — not an app-error-dialog id. In the same trace, at 20:59:35.689, after
aerr_wait and aerr_close both missed, button1 was found and clicked at (927, 2274), and
MainActivity resumed 339 ms later: dismissASystemErrorDialog clicked a button inside
DocumentsUI's own save dialog
, believing it to be a system error dialog. It happened to complete
the save the walk had just failed to complete. There is no reason it always would.


Method notes, since this ticket is where people land

  • A re-run destroys the log. gh run view --job <id> --log resolves by run and serves the
    latest attempt. Read a historical attempt with
    gh api --allow-escape-sequences /repos/{owner}/{repo}/actions/jobs/{job_id}/logs, taking the
    job id from /actions/runs/{run}/attempts/{n}/jobs. That is the correction already on this
    ticket, and it still holds.
  • Artifacts survive a re-run; the API just makes them look like they do not.
    /actions/runs/{run}/artifacts returns every attempt's upload under the same name with
    different ids and created_at. Take the id whose created_at falls in the attempt's window —
    gh run download takes the newest, which on a re-run-to-green is the green one.
  • The per-test logcats are the evidence, not the job log. e2e-report-apiNN carries
    outputs/androidTest-results/connected/debug/<device>/logcat-<class>-<method>.txt, one file per
    test, plus the JUnit XML with the untruncated stack. The job log truncates a stack to its first
    frame. Everything above came out of those files.
## This ticket's own mode has a measured cause, and it is not starvation **`transcodesH264ToH265AndReportsProgress` fails because the emulator's Codec2 HAL process segfaults.** Read out of the per-test logcat in `e2e-report-api34` for run `34000816016` attempt 1 — the API 34 gating leg of 2026-09-06, 62 ms after the test starts: ``` 00:20:39.752 I TestRunner: started: transcodesH264ToH265AndReportsProgress 00:20:39.757 I TransformerInternal: Init [AndroidXMedia3/1.11.0] [emu64xa, sdk_gphone64_x86_64, 34] 00:20:39.814 D MediaCodec: MediaCodec::reclaim(0x7f01432ddb40) c2.goldfish.h264.decoder 00:20:39.822 F DEBUG : Cmdline: /vendor/bin/hw/android.hardware.media.c2@1.0-service-goldfish 00:20:39.822 F DEBUG : pid: 349, tid: 6079, name: sh.h264.decoder 00:20:39.822 F DEBUG : signal 0 (SIGSEGV), code 1 (SEGV_MAPERR) 00:20:39.822 F DEBUG : Cause: null pointer dereference #00 C2Block2D::handle() const+4 /vendor/lib64/libcodec2_vndk.so #01 getClientUsage(std::shared_ptr<C2BlockPool> const&) /vendor/lib64/libcodec2_goldfish_common.so #02 android::C2GoldfishAvcDec::process(...) /vendor/lib64/libcodec2_goldfish_avcdec.so #03 android::SimpleC2Component::processQueue()+2888 /vendor/lib64/libcodec2_goldfish_common.so 00:20:39.830 E tombstoned: Tombstone written to: tombstone_00 00:20:39.839 E CCodec : Codec2 component "c2.goldfish.h264.decoder" died. 00:20:39.846 E MediaCodec: Codec reported err 0xffffffe0/DEAD_OBJECT ... while in state 10/RELEASING 00:20:39.911 D android.hardware.media.c2@1.0-service-goldfish: Goldfish C2 Service starting... ``` The decoder HAL dies and respawns; Media3 is left with a `DEAD_OBJECT` codec, writes no output sample, and its 25-second export watchdog aborts the export — which is the `ExportException: Muxer error` this ticket was filed on. **Nothing in this app decided anything, and it is not the runner starving a working codec: it is a null-pointer dereference in the emulator's own vendor codec library.** ### It is 6 for 6 Every gating leg-attempt that has ever failed this way carries the same HAL crash in the same job's log — grepped for `c2@1.0-service-goldfish` in the `--- native crashes (tail 60) ---` dump that `e2e-run.sh` writes on failure: | run / attempt | leg | date | |---|---|---| | `32855014836` a1 | API 36 | 2026-08-25 | | `32857067112` a1 | API 34 | 2026-08-25 | | `32919928048` a1 | API 36 | 2026-08-26 | | `33261618358` a1 | API 34 | 2026-08-29 | | `33588264439` a1 | API 36 | 2026-09-02 | | `34000816016` a1 | API 34 | 2026-09-06 | Six occurrences in **1210** gating leg-attempts on API 33-36 across 2026-08-20..09-07 — **0.5%**, and 0/301 on API 33, 3/302 on API 34, 0/304 on API 35, 3/303 on API 36. (The API 37 row does not run this test, so it is not in the denominator. Two other failures of this test are excluded and said so below.) ### What this settles, and what it does not **Settles:** the comment above that called this "one emulator-codec weakness expressed at two severities: deterministic on 37, load-dependent below it" was **right about the subsystem**, and the `@FailsOnEmulatorApi37` marker's stated reason — "fail inside the emulator's own `c2.goldfish.h264.decoder`" — is the same vendor HAL. It is one weakness, and this is its 33-36 expression. **Does not settle:** *why* the HAL dereferences null. `MediaCodec::reclaim` is logged 8 ms before the crash, and a reclaim is the resource manager taking a codec instance away from a client — so "a reclaim races `C2GoldfishAvcDec::process` and the block pool is gone under it" is the obvious hypothesis and **is untested**. I am recording it as a hypothesis and not acting on it, because this ticket has twice published a causal claim its own data did not support. **Not a candidate for a fix here.** The crashing code is `/vendor/lib64/*` inside the system image. This is environmental, in the same sense `@FailsOnEmulatorApi37` is, and now with the same kind of evidence behind it. ### Two failures of the same test that are *not* this mode Anchoring on the test name alone would have counted eight. Excluded, with the reason: - `32545625459` a1 (API 37, 2026-08-22) — 37 tests failed together, framework down, gralloc assertion present. That is #108, and this test was collateral. - `32669190757` a1 (API 35, 2026-08-23) — predates #111's shape report, and the log carries the bare `FAILED` marker with no message at all. **Unclassifiable, not classified.** --- ## Live census of every gating leg-attempt, 2026-08-20 .. 2026-09-07 Per leg-attempt, as this ticket requires; cancelled legs excluded (those are the `cancel-in-progress` cancellations the tally above warns about, not runs). **1489 gating leg-attempts, 129 failures, 8.7%.** Classified by anchoring the failing test name to the `FAILED` marker in each failing job's own log, then by the message under it. | mode | all history | since 2026-08-27 | last 2 days | |---|---|---|---| | SAF picker (will not close / never showed / rotation) | 52 | 17 | 7 | | framework down, no test named or test as collateral | 25 | 11 | 4 | | other named-test failure | 16 | 2 | 1 | | wedge — gradle never returned | 9 | 6 | 1 | | **Media3 export watchdog (this ticket)** | **6** | **3** | **1** | | SAF *save* tests — Activity gone / composition unreachable | 7 | 7 | 7 | | emulator never came up (`adb` 224) | 3 | 3 | 1 | | unclassified (no test named, no signature) | 9 | 2 | 0 | | denominator (gating leg-attempts) | 1489 | 796 | 387 | The mode column sums to 127, not 129: the two transcode failures excluded above are left out of it rather than filed under a mode they are not. **Per-mode denominators are not the same denominator**, which is the counting lesson from doing this: three of these modes cannot occur on all five rows. The save tests carry `@FailsOnEmulatorApi37` and only existed from 2026-09-06T15:12, so their rate is 7 in **88** leg-attempts on API 33-36 — not 7 in 1489. The wedge has only ever happened on API 33/34. ### Per-mode disposition | mode | disposition | evidence | |---|---|---| | Media3 export watchdog | **environmental, measured** — goldfish Codec2 HAL SIGSEGV, 6/6 | above | | API 37 picker + framework-down | **#108** — 24 of the 30 API 37 failures since 2026-08-27 carry `hasReadColorBufferDma`; of the 6 that do not, 3 are `adb` 224, 2 have no signature at all, and 1 has the framework-down signature without the assertion reaching the crash tail | job logs | | wedge | **consistent with #219 having fixed it, not demonstrated** — 9 occurrences in 497 API 33/34 leg-attempts before 2026-09-06T02:00Z (1.8%), **0 in the 106 since**. The last one (`34001741668`, 2026-09-06T00:36) is `thePickedInputSurvivesARealRotation` again, and `git merge-base --is-ancestor 32ab54d <head>` says that head **does not contain** #219's fix. At the prior rate, P(0 in 106) ≈ 0.15, so this is suggestive and not yet evidence | `e2e-wedge-api34` of `34001741668`: `started:` the rotation test, no `finished:`, every binder service `found` | | `adb` 224 | **infra, before the suite starts** — 3 occurrences, all API 37, nothing to attribute | job logs | | SAF save tests | **at least four distinct shapes, three of them already fixed or pre-fix** — split below | below | --- ## The newest mode splits into four, and only one of the seven is on today's `main` The `@FailsOnEmulatorApi37`-marked save tests (#226, #250) landed on 2026-09-06 and account for every gating failure on API 33-36 since. Filed as one mode, they are not one: | run / attempt | head | message | shape | |---|---|---|---| | `34041593697` a1 (35) | `fa10d94` — the commit that **added** the test | `IllegalStateException: No compose hierarchies found`, thrown directly | pre-`b23ff0f` | | `34056545386` a1 (35) | `cbbaf74` (#250) | `ComposeTimeoutException ... after 120000 ms` | the `CONVERSION_TIMEOUT_MS` case `19e3539` fixed. **Not a system-service failure at all** — API 35's software encode measured 134.8 s against a 120 s bound | | `34057706195` a1 (34) | `19e3539` | `No compose hierarchies found` ×2, thrown directly | **contains `b23ff0f`**, so that fix did not close this shape | | `34067653670` a1 (35), `34146936252` a1 (35) | `d45abe7`, `73482520` | `waited 300000ms for a node tagged action.saveFile`, **no** composition error | `awaitNode` only appends the composition error when `fetchSemanticsNodes` threw, so the composition was readable throughout. Confirmed in `34067653670`'s per-test logcat: `MainActivity` is `RESUMED` at 23:49:52.981 and stays resumed until the rule tears it down 300 s later, and **no conversion runs at all** in that window. Open | | `34146936252` a2 (35) | `73482520` | `waited 300000ms ...; last composition error: No compose hierarchies` | app `PAUSED` at 17:38:30.522 and never resumes; the back press at 17:38:32.479 follows `Waiting 5000ms for ... com.google.android.permissioncontroller`, i.e. the permission-dialog helper #269 replaced with `pm grant` | | **`34161043035` a1 (35)** | **`9fd96d0` — #269 itself** | `NPE: Cannot run onActivity since Activity has been destroyed already` | **traced to a mechanism; see below** | **So "mode 5 at 8% of leg-attempts" would be wrong** — most of these are shapes their own follow-up commit had already answered, on heads that predate it. ### The one on current `main` has a mechanism, and it is #93's launcher ANR with a new victim From the per-test logcat in `e2e-report-api35` of `34161043035` **attempt 1** (attempt 2 was green; the attempt-1 artifact survives with its own id, which is how it can still be read): ``` 20:59:36.033 MainActivity RESUMED <- the save picker has returned; the app is in front 20:59:37.068 UiDevice: Pressing back button. 20:59:38.092 InteractionController: Timed out waiting 1000ms for command and events. 20:59:41.094 UiDevice: Retrieving node with selector: BySelector [RES='android:id/aerr_wait'] 20:59:41.100 UiObject2: Clicking on (540, 1359). 20:59:41.169 Input channel object '9128b0b Application Not Responding: com.google.android.apps.nexuslauncher (client)' was disposed 20:59:41.713 UiDevice: Pressing back button. 20:59:41.754 TopTaskTracker: onTaskMovedToFront: ... cmp=com.google.android.apps.nexuslauncher 20:59:41.755 MainActivity PAUSED 20:59:42.224 MainActivity STOPPED 20:59:42.278 MainActivity DESTROYED 20:59:42.599 TestRunner: failed ``` `dismissThePicker`'s loop is: ```kotlin repeat(BACK_PRESSES) { if (awaitAppFocus()) return dismissASystemErrorDialog() device.pressBack() } ``` `awaitAppFocus()` is `composeRule.activity.hasWindowFocus()`. **The launcher's ANR dialog makes that false too** — it is a fullscreen `system_server` window, which is the whole reason `dismissASystemErrorDialog` exists. So the guard reads false, the dialog is removed at `41.169`, and the back press at `41.713` is then sent on a reading taken *before* the occluder was removed — into an app that is already in front with nothing to go back to. It finishes `MainActivity`, the launcher comes to the front, and the rest of the test has no Activity. The method's own KDoc names this hazard — *"a third from Recent would finish `MainActivity` and take the rest of the test with it"* — and the guard is what is supposed to prevent it. It cannot, because it cannot tell "the picker is still up" from "a system dialog is on top". `requireAReadableScreen`, twenty lines away, already does the right thing: it re-reads after `dismissASystemErrorDialog()` before doing anything else. `dismissThePicker` does not. PR to follow — the fix is to re-read the focus after removing a dialog and skip the back press if the app already has it, which **removes** an action rather than retrying one. ### One more finding, recorded rather than acted on `ERROR_DIALOG_BUTTONS` ends with `android:id/button1`, the framework's generic AlertDialog positive button — not an app-error-dialog id. In the same trace, at `20:59:35.689`, after `aerr_wait` and `aerr_close` both missed, `button1` was found and clicked at (927, 2274), and `MainActivity` resumed 339 ms later: **`dismissASystemErrorDialog` clicked a button inside DocumentsUI's own save dialog**, believing it to be a system error dialog. It happened to complete the save the walk had just failed to complete. There is no reason it always would. --- ## Method notes, since this ticket is where people land - **A re-run destroys the log.** `gh run view --job <id> --log` resolves by *run* and serves the latest attempt. Read a historical attempt with `gh api --allow-escape-sequences /repos/{owner}/{repo}/actions/jobs/{job_id}/logs`, taking the job id from `/actions/runs/{run}/attempts/{n}/jobs`. That is the correction already on this ticket, and it still holds. - **Artifacts survive a re-run; the API just makes them look like they do not.** `/actions/runs/{run}/artifacts` returns *every* attempt's upload under the same name with different ids and `created_at`. Take the id whose `created_at` falls in the attempt's window — `gh run download` takes the newest, which on a re-run-to-green is the green one. - **The per-test logcats are the evidence, not the job log.** `e2e-report-apiNN` carries `outputs/androidTest-results/connected/debug/<device>/logcat-<class>-<method>.txt`, one file per test, plus the JUnit XML with the untruncated stack. The job log truncates a stack to its first frame. Everything above came out of those files.
JMR-dev commented 2026-09-07 22:15:45 +00:00 (Migrated from github.com)

Where this leaves the ticket, plus one correction to the comment above

Correction first

That comment's heading says "The newest mode splits into four". It is five, and the table
under it is what shows so: No compose hierarchies thrown directly, ComposeTimeoutException at
120 s, waited 300000ms with no composition error, waited 300000ms with one, and the NPE — five
distinct messages across seven leg-attempts. The table is right; the heading undercounted by
folding the NPE in with the rest, which is the one thing that comment was arguing against doing.
docs/ci-failure-modes.md says five.

The work

PR #272 — the dismissThePicker fix, docs/ci-failure-modes.md (the census and per-mode
disposition), and a pointer to it from CLAUDE.md. Local gate green: the instrumented suite is
72 tests / 0 failures on API 33, 34, 35 and 36, plus the JVM gate; API 37 reported
NOT COVERED LOCALLY, no device attached.

#270 — the two SAF-save failure shapes that still have no mechanism. Both predate #269, and the
first move is a re-count over post-#269 leg-attempts rather than a fix.

#271 — ERROR_DIALOG_BUTTONS ends with android:id/button1, which is any AlertDialog's
positive button, and it was measured clicking one inside DocumentsUI.

#190 — has the marker-width measurement it asked for: 0.5% below API 37 against 100% on it, so
the marker is the right width, and the second signature (c2@1.0-service-goldfish) that its
proposed report would need in order to have caught its own row 4.

What I would do with #102, leaving the call to you

Close it as an umbrella. Its "done means" is met for the mode it was filed on — a cause with a
measurement, and one that says plainly there is nothing here to fix. Nothing was retried, no timeout
was widened, and the one fix in #272 removes an action rather than adding a tolerance.

The four modes it accumulated all have somewhere better to live now: #108 and #122 are closed, #190
carries the reporting question, #270 and #271 carry what is genuinely still open, and
docs/ci-failure-modes.md is the standing triage reference an umbrella ticket was being used as.

Two things I could not answer and am not going to pretend otherwise:

  • The wedge is "consistent with #219 having fixed it", not fixed. 0 in 106 API 33/34
    leg-attempts against a prior 1.8%; P(0 in 106) ≈ 0.15 if nothing changed. Worth a re-count in a
    week, not a claim today.
  • There are zero post-#269 gating leg-attempts on main. The only run on #269's head is
    34161043035, whose attempt 1 is the mode #272 fixes. So "the picker tests are fine now" is
    currently unmeasured in either direction, and #270 says so as its first step.
## Where this leaves the ticket, plus one correction to the comment above ### Correction first That comment's heading says *"The newest mode splits into four"*. **It is five**, and the table under it is what shows so: `No compose hierarchies` thrown directly, `ComposeTimeoutException` at 120 s, `waited 300000ms` with no composition error, `waited 300000ms` *with* one, and the NPE — five distinct messages across seven leg-attempts. The table is right; the heading undercounted by folding the NPE in with the rest, which is the one thing that comment was arguing against doing. `docs/ci-failure-modes.md` says five. ### The work **PR #272** — the `dismissThePicker` fix, `docs/ci-failure-modes.md` (the census and per-mode disposition), and a pointer to it from `CLAUDE.md`. Local gate green: the instrumented suite is 72 tests / 0 failures on API 33, 34, 35 and 36, plus the JVM gate; API 37 reported `NOT COVERED LOCALLY`, no device attached. **#270** — the two SAF-save failure shapes that still have no mechanism. Both predate #269, and the first move is a re-count over post-#269 leg-attempts rather than a fix. **#271** — `ERROR_DIALOG_BUTTONS` ends with `android:id/button1`, which is any `AlertDialog`'s positive button, and it was measured clicking one inside DocumentsUI. **#190** — has the marker-width measurement it asked for: 0.5% below API 37 against 100% on it, so the marker is the right width, and the second signature (`c2@1.0-service-goldfish`) that its proposed report would need in order to have caught its own row 4. ### What I would do with #102, leaving the call to you **Close it as an umbrella.** Its "done means" is met for the mode it was filed on — a cause with a measurement, and one that says plainly there is nothing here to fix. Nothing was retried, no timeout was widened, and the one fix in #272 *removes* an action rather than adding a tolerance. The four modes it accumulated all have somewhere better to live now: #108 and #122 are closed, #190 carries the reporting question, #270 and #271 carry what is genuinely still open, and `docs/ci-failure-modes.md` is the standing triage reference an umbrella ticket was being used as. Two things I could not answer and am not going to pretend otherwise: - **The wedge is "consistent with #219 having fixed it", not fixed.** 0 in 106 API 33/34 leg-attempts against a prior 1.8%; P(0 in 106) ≈ 0.15 if nothing changed. Worth a re-count in a week, not a claim today. - **There are zero post-#269 gating leg-attempts on `main`.** The only run on #269's head is `34161043035`, whose attempt 1 is the mode #272 fixes. So "the picker tests are fine now" is currently unmeasured in either direction, and #270 says so as its first step.
JMR-dev commented 2026-09-07 22:43:46 +00:00 (Migrated from github.com)

Retracting my sighting above: that NPE was not environmental

My comment of 2026-09-07T21:20 filed the API 35 Cannot run onActivity since Activity has been destroyed already under this ticket's environmental class. That attribution is wrong, and the
investigation behind #272 has the trace.

What I offered as evidence — no test-file frame in the stack, ColorBuffer errors seven seconds
either side, green on re-run, green on the other four levels — was all true and all beside the point.
The Activity was destroyed by the test's own back press:

20:59:35.689  UiObject2: Clicking on (927, 2274)   <- dismissASystemErrorDialog falls through to
                                                      android:id/button1 and hits DocumentsUI's
                                                      own positive button (#271)
20:59:36.033  MainActivity RESUMED                  <- the picker is now gone, by our own hand
20:59:37.068  UiDevice: Pressing back button        <- pressed anyway, on a stale focus reading
20:59:41.713  UiDevice: Pressing back button
20:59:42.278  MainActivity DESTROYED

dismissThePicker guarded its back presses on Activity.hasWindowFocus, and a fullscreen
system_server ANR dialog makes that false too — so it could not tell "the picker is still up" from
"a dialog is on top of an app that is already in front". #272 re-reads the focus after a dialog is
actually dismissed, which removes the press rather than retrying it.

The part worth keeping

My own caveat at the time was the right instinct and I did not follow it far enough:

a pattern this well known is exactly what a real regression could hide behind.

That is precisely what happened. Four independent signals agreed, the failing test was one the PR
had modified, and the pattern still lost to a logcat trace. A no-test-frame stack is evidence that
the test did not throw, not evidence that the test did not cause it
— the back press that killed
the Activity was ours, and it left no frame because the death was asynchronous.

The re-run that made it green is what nearly buried it, exactly as the "a re-run destroys the log"
note in the same comment warns. That note stands; so does the practice of saving the log first,
which is the only reason the trace above still exists.

## Retracting my sighting above: that NPE was not environmental My comment of 2026-09-07T21:20 filed the API 35 `Cannot run onActivity since Activity has been destroyed already` under this ticket's environmental class. **That attribution is wrong**, and the investigation behind #272 has the trace. What I offered as evidence — no test-file frame in the stack, `ColorBuffer` errors seven seconds either side, green on re-run, green on the other four levels — was all true and all beside the point. The Activity was destroyed **by the test's own back press**: ``` 20:59:35.689 UiObject2: Clicking on (927, 2274) <- dismissASystemErrorDialog falls through to android:id/button1 and hits DocumentsUI's own positive button (#271) 20:59:36.033 MainActivity RESUMED <- the picker is now gone, by our own hand 20:59:37.068 UiDevice: Pressing back button <- pressed anyway, on a stale focus reading 20:59:41.713 UiDevice: Pressing back button 20:59:42.278 MainActivity DESTROYED ``` `dismissThePicker` guarded its back presses on `Activity.hasWindowFocus`, and a fullscreen `system_server` ANR dialog makes that false too — so it could not tell "the picker is still up" from "a dialog is on top of an app that is already in front". #272 re-reads the focus after a dialog is actually dismissed, which removes the press rather than retrying it. ### The part worth keeping My own caveat at the time was the right instinct and I did not follow it far enough: > a pattern this well known is exactly what a real regression could hide behind. That is precisely what happened. Four independent signals agreed, the failing test was one the PR had modified, and the pattern still lost to a logcat trace. **A no-test-frame stack is evidence that the test did not throw, not evidence that the test did not cause it** — the back press that killed the Activity was ours, and it left no frame because the death was asynchronous. The re-run that made it green is what nearly buried it, exactly as the "a re-run destroys the log" note in the same comment warns. That note stands; so does the practice of saving the log first, which is the only reason the trace above still exists.
JMR-dev commented 2026-09-07 22:48:07 +00:00 (Migrated from github.com)

Closing as an umbrella: every mode now has a home

This ticket asked for "either a cause with a measurement behind it, or a documented decision to
treat it as environmental with the same honesty @FailsOnEmulatorApi37 gets."
Both halves are now
answered, and the modes it accumulated have outgrown one ticket.

Its own mode — the Media3 export watchdog — has a measured cause, and it is environmental. The
emulator's Codec2 HAL null-derefs at C2Block2D::handle()+4 via getClientUsage in
libcodec2_goldfish_common.so, killing c2.goldfish.h264.decoder; Media3's 25 s watchdog then
aborts. 6 of 6 occurrences carry that signature. The crashing code is /vendor/lib64/* — there
is nothing in this repo to fix. Rate: 6 in 1210 API 33–36 gating leg-attempts, 0.5%.

The census that settles the counting question. 1489 gating leg-attempts, 2026-08-20 → 09-07,
129 failures (8.7%). Those 129 sit in 83 runs, 45 of which ended green after a re-run — so a
run-level count reports 38 where there were 129, and deletes exactly the ones somebody already
judged to be noise. Count per leg-attempt. Full method, per-mode dispositions and the
evidence-reading procedure are now in docs/ci-failure-modes.md.

Where the modes went:

mode disposition
Media3 export watchdog environmental, measured — this ticket, closed
API 37 picker + framework-down #108 (gralloc assertion; 24 of 30 API 37 failures since 08-27)
Wedge #122 — 0 in 106 leg-attempts since, but "consistent with", not demonstrated
SAF picker back-press killing MainActivity a real defect, fixed in #272
dismissASystemErrorDialog clicking any button1 #271
Two unexplained save-test shapes #270
Overall gating flake rate #190, now carrying the census measurement
adb 224 infra, pre-suite, nothing to attribute

The lesson this ticket paid for twice

It retracted two causal claims its own data did not support, and then a third — mine. I filed the
2026-09-07 API 35 Activity has been destroyed NPE here as environmental on four agreeing signals,
and it was a real defect the whole time; the retraction above has the trace. A no-test-frame stack
is evidence the test did not throw, not evidence the test did not cause it.

The other durable finding: a re-run destroys the log. The trace that overturned my attribution
survives only because someone saved it before pressing re-run. That practice is now written into
docs/ci-failure-modes.md.

Closing. New sightings should go to the specific ticket for their mode, or to a new one citing the
doc — not here.

## Closing as an umbrella: every mode now has a home This ticket asked for *"either a cause with a measurement behind it, or a documented decision to treat it as environmental with the same honesty `@FailsOnEmulatorApi37` gets."* Both halves are now answered, and the modes it accumulated have outgrown one ticket. **Its own mode — the Media3 export watchdog — has a measured cause, and it is environmental.** The emulator's Codec2 HAL null-derefs at `C2Block2D::handle()+4` via `getClientUsage` in `libcodec2_goldfish_common.so`, killing `c2.goldfish.h264.decoder`; Media3's 25 s watchdog then aborts. **6 of 6** occurrences carry that signature. The crashing code is `/vendor/lib64/*` — there is nothing in this repo to fix. Rate: 6 in 1210 API 33–36 gating leg-attempts, **0.5%**. **The census that settles the counting question.** 1489 gating leg-attempts, 2026-08-20 → 09-07, 129 failures (8.7%). Those 129 sit in 83 runs, **45 of which ended green after a re-run** — so a run-level count reports 38 where there were 129, and deletes exactly the ones somebody already judged to be noise. Count per leg-attempt. Full method, per-mode dispositions and the evidence-reading procedure are now in [`docs/ci-failure-modes.md`](../blob/main/docs/ci-failure-modes.md). **Where the modes went:** | mode | disposition | |---|---| | Media3 export watchdog | environmental, measured — this ticket, closed | | API 37 picker + framework-down | #108 (gralloc assertion; 24 of 30 API 37 failures since 08-27) | | Wedge | #122 — 0 in 106 leg-attempts since, but **"consistent with", not demonstrated** | | SAF picker back-press killing `MainActivity` | **a real defect**, fixed in #272 | | `dismissASystemErrorDialog` clicking any `button1` | #271 | | Two unexplained save-test shapes | #270 | | Overall gating flake rate | #190, now carrying the census measurement | | `adb` 224 | infra, pre-suite, nothing to attribute | ### The lesson this ticket paid for twice It retracted two causal claims its own data did not support, and then a third — mine. I filed the 2026-09-07 API 35 `Activity has been destroyed` NPE here as environmental on four agreeing signals, and it was a real defect the whole time; the retraction above has the trace. **A no-test-frame stack is evidence the test did not throw, not evidence the test did not cause it.** The other durable finding: **a re-run destroys the log.** The trace that overturned my attribution survives only because someone saved it before pressing re-run. That practice is now written into `docs/ci-failure-modes.md`. Closing. New sightings should go to the specific ticket for their mode, or to a new one citing the doc — not here.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: JMR-dev/LibreMediaConverter#102