CI legs fail on system services, not assertions: three modes today that look like one cause #102
Closed
opened 2026-08-25 14:25:31 +00:00 by JMR-dev
·
19 comments
No Branch/Tag Specified
main
fix/102-picker-back-press-overshoot
fix/268-saf-picker-determinism
feat/ogg-vorbis-libvorbis
feat/expedited-conversion-work
fix/gate-cache-in-worktrees
chore/gate-runs-shellcheck
test/publish-delete-arm-real-provider
docs/e8-instrumented-coverage
docs/api37-carrier-count-drift
docs/e7-second-constraint
test/publish-to-a-real-saf-destination
fix/api37-report-match-line
fix/api37-task-snapshot-crash
test/join-failure-message-on-device
docs/e2e-read-findings-e7
test/cancelling-a-running-export
test/reattach-to-a-running-job
test/content-uri-reaches-ffmpeg
fix/launcher-wiring-waits-for-the-pick
fix/cancel-tests-need-a-slower-encode
test/cancelling-a-running-join
test/cancelling-a-running-session
test/notification-cancel-action
test/ffmpeg-progress-is-observed
test/fallback-asserts-the-path
test/flac-and-opus-assert-their-format
docs/e2e-read-findings
docs/wave4-coverage-numbers
fix/injectable-startup-sweep-scope
test/session-outcome-seam
test/launcher-callback-identity
test/theme-follows-system-dark
test/audio-drop-arm
fix/rotation-waits-for-recreation
fix/convert-guards-on-ready
test/retry-save-mime
test/hardware-progress-reaches-workmanager
test/ffprobe-mapping-seam
test/device-codec-enumeration-seam
test/unknown-container-row
test/null-message-fallbacks
test/cancel-reaches-workmanager
docs/coverage-wave3-recovery
test/concat-engine-seam
docs/coverage-wave3
test/mediaprobe-merge-seam
test/adaptive-shell-wiring
test/aac-audio-args
test/notification-progress-text
test/media3-muxer-guard
test/hardware-fallback-and-cancellation
test/unprobeable-join-clip
test/one-branch-outcomes
test/foreground-type-regimes
fix/bound-wedge-diagnostics
docs/coverage-wave2
test/screen-wiring
test/viewmodel-setters
test/join-state-mapping
test/conversion-state-mapping
test/dedupe-user-messages
fix/restore-stack-merges
test/refused-jobs
test/concatworker-failure-arms
test/container-capabilities-audio
test/readspec-enum-fallbacks
test/outputpublisher-seams
test/mediaprobe-track-seam
test/outputpublisher-partial-branches
test/fake-provider-scaffolding
docs/coverage-read-findings
chore/gitignore-kotlin
test/bound-the-hangs
docs/coverage-remeasure
ci/baseline-counter-precision
fix/invalid-suggestion-chip
ci/wedged-leg-report
fix/reattachment-overwrites-pick
test/theme-live-branches
fix/failed-save-retry
fix/empty-composition-crash
ci/advisory-failure-report
docs/seven-run-counts
test/release-permission-guard
ci/build-workflow-permissions
docs/api37-point-release
docs/benchmark-populate-path
fix/dead-assertion-probe-test
ci/actionlint
test/device-codecs-encode-consequence
fix/sdkmanager-pipefail
docs/readme-restart-claim
fix/saf-picker-root-discovery
fix/probe-dispatcher-seam
test/media3engine-mime-tables
test/mediaprobe-pure-helpers
fix/codec-vocabulary-drift
docs/robolectric-rationale-correction
docs/api37-advisory-counts
test/r38-8-saf-e2e
test/r38-7-join-states
test/r38-6-conversion-states
test/r38-5-state-seam
fix/jacoco-robolectric-coverage
docs/instrumented-tests-correction
test/r38-2-filecard
test/r38-4-advanced-picker
test/r38-3-pickers
tools/file-issue-script
tools/api-37-emulator
fix/review-app-gaps
docs/review-corrections
No results found.
Labels
Clear labels
above-cut
accessibility
backlog
bug
confirmed
documentation
duplicate
enhancement
good first issue
help wanted
invalid
plausible
question
sev:high
sev:low
sev:medium
wontfix
Worked autonomously overnight: local, JVM-verifiable, no product decision
Barrier affecting people with disabilities
Held for manual review: product/UX call, CI/workflow, hardware, or unverifiable here
Something isn't working
Reviewer demonstrated the defect
Improvements or additions to documentation
This issue or pull request already exists
New feature or request
Good for newcomers
Extra attention is needed
This doesn't seem right
Reviewer could not fully demonstrate it; treat as unproven
Further information is requested
High severity
Low severity
Medium severity
This will not be worked on
Milestone
No items
No Milestone
Projects
Clear projects
No projects
Notifications
Due Date
No due date set.
Dependencies
No dependencies set.
Reference: JMR-dev/LibreMediaConverter#102
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Split out of #101, where I first mis-attributed it to
RealMediaBenchmark. This is the test that actually failed.The failure
Media3EngineTest.transcodesH264ToH265AndReportsProgressfails intermittently on the gating API 34 leg:A sibling signature on PR #95's API 36 leg in the same window:
Both are the hardware transcode failing to make progress, on levels where it is expected to work.
Why this is awkward, and interesting
This is one of the two tests that already carries
@FailsOnEmulatorApi37. It is excluded fromthe API 37 gating leg because it dies inside that image's
c2.goldfish.h264.decoder, and it runs inthe advisory job there instead. On API 33-36 it is a normal gating test and has been reliable all
session — dozens of green runs.
So the same test is: known-broken on one emulator image, reliable on four, and now intermittently
failing on two of those four.
The hypothesis worth testing first
no output sample written in the last 25000 millisecondsis a starvation signature, not acorrectness one.
#93's root cause, established in PR #96, was that the launcher ANRs on a loaded runner — loaded
enough that
system_serverleaves an "Application Not Responding" dialog that never clears. A runnerunder that much pressure would also starve a hardware transcode of its 25-second sample budget.
If that is the same illness, then #96 fixed the accessibility symptom of runner load and this is
the same cause surfacing in the codec path. That would also explain why both appeared in the same
few hours rather than gradually.
Check that before writing a Media3-specific theory. The measurement is whether these failures
correlate with the ANR signature in the same job's logcat.
Frequency, honestly
Two sightings in one window (#95 API 36, #97 API 34). That is not yet a rate, and I overstated a
different flake earlier today on four sightings, so this ticket deliberately does not claim it is a
blocker. It is filed so the next sighting has somewhere to land and so the connection to #93 is on
record rather than rediscovered.
Done means
Either a cause with a measurement behind it — runner load, a Media3 version behaviour, an emulator
configuration — or a documented decision to treat it as environmental with the same honesty
@FailsOnEmulatorApi37gets. Not a retry wrapper added because the test is annoying: the point ofthis test is that a hardware transcode completes, and retrying until it does would delete the
assertion.
A third failure mode, and the reason to widen this ticket rather than file a fourth.
PR #99's gating E2E API 37 leg failed today with the tests all passing:
Zero test failures.
am get-current-userreturning 20 meanssystem_serveris not answering — the framework went down during teardown, after every test had already passed. ThenotAnnotationfilter was intact (56 tests, the right number), so this is not a filtering regression from that PR's workflow edit.Three modes, one shape
AccessibilityWindowManagerdrops every app windowno output sample written in the last 25000 millisecondsam get-current-userfails; framework down, 0 test failuresAll three are system services failing to answer under load, not assertions failing. #96 established mode 1's root cause as a runner loaded enough that the launcher ANRs and
system_serverleaves a dialog that never clears. Modes 2 and 3 are what that same pressure would look like in the codec path and in teardown.They also arrived together. None of these three was seen before today; all three appeared within a few hours, on diffs that cannot cause them (KDoc comments, a lookup table, a workflow step).
What that changes about this ticket
It is probably not a Media3 ticket. Test whether the three correlate with runner load before writing a codec-specific theory — #96 already proved the load hypothesis once, and the measurement is whether these failures carry the ANR signature in the same job's logcat.
If they do, the useful fix is at the harness level — fewer things competing on the runner, a longer settle before the suite starts, or a documented acceptance that these legs are load-sensitive — rather than three separate per-test patches.
Frequency, still honest
Mode 2: two sightings. Mode 3: one. Not claiming a rate, and deliberately not escalating severity on three data points — the correction to #49 earlier today came from doing exactly that on four.
The three modes, now counted — and one of them moved
Census of every gating E2E leg-attempt since 2026-08-24 (advisory job excluded), classified by
anchoring the test name to the
FAILEDmarker in each failing job's own log. 45 failures across400 leg-attempts. Counting runs rather than leg-attempts sees only 19 of those 45, because a
re-run to green erases the evidence — every figure here is per leg-attempt.
SafPickerRoundTripTest— picker and/or rotation (#93)@FailsOnEmulatorApi37probe for #83doesNotOverwriteAPickTheUserHasAlreadyMade(#49)transcodesH264ToH265AndReportsProgress(#102 Media3 mode)routesAFastMp4JobByDeviceCapabilityThe finding that matters for this ticket
#96 fixed the SAF picker on API 33–36 and did not fix it on API 37.
Since #96 merged (2026-08-25T13:42Z) there have been zero picker failures on 33/34/35/36 and
two on API 37 — runs
32865281555and32899061992. I checked withgit merge-base --is-ancestorthat both head commits actually contain #96's fix, so this is not a stale branch.API 37 is also the only leg that aborts in
system_server(#108). Cross-referencing the two:5 of the 8 API 37 legs carrying the
TaskSnapshotPersisterabort trace also had a test fail, andin 4 of those 5 the failing test was the picker.
That is the concrete shape of this ticket's "one cause" hypothesis: on API 37 the picker test is
not failing for #93's original reason (an ANR dialog hiding every window from UiAutomator — that
cause was fixed and stayed fixed on four other API levels). It is failing because
system_serverstops answering, which is the same event that produces #108's abort. The picker test is simply the
most system-service-dependent test in the suite, so it is the first to notice.
Consequence for triage: #108 should not be treated as a cosmetic post-run crash, and the API 37
picker failures should not be filed as a #93 regression. They are one problem, and #108 is the
better handle on it.
Method note
Historical attempts cannot be read with
gh run view --job <id> --log— it resolves by run andserves the latest attempt, so a re-run hands you a green log for a red attempt. Use
gh api --allow-escape-sequences /repos/{owner}/{repo}/actions/jobs/{job_id}/logs, with the job idfrom
/actions/runs/{run}/attempts/{n}/jobs. Without--allow-escape-sequences, gh writes nothingand exits 0.
Correcting the causal claim in my previous comment
Two sentences above are not supported by the data in the same comment, and I am retracting them:
The discriminator is in my own table. The two post-#96 API 37 picker failures are:
32865281555 a232899061992 a1One of the two has no
TaskSnapshotPersistertrace at all. (32865281555in fact failed on API 37twice — a1 with neither an abort trace nor a named test failure, then a2 with the picker.) The
"4 of 5 abort legs also failed the picker" figure does not rescue the claim, because those
co-occurrences are almost all pre-#96, when the picker was failing on every API level for #93's
own cause. Only the post-#96 API 37 failures discriminate, and that is n=2 with one against.
The second sentence is the harmful one: it would tell a future triager to discard the possibility of
a #93 regression on 1-for-2 evidence. Withdrawn.
A third mechanism I failed to enumerate
SafPickerRoundTripTest's own KDoc already documents an API-37-specific failure mode, and it isnot #108's:
So there are at least three candidates for the post-#96 API 37 picker failures — #93's cause
surviving on 37, #108's
system_serverabort, and this surfaceflinger/Gralloc5 restart — and Ipicked one without ruling out the other two.
I also checked whether #96's fix might be version-sensitive:
dismissASystemErrorDialog(theaerr_waitpath) is not API-gated and runs identically on every leg, so a gated remedy is notthe explanation.
What actually survives, and is worth keeping
#96 fixed the picker on API 33–36 and API 37 has failed twice since. Zero picker failures on
33/34/35/36 post-fix; two on 37, both on head commits verified with
git merge-base --is-ancestorto contain the fix.
Which of the three mechanisms is responsible is unresolved. The abort trace being absent from one
of the two failures is precisely why it is unresolved, rather than being the reason to pick #108.
Partial answer to this ticket's question, from the API 37 side
Two of the API 37 aborts are provably one bug, not two — details and evidence on #108.
The crash the
E2E_DISABLE_SYSTEM_UImitigation already handles and the crash in this ticket'sthird mode are the same assertion in the same emulator mapper:
They differ only in who calls it —
RegionSamplingThread(SystemUI's nav-bar luma sampler, whichthe mitigation removes by deleting the package) versus
TaskSnapshotPer(WindowManager, insidesystem_server, where there is no package to remove). All 8 recorded #108 legs showRegionSamplingThread=0, so they are entirely the second caller.What this does and does not settle. It supports "one cause" for the API 37 aborts. It says
nothing about the picker failures — I retracted that claim above, and the reason still stands: one
of the two post-#96 API 37 picker failures had no abort trace at all. A shared root for the aborts
does not retroactively make the picker a symptom of them.
A fourth mode, on a docs-only PR
2026-08-26T02:07-02:30Z, run32922652540, E2E API 34, on #115 — a change toCLAUDE.mdandnothing else, so no test in the suite could have been affected by the diff.
Not an abort, not an assertion. All 59 tests ran, and then the leg hung in teardown until
WEDGE_TIMEOUTkilled it at ~22 minutes. Gradle never printed a summary line, which is why the countis
unknown.bad color buffer handle 375two seconds later is the emulator's own complaint, and it puts this inthe same family as the #108 aborts — the
hasReadColorBufferDmaassertion is also a colour-bufferreadback fault in the goldfish mapper. I am not claiming they are the same bug; I made that
mistake once on this ticket already. What is fair to say is that three of this ticket's now-four
modes name the emulator's colour-buffer path, and none of them names anything in this app.
So the running list for this ticket is:
system_serverabortThe only reason this one was legible is the report from #111 — before it, a wedged leg was an
anonymous red with no counts. The table did overstate it (
completed cleanly: yeson a leg the wedgehad just killed), which is #118 and is a reporting fix, not a pass/fail change.
The wedge mode has recurred, on a different API level
Second sighting:
2026-08-26T03:23Z, PR #117, E2E API 33 (the first was API 34 on #115). Sameshape —
expected: 60, received: 60, failed: unknown, thenWEDGED. No abort trace at all thistime, so it is not #108 in disguise.
That makes the mode table for this ticket:
system_serverabortThe wedge is now the second mode confirmed on more than one API level, and the second whose evidence
came from a PR that could not plausibly have caused it (#115 was docs-only; #117 touches validation
suggestions and nothing near teardown).
Still not claiming a single cause. Three of the four modes name the emulator's colour-buffer path
and this one does not name anything yet — the wedge diagnostics artifact is uploaded per occurrence
(
e2e-wedge-api33,e2e-wedge-api34) and nobody has read one. That is the obvious next evidence,and it is cheap: two artifacts already exist.
The fourth mode now has a name, and my description of it was wrong
I read the two wedge diagnostics artifacts. Correcting what I wrote above: I said "every test
ran, then the leg hung in teardown", trusting
received: Nfrom the run-shape table. Both wedgesactually end on a test that never returned:
Same test, both times, on two PRs that could not have caused it (#115 docs-only, #117 validation
suggestions). Filed as #122.
And it is not this ticket's family. In both artifacts every binder service the wedge probe checks
is
found—input,window,activity,media.player. Nothing crashed. That separates itcleanly from the three colour-buffer modes:
system_serverabort (#108)hasReadColorBufferDmaSo this ticket's premise — "three modes that look like one cause" — is now two answered questions
and two open ones. #108 is the gralloc assertion with a second caller. #122 is a hanging test, not an
emulator fault at all. What remains genuinely open is the Media3 transcode timeout and why the SAF
picker still fails on 37.
That is a narrowing, not a closure: I would not close this until those two have the same treatment.
The API 37 picker question is answered — and my retraction above was wrong, for an instructive reason
Every
E2E API 37gating failure since #96 merged, classified by whether the picker test failed andwhether the gralloc assertion (
hasReadColorBufferDma) appears — 14 leg-attempts:The
picker failed / no abortcell is empty. Six for six.Why I previously said one had no abort
I grepped for
TaskSnapshotPer— the thread name from #108 — rather than for the assertion itself.Re-checking the case I cited as the counterexample (run
32865281555attempt 2, job97952658829):It aborted through the other caller — the SystemUI region-sampling path that
E2E_DISABLE_SYSTEM_UIexists to remove, and which evidently did not take on that run. So the abortwas there; my filter could not see it.
That is a real methodological lesson and I would rather write it down than bury it: grepping the
thread name instead of the fault split one bug into two. #108's title names
TaskSnapshotPersister, and I let the title become the search term. The assertion is the invariant;the thread is just which caller happened to trip it.
What this settles, and what it does not
Settles: the post-#96 API 37 picker failures are not a #93 regression. #93's cause was an ANR
dialog hiding every window from UiAutomator, fixed in #96 and holding on API 33-36 with zero picker
failures there since. On 37 the picker fails only in the presence of a gralloc abort, which takes the
framework down underneath it —
UiAutomation.getWindows()empty,Can't find service: package.Does not settle: it remains an association, six for six, not a proved mechanism. And it does not
resurrect the causal claim I withdrew earlier in the stronger form I first wrote it: the aborts are
not triggered by the picker test. The diagnostics artifact shows three aborts firing before the
first test starts, and one that the picker test survived and finished 1.1 s later. Ambient fault,
occasional victim — the picker is simply the largest target, being the slowest test in the suite.
Remaining open on this ticket
One mode, not two: the Media3 transcode timeout (
transcodesH264ToH265AndReportsProgress, seen onAPI 34 and 36, no abort trace). The other three now have names — #108 (gralloc, two callers), #122
(rotation hang, framework fully alive), and this picker association.
The last unnamed mode has a name too
The transcode failure is Media3's own export watchdog, not an assertion of ours:
Transformerarms a 25-second watchdog and aborts the export when the encoder produces no outputsample for that long. So the emulator's encode stalls; Media3 notices and gives up. Nothing in this
app decided anything.
And the test it hits is already known to be broken on one image
transcodesH264ToH265AndReportsProgressis one of the two tests carrying@FailsOnEmulatorApi37(Media3EngineTest.kt:72). CLAUDE.md records why: two Media3 hardwaretranscodes fail inside the emulator's own
c2.goldfish.h264.decoderon the android-37 images.So the same test that always fails on API 37's emulator codec intermittently fails on 34 and
36 — twice in the census window — with a stall long enough to trip a 25-second watchdog. The
straightforward reading is that this is one emulator-codec weakness expressed at two severities:
deterministic on 37, load-dependent below it. That is an inference from two data points and the
shared marker, not a proof.
What would confirm it: the stall should be visible in logcat as the
c2.goldfishencoderstarving, in the
e2e-diagnostics-apiNNartifact for a run that hit it. Two such runs exist. That isthe same cheap artifact read that settled #122.
This ticket's question is now answered
All four modes have names, and none of them is an assertion failing about this app's behaviour:
system_serveraborthasReadColorBufferDmain the goldfish mapper, two callers, ambientThe premise this was filed under — "three modes today that look like one cause" — turned out to be
half right in an unhelpful way. They are not one cause. But they are all the emulator failing
underneath the suite, in four unrelated subsystems: gralloc, whatever hangs the rotation, and the
goldfish codec. Not one bug; one class of bug.
I would close this in favour of #108 and #122 plus a new ticket for the codec stall if anyone wants
it chased, rather than keep a four-way umbrella open. Leaving that call to the maintainer.
Tally from a batch of eight PRs opened today (#131, #144–#151) — offered because this ticket asks whether the modes share a cause, and a single day's batch is a decent sample of one workload.
Roughly fifty leg-runs. Every failure was on a diff that was docs-only or JVM-test-only, so none of them can be attributed to the change under test.
received: Nthen gradle never returns; killed at 1200s;failed: unknowndocumentsuiin front after 4 back presses;UiAutomation.getWindows()emptyadb … failed with exit code 224during setup, before any test ranCLAUDE.mdWedges split API 33 ×1, API 34 ×2 — so not one level, which also corrects something I wrote on #122 earlier from a smaller sample.
What separates them, and it is not subtle: the wedge happens after every test has reported (
received: 60), the picker failure is a real assertion inside a running suite, and the adb-224 case happens before the suite starts at all. Those are three different points in the lifecycle, which argues against the single shared cause this ticket floats — unless the shared cause is simply "the emulator on this runner is unreliable in several independent ways".The practically useful finding: every one of the five retried clean on a plain re-run, no change to the diff. And on a wedged leg the tests themselves completed — what is lost is the verdict, not the coverage.
One trap worth putting in this ticket, since this is where people land when a leg fails:
status_check.ymlsetsconcurrency: cancel-in-progress: true, so re-running a job on an older run for the same ref cancels whatever newer run is in flight.gh pr checksthen reports every cancelled job asfail, which reads exactly like a build break — nine jobs "failing" in 42 seconds on a one-file docs diff. Check that nothing newer is queued for the ref before retrying, or re-run the newest run instead.Correcting my own tally above. The line "every one of the five retried clean on a plain re-run" was wrong when I posted it, and it was the one practically useful sentence in the comment, so it is worth fixing rather than leaving.
#146 (
test/container-capabilities-audio) needed two re-runs, not one. Run33036437394reached attempt 3, and the gatingE2E API 37leg failed on attempts 1 and 2:E2E API 37adb … failed with exit code 224, 03:45:52Z — emulator never came upSo the "Emulator never came up" row in my table is 2, not 1, and that mode has now repeated on a re-run of the same ref — which is a mildly interesting datum for this ticket, because it is the one mode that happens before the suite starts and therefore cannot be blamed on anything the tests do.
What survives unchanged: every failure was still on a docs-only or JVM-test-only diff, the three modes still sit at three different points in the lifecycle, and all eight branches are green now. What does not survive is "one re-run clears it" — for one branch it took two, and the retry reproduced the original mode rather than a new one.
(#150's single retry did come back clean; attempt 2 succeeded. I checked both rather than assume the correction generalised.)
One more data point, same batch. #148's re-push (a one-file, JVM-test-only diff) hit the SAF picker mode again on the gating
E2E API 37leg:Attempt 2 came back clean, so the running tally for this batch is now:
adb224)Each of the three modes has now repeated within a single day's batch, which is the part worth recording. Whatever this ticket concludes about a shared cause, none of the three is a one-off.
The shape row is doing its job, incidentally:
received: 57, failed: 1, completed cleanly: yessaid in one line that all 57 gating tests ran and exactly one asserted false — which is what separates this mode from the wedge (where the tests complete but the verdict is lost) and from adb-224 (where the suite never starts). Reading the marker rather than the job name, perCLAUDE.md.A mode this tally has not recorded before: a Media3 hardware transcode failing on API 34.
From #165, run
33261618358attempt 1:Two legs, two different modes, one run.
Why the API 34 one is worth a line here.
CLAUDE.mdrecords Media3 hardware transcodes failing "inside the emulator's ownc2.goldfish.h264.decoder" as an API 37 property — two of the three@FailsOnEmulatorApi37tests are exactly that. This is the same class of failure on API 34, which is a gating leg with no marker and no allowance.If it recurs, the interesting question is whether the marker's premise is level-specific or whether API 34 has simply been lucky. One occurrence is not that evidence, which is why this is a note rather than a claim.
Attribution, checked rather than assumed: #165's diff extracts an actions builder from two screen composables and adds two JVM test files. Neither screen references
Media3EngineorTransformer, andUnit testsandStatic analysisboth passed on the same attempt. It was worth checking regardless — this is the first change in wave 2 to touchmainsource that the instrumented suite actually drives, so "test-only diff" was no longer available as a shortcut.Running tally for the two batches: wedge ×4, SAF picker ×3, adb-224 ×3, Media3-on-34 ×1.
Still live, on a PR whose production diff cannot cause it
PR #218, run
34000816016, API 34 gating leg, job101399258547:Same test, same exception, same API 34 gating leg as this ticket's opening report.
and, later in the same log,
ERROR | Failed to find ColorBuffer: 288— the graphics-stack noisethat keeps showing up around these.
Why this occurrence is worth adding
The production diff on #218 is: making
LibreMediaConverterAppopen, adding aprotected open val sweepScopethat still resolves toDispatchers.IOin production, and publishing theJobthatonCreatealready started. Nothing there touches Media3, a codec, a muxer, or a surface. It is aboutas clean a demonstration as this ticket is going to get that the failure is independent of the change
under test — which is the thing #190 says costs a human 40 minutes per occurrence to re-establish.
Also worth noting against the "either the marker is too narrow or the API 36 occurrence is rarer"
question in #190: this is API 34, so the spread is now 34, 36 and 37 for the same test.
Sighting 2026-09-07 — API 35, PR #269,
NPE: Activity has been destroyedRecording this because it cost a re-run and would otherwise be re-diagnosed from scratch. Evidence
re-read from the raw log rather than relayed.
Run
34161043035, job101862812172, E2E API 35. Red once, green on re-run(
101864620664) — same commit, no code change. APIs 33, 34, 36 and 37 all passed on that samecommit, and a local
run-e2e.sh 35on the same tree was 72/72/0.The stack contains no frame from any test file. It is the main looper's, and the Activity was
already destroyed when
onActivityasked for it — the shape this issue's title describes as failingon system services rather than assertions, bracketed seven seconds either side by the colour-buffer
path this issue already names in three of its four modes. The suite ran to completion:
expected 72, received 72, failed 1, completed cleanly: yes.The honest caveat
The failing test was
SafPickerRoundTripTest.aFailedSaveDeletesTheDocumentItCouldNotWrite, whichPR #269 modifies. So this is not a clean "unrelated test" sighting, and it should not be counted as
one. What argues it is this issue's class rather than that diff: no test frame in the stack, the
colour-buffer brackets, green on re-run of the identical commit, green on the other four API levels,
and green locally at API 35. What argues for caution: a pattern this well known is exactly what a
real regression could hide behind. It was re-run for that reason rather than waved through.
One process note worth recording
A re-run destroys the evidence. Job
101862812172now serves the re-run's log — the original20:59 failure is not retrievable at that job id any more. Anyone diagnosing one of these should
capture the log before pressing re-run, or the sighting cannot be checked afterwards. The excerpt
above survives only because it was saved locally at the time.
This ticket's own mode has a measured cause, and it is not starvation
transcodesH264ToH265AndReportsProgressfails because the emulator's Codec2 HAL processsegfaults. Read out of the per-test logcat in
e2e-report-api34for run34000816016attempt 1 — the API 34 gating leg of 2026-09-06, 62 ms after the test starts:
The decoder HAL dies and respawns; Media3 is left with a
DEAD_OBJECTcodec, writes no outputsample, and its 25-second export watchdog aborts the export — which is the
ExportException: Muxer errorthis ticket was filed on. Nothing in this app decided anything,and it is not the runner starving a working codec: it is a null-pointer dereference in the
emulator's own vendor codec library.
It is 6 for 6
Every gating leg-attempt that has ever failed this way carries the same HAL crash in the same
job's log — grepped for
c2@1.0-service-goldfishin the--- native crashes (tail 60) ---dumpthat
e2e-run.shwrites on failure:32855014836a132857067112a132919928048a133261618358a133588264439a134000816016a1Six occurrences in 1210 gating leg-attempts on API 33-36 across 2026-08-20..09-07 — 0.5%,
and 0/301 on API 33, 3/302 on API 34, 0/304 on API 35, 3/303 on API 36. (The API 37 row does not
run this test, so it is not in the denominator. Two other failures of this test are excluded and
said so below.)
What this settles, and what it does not
Settles: the comment above that called this "one emulator-codec weakness
expressed at two severities: deterministic on 37, load-dependent below it" was right about the
subsystem, and the
@FailsOnEmulatorApi37marker's stated reason — "fail inside the emulator'sown
c2.goldfish.h264.decoder" — is the same vendor HAL. It is one weakness, and this is its33-36 expression.
Does not settle: why the HAL dereferences null.
MediaCodec::reclaimis logged 8 ms beforethe crash, and a reclaim is the resource manager taking a codec instance away from a client — so
"a reclaim races
C2GoldfishAvcDec::processand the block pool is gone under it" is the obvioushypothesis and is untested. I am recording it as a hypothesis and not acting on it, because
this ticket has twice published a causal claim its own data did not support.
Not a candidate for a fix here. The crashing code is
/vendor/lib64/*inside the system image.This is environmental, in the same sense
@FailsOnEmulatorApi37is, and now with the same kind ofevidence behind it.
Two failures of the same test that are not this mode
Anchoring on the test name alone would have counted eight. Excluded, with the reason:
32545625459a1 (API 37, 2026-08-22) — 37 tests failed together, framework down, grallocassertion present. That is #108, and this test was collateral.
32669190757a1 (API 35, 2026-08-23) — predates #111's shape report, and the log carries thebare
FAILEDmarker with no message at all. Unclassifiable, not classified.Live census of every gating leg-attempt, 2026-08-20 .. 2026-09-07
Per leg-attempt, as this ticket requires; cancelled legs excluded (those are the
cancel-in-progresscancellations the tally above warns about, not runs). 1489 gatingleg-attempts, 129 failures, 8.7%. Classified by anchoring the failing test name to the
FAILEDmarker in each failing job's own log, then by the message under it.
adb224)The mode column sums to 127, not 129: the two transcode failures excluded above are left out of it rather than filed under a mode they are not.
Per-mode denominators are not the same denominator, which is the counting lesson from doing
this: three of these modes cannot occur on all five rows. The save tests carry
@FailsOnEmulatorApi37and only existed from 2026-09-06T15:12, so their rate is 7 in 88leg-attempts on API 33-36 — not 7 in 1489. The wedge has only ever happened on API 33/34.
Per-mode disposition
hasReadColorBufferDma; of the 6 that do not, 3 areadb224, 2 have no signature at all, and 1 has the framework-down signature without the assertion reaching the crash tail34001741668, 2026-09-06T00:36) isthePickedInputSurvivesARealRotationagain, andgit merge-base --is-ancestor 32ab54d <head>says that head does not contain #219's fix. At the prior rate, P(0 in 106) ≈ 0.15, so this is suggestive and not yet evidencee2e-wedge-api34of34001741668:started:the rotation test, nofinished:, every binder servicefoundadb224The newest mode splits into four, and only one of the seven is on today's
mainThe
@FailsOnEmulatorApi37-marked save tests (#226, #250) landed on 2026-09-06 and account forevery gating failure on API 33-36 since. Filed as one mode, they are not one:
34041593697a1 (35)fa10d94— the commit that added the testIllegalStateException: No compose hierarchies found, thrown directlyb23ff0f34056545386a1 (35)cbbaf74(#250)ComposeTimeoutException ... after 120000 msCONVERSION_TIMEOUT_MScase19e3539fixed. Not a system-service failure at all — API 35's software encode measured 134.8 s against a 120 s bound34057706195a1 (34)19e3539No compose hierarchies found×2, thrown directlyb23ff0f, so that fix did not close this shape34067653670a1 (35),34146936252a1 (35)d45abe7,73482520waited 300000ms for a node tagged action.saveFile, no composition errorawaitNodeonly appends the composition error whenfetchSemanticsNodesthrew, so the composition was readable throughout. Confirmed in34067653670's per-test logcat:MainActivityisRESUMEDat 23:49:52.981 and stays resumed until the rule tears it down 300 s later, and no conversion runs at all in that window. Open34146936252a2 (35)73482520waited 300000ms ...; last composition error: No compose hierarchiesPAUSEDat 17:38:30.522 and never resumes; the back press at 17:38:32.479 followsWaiting 5000ms for ... com.google.android.permissioncontroller, i.e. the permission-dialog helper #269 replaced withpm grant34161043035a1 (35)9fd96d0— #269 itselfNPE: Cannot run onActivity since Activity has been destroyed alreadySo "mode 5 at 8% of leg-attempts" would be wrong — most of these are shapes their own
follow-up commit had already answered, on heads that predate it.
The one on current
mainhas a mechanism, and it is #93's launcher ANR with a new victimFrom the per-test logcat in
e2e-report-api35of34161043035attempt 1 (attempt 2 wasgreen; the attempt-1 artifact survives with its own id, which is how it can still be read):
dismissThePicker's loop is:awaitAppFocus()iscomposeRule.activity.hasWindowFocus(). The launcher's ANR dialog makesthat false too — it is a fullscreen
system_serverwindow, which is the whole reasondismissASystemErrorDialogexists. So the guard reads false, the dialog is removed at41.169, and the back press at41.713is then sent on a reading taken before the occluderwas removed — into an app that is already in front with nothing to go back to. It finishes
MainActivity, the launcher comes to the front, and the rest of the test has no Activity.The method's own KDoc names this hazard — "a third from Recent would finish
MainActivityandtake the rest of the test with it" — and the guard is what is supposed to prevent it. It cannot,
because it cannot tell "the picker is still up" from "a system dialog is on top".
requireAReadableScreen, twenty lines away, already does the right thing: it re-reads afterdismissASystemErrorDialog()before doing anything else.dismissThePickerdoes not. PR tofollow — the fix is to re-read the focus after removing a dialog and skip the back press if the
app already has it, which removes an action rather than retrying one.
One more finding, recorded rather than acted on
ERROR_DIALOG_BUTTONSends withandroid:id/button1, the framework's generic AlertDialogpositive button — not an app-error-dialog id. In the same trace, at
20:59:35.689, afteraerr_waitandaerr_closeboth missed,button1was found and clicked at (927, 2274), andMainActivityresumed 339 ms later:dismissASystemErrorDialogclicked a button insideDocumentsUI's own save dialog, believing it to be a system error dialog. It happened to complete
the save the walk had just failed to complete. There is no reason it always would.
Method notes, since this ticket is where people land
gh run view --job <id> --logresolves by run and serves thelatest attempt. Read a historical attempt with
gh api --allow-escape-sequences /repos/{owner}/{repo}/actions/jobs/{job_id}/logs, taking thejob id from
/actions/runs/{run}/attempts/{n}/jobs. That is the correction already on thisticket, and it still holds.
/actions/runs/{run}/artifactsreturns every attempt's upload under the same name withdifferent ids and
created_at. Take the id whosecreated_atfalls in the attempt's window —gh run downloadtakes the newest, which on a re-run-to-green is the green one.e2e-report-apiNNcarriesoutputs/androidTest-results/connected/debug/<device>/logcat-<class>-<method>.txt, one file pertest, plus the JUnit XML with the untruncated stack. The job log truncates a stack to its first
frame. Everything above came out of those files.
Where this leaves the ticket, plus one correction to the comment above
Correction first
That comment's heading says "The newest mode splits into four". It is five, and the table
under it is what shows so:
No compose hierarchiesthrown directly,ComposeTimeoutExceptionat120 s,
waited 300000mswith no composition error,waited 300000mswith one, and the NPE — fivedistinct messages across seven leg-attempts. The table is right; the heading undercounted by
folding the NPE in with the rest, which is the one thing that comment was arguing against doing.
docs/ci-failure-modes.mdsays five.The work
PR #272 — the
dismissThePickerfix,docs/ci-failure-modes.md(the census and per-modedisposition), and a pointer to it from
CLAUDE.md. Local gate green: the instrumented suite is72 tests / 0 failures on API 33, 34, 35 and 36, plus the JVM gate; API 37 reported
NOT COVERED LOCALLY, no device attached.#270 — the two SAF-save failure shapes that still have no mechanism. Both predate #269, and the
first move is a re-count over post-#269 leg-attempts rather than a fix.
#271 —
ERROR_DIALOG_BUTTONSends withandroid:id/button1, which is anyAlertDialog'spositive button, and it was measured clicking one inside DocumentsUI.
#190 — has the marker-width measurement it asked for: 0.5% below API 37 against 100% on it, so
the marker is the right width, and the second signature (
c2@1.0-service-goldfish) that itsproposed report would need in order to have caught its own row 4.
What I would do with #102, leaving the call to you
Close it as an umbrella. Its "done means" is met for the mode it was filed on — a cause with a
measurement, and one that says plainly there is nothing here to fix. Nothing was retried, no timeout
was widened, and the one fix in #272 removes an action rather than adding a tolerance.
The four modes it accumulated all have somewhere better to live now: #108 and #122 are closed, #190
carries the reporting question, #270 and #271 carry what is genuinely still open, and
docs/ci-failure-modes.mdis the standing triage reference an umbrella ticket was being used as.Two things I could not answer and am not going to pretend otherwise:
leg-attempts against a prior 1.8%; P(0 in 106) ≈ 0.15 if nothing changed. Worth a re-count in a
week, not a claim today.
main. The only run on #269's head is34161043035, whose attempt 1 is the mode #272 fixes. So "the picker tests are fine now" iscurrently unmeasured in either direction, and #270 says so as its first step.
Retracting my sighting above: that NPE was not environmental
My comment of 2026-09-07T21:20 filed the API 35
Cannot run onActivity since Activity has been destroyed alreadyunder this ticket's environmental class. That attribution is wrong, and theinvestigation behind #272 has the trace.
What I offered as evidence — no test-file frame in the stack,
ColorBuffererrors seven secondseither side, green on re-run, green on the other four levels — was all true and all beside the point.
The Activity was destroyed by the test's own back press:
dismissThePickerguarded its back presses onActivity.hasWindowFocus, and a fullscreensystem_serverANR dialog makes that false too — so it could not tell "the picker is still up" from"a dialog is on top of an app that is already in front". #272 re-reads the focus after a dialog is
actually dismissed, which removes the press rather than retrying it.
The part worth keeping
My own caveat at the time was the right instinct and I did not follow it far enough:
That is precisely what happened. Four independent signals agreed, the failing test was one the PR
had modified, and the pattern still lost to a logcat trace. A no-test-frame stack is evidence that
the test did not throw, not evidence that the test did not cause it — the back press that killed
the Activity was ours, and it left no frame because the death was asynchronous.
The re-run that made it green is what nearly buried it, exactly as the "a re-run destroys the log"
note in the same comment warns. That note stands; so does the practice of saving the log first,
which is the only reason the trace above still exists.
Closing as an umbrella: every mode now has a home
This ticket asked for "either a cause with a measurement behind it, or a documented decision to
treat it as environmental with the same honesty
@FailsOnEmulatorApi37gets." Both halves are nowanswered, and the modes it accumulated have outgrown one ticket.
Its own mode — the Media3 export watchdog — has a measured cause, and it is environmental. The
emulator's Codec2 HAL null-derefs at
C2Block2D::handle()+4viagetClientUsageinlibcodec2_goldfish_common.so, killingc2.goldfish.h264.decoder; Media3's 25 s watchdog thenaborts. 6 of 6 occurrences carry that signature. The crashing code is
/vendor/lib64/*— thereis nothing in this repo to fix. Rate: 6 in 1210 API 33–36 gating leg-attempts, 0.5%.
The census that settles the counting question. 1489 gating leg-attempts, 2026-08-20 → 09-07,
129 failures (8.7%). Those 129 sit in 83 runs, 45 of which ended green after a re-run — so a
run-level count reports 38 where there were 129, and deletes exactly the ones somebody already
judged to be noise. Count per leg-attempt. Full method, per-mode dispositions and the
evidence-reading procedure are now in
docs/ci-failure-modes.md.Where the modes went:
MainActivitydismissASystemErrorDialogclicking anybutton1adb224The lesson this ticket paid for twice
It retracted two causal claims its own data did not support, and then a third — mine. I filed the
2026-09-07 API 35
Activity has been destroyedNPE here as environmental on four agreeing signals,and it was a real defect the whole time; the retraction above has the trace. A no-test-frame stack
is evidence the test did not throw, not evidence the test did not cause it.
The other durable finding: a re-run destroys the log. The trace that overturned my attribution
survives only because someone saved it before pressing re-run. That practice is now written into
docs/ci-failure-modes.md.Closing. New sightings should go to the specific ticket for their mode, or to a new one citing the
doc — not here.