The gating emulator legs are flaky enough that a ten-PR stack cannot land without repeated retries #190
Open
opened 2026-09-02 04:03:39 +00:00 by JMR-dev
·
5 comments
No Branch/Tag Specified
main
fix/102-picker-back-press-overshoot
fix/268-saf-picker-determinism
feat/ogg-vorbis-libvorbis
feat/expedited-conversion-work
fix/gate-cache-in-worktrees
chore/gate-runs-shellcheck
test/publish-delete-arm-real-provider
docs/e8-instrumented-coverage
docs/api37-carrier-count-drift
docs/e7-second-constraint
test/publish-to-a-real-saf-destination
fix/api37-report-match-line
fix/api37-task-snapshot-crash
test/join-failure-message-on-device
docs/e2e-read-findings-e7
test/cancelling-a-running-export
test/reattach-to-a-running-job
test/content-uri-reaches-ffmpeg
fix/launcher-wiring-waits-for-the-pick
fix/cancel-tests-need-a-slower-encode
test/cancelling-a-running-join
test/cancelling-a-running-session
test/notification-cancel-action
test/ffmpeg-progress-is-observed
test/fallback-asserts-the-path
test/flac-and-opus-assert-their-format
docs/e2e-read-findings
docs/wave4-coverage-numbers
fix/injectable-startup-sweep-scope
test/session-outcome-seam
test/launcher-callback-identity
test/theme-follows-system-dark
test/audio-drop-arm
fix/rotation-waits-for-recreation
fix/convert-guards-on-ready
test/retry-save-mime
test/hardware-progress-reaches-workmanager
test/ffprobe-mapping-seam
test/device-codec-enumeration-seam
test/unknown-container-row
test/null-message-fallbacks
test/cancel-reaches-workmanager
docs/coverage-wave3-recovery
test/concat-engine-seam
docs/coverage-wave3
test/mediaprobe-merge-seam
test/adaptive-shell-wiring
test/aac-audio-args
test/notification-progress-text
test/media3-muxer-guard
test/hardware-fallback-and-cancellation
test/unprobeable-join-clip
test/one-branch-outcomes
test/foreground-type-regimes
fix/bound-wedge-diagnostics
docs/coverage-wave2
test/screen-wiring
test/viewmodel-setters
test/join-state-mapping
test/conversion-state-mapping
test/dedupe-user-messages
fix/restore-stack-merges
test/refused-jobs
test/concatworker-failure-arms
test/container-capabilities-audio
test/readspec-enum-fallbacks
test/outputpublisher-seams
test/mediaprobe-track-seam
test/outputpublisher-partial-branches
test/fake-provider-scaffolding
docs/coverage-read-findings
chore/gitignore-kotlin
test/bound-the-hangs
docs/coverage-remeasure
ci/baseline-counter-precision
fix/invalid-suggestion-chip
ci/wedged-leg-report
fix/reattachment-overwrites-pick
test/theme-live-branches
fix/failed-save-retry
fix/empty-composition-crash
ci/advisory-failure-report
docs/seven-run-counts
test/release-permission-guard
ci/build-workflow-permissions
docs/api37-point-release
docs/benchmark-populate-path
fix/dead-assertion-probe-test
ci/actionlint
test/device-codecs-encode-consequence
fix/sdkmanager-pipefail
docs/readme-restart-claim
fix/saf-picker-root-discovery
fix/probe-dispatcher-seam
test/media3engine-mime-tables
test/mediaprobe-pure-helpers
fix/codec-vocabulary-drift
docs/robolectric-rationale-correction
docs/api37-advisory-counts
test/r38-8-saf-e2e
test/r38-7-join-states
test/r38-6-conversion-states
test/r38-5-state-seam
fix/jacoco-robolectric-coverage
docs/instrumented-tests-correction
test/r38-2-filecard
test/r38-4-advanced-picker
test/r38-3-pickers
tools/file-issue-script
tools/api-37-emulator
fix/review-app-gaps
docs/review-corrections
No results found.
Labels
Clear labels
above-cut
accessibility
backlog
bug
confirmed
documentation
duplicate
enhancement
good first issue
help wanted
invalid
plausible
question
sev:high
sev:low
sev:medium
wontfix
Worked autonomously overnight: local, JVM-verifiable, no product decision
Barrier affecting people with disabilities
Held for manual review: product/UX call, CI/workflow, hardware, or unverifiable here
Something isn't working
Reviewer demonstrated the defect
Improvements or additions to documentation
This issue or pull request already exists
New feature or request
Good for newcomers
Extra attention is needed
This doesn't seem right
Reviewer could not fully demonstrate it; treat as unproven
Further information is requested
High severity
Low severity
Medium severity
This will not be worked on
Milestone
No items
No Milestone
Projects
Clear projects
No projects
Notifications
Due Date
No due date set.
Dependencies
No dependencies set.
Reference: JMR-dev/LibreMediaConverter#190
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
The gating emulator legs are flaky enough that a stack cannot land without repeated retries
Measured while shepherding the wave-3 stack (#179-#189) on 2026-09-02. Every PR in that stack
whose diff is a JVM test file and nothing else — no production code, no
androidTestchange —and yet four separate gating-leg failures happened, each naming a different instrumented test.
expected 57, received 57, failed 0ConversionWorkerTest.routesAFastMp4JobByDeviceCapabilityfailed 1, completed cleanly: no,INSTRUMENTATION_ABORTED: System has crashedSafPickerRoundTripTest.pickingAFileThroughTheSystemPickerFillsInTheFileCardfailed 1, completed cleanly: yesMedia3EngineTest.transcodesH264ToH265AndReportsProgressfailed 1, completed cleanly: yesAll four carried the same native abort in the log:
Row 2 also showed the system going down around it —
am get-current-userfailing with exit 20,and both test APKs failing to uninstall.
Why this is not four bugs
Each PR's diff was checked before retrying. #183 is 99 lines of one JVM test file; #184 is 89 lines
of another. No production file changed on
mainacross the whole first half of the stack —git diff d354f64..origin/main -- app/src/mainwas empty at the point rows 2-4 happened. A JVMtest file cannot make a Media3 hardware transcode fail on an emulator.
Every one passed on retry. #183 needed three attempts, drawing a different unrelated test each
time.
Why it is worth a ticket rather than a shrug
docs/api-37-emulator-crash.mdand #102 already describe legs failing on system services ratherthan assertions, and #108 records the API 37 gating leg aborting after every test passes. What is
new here is the cost at scale: this is the first time a ten-PR stack has been landed, and the
per-PR retry rate turned a mechanical merge into an hour of adjudication. Each halt needed a human
reading a log to decide whether the red was real, because the only thing distinguishing "flaky
emulator" from "your change broke it" is the run shape plus the diff.
Two things would have made it cheap, both cheaper than fixing the emulator:
hasReadColorBufferDmaabort is a recognisable signature.e2e-report-shape.shalreadyparses the run and knows the failure count; it could also say "a native graphics abort was
present in this run", which is the single fact that turns 40 minutes of log reading into a
glance. It would not change any conclusion — the leg should still be red — but it would put the
evidence where the decision is made.
Media3EngineTest.transcodesH264ToH265AndReportsProgressfails on API 36 too, not only onAPI 37 where it carries
@FailsOnEmulatorApi37. Either the marker is too narrow or the API 36occurrence is rarer and nobody has hit it before. Worth measuring before deciding, because
widening the marker moves a test out of the gating set and that is not free.
Not proposed
Retrying automatically on this signature. A leg that goes green on the third attempt is still
telling you something, and a retry rule keyed on "a native abort was present" would eventually
swallow a real regression that happened to run alongside one. The report should surface the
signature; a person should still decide.
A fifth mode, from the same shepherding run — and the one that matters most, because it is the only one that does not produce a verdict at all.
#187, E2E API 33 — the wedge (#122).
All sixty tests reported, so the API 33 regime was exercised. But
failed:readsunknown, so that leg could not have said whether anything broke — exactly whatdocs/coverage-read-findings.mdrecords about the wedge costing the verdict rather than the execution.Why this one is different from the four in the issue body. Those four were on PRs whose entire diff was a JVM test file, so "unrelated flake" was cheap to establish from the diff alone. #187 is the first PR in the stack that changes production code — it cuts
MediaProbe.probe's two-probe merge into a seam, andMediaProbeis exercised by the instrumented suite. A red leg there deserves suspicion, and a leg that answersunknowngives you nothing to be suspicious with. Retried for a real verdict rather than reasoned around.Note also
bad color buffer handleand ~1 GB of 2.4 GB memory available — the same graphics-stack neighbourhood as thehasReadColorBufferDmaabort in the other four, which is worth recording in case the five modes turn out to have one cause. #102 already suspects as much.This does not change the recommendation in the body. It sharpens it: the report already carries
wedged:(from #118), and that row is the reason this was diagnosable in one glance while the other four each took a log read. Whatever is done for the native-abort signature should follow the wedge row's precedent — put the fact in the table, leave the decision to a person.Wave-4 data point: six PRs, three unrelated red legs, on changes that cannot have caused them.
Filing this because the ticket's premise — "a ten-PR stack cannot land without repeated retries" — is now measured against a second wave rather than inferred from the first.
PRs #206–#211 (wave 4, #192–#195/#198/#199). Every failure so far has been in a test the PR does not touch:
Media3EngineTest.anUnwritableOutputPathFailsInsteadOfHanging, then "Instrumentation run failed due to Process crashed" — green on re-runFileCardTestOutputPublisherStagingTestat:184— failed twice consecutively, see belowAndroidDeviceCodecsseamSafPickerRoundTripTest: "the system picker would not close: after 4 back presses the app still does not have the window focus, and com.google.android.documentsui is in front"Three different mechanisms, three different legs, none touching the diff under test.
The
#209one is not the same kind of flake and is worth separating. It is the JVM suite, it failed on both the original run and the re-run, and it does not reproduce locally — three consecutive full-suite--rerun-tasksruns on that exact branch are green. That is #159, which predicts precisely this ("a loaded CI runner is where it shows"); I have added the new occurrence there.What makes it more than a retry problem: adding a Robolectric test adds an
Application, and everyApplicationschedules another backgroundsweepStaging(). So the pressure on that race grows with the suite, which means it gets worse exactly as this kind of work proceeds. #209's diff is one test case, and it appears to have been enough to flip that race from never-seen to twice-in-a-row on CI.So the two tickets interact: #190 is the retry cost, and #159 is a component of it that is growing rather than static. Fixing #159 would remove one of the three mechanisms above outright.
Wave 4 finished: twelve PRs, six distinct flake mechanisms, and one PR that needed three attempts on a single leg. Closing out the data I have been adding here, since the wave is now complete and the numbers are final.
Every failure across #206–#217 was in a test the PR does not touch. Six distinct mechanisms:
Media3EngineTest+ "Instrumentation run failed due to Process crashed"OutputPublisherStagingTeststaging raceSafPickerRoundTripTest— "the system picker would not close"DeadSystemRuntimeExceptionatPowerManager$WakeLock.release, thencmd: Can't find service: activityCan't find service: settings/inputduring setup#215 is the sharpest illustration of this ticket's premise. It is a test-only change sitting directly on
main— no production diff at all — and its API 37 leg has failed three times for three different reasons: system server death, emulator never coming up, and the SAF picker refusing to close. Its unit tests, static analysis and five other emulator legs pass every time.Two observations that may be useful for whatever fix this ticket eventually gets:
wedged:row from #118 is doing its job. #217's API 33 leg reportedwedged: yes -- gradle was killed after 1200s and never returnedrather than the misleadingreceived: N, completed cleanly: yesthat CLAUDE.md records as the pre-#118 behaviour. The diagnosis took one line of log.Related and worth reading together: #159 (which I have separately upgraded — it now reproduces on the development host, not only on CI) and #122.
A fifth mode, and this one makes the shape table print a fully-green row for a red leg
PR #215, run
33697801472retried, API 37 gating leg, job101397719644:and the leg is red:
Every test ran and every test passed. The process then died in teardown:
Instrumentation.finish()is after the last test. AGP failsDeviceProviderInstrumentTestTaskbecause the runner process aborted; the XML was already writtenand records nothing, because nothing test-shaped went wrong.
Why this one matters more than the count
The four modes in the table above all leave a mark the shape report can see — a non-zero
failed, orcompleted cleanly: no, or awedged:row. This one leaves none.e2e-report-shape.shreadsthe test XML, the XML is complete and clean, so the table prints:
next to a red leg. That is the worst possible output for the decision the report exists to support —
it does not merely fail to help, it actively says the run was fine. Someone reading the summary to
decide "flaky emulator or my change?" gets an answer that fits neither.
Note this is not #108. There the abort truncates the run and
receivedis short. Here the run iscomplete and the abort is after it.
What would fix it, in the same place point 1 already proposes
e2e-report-shape.shknows the gradle exit status. A row likewhenever the exit is non-zero and
failed: 0costs nothing and turns this into a glance. ThehasReadColorBufferDmasignature line proposed in point 1 would also have fired here — there is anative crash in the artifact — but the cheaper and more general fact is that the table currently
never mentions whether gradle actually succeeded.
Occurrence log for this leg on #215
received 55/57, failed 1, completed cleanly: noINSTRUMENTATION_ABORTED: System has crashed, tombstoneAssertion failed: !rcEnc->featureInfo()->hasReadColorBufferDma— row-2 shape, #102received 57/57, failed 0, completed cleanly: yesInstrumentation.finish()— this commentTwo attempts, two different infra failures, on a PR whose diff is one JVM test file and one KDoc
edit. That is the cost this ticket is about, still being paid.
Filed while shepherding the wave-4 PRs (#206-#217). Ten of the twelve are green on every gating leg;
#217's API 33 rotation failure passed on retry and is written up on #122 as a test-synchronisation
bug rather than an image one.
Both of this ticket's "worth measuring before deciding" items now have the measurement
From the #102 census — every gating E2E leg-attempt in the repo's history, 2026-08-20..09-07,
1489 leg-attempts and 129 failures, counted per leg-attempt because a re-run to green replaces
the run's conclusion. Full write-up in
docs/ci-failure-modes.md; the two halves that belong here:2. The marker should not be widened, and now there is a rate rather than an impression
Media3EngineTest.transcodesH264ToH265AndReportsProgressreally does fail below API 37, and thisticket's row 4 (#184, API 36) is one of six:
32855014836a132857067112a132919928048a133261618358a133588264439a134000816016a1Six in 1210 gating leg-attempts on API 33-36 — 0.5%. Split 3 on API 34, 3 on API 36, and
zero on API 33 and API 35 across 301 and 304 leg-attempts respectively.
So the answer to "either the marker is too narrow or the API 36 occurrence is rarer" is rarer,
by a lot: on the android-37 images this test fails every time, and below them once in two
hundred legs. Moving it out of the gating set on 33-36 would trade a 1-in-200 red for permanently
not testing the hardware transcode on four API levels. The marker is the right width.
1. There is a second recognisable signature, and it is a better one than the abort
The reason this test fails is not load. It is that the emulator's Codec2 HAL process segfaults.
From the per-test logcat in
e2e-report-api34of34000816016a1, 62 ms after the test starts:The HAL dies and respawns; Media3 is left with a dead codec and its own 25-second export watchdog
aborts the export, which is the
Muxer errorin the job log. Six of six of the occurrencesabove carry that crash in the same job's
--- native crashes (tail 60) ---dump — grepc2@1.0-service-goldfish.That matters for this ticket's proposal #1 in a specific way: the
hasReadColorBufferDmaabortwould not have flagged row 4. It is an API 37 signature — 24 of the 30 API 37 failures since
2026-08-27 carry it, and it appears on none of the six above. A report that says only "a native
graphics abort was present" would have left #184's API 36 red looking exactly like a product
failure, which is the 40 minutes of log reading this ticket is about.
If the run-shape report is going to name signatures, the useful set looks like two, not one:
hasReadColorBufferDma— the gralloc assertion (API 37's, ex-#108)c2@1.0-service-goldfish— the codec HAL crashGrep that exact string and not the friendlier one.
Codec2 component "c2.goldfish.h264.decoder" diedis aCCodecline in the main buffer, and it is in none of the six job logs; what reachesadb logcat -d -b crashis the tombstone, whoseCmdline:names/vendor/bin/hw/android.hardware.media.c2@1.0-service-goldfish. Measured across all six.Both signatures are already in the crash tail
e2e-run.shwrites on failure, so this is a grepover a file the script already has, in the same shape as the
wedged:row #118 added. And it stays within thisticket's own "Not proposed": neither changes a conclusion or retries anything — the leg is still
red, it just says which of the emulator's subsystems went down underneath it.
One correction to this ticket's framing, offered rather than assumed
The row-4 failure is not in the same family as rows 2 and 3. Rows 2 and 3 are API 37 and
hasReadColorBufferDma; row 4 is API 36 and a codec HAL segfault, with no gralloc abort anywherein it. "All four carried the same native abort in the log" holds for rows 1-3. For row 4 it does
not: I grepped
hasReadColorBufferDmain the job log of all six occurrences above, including33588264439which is row 4, and it is absent from every one — whilec2@1.0-service-goldfishispresent in every one. Two emulator faults, not one, which is why the two-signature list above is
worth having.