The gating emulator legs are flaky enough that a ten-PR stack cannot land without repeated retries #190

Open
opened 2026-09-02 04:03:39 +00:00 by JMR-dev · 5 comments
JMR-dev commented 2026-09-02 04:03:39 +00:00 (Migrated from github.com)

The gating emulator legs are flaky enough that a stack cannot land without repeated retries

Measured while shepherding the wave-3 stack (#179-#189) on 2026-09-02. Every PR in that stack
whose diff is a JVM test file and nothing else
— no production code, no androidTest change —
and yet four separate gating-leg failures happened, each naming a different instrumented test.

# PR leg failing test run shape
1 #181 API 37 (none — the suite passed) expected 57, received 57, failed 0
2 #183 attempt 1 API 37 ConversionWorkerTest.routesAFastMp4JobByDeviceCapability failed 1, completed cleanly: no, INSTRUMENTATION_ABORTED: System has crashed
3 #183 attempt 2 API 37 SafPickerRoundTripTest.pickingAFileThroughTheSystemPickerFillsInTheFileCard failed 1, completed cleanly: yes
4 #184 API 36 Media3EngineTest.transcodesH264ToH265AndReportsProgress failed 1, completed cleanly: yes

All four carried the same native abort in the log:

F DEBUG : Abort message: 'Assertion failed: !rcEnc->featureInfo()->hasReadColorBufferDma'

Row 2 also showed the system going down around it — am get-current-user failing with exit 20,
and both test APKs failing to uninstall.

Why this is not four bugs

Each PR's diff was checked before retrying. #183 is 99 lines of one JVM test file; #184 is 89 lines
of another. No production file changed on main across the whole first half of the stack —
git diff d354f64..origin/main -- app/src/main was empty at the point rows 2-4 happened. A JVM
test file cannot make a Media3 hardware transcode fail on an emulator.

Every one passed on retry. #183 needed three attempts, drawing a different unrelated test each
time.

Why it is worth a ticket rather than a shrug

docs/api-37-emulator-crash.md and #102 already describe legs failing on system services rather
than assertions, and #108 records the API 37 gating leg aborting after every test passes. What is
new here is the cost at scale: this is the first time a ten-PR stack has been landed, and the
per-PR retry rate turned a mechanical merge into an hour of adjudication. Each halt needed a human
reading a log to decide whether the red was real, because the only thing distinguishing "flaky
emulator" from "your change broke it" is the run shape plus the diff.

Two things would have made it cheap, both cheaper than fixing the emulator:

  1. The hasReadColorBufferDma abort is a recognisable signature. e2e-report-shape.sh already
    parses the run and knows the failure count; it could also say "a native graphics abort was
    present in this run"
    , which is the single fact that turns 40 minutes of log reading into a
    glance. It would not change any conclusion — the leg should still be red — but it would put the
    evidence where the decision is made.
  2. Media3EngineTest.transcodesH264ToH265AndReportsProgress fails on API 36 too, not only on
    API 37 where it carries @FailsOnEmulatorApi37. Either the marker is too narrow or the API 36
    occurrence is rarer and nobody has hit it before. Worth measuring before deciding, because
    widening the marker moves a test out of the gating set and that is not free.

Not proposed

Retrying automatically on this signature. A leg that goes green on the third attempt is still
telling you something, and a retry rule keyed on "a native abort was present" would eventually
swallow a real regression that happened to run alongside one. The report should surface the
signature; a person should still decide.

## The gating emulator legs are flaky enough that a stack cannot land without repeated retries Measured while shepherding the wave-3 stack (#179-#189) on 2026-09-02. **Every PR in that stack whose diff is a JVM test file and nothing else** — no production code, no `androidTest` change — and yet four separate gating-leg failures happened, each naming a *different* instrumented test. | # | PR | leg | failing test | run shape | |---|---|---|---|---| | 1 | #181 | API 37 | *(none — the suite passed)* | `expected 57, received 57, failed 0` | | 2 | #183 attempt 1 | API 37 | `ConversionWorkerTest.routesAFastMp4JobByDeviceCapability` | `failed 1, completed cleanly: no`, `INSTRUMENTATION_ABORTED: System has crashed` | | 3 | #183 attempt 2 | API 37 | `SafPickerRoundTripTest.pickingAFileThroughTheSystemPickerFillsInTheFileCard` | `failed 1, completed cleanly: yes` | | 4 | #184 | API 36 | `Media3EngineTest.transcodesH264ToH265AndReportsProgress` | `failed 1, completed cleanly: yes` | All four carried the same native abort in the log: ``` F DEBUG : Abort message: 'Assertion failed: !rcEnc->featureInfo()->hasReadColorBufferDma' ``` Row 2 also showed the system going down around it — `am get-current-user` failing with exit 20, and both test APKs failing to uninstall. ### Why this is not four bugs Each PR's diff was checked before retrying. #183 is 99 lines of one JVM test file; #184 is 89 lines of another. **No production file changed on `main` across the whole first half of the stack** — `git diff d354f64..origin/main -- app/src/main` was empty at the point rows 2-4 happened. A JVM test file cannot make a Media3 hardware transcode fail on an emulator. Every one passed on retry. #183 needed **three** attempts, drawing a different unrelated test each time. ### Why it is worth a ticket rather than a shrug `docs/api-37-emulator-crash.md` and #102 already describe legs failing on system services rather than assertions, and #108 records the API 37 gating leg aborting after every test passes. What is new here is the **cost at scale**: this is the first time a ten-PR stack has been landed, and the per-PR retry rate turned a mechanical merge into an hour of adjudication. Each halt needed a human reading a log to decide whether the red was real, because the only thing distinguishing "flaky emulator" from "your change broke it" is the run shape plus the diff. Two things would have made it cheap, both cheaper than fixing the emulator: 1. **The `hasReadColorBufferDma` abort is a recognisable signature.** `e2e-report-shape.sh` already parses the run and knows the failure count; it could also say *"a native graphics abort was present in this run"*, which is the single fact that turns 40 minutes of log reading into a glance. It would not change any conclusion — the leg should still be red — but it would put the evidence where the decision is made. 2. **`Media3EngineTest.transcodesH264ToH265AndReportsProgress` fails on API 36 too**, not only on API 37 where it carries `@FailsOnEmulatorApi37`. Either the marker is too narrow or the API 36 occurrence is rarer and nobody has hit it before. Worth measuring before deciding, because widening the marker moves a test out of the gating set and that is not free. ### Not proposed Retrying automatically on this signature. A leg that goes green on the third attempt is still telling you something, and a retry rule keyed on "a native abort was present" would eventually swallow a real regression that happened to run alongside one. The report should surface the signature; a person should still decide.
JMR-dev commented 2026-09-02 04:36:36 +00:00 (Migrated from github.com)

A fifth mode, from the same shepherding run — and the one that matters most, because it is the only one that does not produce a verdict at all.

#187, E2E API 33 — the wedge (#122).

expected:          60
received:          60
failed:            unknown
wedged:            yes -- gradle was killed after 1200s and never returned
completed cleanly: no
ERROR | bad color buffer handle 373
MemAvailable:    1061376 kB   (of MemTotal 2526376 kB)

All sixty tests reported, so the API 33 regime was exercised. But failed: reads unknown, so that leg could not have said whether anything broke — exactly what docs/coverage-read-findings.md records about the wedge costing the verdict rather than the execution.

Why this one is different from the four in the issue body. Those four were on PRs whose entire diff was a JVM test file, so "unrelated flake" was cheap to establish from the diff alone. #187 is the first PR in the stack that changes production code — it cuts MediaProbe.probe's two-probe merge into a seam, and MediaProbe is exercised by the instrumented suite. A red leg there deserves suspicion, and a leg that answers unknown gives you nothing to be suspicious with. Retried for a real verdict rather than reasoned around.

Note also bad color buffer handle and ~1 GB of 2.4 GB memory available — the same graphics-stack neighbourhood as the hasReadColorBufferDma abort in the other four, which is worth recording in case the five modes turn out to have one cause. #102 already suspects as much.

This does not change the recommendation in the body. It sharpens it: the report already carries wedged: (from #118), and that row is the reason this was diagnosable in one glance while the other four each took a log read. Whatever is done for the native-abort signature should follow the wedge row's precedent — put the fact in the table, leave the decision to a person.

A fifth mode, from the same shepherding run — and the one that matters most, because it is the only one that does not produce a verdict at all. **#187, E2E API 33 — the wedge (#122).** ``` expected: 60 received: 60 failed: unknown wedged: yes -- gradle was killed after 1200s and never returned completed cleanly: no ERROR | bad color buffer handle 373 MemAvailable: 1061376 kB (of MemTotal 2526376 kB) ``` All sixty tests reported, so the API 33 regime *was* exercised. But `failed:` reads `unknown`, so that leg could not have said whether anything broke — exactly what `docs/coverage-read-findings.md` records about the wedge costing the verdict rather than the execution. **Why this one is different from the four in the issue body.** Those four were on PRs whose entire diff was a JVM test file, so "unrelated flake" was cheap to establish from the diff alone. #187 is the first PR in the stack that changes production code — it cuts `MediaProbe.probe`'s two-probe merge into a seam, and `MediaProbe` is exercised by the instrumented suite. A red leg there deserves suspicion, and a leg that answers `unknown` gives you nothing to be suspicious *with*. Retried for a real verdict rather than reasoned around. Note also `bad color buffer handle` and ~1 GB of 2.4 GB memory available — the same graphics-stack neighbourhood as the `hasReadColorBufferDma` abort in the other four, which is worth recording in case the five modes turn out to have one cause. #102 already suspects as much. This does not change the recommendation in the body. It sharpens it: the report already carries `wedged:` (from #118), and that row is the reason this was diagnosable in one glance while the other four each took a log read. **Whatever is done for the native-abort signature should follow the wedge row's precedent** — put the fact in the table, leave the decision to a person.
JMR-dev commented 2026-09-02 23:29:55 +00:00 (Migrated from github.com)

Wave-4 data point: six PRs, three unrelated red legs, on changes that cannot have caused them.

Filing this because the ticket's premise — "a ten-PR stack cannot land without repeated retries" — is now measured against a second wave rather than inferred from the first.

PRs #206–#211 (wave 4, #192–#195/#198/#199). Every failure so far has been in a test the PR does not touch:

PR what it changes red leg what actually failed
#207 one new JVM test file E2E API 34 Media3EngineTest.anUnwritableOutputPathFailsInsteadOfHanging, then "Instrumentation run failed due to Process crashed" — green on re-run
#209 one case added to FileCardTest Unit tests OutputPublisherStagingTest at :184 — failed twice consecutively, see below
#210 AndroidDeviceCodecs seam E2E API 37 (gating) SafPickerRoundTripTest: "the system picker would not close: after 4 back presses the app still does not have the window focus, and com.google.android.documentsui is in front"

Three different mechanisms, three different legs, none touching the diff under test.

The #209 one is not the same kind of flake and is worth separating. It is the JVM suite, it failed on both the original run and the re-run, and it does not reproduce locally — three consecutive full-suite --rerun-tasks runs on that exact branch are green. That is #159, which predicts precisely this ("a loaded CI runner is where it shows"); I have added the new occurrence there.

What makes it more than a retry problem: adding a Robolectric test adds an Application, and every Application schedules another background sweepStaging(). So the pressure on that race grows with the suite, which means it gets worse exactly as this kind of work proceeds. #209's diff is one test case, and it appears to have been enough to flip that race from never-seen to twice-in-a-row on CI.

So the two tickets interact: #190 is the retry cost, and #159 is a component of it that is growing rather than static. Fixing #159 would remove one of the three mechanisms above outright.

**Wave-4 data point: six PRs, three unrelated red legs, on changes that cannot have caused them.** Filing this because the ticket's premise — "a ten-PR stack cannot land without repeated retries" — is now measured against a second wave rather than inferred from the first. PRs #206–#211 (wave 4, #192–#195/#198/#199). Every failure so far has been in a test the PR does not touch: | PR | what it changes | red leg | what actually failed | |---|---|---|---| | #207 | one new JVM test file | E2E API 34 | `Media3EngineTest.anUnwritableOutputPathFailsInsteadOfHanging`, then *"Instrumentation run failed due to Process crashed"* — **green on re-run** | | #209 | one case added to `FileCardTest` | Unit tests | `OutputPublisherStagingTest` at `:184` — **failed twice consecutively**, see below | | #210 | `AndroidDeviceCodecs` seam | E2E API 37 (gating) | `SafPickerRoundTripTest`: *"the system picker would not close: after 4 back presses the app still does not have the window focus, and com.google.android.documentsui is in front"* | Three different mechanisms, three different legs, none touching the diff under test. **The `#209` one is not the same kind of flake and is worth separating.** It is the JVM suite, it failed on both the original run and the re-run, and it does **not** reproduce locally — three consecutive full-suite `--rerun-tasks` runs on that exact branch are green. That is [#159](https://github.com/JMR-dev/LibreMediaConverter/issues/159), which predicts precisely this ("a loaded CI runner is where it shows"); I have added the new occurrence there. What makes it more than a retry problem: **adding a Robolectric test adds an `Application`, and every `Application` schedules another background `sweepStaging()`**. So the pressure on that race grows with the suite, which means it gets worse exactly as this kind of work proceeds. #209's diff is one test case, and it appears to have been enough to flip that race from never-seen to twice-in-a-row on CI. So the two tickets interact: #190 is the retry cost, and #159 is a component of it that is growing rather than static. Fixing #159 would remove one of the three mechanisms above outright.
JMR-dev commented 2026-09-03 00:42:17 +00:00 (Migrated from github.com)

Wave 4 finished: twelve PRs, six distinct flake mechanisms, and one PR that needed three attempts on a single leg. Closing out the data I have been adding here, since the wave is now complete and the numbers are final.

Every failure across #206–#217 was in a test the PR does not touch. Six distinct mechanisms:

mechanism where filed as
Media3EngineTest + "Instrumentation run failed due to Process crashed" #207, E2E 34 —
OutputPublisherStagingTest staging race #209, Unit tests (×2) #159
SafPickerRoundTripTest — "the system picker would not close" #210 and #215, E2E 37 —
DeadSystemRuntimeException at PowerManager$WakeLock.release, then cmd: Can't find service: activity #215, E2E 37 #102
emulator never booted — Can't find service: settings / input during setup #215, E2E 37 —
wedged — "gradle was killed after 1200s and never returned" #217, E2E 33 #122

#215 is the sharpest illustration of this ticket's premise. It is a test-only change sitting directly on main — no production diff at all — and its API 37 leg has failed three times for three different reasons: system server death, emulator never coming up, and the SAF picker refusing to close. Its unit tests, static analysis and five other emulator legs pass every time.

Two observations that may be useful for whatever fix this ticket eventually gets:

  • The wedged: row from #118 is doing its job. #217's API 33 leg reported wedged: yes -- gradle was killed after 1200s and never returned rather than the misleading received: N, completed cleanly: yes that CLAUDE.md records as the pre-#118 behaviour. The diagnosis took one line of log.
  • The failures are not concentrated in one leg or one API. They span API 33, 34 and 37, the unit-test job, and three unrelated instrumented tests. So a fix targeted at any single test would not have changed the retry cost of this wave.

Related and worth reading together: #159 (which I have separately upgraded — it now reproduces on the development host, not only on CI) and #122.

**Wave 4 finished: twelve PRs, six distinct flake mechanisms, and one PR that needed three attempts on a single leg.** Closing out the data I have been adding here, since the wave is now complete and the numbers are final. Every failure across #206–#217 was in a test the PR does not touch. Six distinct mechanisms: | mechanism | where | filed as | |---|---|---| | `Media3EngineTest` + *"Instrumentation run failed due to Process crashed"* | #207, E2E 34 | — | | `OutputPublisherStagingTest` staging race | #209, Unit tests (×2) | #159 | | `SafPickerRoundTripTest` — *"the system picker would not close"* | #210 and #215, E2E 37 | — | | `DeadSystemRuntimeException` at `PowerManager$WakeLock.release`, then `cmd: Can't find service: activity` | #215, E2E 37 | #102 | | emulator never booted — `Can't find service: settings` / `input` during setup | #215, E2E 37 | — | | **wedged** — *"gradle was killed after 1200s and never returned"* | #217, E2E 33 | #122 | **#215 is the sharpest illustration of this ticket's premise.** It is a test-only change sitting directly on `main` — no production diff at all — and its API 37 leg has failed **three times for three different reasons**: system server death, emulator never coming up, and the SAF picker refusing to close. Its unit tests, static analysis and five other emulator legs pass every time. Two observations that may be useful for whatever fix this ticket eventually gets: - **The `wedged:` row from #118 is doing its job.** #217's API 33 leg reported `wedged: yes -- gradle was killed after 1200s and never returned` rather than the misleading `received: N, completed cleanly: yes` that CLAUDE.md records as the pre-#118 behaviour. The diagnosis took one line of log. - **The failures are not concentrated in one leg or one API.** They span API 33, 34 and 37, the unit-test job, and three unrelated instrumented tests. So a fix targeted at any single test would not have changed the retry cost of this wave. Related and worth reading together: #159 (which I have separately upgraded — it now reproduces on the development host, not only on CI) and #122.
JMR-dev commented 2026-09-06 00:16:24 +00:00 (Migrated from github.com)

A fifth mode, and this one makes the shape table print a fully-green row for a red leg

PR #215, run 33697801472 retried, API 37 gating leg, job 101397719644:

  expected:          57
  received:          57
  failed:            0
  completed cleanly: yes

and the leg is red:

Caused by: org.gradle.api.GradleException: There were failing tests.
##[error]E2E api37 failed (exit 1)

Every test ran and every test passed. The process then died in teardown:

E AndroidRuntime: java.lang.IllegalStateException: Error while unregistering UiTestAutomationService
        at android.app.IUiAutomationConnection$Stub$Proxy.disconnect(IUiAutomationConnection.java:564)
        at android.app.UiAutomation.disconnect(UiAutomation.java:458)
        at android.app.Instrumentation.finish(Instrumentation.java:310)
F libc  : Fatal signal 6 (SIGABRT), code -1 (SI_QUEUE) in tid 7368 (roidJUnitRunner), pid 7354 (emediaconverter)

Instrumentation.finish() is after the last test. AGP fails
DeviceProviderInstrumentTestTask because the runner process aborted; the XML was already written
and records nothing, because nothing test-shaped went wrong.

Why this one matters more than the count

The four modes in the table above all leave a mark the shape report can see — a non-zero failed, or
completed cleanly: no, or a wedged: row. This one leaves none. e2e-report-shape.sh reads
the test XML, the XML is complete and clean, so the table prints:

expected: 57, received: 57, failed: 0, completed cleanly: yes

next to a red leg. That is the worst possible output for the decision the report exists to support —
it does not merely fail to help, it actively says the run was fine. Someone reading the summary to
decide "flaky emulator or my change?" gets an answer that fits neither.

Note this is not #108. There the abort truncates the run and received is short. Here the run is
complete and the abort is after it.

What would fix it, in the same place point 1 already proposes

e2e-report-shape.sh knows the gradle exit status. A row like

gradle: FAILED (no test failure recorded — check for a teardown abort)

whenever the exit is non-zero and failed: 0 costs nothing and turns this into a glance. The
hasReadColorBufferDma signature line proposed in point 1 would also have fired here — there is a
native crash in the artifact — but the cheaper and more general fact is that the table currently
never mentions whether gradle actually succeeded.

Occurrence log for this leg on #215

attempt shape cause
1 received 55/57, failed 1, completed cleanly: no INSTRUMENTATION_ABORTED: System has crashed, tombstone Assertion failed: !rcEnc->featureInfo()->hasReadColorBufferDma — row-2 shape, #102
2 received 57/57, failed 0, completed cleanly: yes teardown SIGABRT in Instrumentation.finish() — this comment

Two attempts, two different infra failures, on a PR whose diff is one JVM test file and one KDoc
edit. That is the cost this ticket is about, still being paid.

Filed while shepherding the wave-4 PRs (#206-#217). Ten of the twelve are green on every gating leg;
#217's API 33 rotation failure passed on retry and is written up on #122 as a test-synchronisation
bug rather than an image one.

## A fifth mode, and this one makes the shape table print a fully-green row for a red leg PR #215, run `33697801472` retried, **API 37 gating leg**, job `101397719644`: ``` expected: 57 received: 57 failed: 0 completed cleanly: yes ``` and the leg is **red**: ``` Caused by: org.gradle.api.GradleException: There were failing tests. ##[error]E2E api37 failed (exit 1) ``` Every test ran and every test passed. The process then died in **teardown**: ``` E AndroidRuntime: java.lang.IllegalStateException: Error while unregistering UiTestAutomationService at android.app.IUiAutomationConnection$Stub$Proxy.disconnect(IUiAutomationConnection.java:564) at android.app.UiAutomation.disconnect(UiAutomation.java:458) at android.app.Instrumentation.finish(Instrumentation.java:310) F libc : Fatal signal 6 (SIGABRT), code -1 (SI_QUEUE) in tid 7368 (roidJUnitRunner), pid 7354 (emediaconverter) ``` `Instrumentation.finish()` is after the last test. AGP fails `DeviceProviderInstrumentTestTask` because the runner process aborted; the XML was already written and records nothing, because nothing test-shaped went wrong. ### Why this one matters more than the count The four modes in the table above all leave a mark the shape report can see — a non-zero `failed`, or `completed cleanly: no`, or a `wedged:` row. **This one leaves none.** `e2e-report-shape.sh` reads the test XML, the XML is complete and clean, so the table prints: > `expected: 57, received: 57, failed: 0, completed cleanly: yes` next to a red leg. That is the worst possible output for the decision the report exists to support — it does not merely fail to help, it actively says the run was fine. Someone reading the summary to decide "flaky emulator or my change?" gets an answer that fits neither. Note this is *not* #108. There the abort truncates the run and `received` is short. Here the run is complete and the abort is after it. ### What would fix it, in the same place point 1 already proposes `e2e-report-shape.sh` knows the gradle exit status. A row like > `gradle: FAILED (no test failure recorded — check for a teardown abort)` whenever the exit is non-zero and `failed: 0` costs nothing and turns this into a glance. The `hasReadColorBufferDma` signature line proposed in point 1 would also have fired here — there is a native crash in the artifact — but the *cheaper* and more general fact is that the table currently never mentions whether gradle actually succeeded. ### Occurrence log for this leg on #215 | attempt | shape | cause | |---|---|---| | 1 | `received 55/57, failed 1, completed cleanly: no` | `INSTRUMENTATION_ABORTED: System has crashed`, tombstone `Assertion failed: !rcEnc->featureInfo()->hasReadColorBufferDma` — row-2 shape, #102 | | 2 | `received 57/57, failed 0, completed cleanly: yes` | teardown SIGABRT in `Instrumentation.finish()` — this comment | Two attempts, two different infra failures, on a PR whose diff is one JVM test file and one KDoc edit. That is the cost this ticket is about, still being paid. *Filed while shepherding the wave-4 PRs (#206-#217). Ten of the twelve are green on every gating leg; #217's API 33 rotation failure passed on retry and is written up on #122 as a test-synchronisation bug rather than an image one.*
JMR-dev commented 2026-09-07 22:07:13 +00:00 (Migrated from github.com)

Both of this ticket's "worth measuring before deciding" items now have the measurement

From the #102 census — every gating E2E leg-attempt in the repo's history, 2026-08-20..09-07,
1489 leg-attempts and 129 failures, counted per leg-attempt because a re-run to green replaces
the run's conclusion. Full write-up in docs/ci-failure-modes.md; the two halves that belong here:

2. The marker should not be widened, and now there is a rate rather than an impression

Media3EngineTest.transcodesH264ToH265AndReportsProgress really does fail below API 37, and this
ticket's row 4 (#184, API 36) is one of six:

run / attempt leg date
32855014836 a1 API 36 2026-08-25
32857067112 a1 API 34 2026-08-25
32919928048 a1 API 36 2026-08-26
33261618358 a1 API 34 2026-08-29
33588264439 a1 API 36 2026-09-02 — this ticket's row 4
34000816016 a1 API 34 2026-09-06

Six in 1210 gating leg-attempts on API 33-36 — 0.5%. Split 3 on API 34, 3 on API 36, and
zero on API 33 and API 35 across 301 and 304 leg-attempts respectively.

So the answer to "either the marker is too narrow or the API 36 occurrence is rarer" is rarer,
by a lot
: on the android-37 images this test fails every time, and below them once in two
hundred legs. Moving it out of the gating set on 33-36 would trade a 1-in-200 red for permanently
not testing the hardware transcode on four API levels. The marker is the right width.

1. There is a second recognisable signature, and it is a better one than the abort

The reason this test fails is not load. It is that the emulator's Codec2 HAL process segfaults.
From the per-test logcat in e2e-report-api34 of 34000816016 a1, 62 ms after the test starts:

00:20:39.814 D MediaCodec: MediaCodec::reclaim(...) c2.goldfish.h264.decoder
00:20:39.822 F DEBUG : Cmdline: /vendor/bin/hw/android.hardware.media.c2@1.0-service-goldfish
00:20:39.822 F DEBUG : signal 0 (SIGSEGV) ... Cause: null pointer dereference
  #00 C2Block2D::handle() const+4                          libcodec2_vndk.so
  #01 getClientUsage(std::shared_ptr<C2BlockPool> const&)  libcodec2_goldfish_common.so
  #02 android::C2GoldfishAvcDec::process(...)              libcodec2_goldfish_avcdec.so
00:20:39.839 E CCodec  : Codec2 component "c2.goldfish.h264.decoder" died.
00:20:39.846 E MediaCodec: Codec reported err 0xffffffe0/DEAD_OBJECT

The HAL dies and respawns; Media3 is left with a dead codec and its own 25-second export watchdog
aborts the export, which is the Muxer error in the job log. Six of six of the occurrences
above carry that crash in the same job's --- native crashes (tail 60) --- dump — grep
c2@1.0-service-goldfish.

That matters for this ticket's proposal #1 in a specific way: the hasReadColorBufferDma abort
would not have flagged row 4.
It is an API 37 signature — 24 of the 30 API 37 failures since
2026-08-27 carry it, and it appears on none of the six above. A report that says only "a native
graphics abort was present" would have left #184's API 36 red looking exactly like a product
failure, which is the 40 minutes of log reading this ticket is about.

If the run-shape report is going to name signatures, the useful set looks like two, not one:

  • hasReadColorBufferDma — the gralloc assertion (API 37's, ex-#108)
  • c2@1.0-service-goldfish — the codec HAL crash

Grep that exact string and not the friendlier one. Codec2 component "c2.goldfish.h264.decoder" died is a CCodec line in the main buffer, and it is in none of the six job logs; what reaches
adb logcat -d -b crash is the tombstone, whose Cmdline: names
/vendor/bin/hw/android.hardware.media.c2@1.0-service-goldfish. Measured across all six.

Both signatures are already in the crash tail e2e-run.sh writes on failure, so this is a grep
over a file the script already has, in the same shape as the wedged: row #118 added. And it stays within this
ticket's own "Not proposed": neither changes a conclusion or retries anything — the leg is still
red, it just says which of the emulator's subsystems went down underneath it.

One correction to this ticket's framing, offered rather than assumed

The row-4 failure is not in the same family as rows 2 and 3. Rows 2 and 3 are API 37 and
hasReadColorBufferDma; row 4 is API 36 and a codec HAL segfault, with no gralloc abort anywhere
in it. "All four carried the same native abort in the log" holds for rows 1-3. For row 4 it does
not: I grepped hasReadColorBufferDma in the job log of all six occurrences above, including
33588264439 which is row 4, and it is absent from every one — while c2@1.0-service-goldfish is
present in every one. Two emulator faults, not one, which is why the two-signature list above is
worth having.

## Both of this ticket's "worth measuring before deciding" items now have the measurement From the #102 census — every gating E2E leg-attempt in the repo's history, 2026-08-20..09-07, **1489 leg-attempts and 129 failures**, counted per leg-attempt because a re-run to green replaces the run's conclusion. Full write-up in `docs/ci-failure-modes.md`; the two halves that belong here: ### 2. The marker should **not** be widened, and now there is a rate rather than an impression `Media3EngineTest.transcodesH264ToH265AndReportsProgress` really does fail below API 37, and this ticket's row 4 (#184, API 36) is one of six: | run / attempt | leg | date | |---|---|---| | `32855014836` a1 | API 36 | 2026-08-25 | | `32857067112` a1 | API 34 | 2026-08-25 | | `32919928048` a1 | API 36 | 2026-08-26 | | `33261618358` a1 | API 34 | 2026-08-29 | | **`33588264439` a1** | **API 36** | **2026-09-02** — this ticket's row 4 | | `34000816016` a1 | API 34 | 2026-09-06 | **Six in 1210 gating leg-attempts on API 33-36 — 0.5%.** Split 3 on API 34, 3 on API 36, and **zero on API 33 and API 35** across 301 and 304 leg-attempts respectively. So the answer to "either the marker is too narrow or the API 36 occurrence is rarer" is **rarer, by a lot**: on the android-37 images this test fails every time, and below them once in two hundred legs. Moving it out of the gating set on 33-36 would trade a 1-in-200 red for permanently not testing the hardware transcode on four API levels. **The marker is the right width.** ### 1. There is a second recognisable signature, and it is a better one than the abort The reason this test fails is not load. It is that **the emulator's Codec2 HAL process segfaults**. From the per-test logcat in `e2e-report-api34` of `34000816016` a1, 62 ms after the test starts: ``` 00:20:39.814 D MediaCodec: MediaCodec::reclaim(...) c2.goldfish.h264.decoder 00:20:39.822 F DEBUG : Cmdline: /vendor/bin/hw/android.hardware.media.c2@1.0-service-goldfish 00:20:39.822 F DEBUG : signal 0 (SIGSEGV) ... Cause: null pointer dereference #00 C2Block2D::handle() const+4 libcodec2_vndk.so #01 getClientUsage(std::shared_ptr<C2BlockPool> const&) libcodec2_goldfish_common.so #02 android::C2GoldfishAvcDec::process(...) libcodec2_goldfish_avcdec.so 00:20:39.839 E CCodec : Codec2 component "c2.goldfish.h264.decoder" died. 00:20:39.846 E MediaCodec: Codec reported err 0xffffffe0/DEAD_OBJECT ``` The HAL dies and respawns; Media3 is left with a dead codec and its own 25-second export watchdog aborts the export, which is the `Muxer error` in the job log. **Six of six** of the occurrences above carry that crash in the same job's `--- native crashes (tail 60) ---` dump — grep `c2@1.0-service-goldfish`. That matters for this ticket's proposal #1 in a specific way: **the `hasReadColorBufferDma` abort would not have flagged row 4.** It is an API 37 signature — 24 of the 30 API 37 failures since 2026-08-27 carry it, and it appears on **none** of the six above. A report that says only "a native graphics abort was present" would have left #184's API 36 red looking exactly like a product failure, which is the 40 minutes of log reading this ticket is about. If the run-shape report is going to name signatures, the useful set looks like two, not one: - `hasReadColorBufferDma` — the gralloc assertion (API 37's, ex-#108) - `c2@1.0-service-goldfish` — the codec HAL crash **Grep that exact string and not the friendlier one.** `Codec2 component "c2.goldfish.h264.decoder" died` is a `CCodec` line in the main buffer, and it is in **none** of the six job logs; what reaches `adb logcat -d -b crash` is the tombstone, whose `Cmdline:` names `/vendor/bin/hw/android.hardware.media.c2@1.0-service-goldfish`. Measured across all six. Both signatures are already in the crash tail `e2e-run.sh` writes on failure, so this is a grep over a file the script already has, in the same shape as the `wedged:` row #118 added. And it stays within this ticket's own "Not proposed": neither changes a conclusion or retries anything — the leg is still red, it just says which of the emulator's subsystems went down underneath it. ### One correction to this ticket's framing, offered rather than assumed The row-4 failure is **not** in the same family as rows 2 and 3. Rows 2 and 3 are API 37 and `hasReadColorBufferDma`; row 4 is API 36 and a codec HAL segfault, with no gralloc abort anywhere in it. "All four carried the same native abort in the log" holds for rows 1-3. For row 4 it does not: I grepped `hasReadColorBufferDma` in the job log of **all six** occurrences above, including `33588264439` which is row 4, and it is absent from every one — while `c2@1.0-service-goldfish` is present in every one. Two emulator faults, not one, which is why the two-signature list above is worth having.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: JMR-dev/LibreMediaConverter#190