Re-derive the API 37 emulator failure, and harden the local sweep #47
+437
-132
@@ -1,127 +1,381 @@
|
||||
# API 37 is not tested in CI: a crash in Google's `android-37.0` emulator image
|
||||
# API 37 on the emulator: a guest gralloc bug that only the host GL renderer triggers
|
||||
|
||||
**Status:** open upstream, worked around by removing API 37 from the E2E matrix.
|
||||
The app itself is verified good on real API 37 hardware — this is an emulator bug only.
|
||||
**Last verified:** 2026-08-21, against emulator `37.1.11.0` and system image revision 6
|
||||
**Status:** the bug is real and still open upstream, but the previous diagnosis in this file was
|
||||
wrong about its most important detail. **The renderer decides whether API 37 boots**, and once it
|
||||
boots, disabling SystemUI collapses the crash rate far enough to run a suite —
|
||||
`tools/local-emulator/run-e2e.sh 37` gets through the whole instrumented suite and comes back with
|
||||
**2 failures, 0 errors and the two by-design skips** (measured 49 / 2 / 0 / 2 at `22c7914`, where
|
||||
the suite was 49 tests — [Reading these totals](#reading-these-totals) before comparing any total
|
||||
with another). The crashes do not stop outright, and the two failures are real; both are quantified
|
||||
below. CI's matrix should still stop at 36 — see
|
||||
[So should CI take API 37?](#so-should-ci-take-api-37).
|
||||
**Last verified:** 2026-08-22, emulator `37.1.11.0` (build 15917651), Fedora 44,
|
||||
against system images `android-37.0` rev 6 **and** `android-37.1` rev 8.
|
||||
|
||||
`minSdk` is 33 and `targetSdk` is 37, and the E2E matrix in
|
||||
[`status_check.yml`](../.github/workflows/status_check.yml) runs API 33 through 36.
|
||||
API 37 is deliberately absent. This is why.
|
||||
## The correction
|
||||
|
||||
## Summary
|
||||
This file previously said, under "What was ruled out":
|
||||
|
||||
The `android-37.0` emulator system image crashes `surfaceflinger` in a loop. The app
|
||||
under test never gets a working framework, so every instrumented test fails regardless
|
||||
of what the app does. The bug is in the emulator image, not in this project.
|
||||
> **GPU mode.** Both `swiftshader_indirect` and `host` crash, with the same assertion and
|
||||
> the same frames. The crash is in the gralloc mapper, below the renderer.
|
||||
|
||||
The crash is an assertion inside the emulator's own gralloc implementation:
|
||||
**That is wrong.** The mapper is below the renderer, but *whether the mapper's bad path is
|
||||
reached* is not. Re-measured on 2026-08-22, seven runs, one variable at a time:
|
||||
|
||||
| # | system image | `-gpu` | GLES the emulator chose | booted? | surfaceflinger aborts |
|
||||
|---|---|---|---|---|---|
|
||||
| r01 | `android-37.0` rev 6 | `host` | host (Mesa Iris Xe) | **no**, 422 s | 71, looping |
|
||||
| r02 | `android-37.1` rev 8 | `host` | host (Mesa Iris Xe) | **no**, 362 s | 65, looping |
|
||||
| r03 | `android-37.0` rev 6 | `swangle_indirect` | ANGLE | **yes, 85 s** | 1 |
|
||||
| r04 | `android-37.0` rev 6 | `host` + `-feature -GLDMA,-GLDMA2,-GLDirectMem` | host | **no**, 363 s | 57, looping |
|
||||
| r05 | `android-37.0` rev 6 | `angle_indirect` | ANGLE | **yes, 112 s** | 2 |
|
||||
| r06 | `android-37.1` rev 8 | `swangle_indirect` | ANGLE | **yes, 285 s** | 23 |
|
||||
| r07 | `android-37.0` rev 6 | `host` + `-feature -HostComposition` | host | **no**, wedged adb at 208 s | not readable |
|
||||
|
||||
The discriminator is exact across all seven: **a run boots if and only if the emulator log says
|
||||
something other than `gles_mode_selected:host`.**
|
||||
|
||||
One caveat about how independent those rows are, because the table flatters itself. `-gpu
|
||||
angle_indirect` (r05) and `-gpu swangle_indirect` (r03) both logged `gles_mode_selected:swangle`
|
||||
and both reported the same adapter, differing only in the Vulkan backend beneath
|
||||
(`vulkan_mode_selected:lavapipe` against `swiftshader`). So they are closer to one GLES path
|
||||
reached two ways than to two renderers agreeing — note that at API 33–36
|
||||
[`docs/local-emulator.md`](local-emulator.md) records `angle_indirect` resolving to ANGLE on
|
||||
*llvmpipe*, a genuinely different adapter, which it did not do here. What is 7-for-7 is the
|
||||
host-GLES-versus-not split, not "two independent renderers both work".
|
||||
|
||||
```
|
||||
# r01, r02, r04, r07 -- never boots
|
||||
INFO | emuglConfig_init: vulkan_mode_selected:host gles_mode_selected:host
|
||||
INFO | Graphics Adapter Android Emulator OpenGL ES Translator (Mesa Intel(R) Iris(R) Xe Graphics (TGL GT2))
|
||||
|
||||
# r03, r05, r06 -- boots
|
||||
INFO | emuglConfig_init: vulkan_mode_selected:swiftshader gles_mode_selected:swangle
|
||||
INFO | Graphics Adapter Android Emulator OpenGL ES Translator (ANGLE (Google, Vulkan 1.2.0
|
||||
| (SwiftShader Device (Subzero) (0x0000C0DE)), SwiftShader driver-5.0.0))
|
||||
```
|
||||
|
||||
### Why the wrong claim looked right
|
||||
|
||||
It rested on two samples of two different things, and neither of them was ANGLE.
|
||||
|
||||
- The **local** `swiftshader_indirect` sample was void. On this workstation *every*
|
||||
SwiftShader-GLES launch segfaults the host emulator before the guest matters at all —
|
||||
SELinux denies `execheap` to SwiftShader's Reactor JIT. That is
|
||||
[`docs/local-emulator.md`](local-emulator.md), and it was not yet understood when this file
|
||||
was written. So "`swiftshader_indirect` crashes" was true, for an entirely unrelated reason,
|
||||
and told you nothing about the gralloc assertion.
|
||||
- The **CI** sample was one `swiftshader_indirect` run on a GPU-less `ubuntu-latest`, and the
|
||||
**local** sample was one `-gpu host` run. Two renderers, one measurement each, and the pair
|
||||
written up as "both GPU modes".
|
||||
|
||||
`angle_indirect` and `swangle_indirect` — the two modes that work — had never been tried on
|
||||
API 37. Neither had a second system image.
|
||||
|
||||
The lesson is the same one `docs/local-emulator.md` ends on, which makes it worth repeating:
|
||||
"both backends fail" is a claim about a matrix, and a matrix needs cells, not inference. Two
|
||||
observations of two different configurations do not establish anything about a third.
|
||||
|
||||
## What the bug actually is
|
||||
|
||||
`surfaceflinger` aborts inside the emulator's own gralloc mapper:
|
||||
|
||||
```
|
||||
Executable: /system/bin/surfaceflinger
|
||||
signal 6 (SIGABRT), code -1 (SI_QUEUE), tid: RegionSampling
|
||||
Abort message: 'Assertion failed: !rcEnc->featureInfo()->hasReadColorBufferDma'
|
||||
|
||||
#03 mapper.ranchu.so GoldfishMapper::readFromHost(cb_handle_t const&) const
|
||||
#04 mapper.ranchu.so GoldfishMapper::GoldfishMapper()::'lambda'(...)::__invoke
|
||||
#05 libui.so android::Gralloc5Mapper::lock(...)
|
||||
#06 libui.so android::GraphicBufferMapper::lock(...)
|
||||
#07 libui.so android::GraphicBuffer::lockAsync(...)
|
||||
#08 libui.so android::GraphicBuffer::lock(...)
|
||||
#09 surfaceflinger android::RegionSamplingThread::threadMain()
|
||||
#03 /vendor/lib64/hw/mapper.ranchu.so GoldfishMapper::readFromHost(cb_handle_t const&) const+543
|
||||
#04 /vendor/lib64/hw/mapper.ranchu.so GoldfishMapper::GoldfishMapper()::'lambda'(...)::__invoke+704
|
||||
#05 /system/lib64/libui.so android::Gralloc5Mapper::lock(...)+63
|
||||
#06 /system/lib64/libui.so android::GraphicBufferMapper::lock(...)+198
|
||||
#07 /system/lib64/libui.so android::GraphicBuffer::lockAsync(...)+545
|
||||
#08 /system/lib64/libui.so android::GraphicBuffer::lock(...)+67
|
||||
#09 /system/bin/surfaceflinger android::RegionSamplingThread::threadMain()+2571
|
||||
```
|
||||
|
||||
`RegionSamplingThread` is SystemUI's navigation-bar luma sampling. It calls
|
||||
`GraphicBuffer::lock`, which routes into `GoldfishMapper::readFromHost`, which asserts
|
||||
that the host has *not* negotiated the `ReadColorBufferDma` capability. On this image
|
||||
the host has, so the assertion fails and `surfaceflinger` aborts. It restarts and
|
||||
aborts again.
|
||||
`RegionSamplingThread` is SystemUI's nav-bar luma sampling. It locks a `GraphicBuffer` for CPU
|
||||
read; that routes through the Gralloc5 mapper into `GoldfishMapper::readFromHost`, which is the
|
||||
*non-DMA* readback path and asserts that the host has not negotiated `ReadColorBufferDma`. The
|
||||
host always has, so the assert fires whenever that path is taken.
|
||||
|
||||
## Impact
|
||||
Two facts pin down what "always" means:
|
||||
|
||||
The failure surfaces in two different ways depending on how far the job gets, which is
|
||||
why it took several rounds to identify:
|
||||
- **The capability is negotiated regardless of renderer.** The evidence is the aborts
|
||||
themselves: the assertion that fires is `!hasReadColorBufferDma`, and it fires under ANGLE
|
||||
(r03/r05/r06) as well as under the host translator — just far less often. That is a direct
|
||||
observation of the guest having negotiated DMA readback under both, and it stands alone.
|
||||
(Supporting only, and weaker than it first looks: `ANDROID_EMU_read_color_buffer_dma` appears
|
||||
in exactly one file in the SDK, `emulator/lib64/libgfxstream_backend.so`, which every `-gpu`
|
||||
mode goes through. A string search establishes where the extension is implemented, not that
|
||||
it is negotiated on every path.)
|
||||
- **It is not gated by any feature flag the emulator exposes.** See the ruled-out list below.
|
||||
|
||||
| Guest RAM | Where it dies | What CI reports |
|
||||
|---|---|---|
|
||||
| 1536 MB | during APK install | `Unknown failure: cmd: Can't find service: package` |
|
||||
| 2560 MB | during the test run | `There were failing tests` — all of them |
|
||||
So the renderer does not decide whether the guest *believes* DMA readback exists. It decides how
|
||||
often `RegionSamplingThread` ends up in `readFromHost` — which under the host GL translator is
|
||||
constantly, and under ANGLE is occasionally.
|
||||
|
||||
At 2560 MB the install succeeds and the tests actually execute, then fail wholesale.
|
||||
The first failure in the report is misleading:
|
||||
### Why one abort takes down the whole device
|
||||
|
||||
`surfaceflinger` is a critical service. When it dies, `init` kills the framework with it:
|
||||
|
||||
```
|
||||
kotlin.UninitializedPropertyAccessException: lateinit property output has not
|
||||
been initialized
|
||||
at Media3EngineTest.tearDown(Media3EngineTest.kt:53)
|
||||
|
||||
java.lang.IllegalStateException: WorkManager is not initialized properly.
|
||||
You have explicitly disabled WorkManagerInitializer in your manifest, ...
|
||||
08-22 21:40:28.253 I/init: Sending SIGKILL to service 'zygote' (pid 470) process group...
|
||||
08-22 21:40:28.260 I/init: Service 'zygote' (pid 470) received SIGKILL
|
||||
```
|
||||
|
||||
Neither is a real defect in this project. `tearDown` throws because `setUp` never got
|
||||
far enough to assign `output`, and WorkManager's `InitializationProvider` never runs
|
||||
because content-provider installation fails on a framework whose `surfaceflinger` is
|
||||
crash-looping. The same tests pass at API 33, 34, 35, and 36 in the same CI run, and
|
||||
the first `surfaceflinger` abort is timestamped *before* the test results are reported.
|
||||
|
||||
This is not inference. The full suite was run against a physical API 37 device and
|
||||
passed — see [Verified on real API 37 hardware](#verified-on-real-api-37-hardware)
|
||||
below. `ConversionWorkerTest` and `ConcatWorkerTest`, which drive a real WorkManager
|
||||
round trip and are among the tests that failed this way in CI, both pass there.
|
||||
Everything above zygote goes with it, which is why the symptoms look nothing like a graphics
|
||||
bug. Under `-gpu host` the cycle repeats every five to seven seconds forever and
|
||||
`sys.boot_completed` is never set. Under ANGLE the aborts are sparse enough that the boot
|
||||
usually completes between them — but they do not stop, and each one is a framework restart.
|
||||
That is the difference between "boots" and "is usable", and it is the reason this is not simply
|
||||
fixed by changing the renderer. See [Can the suite run on it?](#can-the-suite-run-on-it) below.
|
||||
|
||||
## Environment
|
||||
|
||||
Reproduced identically in two unrelated environments, so it is not specific to a host
|
||||
GPU, driver, or CI runner.
|
||||
|
||||
| | GitHub Actions | Local workstation |
|
||||
|---|---|---|
|
||||
| Host | `ubuntu-latest`, no GPU | Fedora, Intel Iris Xe (TGL GT2) |
|
||||
| Host | `ubuntu-latest`, no GPU | Fedora 44, Intel Iris Xe (TGL GT2), kernel `7.1.8-200.fc44` |
|
||||
| Emulator | `37.1.11.0` (build 15917651) | `37.1.11.0` (build 15917651) |
|
||||
| GPU mode | `swiftshader_indirect` | `host` |
|
||||
| Result | boots, aborts during tests | aborts before boot completes |
|
||||
| GPU mode measured | `swiftshader_indirect` | `host`, `angle_indirect`, `swangle_indirect` |
|
||||
|
||||
System image: `system-images;android-37.0;google_apis;x86_64`, `Pkg.Revision=6`,
|
||||
`AndroidVersion.ApiLevel=37.0`, `AndroidVersion.ExtensionLevel=22`
|
||||
Images, both reproducing it:
|
||||
|
||||
```
|
||||
Build fingerprint: google/sdk_gphone64_x86_64/emu64xa:17/CE2A.260420.019/15611780:userdebug/dev-keys
|
||||
Kernel Release: 6.12.58-android16-6-gccafb60de224-ab14828483
|
||||
system-images;android-37.0;google_apis;x86_64 Pkg.Revision=6 ApiLevel=37.0 ExtensionLevel=22
|
||||
fingerprint google/sdk_gphone64_x86_64/emu64xa:17/CE2A.260420.019/15611780:userdebug/dev-keys
|
||||
system-images;android-37.1;google_apis_ps16k;x86_64 Pkg.Revision=8 ApiLevel=37.1 ExtensionLevel=23
|
||||
ro.build.version.codename=REL (a release image, not a preview)
|
||||
```
|
||||
|
||||
**Note the `ps16k` in the second one — it is not optional, and it is why the 37.1 result is
|
||||
interpretable.** From API 37.1 onward Google ships *only* 16 KB-page x86_64 images; there is no
|
||||
plain `google_apis` variant to pick. `sdkmanager --list` for 37.1 and 37.2-beta* offers nothing
|
||||
but `google_apis_ps16k` and `google_apis_playstore_ps16k`. That makes page-size alignment a
|
||||
prerequisite rather than a detail: a `.so` that is not 16 KB aligned will not load on such a
|
||||
guest, and the resulting failure looks like an app bug. Checked before the first `ps16k` boot,
|
||||
using the same test `build.yml` applies to release APKs — all 20 libraries in the committed
|
||||
`bin/ffmpeg-kit-next-8.1.1.aar`, both ABIs, report `0x4000`:
|
||||
|
||||
```
|
||||
$ for f in jni/*/*.so; do readelf -lW "$f" | awk '$1=="LOAD"{print $NF}' | sort -u; done
|
||||
0x4000 (x20: libavcodec, libavdevice, libavfilter, libavformat, libavutil,
|
||||
libc++_shared, libffmpegkit, libffmpegkit_abidetect, libswresample, libswscale
|
||||
-- arm64-v8a and x86_64)
|
||||
```
|
||||
|
||||
So when `android-37.1` reproduced the abort, that was the gralloc bug and not a page-size
|
||||
mismatch. `image_pkg_for_api` in `tools/local-emulator/run-e2e.sh` encodes the `ps16k` tag for
|
||||
37.1; if this ever fails after an FFmpeg rebuild, re-run the alignment check first.
|
||||
|
||||
## What was ruled out, and how
|
||||
|
||||
Each of these was tested rather than reasoned about, because the first three attempts
|
||||
at this bug were plausible fixes that turned out to address earlier, unrelated failures.
|
||||
**A newer system image.** This file's own revisit trigger was "a new `android-37.0` system image
|
||||
revision ships (this was revision 6)". That trigger was written too narrowly and would never have
|
||||
fired: `android-37.0` is *still* revision 6, but Google shipped a whole new minor level.
|
||||
`android-37.1` `google_apis_ps16k` revision 8 — a `REL` build, not a beta — was installed and
|
||||
tested (r02, r06) and **behaves identically**: same assertion, same frames, never boots under
|
||||
`-gpu host`, and *worse* under ANGLE (23 aborts to `37.0`'s 1). `android-37.2-beta3` exists too
|
||||
but was not needed; two independent images agreeing settles it, and a beta could not be used by
|
||||
CI anyway.
|
||||
|
||||
**Guest memory.** The emulator raises an undersized guest to a minimum on its own, but
|
||||
only for API levels it recognises, and it does not recognise `"37.0"`. API 33 bumps to
|
||||
2048 MB and 34/35/36 to 2560 MB, while API 37 logged no bump at all and ran at the
|
||||
`pixel_6` default of 1536 MB. Setting `ram-size: 2560M` explicitly fixed that asymmetry
|
||||
and did change the outcome — the job got past install and into the test run — but it is
|
||||
not the underlying bug. At the moment of failure the guest reported `MemTotal 2527392
|
||||
kB` with `MemAvailable 1507104 kB`: 1.5 GB free, and no OOM kills.
|
||||
**An ATD image.** Still does not exist for API 37. `sdkmanager --list` offers `aosp_atd` and
|
||||
`google_atd` for API 30 through 36 and nothing above:
|
||||
|
||||
**GPU mode.** Both `swiftshader_indirect` and `host` crash, with the same assertion and
|
||||
the same frames. The crash is in the gralloc mapper, below the renderer.
|
||||
```
|
||||
system-images;android-36;google_atd;x86_64 | 1 | Google APIs ATD Intel x86_64 Atom System Image
|
||||
(no android-37 ATD of any kind)
|
||||
```
|
||||
|
||||
**Disabling the DMA feature.** `GLDMA` is the host feature that most plausibly backs the
|
||||
guest's `hasReadColorBufferDma`. Launching with `-feature -GLDMA` was accepted by the
|
||||
emulator — the log confirms `Feature 'GLDMA' (51) is overridden to 'disabled'` — and
|
||||
`surfaceflinger` still aborted 13 times and the device never finished booting. Whatever
|
||||
sets that guest capability, it is not this flag.
|
||||
For API 37 the only x86_64 images are `google_apis`, `google_apis_playstore`, their `ps16k`
|
||||
16 KB-page variants, and Wear OS. Check again when revisiting.
|
||||
|
||||
**An ATD image.** `google_atd` / `aosp_atd` images are built for automated testing and
|
||||
ship without the SystemUI package set, which is what drives `RegionSamplingThread` in
|
||||
the first place. That would likely sidestep the bug class entirely, but **no ATD image
|
||||
exists for `android-37.0`** — only `google_apis`, `google_apis_playstore`, the `ps16k`
|
||||
16 KB-page variants, and Wear OS. Check again when revisiting; if an ATD image appears,
|
||||
try it before anything else here.
|
||||
**The DMA feature flags.** `GLDMA` alone was ruled out previously; `GLDMA2` and `GLDirectMem`
|
||||
were not, and the per-image `advancedFeatures.ini` turns all three on. Disabling all three
|
||||
together (r04) is accepted by the emulator and changes nothing:
|
||||
|
||||
```
|
||||
INFO | Feature 'GLDMA' (51) is overridden to 'disabled'
|
||||
INFO | Feature 'GLDMA2' (52) is overridden to 'disabled'
|
||||
INFO | Feature 'GLDirectMem' (53) is overridden to 'disabled'
|
||||
... 57 surfaceflinger aborts, device never boots
|
||||
```
|
||||
|
||||
**Host composition.** `-feature -HostComposition` (r07) was the best remaining guess at what
|
||||
forces the readback. It did not help; it made things worse, wedging adb entirely at 208 s so the
|
||||
crash buffer could not even be read. Recorded as inconclusive rather than ruled out, because no
|
||||
evidence came back from it.
|
||||
|
||||
**Guest feature negotiation differing from API 36.** It does not. The image-level
|
||||
`advancedFeatures.ini` for `android-37.0` is byte-identical to `android-36`'s except for one
|
||||
unrelated line:
|
||||
|
||||
```
|
||||
$ diff android-36/google_apis/x86_64/advancedFeatures.ini android-37.0/google_apis/x86_64/advancedFeatures.ini
|
||||
+QemuCameraSensorOrientation = on
|
||||
```
|
||||
|
||||
`GLDMA`, `GLDMA2`, `GLDirectMem`, `GrallocSync`, `HostComposition` and `YUVCache` are on in
|
||||
both. API 36 boots and passes. So nothing about the host/guest feature handshake changed — the
|
||||
regression is in the guest's Gralloc5 mapper or in what API 37's `RegionSamplingThread` asks of
|
||||
it, not in what the emulator advertises.
|
||||
|
||||
**Guest memory.** Ruled out previously and not revisited; every run above used
|
||||
`hw.ramSize=2560`, the same value the E2E matrix pins, and none of them OOMed.
|
||||
|
||||
**A host-side crash.** Not this bug, and worth stating because the other emulator failure on this
|
||||
workstation *is* host-side. Every run above left `coredumpctl` empty and produced zero
|
||||
`avc: denied` lines, and the qemu process was still alive at the end of the ones that never
|
||||
booted (`emulator_alive=yes`). The host emulator is fine; the guest is not.
|
||||
|
||||
## Can the suite run on it?
|
||||
|
||||
**Almost.** `tools/local-emulator/run-e2e.sh 37` now runs the whole suite locally, and all of it
|
||||
passes except two tests. Measured at `22c7914`: **49 tests, 2 failures, 0 errors, 2 skipped** — 45
|
||||
passed, the two `Media3EngineTest` failures dissected below, and the two `assumeTrue` skips every
|
||||
level has. It costs two deviations from how every other level is run, and both are worth
|
||||
understanding before trusting the leg.
|
||||
|
||||
Two things about that total before it is compared with anything. It is the size of the suite on
|
||||
the checkout that ran, not a property of API 37 — `app/src/androidTest` held 49 `@Test` methods at
|
||||
`22c7914`, and a newer checkout reports its own count; see
|
||||
[Reading these totals](#reading-these-totals). And **the Pixel has never run 49**: its green run
|
||||
was 40 / 0 / 0 / 2 at `edd6385`, the same suite nine tests earlier. What compares across the two
|
||||
is two failures against none, and the same two skips — not the totals.
|
||||
|
||||
The same numbers and the same two test names came back twice, which is real corroboration — but
|
||||
by two different routes, and only one of them is the harness. The first was driven by hand
|
||||
(`pm disable-user`, then several minutes of incidental framework restarts, then `e2e-run.sh`
|
||||
directly); the second went through `disable_region_sampling`'s `stop; start`. **The harness path
|
||||
itself has one green measurement.** What would make this routine is a second consecutive
|
||||
`run-e2e.sh 37` whose only failures are the same two.
|
||||
|
||||
### Booting is not the same as being usable
|
||||
|
||||
Changing the renderer gets the device to `sys.boot_completed=1`, and that is all it gets you. The
|
||||
aborts do not stop, and each one is a framework restart. A five-minute test run does not survive
|
||||
that. What it looks like from Gradle:
|
||||
|
||||
```
|
||||
Shell command failed (1): rm -rf "/sdcard/Android/media/org.libremediaconverter/..."
|
||||
rm: ...: Transport endpoint is not connected
|
||||
Starting 0 tests on lmc_e2e_api37(AVD) - 17
|
||||
Shell command failed (20): am get-current-user
|
||||
cmd: Can't find service: activity
|
||||
Device emulator-5572 failed to uninstall test APK org.libremediaconverter.
|
||||
[cmd: Can't find service: package]
|
||||
Test run failed to complete. No test results.
|
||||
onError: commandError=false message=INSTRUMENTATION_ABORTED: System has crashed.
|
||||
```
|
||||
|
||||
Measured idle rate on `android-37.0` under `swangle_indirect`: **10 aborts in 150 s, then 11 more
|
||||
in the next 150 s**. Steady, not a start-up transient.
|
||||
|
||||
### The fix is to remove the region-sampling listener, not to survive it
|
||||
|
||||
`RegionSamplingThread` exists only because SystemUI registers a nav-bar luma-sampling listener.
|
||||
Take SystemUI away and the thread is never started, so the mapper's bad path is never called:
|
||||
|
||||
```
|
||||
$ adb shell pm disable-user --user 0 com.android.systemui
|
||||
Package com.android.systemui new state: disabled-user
|
||||
|
||||
=== aborts at start of measurement: 36
|
||||
=== idle 180s with SystemUI disabled ===
|
||||
=== aborts after: 36 NEW IN WINDOW: 0
|
||||
--- services still up? ---
|
||||
activity Service activity: found
|
||||
package Service package: found
|
||||
window Service window: found
|
||||
```
|
||||
|
||||
**Zero in 180 s, against 10–11 per 150 s.** That is the strongest evidence that region sampling
|
||||
is the dominant trigger, and it is worth recording even by someone who never wants the workaround.
|
||||
It does not establish it as the *only* trigger: the paragraph below has an abort surviving the
|
||||
disable, and nothing measured here says whether that residue is a second caller of the readback
|
||||
path or a disable that did not fully take.
|
||||
|
||||
Do not read that as "the crashes stop", though, because the harness path does not reproduce a
|
||||
clean zero. Its own post-disable check on the run recorded below printed
|
||||
|
||||
```
|
||||
quiet check: 1 new surfaceflinger aborts in 45 s (want 0)
|
||||
surfaceflinger hasReadColorBufferDma aborts: 4 (whole run)
|
||||
```
|
||||
|
||||
So what is reliably achieved is a **rate collapse** — from roughly one abort every fourteen
|
||||
seconds to one every forty-five — which a 47-second Gradle run survives and a five-minute one
|
||||
might not. The 180-second zero above is one measurement on a device that had been up for twelve
|
||||
minutes and had already cycled its framework several times. The harness prints the quiet-check
|
||||
delta on every run precisely so this is visible rather than assumed.
|
||||
|
||||
One ordering detail cost a whole run and is now encoded in `disable_region_sampling`: by the time
|
||||
`sys.boot_completed` flips, SystemUI has **already registered**, and `pm disable-user` does not
|
||||
retract an existing registration — it only stops the package being started again. Disabling it
|
||||
and proceeding straight to the tests fails exactly as before. The harness therefore does
|
||||
`stop; start` afterwards, so the framework that comes back never starts SystemUI at all.
|
||||
|
||||
### The two deviations, stated plainly
|
||||
|
||||
1. **The renderer is ANGLE, not the host GPU.** Shared with nothing else in the matrix — API
|
||||
33–36 run `-gpu host` locally, and CI runs `swiftshader_indirect`.
|
||||
2. **SystemUI is disabled.** The API 37 leg does not run the same device configuration as any
|
||||
other leg or as the Pixel. It is defensible here only because nothing in this suite touches
|
||||
system UI — these are Media3, FFmpeg and WorkManager tests — and because the alternative is no
|
||||
local API 37 coverage at all. **Anything that ever does depend on system UI must not trust
|
||||
this leg.**
|
||||
|
||||
### The two remaining failures are the same bug, one layer down
|
||||
|
||||
```
|
||||
org.libremediaconverter.convert.Media3EngineTest > runsFromAThreadWithNoLooper FAILED
|
||||
org.libremediaconverter.convert.Media3EngineTest > transcodesH264ToH265AndReportsProgress FAILED
|
||||
|
||||
androidx.media3.transformer.ExportException: Codec exception:
|
||||
CodecInfo{type=VideoDecoder, ..., mime=video/avc, name=c2.goldfish.h264.decoder}
|
||||
at androidx.media3.transformer.DefaultCodec.maybeDequeueOutputBuffer(DefaultCodec.java:398)
|
||||
Caused by: android.media.MediaCodec$CodecException:
|
||||
at android.media.MediaCodec.native_dequeueOutputBuffer(Native Method)
|
||||
```
|
||||
|
||||
Three measurements say this is the emulator image and not this app, and not the software
|
||||
renderer:
|
||||
|
||||
- **Control at API 35 under the identical renderer.** `GPU_MODE=swangle_indirect
|
||||
tools/local-emulator/run-e2e.sh 35` → **49 / 0 / 0 / 2** at `22c7914`, green.
|
||||
`c2.goldfish.h264.decoder` is perfectly happy under ANGLE one API level down, so the renderer is
|
||||
not what breaks it.
|
||||
- **Real API 37 hardware passes**, see below. There is no `c2.goldfish.*` codec on a Pixel.
|
||||
- The failing call is `dequeueOutputBuffer` on the *goldfish* decoder — the emulator's own codec,
|
||||
which like `RegionSamplingThread` gets its frames out of a host-side colour buffer. Same
|
||||
readback machinery, one layer down. This is inference rather than a measurement, and is flagged
|
||||
as such; what is measured is the first two bullets.
|
||||
|
||||
**Do not try `-feature -HardwareDecoder`.** It is the obvious next idea and it is much worse:
|
||||
forcing the guest onto software decoders took the run from 2 failures to **46**, across
|
||||
`RemuxTest`, `ForcedFailureTest`, `HardwareFallbackTest` and `UnopenableUriTest` as well. The
|
||||
suite depends on those decoders existing.
|
||||
|
||||
### So should CI take API 37?
|
||||
|
||||
**No, and the matrix should still stop at 36.** Three reasons, in order of weight:
|
||||
|
||||
1. CI runs `swiftshader_indirect` on a GPU-less runner. The aborts happen under ANGLE too — they
|
||||
are merely sparser — so nothing here says a runner would be stable.
|
||||
2. The working configuration needs SystemUI disabled and a framework restart mid-job. That is a
|
||||
lot of bespoke device surgery to put behind a merge gate, and it silently weakens what the leg
|
||||
proves.
|
||||
3. Even at its best two tests fail, so the leg would be permanently red or permanently
|
||||
allow-listed. Neither is a gate worth having.
|
||||
|
||||
What has changed is the *local* story: API 37 is no longer a level nobody can look at. A
|
||||
regression that shows up at 37 and not at 36 can now be reproduced on this workstation in about
|
||||
four minutes, which is what the missing matrix row was really costing.
|
||||
|
||||
## Verified on real API 37 hardware
|
||||
|
||||
The bug is confined to the emulator image. On 2026-08-21 the whole instrumented suite
|
||||
was run against a physical device and passed:
|
||||
Unchanged and still true. On 2026-08-21 the whole instrumented suite ran green on a physical
|
||||
device:
|
||||
|
||||
```
|
||||
Device: Pixel 10 Pro XL (mustang), arm64-v8a
|
||||
@@ -132,65 +386,94 @@ API: 37 (Android 17, codename REL -- a release build, not a preview)
|
||||
40 tests, 0 failures, 0 errors, 2 skipped BUILD SUCCESSFUL
|
||||
```
|
||||
|
||||
**That "40" is a measurement of the tree it ran on, not a baseline for today**, and it is not a
|
||||
contradiction of the totals in [`docs/local-emulator.md`](local-emulator.md) either.
|
||||
|
||||
### Reading these totals
|
||||
|
||||
Every total in this file and in [`docs/local-emulator.md`](local-emulator.md) is the size of
|
||||
`app/src/androidTest` on the checkout that produced it, and nothing else. The reported total has
|
||||
equalled that checkout's `@Test` count everywhere it has been checked:
|
||||
|
||||
| checkout | `@Test` methods | total the run reported |
|
||||
|---|---|---|
|
||||
| `edd6385` | 40 | 40 — the Pixel run above |
|
||||
| `22c7914` | 49 | 49 — the four local levels, and API 37 |
|
||||
| `18c53a3` | 57 | not run |
|
||||
|
||||
So the number to expect is not written down here. It is derived from the checkout in front of
|
||||
you, which is the only thing that cannot go stale:
|
||||
|
||||
```bash
|
||||
grep -rho '@Test' app/src/androidTest | wc -l
|
||||
```
|
||||
|
||||
**Before each release, run the suite on the Pixel 10 Pro XL and expect that many tests, 0
|
||||
failures, 0 errors, 2 skipped.** The failure, error and skip counts are the invariant; the total
|
||||
is not. A total that disagrees with your own checkout's count is the signal — an old checkout, a
|
||||
stale build, or tests that never ran — and it is worth stopping on either way.
|
||||
|
||||
The two skips are `RealMediaBenchmark.hardwareVersusSoftwareOnRealVideo` and
|
||||
`av1InputRoutesAccordingToDeviceDecodeSupport`, which `assumeTrue` their sample files
|
||||
are present and skip when they are not. That is by design and unrelated to API level.
|
||||
`av1InputRoutesAccordingToDeviceDecodeSupport`, which `assumeTrue` their sample files are present
|
||||
and skip when they are not. That is by design and unrelated to API level.
|
||||
|
||||
One harmless warning appears during the run and can be ignored:
|
||||
`No UID for androidx.test.services in user 0`, from an `appops` call the test services
|
||||
package makes before it is fully registered.
|
||||
|
||||
So the app is correct on Android 17. What is missing is only *automated* coverage in
|
||||
CI. Until the image is fixed, run the suite on a physical API 37 device before release;
|
||||
that is the substitute for the missing matrix row.
|
||||
`No UID for androidx.test.services in user 0`, from an `appops` call the test services package
|
||||
makes before it is fully registered.
|
||||
|
||||
## Reproducing it
|
||||
|
||||
Locally, with `-gpu host` so the emulator itself does not segfault on Intel graphics:
|
||||
Both halves, so the renderer claim can be checked rather than taken on trust:
|
||||
|
||||
```bash
|
||||
export ANDROID_HOME="$HOME/Android/Sdk"
|
||||
export PATH="$ANDROID_HOME/platform-tools:$ANDROID_HOME/emulator:$ANDROID_HOME/cmdline-tools/latest/bin:$PATH"
|
||||
|
||||
sdkmanager --install "system-images;android-37.0;google_apis;x86_64"
|
||||
echo no | avdmanager create avd -n api37_repro \
|
||||
-k "system-images;android-37.0;google_apis;x86_64" -d pixel_6 --force
|
||||
|
||||
$ANDROID_HOME/emulator/emulator -avd api37_repro \
|
||||
-no-window -gpu host -noaudio -no-boot-anim -camera-back none -no-snapshot &
|
||||
# never boots -- surfaceflinger aborts every ~6 s, forever
|
||||
emulator -avd api37_repro -no-window -gpu host \
|
||||
-noaudio -no-boot-anim -camera-back none -no-snapshot &
|
||||
|
||||
# Boot never completes. Count the aborts:
|
||||
adb logcat -d -b crash | grep -c hasReadColorBufferDma
|
||||
# boots in ~85 s, having aborted once or twice on the way
|
||||
emulator -avd api37_repro -no-window -gpu swangle_indirect \
|
||||
-noaudio -no-boot-anim -camera-back none -no-snapshot &
|
||||
```
|
||||
|
||||
`sys.boot_completed` never reaches `1`, `pgrep -f system_server` stays empty, and
|
||||
`keystore2`'s watchdog logs `await_boot_completed ... Overdue` indefinitely.
|
||||
Count the aborts either way:
|
||||
|
||||
To see the CI-side form instead, restore the API 37 row in the E2E matrix of
|
||||
`status_check.yml` (`api-level: "37.0"` — a bare `37` fails earlier still, during SDK
|
||||
setup, because there is no `platforms;android-37`).
|
||||
```bash
|
||||
adb -s emulator-5554 logcat -d -b crash | grep -c hasReadColorBufferDma
|
||||
```
|
||||
|
||||
Under `-gpu host`, `sys.boot_completed` never reaches `1`, `pgrep -f system_server` stays empty,
|
||||
and `keystore2`'s watchdog logs `await_boot_completed ... Overdue` indefinitely.
|
||||
|
||||
`tools/local-emulator/run-e2e.sh 37` does all of this, with the working renderer picked
|
||||
automatically — see `gpu_for_api` in that file.
|
||||
|
||||
## Filing this upstream
|
||||
|
||||
Not yet filed. To file it:
|
||||
Not yet filed. The report is stronger than it was, because the renderer dependency narrows it:
|
||||
|
||||
1. Go to <https://issuetracker.google.com/> and sign in with a Google account.
|
||||
2. Choose **Report an issue**, then pick the component for the Android emulator — search
|
||||
the component picker for "Emulator"; it sits under the Android Studio component tree.
|
||||
If the picker is unclear, Android Studio's **Help → Submit Feedback** opens the same
|
||||
tracker with the component preselected, and the emulator's own **Extended controls →
|
||||
Help → File a bug** does likewise.
|
||||
3. Title it for the mechanism, not the symptom, so it is searchable — for example:
|
||||
`surfaceflinger aborts in GoldfishMapper::readFromHost (hasReadColorBufferDma) on
|
||||
android-37.0 google_apis x86_64`.
|
||||
4. Paste the assertion and backtrace from the top of this document, the environment
|
||||
table, and the reproduction steps above. State explicitly that it reproduces on two
|
||||
unrelated hosts under both GPU modes — that is the detail that stops it being closed
|
||||
as a local graphics problem.
|
||||
5. List what was ruled out. Bugs that arrive with `-feature -GLDMA` already eliminated
|
||||
tend not to bounce back asking for it.
|
||||
6. Attach:
|
||||
- the guest tombstone, via `adb pull /data/tombstones` (or the `pbtombstone` output
|
||||
the crash log names)
|
||||
1. Go to <https://issuetracker.google.com/>, **Report an issue**, and pick the Android emulator
|
||||
component (search the component picker for "Emulator"; Android Studio's **Help → Submit
|
||||
Feedback** opens the same tracker with it preselected).
|
||||
2. Title it for the mechanism: `surfaceflinger aborts in GoldfishMapper::readFromHost
|
||||
(hasReadColorBufferDma) on android-37.0 and android-37.1 x86_64 -- fatal under -gpu host,
|
||||
intermittent under ANGLE`.
|
||||
3. Paste the assertion and backtrace, the environment block, and the seven-row matrix. The
|
||||
matrix is the valuable part: it shows the abort is not renderer-specific but its *frequency*
|
||||
is, which points at the readback path rather than at any one GL implementation.
|
||||
4. State that it reproduces on two independent system images (`37.0` rev 6 and `37.1` rev 8) and
|
||||
on two unrelated hosts, and that `-feature -GLDMA,-GLDMA2,-GLDirectMem` does not suppress it.
|
||||
5. Attach:
|
||||
- the guest tombstone, via `adb pull /data/tombstones` (or the `pbtombstone` output the crash
|
||||
log names)
|
||||
- `adb logcat -d -b crash > crash.txt`
|
||||
- the emulator's own stdout log, captured by redirecting the launch command
|
||||
- the emulator's own stdout log (`-verbose -debug all`, redirected)
|
||||
- the AVD's `config.ini`
|
||||
- a link to a failing CI job, which shows it on hardware you do not control:
|
||||
<https://github.com/JMR-dev/LibreMediaConverter/actions/runs/32545625459/job/96963461184>
|
||||
@@ -199,16 +482,38 @@ Record the issue number here once filed.
|
||||
|
||||
## When to revisit
|
||||
|
||||
Re-add the API 37 row when any of these happens:
|
||||
The old trigger list named "a new `android-37.0` revision", which is why nothing ever fired even
|
||||
though a new API level shipped. Watch for these instead:
|
||||
|
||||
- a new `android-37.0` system image revision ships (this was revision 6)
|
||||
- an ATD image appears for API 37
|
||||
- the upstream issue is marked fixed
|
||||
- **any new API 37.x system image**, not just a new revision of `37.0` — `37.1` rev 8 and
|
||||
`37.2-beta*` already exist, and more will. Test with `-gpu host`: if it boots, the guest mapper
|
||||
is fixed.
|
||||
- **an ATD image for API 37.** Still none as of 2026-08-22. ATD images ship without SystemUI,
|
||||
which is what drives `RegionSamplingThread`, so one would very likely sidestep the bug
|
||||
entirely. Try it before anything else here.
|
||||
- **the upstream issue being marked fixed.**
|
||||
|
||||
Until then the gap is narrower than the missing row suggests. `targetSdk` is 37, so the
|
||||
app is compiled and unit-tested against it; the API-dependent behaviour this matrix
|
||||
exists to exercise — the foreground service type, absent below 34, `dataSync` at 34,
|
||||
`mediaProcessing` from 35 — is covered at 35 and 36; and the full instrumented suite has
|
||||
been run green on real API 37 hardware. What is missing is *automated* API 37 coverage,
|
||||
so a regression there would not be caught by a pull request. Run the suite on a physical
|
||||
API 37 device before each release for as long as this row is absent.
|
||||
## Correction owed to `CLAUDE.md`
|
||||
|
||||
`CLAUDE.md` currently says:
|
||||
|
||||
> - **The API 37 image is broken.** `android-37.0` crash-loops surfaceflinger inside its own
|
||||
> gralloc mapper, so every test fails there regardless of this app.
|
||||
> `docs/api-37-emulator-crash.md` records the evidence and the ruled-out fixes; CI's matrix
|
||||
> therefore stops at API 36 even though targetSdk is 37.
|
||||
|
||||
The first sentence is right, and now under-specified in one direction and over-specified in the
|
||||
other: it is not only `android-37.0` (it is `37.1` too), and it does not crash-loop under every
|
||||
renderer. Proposed replacement, offered for review rather than applied here — `CLAUDE.md` is left
|
||||
alone deliberately, because several branches touch it:
|
||||
|
||||
> - **The API 37 images crash-loop surfaceflinger under the host GL renderer.** Both
|
||||
> `android-37.0` and `android-37.1` abort inside their own gralloc mapper
|
||||
> (`RegionSamplingThread` → `GoldfishMapper::readFromHost`), and when surfaceflinger dies init
|
||||
> SIGKILLs zygote, so the framework restarts under the test run. Under `-gpu host` it never
|
||||
> boots at all; under `-gpu swangle_indirect` it boots and the aborts merely become
|
||||
> intermittent. `docs/api-37-emulator-crash.md` has the seven-run matrix and the ruled-out
|
||||
> list, and `tools/local-emulator/run-e2e.sh` picks the working renderer per API level.
|
||||
> CI's matrix stops at 36 because CI runs `swiftshader_indirect` on a GPU-less runner and the
|
||||
> aborts continue there too. **API 37 still needs a manual check on the Pixel 10 Pro XL before
|
||||
> each release.**
|
||||
|
||||
+52
-17
@@ -1,8 +1,10 @@
|
||||
# Emulators do run on this host: the segfault is SwiftShader's JIT against SELinux
|
||||
|
||||
**Status:** solved. Local instrumented runs work with `-gpu host`, and the suite is green
|
||||
on API 33–36 — 49 tests, 0 failures, 0 errors, 2 skipped on every level. See
|
||||
[The sweep, run](#the-sweep-run).
|
||||
**Status:** solved. Local instrumented runs work with `-gpu host`, and the whole suite is green
|
||||
on API 33–36 — 0 failures, 0 errors and the two by-design skips on every level, measured as
|
||||
49 / 0 / 0 / 2 at `22c7914`, where the suite was 49 tests. See [The sweep, run](#the-sweep-run),
|
||||
and [Reading these totals](api-37-emulator-crash.md#reading-these-totals) before comparing any
|
||||
total with another checkout's.
|
||||
**Last verified:** 2026-08-22, emulator `37.1.11.0` (build 15917651), Fedora 44,
|
||||
kernel `7.1.8-200.fc44`, `selinux-policy-44.6-1.fc44`
|
||||
|
||||
@@ -190,11 +192,28 @@ emulator -avd <name> -no-window -gpu host \
|
||||
`.github/scripts/e2e-run.sh` for the run itself. Use it rather than the raw command:
|
||||
|
||||
```bash
|
||||
tools/local-emulator/run-e2e.sh # API 33 34 35 36
|
||||
tools/local-emulator/run-e2e.sh # API 33 34 35 36 37
|
||||
tools/local-emulator/run-e2e.sh 35 # one level
|
||||
tools/local-emulator/run-e2e.sh 37 37.1 # both API 37 images
|
||||
GPU_MODE=swangle_indirect tools/local-emulator/run-e2e.sh 35
|
||||
```
|
||||
|
||||
Levels are the labels above, not SDK ints: API 37's SDK directories are dotted
|
||||
(`android-37.0`, `android-37.1`) and there is no `android-37`, so `37` is accepted as a
|
||||
spelling of `37.0`. Setting `GPU_MODE` forces one renderer on every level, which is what
|
||||
you want when measuring a mode; leaving it unset lets `gpu_for_api` pick, which is what
|
||||
you want when running the suite — 33–36 need `host` and 37 must not have it.
|
||||
|
||||
**A bare `run-e2e.sh` exits 1, and that is the design.** API 37 is in the default list
|
||||
deliberately — leaving it out is what left the level unlooked-at for as long as it was — and
|
||||
it is permanently two failures short of green: `Media3EngineTest` cannot drive the emulator's
|
||||
`c2.goldfish.h264.decoder` on those images, which
|
||||
[`api-37-emulator-crash.md`](api-37-emulator-crash.md) pins on the image and not on this app
|
||||
(API 35 under the same renderer is green). The summary row names the two expected failures so
|
||||
that a third is visibly new, and the script repeats the point on the way out. Anything that
|
||||
treats a non-zero exit as breakage — a wrapper, a hook, a habit — should name the levels it
|
||||
wants: `run-e2e.sh 33 34 35 36` is the sweep that can be green.
|
||||
|
||||
`swangle_indirect` is the fallback worth knowing about. It is entirely software, so it
|
||||
does not depend on reaching the session's GPU — useful over plain SSH, where `-gpu host`
|
||||
has not been tested and may not find a device. It is also the closest local analogue to
|
||||
@@ -236,8 +255,9 @@ after an AGP upgrade.
|
||||
## The sweep, run
|
||||
|
||||
`tools/local-emulator/run-e2e.sh`, one invocation per level so each got a freshly created
|
||||
AVD, `-gpu host` throughout, 2026-08-22 19:42–19:56. Every level matches the physical
|
||||
Pixel 10 Pro XL (API 37) baseline of 49 / 0 / 0 / 2 exactly:
|
||||
AVD, `-gpu host` throughout, 2026-08-22 19:42–19:56, on `22c7914`. All four levels agree exactly,
|
||||
and 49 is that checkout's whole suite — every `@Test` in `app/src/androidTest`, two of which skip
|
||||
by design everywhere:
|
||||
|
||||
| API | Android | AVD | Boot | `connectedDebugAndroidTest` | Tests | Failures | Errors | Skipped |
|
||||
|---|---|---|---|---|---|---|---|---|
|
||||
@@ -246,6 +266,12 @@ Pixel 10 Pro XL (API 37) baseline of 49 / 0 / 0 / 2 exactly:
|
||||
| 35 | 15 | `lmc_e2e_api35` | 40 s | 3 m 46 s | 49 | 0 | 0 | 2 |
|
||||
| 36 | 16 | `lmc_e2e_api36` | 90 s | 2 m 18 s | 49 | 0 | 0 | 2 |
|
||||
|
||||
The physical Pixel has never reported 49, and an earlier version of this paragraph said the
|
||||
sweep matched it exactly. Its green API 37 run was 40 / 0 / 0 / 2, at `edd6385` — the same suite
|
||||
nine tests earlier. What matches is 0 failures, 0 errors and the same two skips; totals only ever
|
||||
match between runs of one checkout, which
|
||||
[`api-37-emulator-crash.md`](api-37-emulator-crash.md#reading-these-totals) sets out.
|
||||
|
||||
Thirteen and a half minutes for the four levels, AVD creation and cold boots included;
|
||||
fifteen with the pre-warm build in front of them. Nothing needed a retry, and no level
|
||||
produced a `diagnostics-api*.txt` — `e2e-run.sh` writes that only on the failure path, so
|
||||
@@ -352,9 +378,11 @@ and the same binaries against the same kernel boot fine under `-gpu host`.
|
||||
before `sys.boot_completed` is ever set. Nothing in the guest — system image variant,
|
||||
RAM, disk size, ATD versus `google_apis` — can influence a host-side `mprotect` denial,
|
||||
so none of those axes was varied. (The API 37 failure in
|
||||
[`api-37-emulator-crash.md`](api-37-emulator-crash.md) is genuinely guest-side and
|
||||
genuinely unrelated: there the host emulator survives and the guest's `surfaceflinger`
|
||||
aborts.)
|
||||
[`api-37-emulator-crash.md`](api-37-emulator-crash.md) is genuinely guest-side — there the host
|
||||
emulator survives and the guest's `surfaceflinger` aborts — but it is **not** unrelated, as this
|
||||
paragraph originally claimed. Both are decided by the renderer, in opposite directions: below 37
|
||||
you must avoid SwiftShader GLES and `-gpu host` is the answer; at 37 you must avoid the *host* GL
|
||||
translator and `-gpu host` is the thing that never boots.)
|
||||
|
||||
**Turning the SELinux boolean on** — deliberately *not* done, though it would almost
|
||||
certainly work:
|
||||
@@ -401,8 +429,13 @@ are easy to forget to look at.
|
||||
"SwiftShader 4.0.0.1" as reported by the GLES translator.
|
||||
- **If `-gpu host` regresses** after a Mesa or kernel update, fall back to
|
||||
`GPU_MODE=swangle_indirect`, which needs no GPU at all.
|
||||
- **This changes nothing about API 37.** That image is broken for a different reason and
|
||||
still must be checked on the physical Pixel 10 Pro XL before each release.
|
||||
- **API 37 needs the opposite renderer, and this file used to say it needed nothing.** The
|
||||
original bullet here read "This changes nothing about API 37"; that turned out to be wrong.
|
||||
The API 37 images abort `surfaceflinger` under the *host* GL translator and boot under ANGLE —
|
||||
the exact mirror of the rule above — and `run-e2e.sh` therefore picks the renderer per API
|
||||
level. See [`api-37-emulator-crash.md`](api-37-emulator-crash.md), which was rewritten on
|
||||
2026-08-22 with the seven-run matrix. API 37 still must be checked on the physical Pixel 10 Pro
|
||||
XL before each release.
|
||||
|
||||
## Correction owed to `CLAUDE.md`
|
||||
|
||||
@@ -432,12 +465,14 @@ Proposed replacement for the section, offered for review rather than applied her
|
||||
> `auto` (the default), `off` and `guest` do when headless. `-gpu host` works, and the
|
||||
> harness both picks it and refuses the others. `docs/local-emulator.md` has the
|
||||
> backtrace and the mode matrix.
|
||||
> - **The API 37 image is broken.** `android-37.0` crash-loops surfaceflinger inside its
|
||||
> own gralloc mapper, so every test fails there regardless of this app —
|
||||
> `docs/api-37-emulator-crash.md` records the evidence and the ruled-out fixes. This is
|
||||
> unrelated to the renderer above: it is a guest-side bug that CI hits too, which is why
|
||||
> the matrix stops at API 36 even though targetSdk is 37. **API 37 needs a manual check
|
||||
> on the Pixel 10 Pro XL before each release.**
|
||||
> - **API 37 needs the opposite renderer, and SystemUI turned off.** Both `android-37.0` and
|
||||
> `android-37.1` abort surfaceflinger inside their own gralloc mapper, and init SIGKILLs
|
||||
> zygote each time. Under `-gpu host` they never boot; under `-gpu swangle_indirect` they
|
||||
> boot, and disabling SystemUI removes the trigger. `run-e2e.sh` does all of that per level,
|
||||
> and the local API 37 result is two failures and the two usual skips, not a clean run. CI's
|
||||
> matrix still stops at 36.
|
||||
> `docs/api-37-emulator-crash.md` has the matrix and the reasoning. **API 37 needs a manual
|
||||
> check on the Pixel 10 Pro XL before each release.**
|
||||
|
||||
The wording is worth getting right rather than merely correcting, because the original was
|
||||
not a careless sentence — it was a reasonable inference from three crashes, written down
|
||||
|
||||
+344
-48
@@ -3,13 +3,24 @@
|
||||
# Runs the instrumented suite on a local emulator, on this workstation, for one or more
|
||||
# API levels.
|
||||
#
|
||||
# Usage: tools/local-emulator/run-e2e.sh [API ...] # default: 33 34 35 36
|
||||
# Usage: tools/local-emulator/run-e2e.sh [API ...] # default: 33 34 35 36 37
|
||||
#
|
||||
# GPU_MODE=host renderer to use; see the refusal list below
|
||||
# API levels are the labels below, not SDK ints: 33-36, plus `37` (= `37.0`) and `37.1`.
|
||||
#
|
||||
# GPU_MODE= force one renderer on every level; unset means per-API (gpu_for_api)
|
||||
# EMULATOR_PORT=5560 console port, so the serial is deterministic
|
||||
# BOOT_TIMEOUT=300 seconds to wait for sys.boot_completed
|
||||
# KEEP_AVD=1 do not delete an AVD this script created
|
||||
#
|
||||
# EXIT CODE: 0 only if every level was green; 1 if any level failed, wedged or could not be
|
||||
# set up; 2 if it refused to start at all. **A bare `run-e2e.sh` therefore exits 1 by design.**
|
||||
# API 37 is in the default list on purpose -- leaving it out is what left the level unlooked-at
|
||||
# for as long as it was -- and it is permanently two failures short of green, on the emulator's
|
||||
# own c2.goldfish.h264.decoder rather than on anything this app does. The summary names the two,
|
||||
# so a third is visibly new, and the last line printed says the same thing. Anything that reads a
|
||||
# non-zero exit as breakage should name the levels it wants: `run-e2e.sh 33 34 35 36` is the
|
||||
# sweep that can be green. docs/api-37-emulator-crash.md has the measurements.
|
||||
#
|
||||
# WHY THIS EXISTS, AND WHAT IT DELIBERATELY DOES NOT DO
|
||||
#
|
||||
# It is a *launcher*, not a second test harness. The diagnostics -- the FAILED-vs-WEDGED
|
||||
@@ -34,6 +45,14 @@
|
||||
# those modes was measured crashing. docs/local-emulator.md has the backtrace, the faulting
|
||||
# page's RW-without-E segment flags, and the full mode matrix.
|
||||
#
|
||||
# AND THE ONE THING API 37 NEEDS THAT 33-36 DO NOT: the opposite renderer. On the API 37
|
||||
# images the guest's Gralloc5 mapper aborts surfaceflinger from RegionSamplingThread
|
||||
# (`Assertion failed: !rcEnc->featureInfo()->hasReadColorBufferDma`). Under `-gpu host` that
|
||||
# repeats every few seconds and the device never boots; under ANGLE it fires a handful of
|
||||
# times and the boot survives. So `host` is required below 37 and forbidden at 37, which is
|
||||
# why the renderer is chosen per level in gpu_for_api rather than set once.
|
||||
# docs/api-37-emulator-crash.md has that matrix.
|
||||
#
|
||||
# THE OTHER LOCAL-ONLY HAZARD: a physical Pixel is usually plugged into this machine, so
|
||||
# `adb` is ambiguous in a way it never is on a runner, and an unpinned run would install
|
||||
# and execute this suite on the phone. Every path below pins the emulator serial.
|
||||
@@ -65,12 +84,14 @@ export ANDROID_HOME="${ANDROID_HOME:-$HOME/Android/Sdk}"
|
||||
export ANDROID_SDK_ROOT="$ANDROID_HOME"
|
||||
export PATH="$ANDROID_HOME/platform-tools:$ANDROID_HOME/emulator:$ANDROID_HOME/cmdline-tools/latest/bin:$PATH"
|
||||
|
||||
GPU_MODE="${GPU_MODE:-host}"
|
||||
# Empty means "let each level pick" -- see gpu_for_api. Setting GPU_MODE forces one renderer
|
||||
# on every level, which is what you want when measuring a mode, not when running the suite.
|
||||
GPU_MODE="${GPU_MODE:-}"
|
||||
EMULATOR_PORT="${EMULATOR_PORT:-5560}"
|
||||
BOOT_TIMEOUT="${BOOT_TIMEOUT:-300}"
|
||||
SERIAL="emulator-${EMULATOR_PORT}"
|
||||
APIS=("$@")
|
||||
[ "${#APIS[@]}" -eq 0 ] && APIS=(33 34 35 36)
|
||||
[ "${#APIS[@]}" -eq 0 ] && APIS=(33 34 35 36 37)
|
||||
|
||||
# Matches CI. `disk-size: 8G` because the FFmpeg libraries do not fit the default userdata
|
||||
# partition; `ram-size: 2560M` because the emulator's own floor varies by API level and
|
||||
@@ -83,21 +104,28 @@ RESULTS_DIR="app/build/outputs/androidTest-results"
|
||||
LOG_DIR="${TMPDIR:-/tmp}/lmc-local-e2e"
|
||||
mkdir -p "$LOG_DIR"
|
||||
|
||||
# The two things that outlive a level, declared here rather than where they are first
|
||||
# assigned, because the cleanup trap below can fire before either has been reached.
|
||||
EMU_PID=""
|
||||
CREATED_AVDS=()
|
||||
|
||||
# ---------------------------------------------------------------- renderer preflight ---
|
||||
case "$GPU_MODE" in
|
||||
swiftshader_indirect | auto | off | guest)
|
||||
echo "REFUSING to launch with -gpu $GPU_MODE."
|
||||
echo "On this host that resolves to SwiftShader's GLES, whose JIT is denied execheap by"
|
||||
echo "SELinux; the emulator segfaults (exit 139) before boot. See docs/local-emulator.md."
|
||||
echo "Working modes: host (default), angle_indirect, swangle_indirect."
|
||||
exit 2
|
||||
;;
|
||||
host | angle_indirect | swangle_indirect) ;;
|
||||
*)
|
||||
echo "Unrecognised GPU_MODE '$GPU_MODE'. Known-good: host, angle_indirect, swangle_indirect."
|
||||
exit 2
|
||||
;;
|
||||
esac
|
||||
if [ -n "$GPU_MODE" ]; then
|
||||
case "$GPU_MODE" in
|
||||
swiftshader_indirect | auto | off | guest)
|
||||
echo "REFUSING to launch with -gpu $GPU_MODE."
|
||||
echo "On this host that resolves to SwiftShader's GLES, whose JIT is denied execheap by"
|
||||
echo "SELinux; the emulator segfaults (exit 139) before boot. See docs/local-emulator.md."
|
||||
echo "Working modes: host, angle_indirect, swangle_indirect."
|
||||
exit 2
|
||||
;;
|
||||
host | angle_indirect | swangle_indirect) ;;
|
||||
*)
|
||||
echo "Unrecognised GPU_MODE '$GPU_MODE'. Known-good: host, angle_indirect, swangle_indirect."
|
||||
exit 2
|
||||
;;
|
||||
esac
|
||||
fi
|
||||
|
||||
# A courtesy, not a gate: the boolean being on means SwiftShader would work too, and the
|
||||
# refusal list above could be relaxed. It is off on a stock Fedora.
|
||||
@@ -121,14 +149,90 @@ host_forensics() {
|
||||
journalctl --since "$since" --no-pager 2> /dev/null | grep -E 'avc: .*denied' | tail -10 || echo " (none)"
|
||||
}
|
||||
|
||||
# The API 37 counterpart of host_forensics. `-gpu host` there aborts surfaceflinger in a loop
|
||||
# and the device never boots; the working renderers abort it a few times and survive. Either way
|
||||
# the count is the number to look at, and the crash buffer is where it lives -- so print it on
|
||||
# every 37 level, not only on the failure path, because a level that passed with 40 aborts is
|
||||
# telling you something a level that passed with 1 is not.
|
||||
guest_forensics() {
|
||||
local api="$1" n
|
||||
case "$api" in 37 | 37.*) ;; *) return 0 ;; esac
|
||||
n="$(emu_adb logcat -d -b crash 2> /dev/null | grep -c 'hasReadColorBufferDma')"
|
||||
echo " surfaceflinger hasReadColorBufferDma aborts: ${n:-?} (docs/api-37-emulator-crash.md)"
|
||||
}
|
||||
|
||||
# API label -> system image. API 33-36 are plain integers with a `google_apis` image. API 37
|
||||
# is not: its SDK directories are dotted minor versions (`android-37.0`, `android-37.1`), there
|
||||
# is no `android-37`, and from 37.1 onwards Google ships only 16 KB-page (`ps16k`) images for
|
||||
# x86_64. `37` is accepted as a spelling of `37.0` because that is what people type.
|
||||
image_pkg_for_api() {
|
||||
case "$1" in
|
||||
37 | 37.0) echo "system-images;android-37.0;google_apis;x86_64" ;;
|
||||
37.1) echo "system-images;android-37.1;google_apis_ps16k;x86_64" ;;
|
||||
*) echo "system-images;android-$1;google_apis;x86_64" ;;
|
||||
esac
|
||||
}
|
||||
|
||||
# The renderer requirement is per-API and the two levels want OPPOSITE things, which is why this
|
||||
# is a function and not a constant.
|
||||
#
|
||||
# 33-36: must NOT be SwiftShader GLES (host-side SELinux/execheap segfault) -- `host` is right.
|
||||
# 37.x: must NOT be the host GL translator. With `-gpu host` the guest's Gralloc5 mapper
|
||||
# aborts surfaceflinger in a loop and the device never boots; under ANGLE the same
|
||||
# assertion fires a handful of times and the boot survives it. Measured, not guessed --
|
||||
# docs/api-37-emulator-crash.md has the matrix.
|
||||
#
|
||||
# `swangle_indirect` rather than `angle_indirect` for 37: both boot, and swangle names its
|
||||
# renderer outright instead of resolving through `auto`'s path.
|
||||
gpu_for_api() {
|
||||
if [ -n "$GPU_MODE" ]; then
|
||||
echo "$GPU_MODE"
|
||||
return
|
||||
fi
|
||||
case "$1" in
|
||||
37 | 37.*) echo "swangle_indirect" ;;
|
||||
*) echo "host" ;;
|
||||
esac
|
||||
}
|
||||
|
||||
# `lmc_e2e_api37.0` would be a legal AVD name but an awkward one to type and to grep for.
|
||||
# `37` and `37.0` therefore give two AVD names (`lmc_e2e_api37`, `lmc_e2e_api37_0`) for the one
|
||||
# image. Harmless -- two AVDs off the same system image cost only disk -- and deliberately not
|
||||
# normalised, so that `run-e2e.sh 37 37.0` does not have both levels fight over one AVD.
|
||||
avd_for_api() { echo "lmc_e2e_api${1//./_}"; }
|
||||
|
||||
# Where avdmanager actually put the AVD. `$HOME/.android/avd` is only the default:
|
||||
# ANDROID_AVD_HOME, ANDROID_USER_HOME and ANDROID_SDK_HOME each move it, and hardcoding the
|
||||
# default meant a machine that sets any of them silently ran every level at stock RAM and
|
||||
# userdata size. Rather than encode a precedence that cannot be verified from here, look in
|
||||
# every location avdmanager honours and let the existence check pick.
|
||||
avd_config_path() {
|
||||
local avd="$1" base cfg
|
||||
for base in "${ANDROID_AVD_HOME:-}" \
|
||||
"${ANDROID_USER_HOME:+$ANDROID_USER_HOME/avd}" \
|
||||
"${ANDROID_SDK_HOME:+$ANDROID_SDK_HOME/.android/avd}" \
|
||||
"$HOME/.android/avd"; do
|
||||
[ -n "$base" ] || continue
|
||||
cfg="$base/${avd}.avd/config.ini"
|
||||
if [ -f "$cfg" ]; then
|
||||
echo "$cfg"
|
||||
return 0
|
||||
fi
|
||||
done
|
||||
return 1
|
||||
}
|
||||
|
||||
ensure_avd() {
|
||||
local api="$1" avd="$2"
|
||||
local pkg="system-images;android-${api};google_apis;x86_64"
|
||||
local pkg
|
||||
pkg="$(image_pkg_for_api "$api")"
|
||||
local img_dir="$ANDROID_HOME/system-images/${pkg#system-images;}"
|
||||
img_dir="${img_dir//;//}"
|
||||
|
||||
if avdmanager list avd -c 2> /dev/null | grep -qx "$avd"; then
|
||||
echo " reusing existing AVD $avd"
|
||||
else
|
||||
if [ ! -d "$ANDROID_HOME/system-images/android-${api}/google_apis/x86_64" ]; then
|
||||
if [ ! -d "$img_dir" ]; then
|
||||
echo " installing $pkg"
|
||||
yes | sdkmanager --install "$pkg" > /dev/null 2>&1 || {
|
||||
echo " FAILED to install $pkg"
|
||||
@@ -144,18 +248,37 @@ ensure_avd() {
|
||||
fi
|
||||
|
||||
# Written into config.ini rather than passed on the command line, which is how
|
||||
# reactivecircus/android-emulator-runner applies the same two settings in CI.
|
||||
local cfg="$HOME/.android/avd/${avd}.avd/config.ini"
|
||||
sed -i -e '/^disk\.dataPartition\.size=/d' -e '/^hw\.ramSize=/d' "$cfg"
|
||||
printf 'disk.dataPartition.size=%s\nhw.ramSize=%s\n' "$DISK_SIZE_BYTES" "$RAM_SIZE_MB" >> "$cfg"
|
||||
# reactivecircus/android-emulator-runner applies the same two settings in CI. On the reuse
|
||||
# path too, so an AVD left over from an older run gets today's pins.
|
||||
#
|
||||
# A level that cannot be pinned FAILS rather than running at the defaults. Unpinned, it
|
||||
# dies much later with "not enough space", which reads as a device problem -- CI's own
|
||||
# history is where that lesson comes from -- and nothing points back to a `sed` that
|
||||
# edited a path this script guessed wrong.
|
||||
local cfg
|
||||
if ! cfg="$(avd_config_path "$avd")"; then
|
||||
echo " FAILED: no config.ini for $avd in any directory avdmanager uses"
|
||||
echo " (ANDROID_AVD_HOME=${ANDROID_AVD_HOME:-unset}, ANDROID_USER_HOME=${ANDROID_USER_HOME:-unset},"
|
||||
echo " ANDROID_SDK_HOME=${ANDROID_SDK_HOME:-unset}, HOME=$HOME)"
|
||||
return 1
|
||||
fi
|
||||
if ! sed -i -e '/^disk\.dataPartition\.size=/d' -e '/^hw\.ramSize=/d' "$cfg"; then
|
||||
echo " FAILED to rewrite $cfg"
|
||||
return 1
|
||||
fi
|
||||
if ! printf 'disk.dataPartition.size=%s\nhw.ramSize=%s\n' \
|
||||
"$DISK_SIZE_BYTES" "$RAM_SIZE_MB" >> "$cfg"; then
|
||||
echo " FAILED to write the RAM/disk pins into $cfg"
|
||||
return 1
|
||||
fi
|
||||
}
|
||||
|
||||
boot_emulator() {
|
||||
local avd="$1" api="$2"
|
||||
local avd="$1" api="$2" gpu="$3"
|
||||
local boot_log="$LOG_DIR/emulator-api${api}.log"
|
||||
|
||||
emulator -avd "$avd" -port "$EMULATOR_PORT" \
|
||||
-no-window -gpu "$GPU_MODE" -noaudio -no-boot-anim -camera-back none -no-snapshot \
|
||||
-no-window -gpu "$gpu" -noaudio -no-boot-anim -camera-back none -no-snapshot \
|
||||
> "$boot_log" 2>&1 &
|
||||
EMU_PID=$!
|
||||
|
||||
@@ -181,6 +304,92 @@ boot_emulator() {
|
||||
return 1
|
||||
}
|
||||
|
||||
# API 37 only, and the reason API 37 can be run at all.
|
||||
#
|
||||
# The abort that breaks these images is reached from SurfaceFlinger's RegionSamplingThread,
|
||||
# which exists only because SystemUI registers a nav-bar luma-sampling listener. Each abort
|
||||
# kills surfaceflinger, and init responds by SIGKILLing zygote -- so the whole framework
|
||||
# restarts underneath the test run, which arrives as `Can't find service: package` and
|
||||
# `INSTRUMENTATION_ABORTED: System has crashed`. Under the host GL renderer that repeats
|
||||
# forever; under ANGLE it is roughly one every fifteen seconds, which a five-minute suite does
|
||||
# not survive either.
|
||||
#
|
||||
# Removing the listener removes the whole chain. Measured on android-37.0 under
|
||||
# swangle_indirect: 10-11 aborts per 150 s idle with SystemUI running, and 0 in 180 s with it
|
||||
# disabled, framework services up throughout.
|
||||
#
|
||||
# THIS IS A DEVIATION, and it is deliberately loud rather than silent. The API 37 leg does not
|
||||
# run the same device configuration as API 33-36 or as the Pixel. It is defensible only
|
||||
# because nothing in this suite touches SystemUI -- these are Media3, FFmpeg and WorkManager
|
||||
# tests -- and because the alternative is no API 37 coverage at all. Anything that ever does
|
||||
# depend on system UI must not trust this leg. docs/api-37-emulator-crash.md explains why.
|
||||
#
|
||||
# The retry loop is not defensive padding: at the moment boot_completed flips, the framework
|
||||
# may be in one of its restarts and `pm` is simply not published yet. The first attempt at this
|
||||
# failed exactly that way, with `cmd: Can't find service: package`.
|
||||
#
|
||||
# The framework restart at the end is not optional, and finding that out cost a run. By the
|
||||
# time `sys.boot_completed` flips, SystemUI has already registered its region-sampling listener,
|
||||
# and `pm disable-user` does not retract a registration that already happened -- it only stops
|
||||
# the package being started again. So the first attempt disabled SystemUI, reported success, and
|
||||
# then died exactly as before with `Starting 0 tests` and four more aborts. `stop; start` cycles
|
||||
# zygote deliberately, and the framework that comes back up does not start SystemUI at all.
|
||||
disable_region_sampling() {
|
||||
local api="$1" out i before after ready
|
||||
case "$api" in 37 | 37.*) ;; *) return 0 ;; esac
|
||||
|
||||
out=""
|
||||
for i in $(seq 1 20); do
|
||||
out="$(emu_adb shell pm disable-user --user 0 com.android.systemui 2>&1 | tr -d '\r')"
|
||||
case "$out" in
|
||||
*"new state: disabled"*)
|
||||
echo " SystemUI disabled on attempt $i"
|
||||
break
|
||||
;;
|
||||
esac
|
||||
out=""
|
||||
sleep 5
|
||||
done
|
||||
if [ -z "$out" ]; then
|
||||
echo " WARNING: could not disable SystemUI after 20 attempts."
|
||||
echo " Expect INSTRUMENTATION_ABORTED -- docs/api-37-emulator-crash.md"
|
||||
return 0
|
||||
fi
|
||||
|
||||
echo " restarting the framework so the region-sampling listener goes with it"
|
||||
emu_adb shell stop > /dev/null 2>&1
|
||||
emu_adb shell start > /dev/null 2>&1
|
||||
# There is no property worth waiting on here, and an earlier version of this only looked
|
||||
# like it was waiting on one: `stop` does not clear sys.boot_completed, so it still reads
|
||||
# `1` throughout the restart and any loop over it returns at once. The loop below is the
|
||||
# wait -- and it polls the better thing anyway, since `Can't find service: package` is the
|
||||
# failure it exists to prevent.
|
||||
ready=0
|
||||
for i in $(seq 1 30); do
|
||||
if emu_adb shell service check package 2> /dev/null | grep -q ': found' \
|
||||
&& emu_adb shell service check activity 2> /dev/null | grep -q ': found'; then
|
||||
ready=1
|
||||
break
|
||||
fi
|
||||
sleep 5
|
||||
done
|
||||
if [ "$ready" -ne 1 ]; then
|
||||
echo " WARNING: package and activity services still absent 150 s after the restart."
|
||||
echo " Expect INSTRUMENTATION_ABORTED -- docs/api-37-emulator-crash.md"
|
||||
fi
|
||||
|
||||
# Prove it worked rather than assume it. Zero new aborts over this window is what makes the
|
||||
# difference between a run that completes and one that reports `Starting 0 tests`.
|
||||
before="$(emu_adb logcat -d -b crash 2> /dev/null | grep -c 'hasReadColorBufferDma')"
|
||||
emu_adb shell 'sleep 45' > /dev/null 2>&1
|
||||
after="$(emu_adb logcat -d -b crash 2> /dev/null | grep -c 'hasReadColorBufferDma')"
|
||||
echo " quiet check: $((after - before)) new surfaceflinger aborts in 45 s (want 0)"
|
||||
if [ "$((after - before))" -ne 0 ]; then
|
||||
echo " WARNING: region sampling is still live; the run may not survive."
|
||||
fi
|
||||
return 0
|
||||
}
|
||||
|
||||
# CI gets this from the action's `disable-animations: true`.
|
||||
disable_animations() {
|
||||
local s
|
||||
@@ -189,17 +398,77 @@ disable_animations() {
|
||||
done
|
||||
}
|
||||
|
||||
# `${EMU_PID:-0}` used to guard these three calls, and it guarded the wrong thing: EMU_PID
|
||||
# is *empty*, not unset, if the background launch never produced a job, and `kill` reads pid
|
||||
# 0 as "the sender's whole process group" -- this script and, on a terminal, everything else
|
||||
# in the foreground group with it. The `kill -0` wait loop had the same shape and would have
|
||||
# spent its full grace period testing the group. Nothing to stop is now a return, never a
|
||||
# guess. (boot_emulator's own `kill -0 "$EMU_PID"` is unguarded and cannot reach that form:
|
||||
# it runs only after the assignment.)
|
||||
#
|
||||
# max_wait is a parameter so the interrupt path need not sit through the full grace period.
|
||||
stop_emulator() {
|
||||
local max_wait="${1:-30}" waited=0
|
||||
[ -n "${EMU_PID:-}" ] || return 0
|
||||
emu_adb emu kill > /dev/null 2>&1
|
||||
local waited=0
|
||||
while kill -0 "${EMU_PID:-0}" 2> /dev/null && [ "$waited" -lt 30 ]; do
|
||||
while kill -0 "$EMU_PID" 2> /dev/null && [ "$waited" -lt "$max_wait" ]; do
|
||||
sleep 2
|
||||
waited=$((waited + 2))
|
||||
done
|
||||
kill -9 "${EMU_PID:-0}" 2> /dev/null
|
||||
wait "${EMU_PID:-0}" 2> /dev/null
|
||||
kill -9 "$EMU_PID" 2> /dev/null
|
||||
wait "$EMU_PID" 2> /dev/null
|
||||
EMU_PID=""
|
||||
}
|
||||
|
||||
delete_created_avds() {
|
||||
local avd
|
||||
[ "${KEEP_AVD:-0}" = "1" ] && return 0
|
||||
for avd in ${CREATED_AVDS[@]+"${CREATED_AVDS[@]}"}; do
|
||||
# A SIGKILLed emulator does not get to remove its own lock files, and avdmanager can
|
||||
# refuse over them. Staying silent there would leak the very thing this exists to clean.
|
||||
avdmanager delete avd -n "$avd" > /dev/null 2>&1 \
|
||||
|| echo " WARNING: could not delete AVD $avd -- 'avdmanager delete avd -n $avd' by hand"
|
||||
done
|
||||
CREATED_AVDS=()
|
||||
}
|
||||
|
||||
# What an interrupted sweep used to leave behind: a headless emulator holding console port
|
||||
# $EMULATOR_PORT, and an lmc_e2e_apiNN AVD. The next run's `emulator -port` then collides
|
||||
# with the orphan, and `emu_adb` can resolve to it -- on a workstation that also has the
|
||||
# Pixel plugged in, exactly the ambiguity the ANDROID_SERIAL pinning exists to prevent. A
|
||||
# sweep is up to five boots long, so the window for one Ctrl-C is not small.
|
||||
#
|
||||
# Idempotent, and called explicitly on the normal path so its output cannot land after the
|
||||
# summary; the EXIT trap then finds nothing left to do. The emulator logs are deliberately
|
||||
# NOT removed -- they live in $LOG_DIR and are the only evidence a failed boot leaves.
|
||||
CLEANED=0
|
||||
cleanup() {
|
||||
[ "$CLEANED" = "1" ] && return 0
|
||||
CLEANED=1
|
||||
stop_emulator "${1:-30}"
|
||||
delete_created_avds
|
||||
}
|
||||
|
||||
# 6 s, not 30: Ctrl-C has already reached the emulator through the foreground process group,
|
||||
# so this is only waiting for it to finish writing, and `kill -9` follows regardless. The
|
||||
# EXIT trap is disarmed before exiting so the status below is the one that survives.
|
||||
#
|
||||
# bash runs a trap only between commands, so this starts when whatever was in the foreground
|
||||
# returns -- which for Ctrl-C is immediately, because the same interrupt reached that command
|
||||
# too. `kill -INT` aimed at this script alone waits for the foreground command to finish.
|
||||
on_signal() {
|
||||
echo
|
||||
echo "interrupted (SIG$1) -- stopping the emulator and removing the AVDs this run created"
|
||||
echo " emulator logs kept in $LOG_DIR"
|
||||
cleanup 6
|
||||
trap - EXIT
|
||||
exit "$2"
|
||||
}
|
||||
|
||||
trap 'on_signal INT 130' INT
|
||||
trap 'on_signal TERM 143' TERM
|
||||
trap cleanup EXIT
|
||||
|
||||
# The XML is authoritative. The console counter double-counts skips, so a run that reports
|
||||
# "42 tests" on stdout can be 40 in the report.
|
||||
#
|
||||
@@ -236,37 +505,44 @@ PY
|
||||
}
|
||||
|
||||
# ------------------------------------------------------------------------------ main ---
|
||||
CREATED_AVDS=()
|
||||
SUMMARY=()
|
||||
overall=0
|
||||
|
||||
for api in "${APIS[@]}"; do
|
||||
if [ "$api" = "37" ] || [ "$api" = "37.0" ]; then
|
||||
echo "SKIPPING API $api: the android-37.0 image crash-loops surfaceflinger."
|
||||
echo " See docs/api-37-emulator-crash.md. Test API 37 on the physical Pixel."
|
||||
continue
|
||||
fi
|
||||
# Whether the red exit is the expected one depends on which level produced it, and only the
|
||||
# loop knows that -- so it is recorded where `overall` is set rather than guessed from the
|
||||
# summary afterwards. A note at the end claiming a genuine API 34 failure was "by design"
|
||||
# would be the same defect it is there to prevent, one layer up.
|
||||
NON37_RED=0
|
||||
mark_red() {
|
||||
overall=1
|
||||
case "$1" in 37 | 37.*) ;; *) NON37_RED=1 ;; esac
|
||||
}
|
||||
|
||||
avd="lmc_e2e_api${api}"
|
||||
for api in "${APIS[@]}"; do
|
||||
avd="$(avd_for_api "$api")"
|
||||
gpu="$(gpu_for_api "$api")"
|
||||
started="$(date '+%Y-%m-%d %H:%M:%S')"
|
||||
echo "=============================================================="
|
||||
echo "API $api (avd=$avd gpu=$GPU_MODE serial=$SERIAL)"
|
||||
echo "API $api (avd=$avd gpu=$gpu serial=$SERIAL)"
|
||||
echo " image: $(image_pkg_for_api "$api")"
|
||||
echo "=============================================================="
|
||||
|
||||
if ! ensure_avd "$api" "$avd"; then
|
||||
SUMMARY+=("API $api: AVD SETUP FAILED")
|
||||
overall=1
|
||||
mark_red "$api"
|
||||
continue
|
||||
fi
|
||||
|
||||
if ! boot_emulator "$avd" "$api"; then
|
||||
if ! boot_emulator "$avd" "$api" "$gpu"; then
|
||||
host_forensics "$started"
|
||||
guest_forensics "$api"
|
||||
SUMMARY+=("API $api: BOOT FAILED")
|
||||
overall=1
|
||||
mark_red "$api"
|
||||
stop_emulator
|
||||
continue
|
||||
fi
|
||||
|
||||
disable_region_sampling "$api"
|
||||
disable_animations
|
||||
rm -rf "$RESULTS_DIR"
|
||||
|
||||
@@ -281,23 +557,43 @@ for api in "${APIS[@]}"; do
|
||||
unset ANDROID_SERIAL E2E_EXTRA_GRADLE_ARGS
|
||||
|
||||
line="$(summarise_results "$api")"
|
||||
guest_forensics "$api"
|
||||
# API 37 is in the default list on purpose, and it is expected to be red. Leaving it out would
|
||||
# put the level back where this whole exercise found it -- untested and unlooked-at -- but a
|
||||
# summary that just says "2 failures" with no explanation trains people to ignore the exit
|
||||
# code. So the row says which two, and a THIRD failure is then obviously new.
|
||||
case "$api" in
|
||||
37 | 37.*)
|
||||
line="$line
|
||||
expected here: 2 failures, both Media3EngineTest, on c2.goldfish.h264.decoder.
|
||||
A third is new -- docs/api-37-emulator-crash.md"
|
||||
;;
|
||||
esac
|
||||
if [ "$rc" -ne 0 ]; then
|
||||
line="$line [gradle exit $rc]"
|
||||
overall=1
|
||||
mark_red "$api"
|
||||
host_forensics "$started"
|
||||
fi
|
||||
SUMMARY+=("$line")
|
||||
stop_emulator
|
||||
done
|
||||
|
||||
if [ "${KEEP_AVD:-0}" != "1" ]; then
|
||||
for avd in ${CREATED_AVDS[@]+"${CREATED_AVDS[@]}"}; do
|
||||
avdmanager delete avd -n "$avd" > /dev/null 2>&1
|
||||
done
|
||||
fi
|
||||
cleanup
|
||||
|
||||
echo
|
||||
echo "===================== LOCAL E2E SUMMARY ======================"
|
||||
printf '%s\n' ${SUMMARY[@]+"${SUMMARY[@]}"}
|
||||
echo "=============================================================="
|
||||
|
||||
# An unexplained red exit trains people to stop reading exit codes, and this one is expected
|
||||
# whenever API 37 is in the sweep -- which the default list makes the common case. Said here
|
||||
# rather than only in the docs, because this is where it is actually read. Only when 37.x is
|
||||
# the ONLY thing that went red: a note calling a real failure elsewhere "by design" would be
|
||||
# worse than no note at all.
|
||||
if [ "$overall" -ne 0 ] && [ "$NON37_RED" -eq 0 ]; then
|
||||
echo "note: the only level that went red is API 37, which exits non-zero by design -- it is"
|
||||
echo " permanently 2 failures short of green. Confirm its row above shows exactly those"
|
||||
echo " two and nothing else; docs/api-37-emulator-crash.md says why they are the image."
|
||||
fi
|
||||
|
||||
exit "$overall"
|
||||
|
||||
Reference in New Issue
Block a user