diff --git a/docs/api-37-emulator-crash.md b/docs/api-37-emulator-crash.md index d70ea69..0cc6ff0 100644 --- a/docs/api-37-emulator-crash.md +++ b/docs/api-37-emulator-crash.md @@ -1,127 +1,381 @@ -# API 37 is not tested in CI: a crash in Google's `android-37.0` emulator image +# API 37 on the emulator: a guest gralloc bug that only the host GL renderer triggers -**Status:** open upstream, worked around by removing API 37 from the E2E matrix. -The app itself is verified good on real API 37 hardware — this is an emulator bug only. -**Last verified:** 2026-08-21, against emulator `37.1.11.0` and system image revision 6 +**Status:** the bug is real and still open upstream, but the previous diagnosis in this file was +wrong about its most important detail. **The renderer decides whether API 37 boots**, and once it +boots, disabling SystemUI collapses the crash rate far enough to run a suite — +`tools/local-emulator/run-e2e.sh 37` gets through the whole instrumented suite and comes back with +**2 failures, 0 errors and the two by-design skips** (measured 49 / 2 / 0 / 2 at `22c7914`, where +the suite was 49 tests — [Reading these totals](#reading-these-totals) before comparing any total +with another). The crashes do not stop outright, and the two failures are real; both are quantified +below. CI's matrix should still stop at 36 — see +[So should CI take API 37?](#so-should-ci-take-api-37). +**Last verified:** 2026-08-22, emulator `37.1.11.0` (build 15917651), Fedora 44, +against system images `android-37.0` rev 6 **and** `android-37.1` rev 8. -`minSdk` is 33 and `targetSdk` is 37, and the E2E matrix in -[`status_check.yml`](../.github/workflows/status_check.yml) runs API 33 through 36. -API 37 is deliberately absent. This is why. +## The correction -## Summary +This file previously said, under "What was ruled out": -The `android-37.0` emulator system image crashes `surfaceflinger` in a loop. The app -under test never gets a working framework, so every instrumented test fails regardless -of what the app does. The bug is in the emulator image, not in this project. +> **GPU mode.** Both `swiftshader_indirect` and `host` crash, with the same assertion and +> the same frames. The crash is in the gralloc mapper, below the renderer. -The crash is an assertion inside the emulator's own gralloc implementation: +**That is wrong.** The mapper is below the renderer, but *whether the mapper's bad path is +reached* is not. Re-measured on 2026-08-22, seven runs, one variable at a time: + +| # | system image | `-gpu` | GLES the emulator chose | booted? | surfaceflinger aborts | +|---|---|---|---|---|---| +| r01 | `android-37.0` rev 6 | `host` | host (Mesa Iris Xe) | **no**, 422 s | 71, looping | +| r02 | `android-37.1` rev 8 | `host` | host (Mesa Iris Xe) | **no**, 362 s | 65, looping | +| r03 | `android-37.0` rev 6 | `swangle_indirect` | ANGLE | **yes, 85 s** | 1 | +| r04 | `android-37.0` rev 6 | `host` + `-feature -GLDMA,-GLDMA2,-GLDirectMem` | host | **no**, 363 s | 57, looping | +| r05 | `android-37.0` rev 6 | `angle_indirect` | ANGLE | **yes, 112 s** | 2 | +| r06 | `android-37.1` rev 8 | `swangle_indirect` | ANGLE | **yes, 285 s** | 23 | +| r07 | `android-37.0` rev 6 | `host` + `-feature -HostComposition` | host | **no**, wedged adb at 208 s | not readable | + +The discriminator is exact across all seven: **a run boots if and only if the emulator log says +something other than `gles_mode_selected:host`.** + +One caveat about how independent those rows are, because the table flatters itself. `-gpu +angle_indirect` (r05) and `-gpu swangle_indirect` (r03) both logged `gles_mode_selected:swangle` +and both reported the same adapter, differing only in the Vulkan backend beneath +(`vulkan_mode_selected:lavapipe` against `swiftshader`). So they are closer to one GLES path +reached two ways than to two renderers agreeing — note that at API 33–36 +[`docs/local-emulator.md`](local-emulator.md) records `angle_indirect` resolving to ANGLE on +*llvmpipe*, a genuinely different adapter, which it did not do here. What is 7-for-7 is the +host-GLES-versus-not split, not "two independent renderers both work". + +``` +# r01, r02, r04, r07 -- never boots +INFO | emuglConfig_init: vulkan_mode_selected:host gles_mode_selected:host +INFO | Graphics Adapter Android Emulator OpenGL ES Translator (Mesa Intel(R) Iris(R) Xe Graphics (TGL GT2)) + +# r03, r05, r06 -- boots +INFO | emuglConfig_init: vulkan_mode_selected:swiftshader gles_mode_selected:swangle +INFO | Graphics Adapter Android Emulator OpenGL ES Translator (ANGLE (Google, Vulkan 1.2.0 + | (SwiftShader Device (Subzero) (0x0000C0DE)), SwiftShader driver-5.0.0)) +``` + +### Why the wrong claim looked right + +It rested on two samples of two different things, and neither of them was ANGLE. + +- The **local** `swiftshader_indirect` sample was void. On this workstation *every* + SwiftShader-GLES launch segfaults the host emulator before the guest matters at all — + SELinux denies `execheap` to SwiftShader's Reactor JIT. That is + [`docs/local-emulator.md`](local-emulator.md), and it was not yet understood when this file + was written. So "`swiftshader_indirect` crashes" was true, for an entirely unrelated reason, + and told you nothing about the gralloc assertion. +- The **CI** sample was one `swiftshader_indirect` run on a GPU-less `ubuntu-latest`, and the + **local** sample was one `-gpu host` run. Two renderers, one measurement each, and the pair + written up as "both GPU modes". + +`angle_indirect` and `swangle_indirect` — the two modes that work — had never been tried on +API 37. Neither had a second system image. + +The lesson is the same one `docs/local-emulator.md` ends on, which makes it worth repeating: +"both backends fail" is a claim about a matrix, and a matrix needs cells, not inference. Two +observations of two different configurations do not establish anything about a third. + +## What the bug actually is + +`surfaceflinger` aborts inside the emulator's own gralloc mapper: ``` Executable: /system/bin/surfaceflinger signal 6 (SIGABRT), code -1 (SI_QUEUE), tid: RegionSampling Abort message: 'Assertion failed: !rcEnc->featureInfo()->hasReadColorBufferDma' - #03 mapper.ranchu.so GoldfishMapper::readFromHost(cb_handle_t const&) const - #04 mapper.ranchu.so GoldfishMapper::GoldfishMapper()::'lambda'(...)::__invoke - #05 libui.so android::Gralloc5Mapper::lock(...) - #06 libui.so android::GraphicBufferMapper::lock(...) - #07 libui.so android::GraphicBuffer::lockAsync(...) - #08 libui.so android::GraphicBuffer::lock(...) - #09 surfaceflinger android::RegionSamplingThread::threadMain() + #03 /vendor/lib64/hw/mapper.ranchu.so GoldfishMapper::readFromHost(cb_handle_t const&) const+543 + #04 /vendor/lib64/hw/mapper.ranchu.so GoldfishMapper::GoldfishMapper()::'lambda'(...)::__invoke+704 + #05 /system/lib64/libui.so android::Gralloc5Mapper::lock(...)+63 + #06 /system/lib64/libui.so android::GraphicBufferMapper::lock(...)+198 + #07 /system/lib64/libui.so android::GraphicBuffer::lockAsync(...)+545 + #08 /system/lib64/libui.so android::GraphicBuffer::lock(...)+67 + #09 /system/bin/surfaceflinger android::RegionSamplingThread::threadMain()+2571 ``` -`RegionSamplingThread` is SystemUI's navigation-bar luma sampling. It calls -`GraphicBuffer::lock`, which routes into `GoldfishMapper::readFromHost`, which asserts -that the host has *not* negotiated the `ReadColorBufferDma` capability. On this image -the host has, so the assertion fails and `surfaceflinger` aborts. It restarts and -aborts again. +`RegionSamplingThread` is SystemUI's nav-bar luma sampling. It locks a `GraphicBuffer` for CPU +read; that routes through the Gralloc5 mapper into `GoldfishMapper::readFromHost`, which is the +*non-DMA* readback path and asserts that the host has not negotiated `ReadColorBufferDma`. The +host always has, so the assert fires whenever that path is taken. -## Impact +Two facts pin down what "always" means: -The failure surfaces in two different ways depending on how far the job gets, which is -why it took several rounds to identify: +- **The capability is negotiated regardless of renderer.** The evidence is the aborts + themselves: the assertion that fires is `!hasReadColorBufferDma`, and it fires under ANGLE + (r03/r05/r06) as well as under the host translator — just far less often. That is a direct + observation of the guest having negotiated DMA readback under both, and it stands alone. + (Supporting only, and weaker than it first looks: `ANDROID_EMU_read_color_buffer_dma` appears + in exactly one file in the SDK, `emulator/lib64/libgfxstream_backend.so`, which every `-gpu` + mode goes through. A string search establishes where the extension is implemented, not that + it is negotiated on every path.) +- **It is not gated by any feature flag the emulator exposes.** See the ruled-out list below. -| Guest RAM | Where it dies | What CI reports | -|---|---|---| -| 1536 MB | during APK install | `Unknown failure: cmd: Can't find service: package` | -| 2560 MB | during the test run | `There were failing tests` — all of them | +So the renderer does not decide whether the guest *believes* DMA readback exists. It decides how +often `RegionSamplingThread` ends up in `readFromHost` — which under the host GL translator is +constantly, and under ANGLE is occasionally. -At 2560 MB the install succeeds and the tests actually execute, then fail wholesale. -The first failure in the report is misleading: +### Why one abort takes down the whole device + +`surfaceflinger` is a critical service. When it dies, `init` kills the framework with it: ``` -kotlin.UninitializedPropertyAccessException: lateinit property output has not - been initialized - at Media3EngineTest.tearDown(Media3EngineTest.kt:53) - -java.lang.IllegalStateException: WorkManager is not initialized properly. - You have explicitly disabled WorkManagerInitializer in your manifest, ... +08-22 21:40:28.253 I/init: Sending SIGKILL to service 'zygote' (pid 470) process group... +08-22 21:40:28.260 I/init: Service 'zygote' (pid 470) received SIGKILL ``` -Neither is a real defect in this project. `tearDown` throws because `setUp` never got -far enough to assign `output`, and WorkManager's `InitializationProvider` never runs -because content-provider installation fails on a framework whose `surfaceflinger` is -crash-looping. The same tests pass at API 33, 34, 35, and 36 in the same CI run, and -the first `surfaceflinger` abort is timestamped *before* the test results are reported. - -This is not inference. The full suite was run against a physical API 37 device and -passed — see [Verified on real API 37 hardware](#verified-on-real-api-37-hardware) -below. `ConversionWorkerTest` and `ConcatWorkerTest`, which drive a real WorkManager -round trip and are among the tests that failed this way in CI, both pass there. +Everything above zygote goes with it, which is why the symptoms look nothing like a graphics +bug. Under `-gpu host` the cycle repeats every five to seven seconds forever and +`sys.boot_completed` is never set. Under ANGLE the aborts are sparse enough that the boot +usually completes between them — but they do not stop, and each one is a framework restart. +That is the difference between "boots" and "is usable", and it is the reason this is not simply +fixed by changing the renderer. See [Can the suite run on it?](#can-the-suite-run-on-it) below. ## Environment -Reproduced identically in two unrelated environments, so it is not specific to a host -GPU, driver, or CI runner. - | | GitHub Actions | Local workstation | |---|---|---| -| Host | `ubuntu-latest`, no GPU | Fedora, Intel Iris Xe (TGL GT2) | +| Host | `ubuntu-latest`, no GPU | Fedora 44, Intel Iris Xe (TGL GT2), kernel `7.1.8-200.fc44` | | Emulator | `37.1.11.0` (build 15917651) | `37.1.11.0` (build 15917651) | -| GPU mode | `swiftshader_indirect` | `host` | -| Result | boots, aborts during tests | aborts before boot completes | +| GPU mode measured | `swiftshader_indirect` | `host`, `angle_indirect`, `swangle_indirect` | -System image: `system-images;android-37.0;google_apis;x86_64`, `Pkg.Revision=6`, -`AndroidVersion.ApiLevel=37.0`, `AndroidVersion.ExtensionLevel=22` +Images, both reproducing it: ``` -Build fingerprint: google/sdk_gphone64_x86_64/emu64xa:17/CE2A.260420.019/15611780:userdebug/dev-keys -Kernel Release: 6.12.58-android16-6-gccafb60de224-ab14828483 +system-images;android-37.0;google_apis;x86_64 Pkg.Revision=6 ApiLevel=37.0 ExtensionLevel=22 + fingerprint google/sdk_gphone64_x86_64/emu64xa:17/CE2A.260420.019/15611780:userdebug/dev-keys +system-images;android-37.1;google_apis_ps16k;x86_64 Pkg.Revision=8 ApiLevel=37.1 ExtensionLevel=23 + ro.build.version.codename=REL (a release image, not a preview) ``` +**Note the `ps16k` in the second one — it is not optional, and it is why the 37.1 result is +interpretable.** From API 37.1 onward Google ships *only* 16 KB-page x86_64 images; there is no +plain `google_apis` variant to pick. `sdkmanager --list` for 37.1 and 37.2-beta* offers nothing +but `google_apis_ps16k` and `google_apis_playstore_ps16k`. That makes page-size alignment a +prerequisite rather than a detail: a `.so` that is not 16 KB aligned will not load on such a +guest, and the resulting failure looks like an app bug. Checked before the first `ps16k` boot, +using the same test `build.yml` applies to release APKs — all 20 libraries in the committed +`bin/ffmpeg-kit-next-8.1.1.aar`, both ABIs, report `0x4000`: + +``` +$ for f in jni/*/*.so; do readelf -lW "$f" | awk '$1=="LOAD"{print $NF}' | sort -u; done +0x4000 (x20: libavcodec, libavdevice, libavfilter, libavformat, libavutil, + libc++_shared, libffmpegkit, libffmpegkit_abidetect, libswresample, libswscale + -- arm64-v8a and x86_64) +``` + +So when `android-37.1` reproduced the abort, that was the gralloc bug and not a page-size +mismatch. `image_pkg_for_api` in `tools/local-emulator/run-e2e.sh` encodes the `ps16k` tag for +37.1; if this ever fails after an FFmpeg rebuild, re-run the alignment check first. + ## What was ruled out, and how -Each of these was tested rather than reasoned about, because the first three attempts -at this bug were plausible fixes that turned out to address earlier, unrelated failures. +**A newer system image.** This file's own revisit trigger was "a new `android-37.0` system image +revision ships (this was revision 6)". That trigger was written too narrowly and would never have +fired: `android-37.0` is *still* revision 6, but Google shipped a whole new minor level. +`android-37.1` `google_apis_ps16k` revision 8 — a `REL` build, not a beta — was installed and +tested (r02, r06) and **behaves identically**: same assertion, same frames, never boots under +`-gpu host`, and *worse* under ANGLE (23 aborts to `37.0`'s 1). `android-37.2-beta3` exists too +but was not needed; two independent images agreeing settles it, and a beta could not be used by +CI anyway. -**Guest memory.** The emulator raises an undersized guest to a minimum on its own, but -only for API levels it recognises, and it does not recognise `"37.0"`. API 33 bumps to -2048 MB and 34/35/36 to 2560 MB, while API 37 logged no bump at all and ran at the -`pixel_6` default of 1536 MB. Setting `ram-size: 2560M` explicitly fixed that asymmetry -and did change the outcome — the job got past install and into the test run — but it is -not the underlying bug. At the moment of failure the guest reported `MemTotal 2527392 -kB` with `MemAvailable 1507104 kB`: 1.5 GB free, and no OOM kills. +**An ATD image.** Still does not exist for API 37. `sdkmanager --list` offers `aosp_atd` and +`google_atd` for API 30 through 36 and nothing above: -**GPU mode.** Both `swiftshader_indirect` and `host` crash, with the same assertion and -the same frames. The crash is in the gralloc mapper, below the renderer. +``` +system-images;android-36;google_atd;x86_64 | 1 | Google APIs ATD Intel x86_64 Atom System Image +(no android-37 ATD of any kind) +``` -**Disabling the DMA feature.** `GLDMA` is the host feature that most plausibly backs the -guest's `hasReadColorBufferDma`. Launching with `-feature -GLDMA` was accepted by the -emulator — the log confirms `Feature 'GLDMA' (51) is overridden to 'disabled'` — and -`surfaceflinger` still aborted 13 times and the device never finished booting. Whatever -sets that guest capability, it is not this flag. +For API 37 the only x86_64 images are `google_apis`, `google_apis_playstore`, their `ps16k` +16 KB-page variants, and Wear OS. Check again when revisiting. -**An ATD image.** `google_atd` / `aosp_atd` images are built for automated testing and -ship without the SystemUI package set, which is what drives `RegionSamplingThread` in -the first place. That would likely sidestep the bug class entirely, but **no ATD image -exists for `android-37.0`** — only `google_apis`, `google_apis_playstore`, the `ps16k` -16 KB-page variants, and Wear OS. Check again when revisiting; if an ATD image appears, -try it before anything else here. +**The DMA feature flags.** `GLDMA` alone was ruled out previously; `GLDMA2` and `GLDirectMem` +were not, and the per-image `advancedFeatures.ini` turns all three on. Disabling all three +together (r04) is accepted by the emulator and changes nothing: + +``` +INFO | Feature 'GLDMA' (51) is overridden to 'disabled' +INFO | Feature 'GLDMA2' (52) is overridden to 'disabled' +INFO | Feature 'GLDirectMem' (53) is overridden to 'disabled' +... 57 surfaceflinger aborts, device never boots +``` + +**Host composition.** `-feature -HostComposition` (r07) was the best remaining guess at what +forces the readback. It did not help; it made things worse, wedging adb entirely at 208 s so the +crash buffer could not even be read. Recorded as inconclusive rather than ruled out, because no +evidence came back from it. + +**Guest feature negotiation differing from API 36.** It does not. The image-level +`advancedFeatures.ini` for `android-37.0` is byte-identical to `android-36`'s except for one +unrelated line: + +``` +$ diff android-36/google_apis/x86_64/advancedFeatures.ini android-37.0/google_apis/x86_64/advancedFeatures.ini ++QemuCameraSensorOrientation = on +``` + +`GLDMA`, `GLDMA2`, `GLDirectMem`, `GrallocSync`, `HostComposition` and `YUVCache` are on in +both. API 36 boots and passes. So nothing about the host/guest feature handshake changed — the +regression is in the guest's Gralloc5 mapper or in what API 37's `RegionSamplingThread` asks of +it, not in what the emulator advertises. + +**Guest memory.** Ruled out previously and not revisited; every run above used +`hw.ramSize=2560`, the same value the E2E matrix pins, and none of them OOMed. + +**A host-side crash.** Not this bug, and worth stating because the other emulator failure on this +workstation *is* host-side. Every run above left `coredumpctl` empty and produced zero +`avc: denied` lines, and the qemu process was still alive at the end of the ones that never +booted (`emulator_alive=yes`). The host emulator is fine; the guest is not. + +## Can the suite run on it? + +**Almost.** `tools/local-emulator/run-e2e.sh 37` now runs the whole suite locally, and all of it +passes except two tests. Measured at `22c7914`: **49 tests, 2 failures, 0 errors, 2 skipped** — 45 +passed, the two `Media3EngineTest` failures dissected below, and the two `assumeTrue` skips every +level has. It costs two deviations from how every other level is run, and both are worth +understanding before trusting the leg. + +Two things about that total before it is compared with anything. It is the size of the suite on +the checkout that ran, not a property of API 37 — `app/src/androidTest` held 49 `@Test` methods at +`22c7914`, and a newer checkout reports its own count; see +[Reading these totals](#reading-these-totals). And **the Pixel has never run 49**: its green run +was 40 / 0 / 0 / 2 at `edd6385`, the same suite nine tests earlier. What compares across the two +is two failures against none, and the same two skips — not the totals. + +The same numbers and the same two test names came back twice, which is real corroboration — but +by two different routes, and only one of them is the harness. The first was driven by hand +(`pm disable-user`, then several minutes of incidental framework restarts, then `e2e-run.sh` +directly); the second went through `disable_region_sampling`'s `stop; start`. **The harness path +itself has one green measurement.** What would make this routine is a second consecutive +`run-e2e.sh 37` whose only failures are the same two. + +### Booting is not the same as being usable + +Changing the renderer gets the device to `sys.boot_completed=1`, and that is all it gets you. The +aborts do not stop, and each one is a framework restart. A five-minute test run does not survive +that. What it looks like from Gradle: + +``` +Shell command failed (1): rm -rf "/sdcard/Android/media/org.libremediaconverter/..." + rm: ...: Transport endpoint is not connected +Starting 0 tests on lmc_e2e_api37(AVD) - 17 +Shell command failed (20): am get-current-user + cmd: Can't find service: activity +Device emulator-5572 failed to uninstall test APK org.libremediaconverter. + [cmd: Can't find service: package] +Test run failed to complete. No test results. + onError: commandError=false message=INSTRUMENTATION_ABORTED: System has crashed. +``` + +Measured idle rate on `android-37.0` under `swangle_indirect`: **10 aborts in 150 s, then 11 more +in the next 150 s**. Steady, not a start-up transient. + +### The fix is to remove the region-sampling listener, not to survive it + +`RegionSamplingThread` exists only because SystemUI registers a nav-bar luma-sampling listener. +Take SystemUI away and the thread is never started, so the mapper's bad path is never called: + +``` +$ adb shell pm disable-user --user 0 com.android.systemui +Package com.android.systemui new state: disabled-user + +=== aborts at start of measurement: 36 +=== idle 180s with SystemUI disabled === +=== aborts after: 36 NEW IN WINDOW: 0 +--- services still up? --- + activity Service activity: found + package Service package: found + window Service window: found +``` + +**Zero in 180 s, against 10–11 per 150 s.** That is the strongest evidence that region sampling +is the dominant trigger, and it is worth recording even by someone who never wants the workaround. +It does not establish it as the *only* trigger: the paragraph below has an abort surviving the +disable, and nothing measured here says whether that residue is a second caller of the readback +path or a disable that did not fully take. + +Do not read that as "the crashes stop", though, because the harness path does not reproduce a +clean zero. Its own post-disable check on the run recorded below printed + +``` + quiet check: 1 new surfaceflinger aborts in 45 s (want 0) + surfaceflinger hasReadColorBufferDma aborts: 4 (whole run) +``` + +So what is reliably achieved is a **rate collapse** — from roughly one abort every fourteen +seconds to one every forty-five — which a 47-second Gradle run survives and a five-minute one +might not. The 180-second zero above is one measurement on a device that had been up for twelve +minutes and had already cycled its framework several times. The harness prints the quiet-check +delta on every run precisely so this is visible rather than assumed. + +One ordering detail cost a whole run and is now encoded in `disable_region_sampling`: by the time +`sys.boot_completed` flips, SystemUI has **already registered**, and `pm disable-user` does not +retract an existing registration — it only stops the package being started again. Disabling it +and proceeding straight to the tests fails exactly as before. The harness therefore does +`stop; start` afterwards, so the framework that comes back never starts SystemUI at all. + +### The two deviations, stated plainly + +1. **The renderer is ANGLE, not the host GPU.** Shared with nothing else in the matrix — API + 33–36 run `-gpu host` locally, and CI runs `swiftshader_indirect`. +2. **SystemUI is disabled.** The API 37 leg does not run the same device configuration as any + other leg or as the Pixel. It is defensible here only because nothing in this suite touches + system UI — these are Media3, FFmpeg and WorkManager tests — and because the alternative is no + local API 37 coverage at all. **Anything that ever does depend on system UI must not trust + this leg.** + +### The two remaining failures are the same bug, one layer down + +``` +org.libremediaconverter.convert.Media3EngineTest > runsFromAThreadWithNoLooper FAILED +org.libremediaconverter.convert.Media3EngineTest > transcodesH264ToH265AndReportsProgress FAILED + +androidx.media3.transformer.ExportException: Codec exception: + CodecInfo{type=VideoDecoder, ..., mime=video/avc, name=c2.goldfish.h264.decoder} + at androidx.media3.transformer.DefaultCodec.maybeDequeueOutputBuffer(DefaultCodec.java:398) +Caused by: android.media.MediaCodec$CodecException: + at android.media.MediaCodec.native_dequeueOutputBuffer(Native Method) +``` + +Three measurements say this is the emulator image and not this app, and not the software +renderer: + +- **Control at API 35 under the identical renderer.** `GPU_MODE=swangle_indirect + tools/local-emulator/run-e2e.sh 35` → **49 / 0 / 0 / 2** at `22c7914`, green. + `c2.goldfish.h264.decoder` is perfectly happy under ANGLE one API level down, so the renderer is + not what breaks it. +- **Real API 37 hardware passes**, see below. There is no `c2.goldfish.*` codec on a Pixel. +- The failing call is `dequeueOutputBuffer` on the *goldfish* decoder — the emulator's own codec, + which like `RegionSamplingThread` gets its frames out of a host-side colour buffer. Same + readback machinery, one layer down. This is inference rather than a measurement, and is flagged + as such; what is measured is the first two bullets. + +**Do not try `-feature -HardwareDecoder`.** It is the obvious next idea and it is much worse: +forcing the guest onto software decoders took the run from 2 failures to **46**, across +`RemuxTest`, `ForcedFailureTest`, `HardwareFallbackTest` and `UnopenableUriTest` as well. The +suite depends on those decoders existing. + +### So should CI take API 37? + +**No, and the matrix should still stop at 36.** Three reasons, in order of weight: + +1. CI runs `swiftshader_indirect` on a GPU-less runner. The aborts happen under ANGLE too — they + are merely sparser — so nothing here says a runner would be stable. +2. The working configuration needs SystemUI disabled and a framework restart mid-job. That is a + lot of bespoke device surgery to put behind a merge gate, and it silently weakens what the leg + proves. +3. Even at its best two tests fail, so the leg would be permanently red or permanently + allow-listed. Neither is a gate worth having. + +What has changed is the *local* story: API 37 is no longer a level nobody can look at. A +regression that shows up at 37 and not at 36 can now be reproduced on this workstation in about +four minutes, which is what the missing matrix row was really costing. ## Verified on real API 37 hardware -The bug is confined to the emulator image. On 2026-08-21 the whole instrumented suite -was run against a physical device and passed: +Unchanged and still true. On 2026-08-21 the whole instrumented suite ran green on a physical +device: ``` Device: Pixel 10 Pro XL (mustang), arm64-v8a @@ -132,65 +386,94 @@ API: 37 (Android 17, codename REL -- a release build, not a preview) 40 tests, 0 failures, 0 errors, 2 skipped BUILD SUCCESSFUL ``` +**That "40" is a measurement of the tree it ran on, not a baseline for today**, and it is not a +contradiction of the totals in [`docs/local-emulator.md`](local-emulator.md) either. + +### Reading these totals + +Every total in this file and in [`docs/local-emulator.md`](local-emulator.md) is the size of +`app/src/androidTest` on the checkout that produced it, and nothing else. The reported total has +equalled that checkout's `@Test` count everywhere it has been checked: + +| checkout | `@Test` methods | total the run reported | +|---|---|---| +| `edd6385` | 40 | 40 — the Pixel run above | +| `22c7914` | 49 | 49 — the four local levels, and API 37 | +| `18c53a3` | 57 | not run | + +So the number to expect is not written down here. It is derived from the checkout in front of +you, which is the only thing that cannot go stale: + +```bash +grep -rho '@Test' app/src/androidTest | wc -l +``` + +**Before each release, run the suite on the Pixel 10 Pro XL and expect that many tests, 0 +failures, 0 errors, 2 skipped.** The failure, error and skip counts are the invariant; the total +is not. A total that disagrees with your own checkout's count is the signal — an old checkout, a +stale build, or tests that never ran — and it is worth stopping on either way. + The two skips are `RealMediaBenchmark.hardwareVersusSoftwareOnRealVideo` and -`av1InputRoutesAccordingToDeviceDecodeSupport`, which `assumeTrue` their sample files -are present and skip when they are not. That is by design and unrelated to API level. +`av1InputRoutesAccordingToDeviceDecodeSupport`, which `assumeTrue` their sample files are present +and skip when they are not. That is by design and unrelated to API level. One harmless warning appears during the run and can be ignored: -`No UID for androidx.test.services in user 0`, from an `appops` call the test services -package makes before it is fully registered. - -So the app is correct on Android 17. What is missing is only *automated* coverage in -CI. Until the image is fixed, run the suite on a physical API 37 device before release; -that is the substitute for the missing matrix row. +`No UID for androidx.test.services in user 0`, from an `appops` call the test services package +makes before it is fully registered. ## Reproducing it -Locally, with `-gpu host` so the emulator itself does not segfault on Intel graphics: +Both halves, so the renderer claim can be checked rather than taken on trust: ```bash +export ANDROID_HOME="$HOME/Android/Sdk" +export PATH="$ANDROID_HOME/platform-tools:$ANDROID_HOME/emulator:$ANDROID_HOME/cmdline-tools/latest/bin:$PATH" + sdkmanager --install "system-images;android-37.0;google_apis;x86_64" echo no | avdmanager create avd -n api37_repro \ -k "system-images;android-37.0;google_apis;x86_64" -d pixel_6 --force -$ANDROID_HOME/emulator/emulator -avd api37_repro \ - -no-window -gpu host -noaudio -no-boot-anim -camera-back none -no-snapshot & +# never boots -- surfaceflinger aborts every ~6 s, forever +emulator -avd api37_repro -no-window -gpu host \ + -noaudio -no-boot-anim -camera-back none -no-snapshot & -# Boot never completes. Count the aborts: -adb logcat -d -b crash | grep -c hasReadColorBufferDma +# boots in ~85 s, having aborted once or twice on the way +emulator -avd api37_repro -no-window -gpu swangle_indirect \ + -noaudio -no-boot-anim -camera-back none -no-snapshot & ``` -`sys.boot_completed` never reaches `1`, `pgrep -f system_server` stays empty, and -`keystore2`'s watchdog logs `await_boot_completed ... Overdue` indefinitely. +Count the aborts either way: -To see the CI-side form instead, restore the API 37 row in the E2E matrix of -`status_check.yml` (`api-level: "37.0"` — a bare `37` fails earlier still, during SDK -setup, because there is no `platforms;android-37`). +```bash +adb -s emulator-5554 logcat -d -b crash | grep -c hasReadColorBufferDma +``` + +Under `-gpu host`, `sys.boot_completed` never reaches `1`, `pgrep -f system_server` stays empty, +and `keystore2`'s watchdog logs `await_boot_completed ... Overdue` indefinitely. + +`tools/local-emulator/run-e2e.sh 37` does all of this, with the working renderer picked +automatically — see `gpu_for_api` in that file. ## Filing this upstream -Not yet filed. To file it: +Not yet filed. The report is stronger than it was, because the renderer dependency narrows it: -1. Go to and sign in with a Google account. -2. Choose **Report an issue**, then pick the component for the Android emulator — search - the component picker for "Emulator"; it sits under the Android Studio component tree. - If the picker is unclear, Android Studio's **Help → Submit Feedback** opens the same - tracker with the component preselected, and the emulator's own **Extended controls → - Help → File a bug** does likewise. -3. Title it for the mechanism, not the symptom, so it is searchable — for example: - `surfaceflinger aborts in GoldfishMapper::readFromHost (hasReadColorBufferDma) on - android-37.0 google_apis x86_64`. -4. Paste the assertion and backtrace from the top of this document, the environment - table, and the reproduction steps above. State explicitly that it reproduces on two - unrelated hosts under both GPU modes — that is the detail that stops it being closed - as a local graphics problem. -5. List what was ruled out. Bugs that arrive with `-feature -GLDMA` already eliminated - tend not to bounce back asking for it. -6. Attach: - - the guest tombstone, via `adb pull /data/tombstones` (or the `pbtombstone` output - the crash log names) +1. Go to , **Report an issue**, and pick the Android emulator + component (search the component picker for "Emulator"; Android Studio's **Help → Submit + Feedback** opens the same tracker with it preselected). +2. Title it for the mechanism: `surfaceflinger aborts in GoldfishMapper::readFromHost + (hasReadColorBufferDma) on android-37.0 and android-37.1 x86_64 -- fatal under -gpu host, + intermittent under ANGLE`. +3. Paste the assertion and backtrace, the environment block, and the seven-row matrix. The + matrix is the valuable part: it shows the abort is not renderer-specific but its *frequency* + is, which points at the readback path rather than at any one GL implementation. +4. State that it reproduces on two independent system images (`37.0` rev 6 and `37.1` rev 8) and + on two unrelated hosts, and that `-feature -GLDMA,-GLDMA2,-GLDirectMem` does not suppress it. +5. Attach: + - the guest tombstone, via `adb pull /data/tombstones` (or the `pbtombstone` output the crash + log names) - `adb logcat -d -b crash > crash.txt` - - the emulator's own stdout log, captured by redirecting the launch command + - the emulator's own stdout log (`-verbose -debug all`, redirected) - the AVD's `config.ini` - a link to a failing CI job, which shows it on hardware you do not control: @@ -199,16 +482,38 @@ Record the issue number here once filed. ## When to revisit -Re-add the API 37 row when any of these happens: +The old trigger list named "a new `android-37.0` revision", which is why nothing ever fired even +though a new API level shipped. Watch for these instead: -- a new `android-37.0` system image revision ships (this was revision 6) -- an ATD image appears for API 37 -- the upstream issue is marked fixed +- **any new API 37.x system image**, not just a new revision of `37.0` — `37.1` rev 8 and + `37.2-beta*` already exist, and more will. Test with `-gpu host`: if it boots, the guest mapper + is fixed. +- **an ATD image for API 37.** Still none as of 2026-08-22. ATD images ship without SystemUI, + which is what drives `RegionSamplingThread`, so one would very likely sidestep the bug + entirely. Try it before anything else here. +- **the upstream issue being marked fixed.** -Until then the gap is narrower than the missing row suggests. `targetSdk` is 37, so the -app is compiled and unit-tested against it; the API-dependent behaviour this matrix -exists to exercise — the foreground service type, absent below 34, `dataSync` at 34, -`mediaProcessing` from 35 — is covered at 35 and 36; and the full instrumented suite has -been run green on real API 37 hardware. What is missing is *automated* API 37 coverage, -so a regression there would not be caught by a pull request. Run the suite on a physical -API 37 device before each release for as long as this row is absent. +## Correction owed to `CLAUDE.md` + +`CLAUDE.md` currently says: + +> - **The API 37 image is broken.** `android-37.0` crash-loops surfaceflinger inside its own +> gralloc mapper, so every test fails there regardless of this app. +> `docs/api-37-emulator-crash.md` records the evidence and the ruled-out fixes; CI's matrix +> therefore stops at API 36 even though targetSdk is 37. + +The first sentence is right, and now under-specified in one direction and over-specified in the +other: it is not only `android-37.0` (it is `37.1` too), and it does not crash-loop under every +renderer. Proposed replacement, offered for review rather than applied here — `CLAUDE.md` is left +alone deliberately, because several branches touch it: + +> - **The API 37 images crash-loop surfaceflinger under the host GL renderer.** Both +> `android-37.0` and `android-37.1` abort inside their own gralloc mapper +> (`RegionSamplingThread` → `GoldfishMapper::readFromHost`), and when surfaceflinger dies init +> SIGKILLs zygote, so the framework restarts under the test run. Under `-gpu host` it never +> boots at all; under `-gpu swangle_indirect` it boots and the aborts merely become +> intermittent. `docs/api-37-emulator-crash.md` has the seven-run matrix and the ruled-out +> list, and `tools/local-emulator/run-e2e.sh` picks the working renderer per API level. +> CI's matrix stops at 36 because CI runs `swiftshader_indirect` on a GPU-less runner and the +> aborts continue there too. **API 37 still needs a manual check on the Pixel 10 Pro XL before +> each release.** diff --git a/docs/local-emulator.md b/docs/local-emulator.md index 6642f7f..57b2e0e 100644 --- a/docs/local-emulator.md +++ b/docs/local-emulator.md @@ -1,8 +1,10 @@ # Emulators do run on this host: the segfault is SwiftShader's JIT against SELinux -**Status:** solved. Local instrumented runs work with `-gpu host`, and the suite is green -on API 33–36 — 49 tests, 0 failures, 0 errors, 2 skipped on every level. See -[The sweep, run](#the-sweep-run). +**Status:** solved. Local instrumented runs work with `-gpu host`, and the whole suite is green +on API 33–36 — 0 failures, 0 errors and the two by-design skips on every level, measured as +49 / 0 / 0 / 2 at `22c7914`, where the suite was 49 tests. See [The sweep, run](#the-sweep-run), +and [Reading these totals](api-37-emulator-crash.md#reading-these-totals) before comparing any +total with another checkout's. **Last verified:** 2026-08-22, emulator `37.1.11.0` (build 15917651), Fedora 44, kernel `7.1.8-200.fc44`, `selinux-policy-44.6-1.fc44` @@ -190,11 +192,28 @@ emulator -avd -no-window -gpu host \ `.github/scripts/e2e-run.sh` for the run itself. Use it rather than the raw command: ```bash -tools/local-emulator/run-e2e.sh # API 33 34 35 36 +tools/local-emulator/run-e2e.sh # API 33 34 35 36 37 tools/local-emulator/run-e2e.sh 35 # one level +tools/local-emulator/run-e2e.sh 37 37.1 # both API 37 images GPU_MODE=swangle_indirect tools/local-emulator/run-e2e.sh 35 ``` +Levels are the labels above, not SDK ints: API 37's SDK directories are dotted +(`android-37.0`, `android-37.1`) and there is no `android-37`, so `37` is accepted as a +spelling of `37.0`. Setting `GPU_MODE` forces one renderer on every level, which is what +you want when measuring a mode; leaving it unset lets `gpu_for_api` pick, which is what +you want when running the suite — 33–36 need `host` and 37 must not have it. + +**A bare `run-e2e.sh` exits 1, and that is the design.** API 37 is in the default list +deliberately — leaving it out is what left the level unlooked-at for as long as it was — and +it is permanently two failures short of green: `Media3EngineTest` cannot drive the emulator's +`c2.goldfish.h264.decoder` on those images, which +[`api-37-emulator-crash.md`](api-37-emulator-crash.md) pins on the image and not on this app +(API 35 under the same renderer is green). The summary row names the two expected failures so +that a third is visibly new, and the script repeats the point on the way out. Anything that +treats a non-zero exit as breakage — a wrapper, a hook, a habit — should name the levels it +wants: `run-e2e.sh 33 34 35 36` is the sweep that can be green. + `swangle_indirect` is the fallback worth knowing about. It is entirely software, so it does not depend on reaching the session's GPU — useful over plain SSH, where `-gpu host` has not been tested and may not find a device. It is also the closest local analogue to @@ -236,8 +255,9 @@ after an AGP upgrade. ## The sweep, run `tools/local-emulator/run-e2e.sh`, one invocation per level so each got a freshly created -AVD, `-gpu host` throughout, 2026-08-22 19:42–19:56. Every level matches the physical -Pixel 10 Pro XL (API 37) baseline of 49 / 0 / 0 / 2 exactly: +AVD, `-gpu host` throughout, 2026-08-22 19:42–19:56, on `22c7914`. All four levels agree exactly, +and 49 is that checkout's whole suite — every `@Test` in `app/src/androidTest`, two of which skip +by design everywhere: | API | Android | AVD | Boot | `connectedDebugAndroidTest` | Tests | Failures | Errors | Skipped | |---|---|---|---|---|---|---|---|---| @@ -246,6 +266,12 @@ Pixel 10 Pro XL (API 37) baseline of 49 / 0 / 0 / 2 exactly: | 35 | 15 | `lmc_e2e_api35` | 40 s | 3 m 46 s | 49 | 0 | 0 | 2 | | 36 | 16 | `lmc_e2e_api36` | 90 s | 2 m 18 s | 49 | 0 | 0 | 2 | +The physical Pixel has never reported 49, and an earlier version of this paragraph said the +sweep matched it exactly. Its green API 37 run was 40 / 0 / 0 / 2, at `edd6385` — the same suite +nine tests earlier. What matches is 0 failures, 0 errors and the same two skips; totals only ever +match between runs of one checkout, which +[`api-37-emulator-crash.md`](api-37-emulator-crash.md#reading-these-totals) sets out. + Thirteen and a half minutes for the four levels, AVD creation and cold boots included; fifteen with the pre-warm build in front of them. Nothing needed a retry, and no level produced a `diagnostics-api*.txt` — `e2e-run.sh` writes that only on the failure path, so @@ -352,9 +378,11 @@ and the same binaries against the same kernel boot fine under `-gpu host`. before `sys.boot_completed` is ever set. Nothing in the guest — system image variant, RAM, disk size, ATD versus `google_apis` — can influence a host-side `mprotect` denial, so none of those axes was varied. (The API 37 failure in -[`api-37-emulator-crash.md`](api-37-emulator-crash.md) is genuinely guest-side and -genuinely unrelated: there the host emulator survives and the guest's `surfaceflinger` -aborts.) +[`api-37-emulator-crash.md`](api-37-emulator-crash.md) is genuinely guest-side — there the host +emulator survives and the guest's `surfaceflinger` aborts — but it is **not** unrelated, as this +paragraph originally claimed. Both are decided by the renderer, in opposite directions: below 37 +you must avoid SwiftShader GLES and `-gpu host` is the answer; at 37 you must avoid the *host* GL +translator and `-gpu host` is the thing that never boots.) **Turning the SELinux boolean on** — deliberately *not* done, though it would almost certainly work: @@ -401,8 +429,13 @@ are easy to forget to look at. "SwiftShader 4.0.0.1" as reported by the GLES translator. - **If `-gpu host` regresses** after a Mesa or kernel update, fall back to `GPU_MODE=swangle_indirect`, which needs no GPU at all. -- **This changes nothing about API 37.** That image is broken for a different reason and - still must be checked on the physical Pixel 10 Pro XL before each release. +- **API 37 needs the opposite renderer, and this file used to say it needed nothing.** The + original bullet here read "This changes nothing about API 37"; that turned out to be wrong. + The API 37 images abort `surfaceflinger` under the *host* GL translator and boot under ANGLE — + the exact mirror of the rule above — and `run-e2e.sh` therefore picks the renderer per API + level. See [`api-37-emulator-crash.md`](api-37-emulator-crash.md), which was rewritten on + 2026-08-22 with the seven-run matrix. API 37 still must be checked on the physical Pixel 10 Pro + XL before each release. ## Correction owed to `CLAUDE.md` @@ -432,12 +465,14 @@ Proposed replacement for the section, offered for review rather than applied her > `auto` (the default), `off` and `guest` do when headless. `-gpu host` works, and the > harness both picks it and refuses the others. `docs/local-emulator.md` has the > backtrace and the mode matrix. -> - **The API 37 image is broken.** `android-37.0` crash-loops surfaceflinger inside its -> own gralloc mapper, so every test fails there regardless of this app — -> `docs/api-37-emulator-crash.md` records the evidence and the ruled-out fixes. This is -> unrelated to the renderer above: it is a guest-side bug that CI hits too, which is why -> the matrix stops at API 36 even though targetSdk is 37. **API 37 needs a manual check -> on the Pixel 10 Pro XL before each release.** +> - **API 37 needs the opposite renderer, and SystemUI turned off.** Both `android-37.0` and +> `android-37.1` abort surfaceflinger inside their own gralloc mapper, and init SIGKILLs +> zygote each time. Under `-gpu host` they never boot; under `-gpu swangle_indirect` they +> boot, and disabling SystemUI removes the trigger. `run-e2e.sh` does all of that per level, +> and the local API 37 result is two failures and the two usual skips, not a clean run. CI's +> matrix still stops at 36. +> `docs/api-37-emulator-crash.md` has the matrix and the reasoning. **API 37 needs a manual +> check on the Pixel 10 Pro XL before each release.** The wording is worth getting right rather than merely correcting, because the original was not a careless sentence — it was a reasonable inference from three crashes, written down diff --git a/tools/local-emulator/run-e2e.sh b/tools/local-emulator/run-e2e.sh index 7ddabe4..811e809 100755 --- a/tools/local-emulator/run-e2e.sh +++ b/tools/local-emulator/run-e2e.sh @@ -3,13 +3,24 @@ # Runs the instrumented suite on a local emulator, on this workstation, for one or more # API levels. # -# Usage: tools/local-emulator/run-e2e.sh [API ...] # default: 33 34 35 36 +# Usage: tools/local-emulator/run-e2e.sh [API ...] # default: 33 34 35 36 37 # -# GPU_MODE=host renderer to use; see the refusal list below +# API levels are the labels below, not SDK ints: 33-36, plus `37` (= `37.0`) and `37.1`. +# +# GPU_MODE= force one renderer on every level; unset means per-API (gpu_for_api) # EMULATOR_PORT=5560 console port, so the serial is deterministic # BOOT_TIMEOUT=300 seconds to wait for sys.boot_completed # KEEP_AVD=1 do not delete an AVD this script created # +# EXIT CODE: 0 only if every level was green; 1 if any level failed, wedged or could not be +# set up; 2 if it refused to start at all. **A bare `run-e2e.sh` therefore exits 1 by design.** +# API 37 is in the default list on purpose -- leaving it out is what left the level unlooked-at +# for as long as it was -- and it is permanently two failures short of green, on the emulator's +# own c2.goldfish.h264.decoder rather than on anything this app does. The summary names the two, +# so a third is visibly new, and the last line printed says the same thing. Anything that reads a +# non-zero exit as breakage should name the levels it wants: `run-e2e.sh 33 34 35 36` is the +# sweep that can be green. docs/api-37-emulator-crash.md has the measurements. +# # WHY THIS EXISTS, AND WHAT IT DELIBERATELY DOES NOT DO # # It is a *launcher*, not a second test harness. The diagnostics -- the FAILED-vs-WEDGED @@ -34,6 +45,14 @@ # those modes was measured crashing. docs/local-emulator.md has the backtrace, the faulting # page's RW-without-E segment flags, and the full mode matrix. # +# AND THE ONE THING API 37 NEEDS THAT 33-36 DO NOT: the opposite renderer. On the API 37 +# images the guest's Gralloc5 mapper aborts surfaceflinger from RegionSamplingThread +# (`Assertion failed: !rcEnc->featureInfo()->hasReadColorBufferDma`). Under `-gpu host` that +# repeats every few seconds and the device never boots; under ANGLE it fires a handful of +# times and the boot survives. So `host` is required below 37 and forbidden at 37, which is +# why the renderer is chosen per level in gpu_for_api rather than set once. +# docs/api-37-emulator-crash.md has that matrix. +# # THE OTHER LOCAL-ONLY HAZARD: a physical Pixel is usually plugged into this machine, so # `adb` is ambiguous in a way it never is on a runner, and an unpinned run would install # and execute this suite on the phone. Every path below pins the emulator serial. @@ -65,12 +84,14 @@ export ANDROID_HOME="${ANDROID_HOME:-$HOME/Android/Sdk}" export ANDROID_SDK_ROOT="$ANDROID_HOME" export PATH="$ANDROID_HOME/platform-tools:$ANDROID_HOME/emulator:$ANDROID_HOME/cmdline-tools/latest/bin:$PATH" -GPU_MODE="${GPU_MODE:-host}" +# Empty means "let each level pick" -- see gpu_for_api. Setting GPU_MODE forces one renderer +# on every level, which is what you want when measuring a mode, not when running the suite. +GPU_MODE="${GPU_MODE:-}" EMULATOR_PORT="${EMULATOR_PORT:-5560}" BOOT_TIMEOUT="${BOOT_TIMEOUT:-300}" SERIAL="emulator-${EMULATOR_PORT}" APIS=("$@") -[ "${#APIS[@]}" -eq 0 ] && APIS=(33 34 35 36) +[ "${#APIS[@]}" -eq 0 ] && APIS=(33 34 35 36 37) # Matches CI. `disk-size: 8G` because the FFmpeg libraries do not fit the default userdata # partition; `ram-size: 2560M` because the emulator's own floor varies by API level and @@ -83,21 +104,28 @@ RESULTS_DIR="app/build/outputs/androidTest-results" LOG_DIR="${TMPDIR:-/tmp}/lmc-local-e2e" mkdir -p "$LOG_DIR" +# The two things that outlive a level, declared here rather than where they are first +# assigned, because the cleanup trap below can fire before either has been reached. +EMU_PID="" +CREATED_AVDS=() + # ---------------------------------------------------------------- renderer preflight --- -case "$GPU_MODE" in - swiftshader_indirect | auto | off | guest) - echo "REFUSING to launch with -gpu $GPU_MODE." - echo "On this host that resolves to SwiftShader's GLES, whose JIT is denied execheap by" - echo "SELinux; the emulator segfaults (exit 139) before boot. See docs/local-emulator.md." - echo "Working modes: host (default), angle_indirect, swangle_indirect." - exit 2 - ;; - host | angle_indirect | swangle_indirect) ;; - *) - echo "Unrecognised GPU_MODE '$GPU_MODE'. Known-good: host, angle_indirect, swangle_indirect." - exit 2 - ;; -esac +if [ -n "$GPU_MODE" ]; then + case "$GPU_MODE" in + swiftshader_indirect | auto | off | guest) + echo "REFUSING to launch with -gpu $GPU_MODE." + echo "On this host that resolves to SwiftShader's GLES, whose JIT is denied execheap by" + echo "SELinux; the emulator segfaults (exit 139) before boot. See docs/local-emulator.md." + echo "Working modes: host, angle_indirect, swangle_indirect." + exit 2 + ;; + host | angle_indirect | swangle_indirect) ;; + *) + echo "Unrecognised GPU_MODE '$GPU_MODE'. Known-good: host, angle_indirect, swangle_indirect." + exit 2 + ;; + esac +fi # A courtesy, not a gate: the boolean being on means SwiftShader would work too, and the # refusal list above could be relaxed. It is off on a stock Fedora. @@ -121,14 +149,90 @@ host_forensics() { journalctl --since "$since" --no-pager 2> /dev/null | grep -E 'avc: .*denied' | tail -10 || echo " (none)" } +# The API 37 counterpart of host_forensics. `-gpu host` there aborts surfaceflinger in a loop +# and the device never boots; the working renderers abort it a few times and survive. Either way +# the count is the number to look at, and the crash buffer is where it lives -- so print it on +# every 37 level, not only on the failure path, because a level that passed with 40 aborts is +# telling you something a level that passed with 1 is not. +guest_forensics() { + local api="$1" n + case "$api" in 37 | 37.*) ;; *) return 0 ;; esac + n="$(emu_adb logcat -d -b crash 2> /dev/null | grep -c 'hasReadColorBufferDma')" + echo " surfaceflinger hasReadColorBufferDma aborts: ${n:-?} (docs/api-37-emulator-crash.md)" +} + +# API label -> system image. API 33-36 are plain integers with a `google_apis` image. API 37 +# is not: its SDK directories are dotted minor versions (`android-37.0`, `android-37.1`), there +# is no `android-37`, and from 37.1 onwards Google ships only 16 KB-page (`ps16k`) images for +# x86_64. `37` is accepted as a spelling of `37.0` because that is what people type. +image_pkg_for_api() { + case "$1" in + 37 | 37.0) echo "system-images;android-37.0;google_apis;x86_64" ;; + 37.1) echo "system-images;android-37.1;google_apis_ps16k;x86_64" ;; + *) echo "system-images;android-$1;google_apis;x86_64" ;; + esac +} + +# The renderer requirement is per-API and the two levels want OPPOSITE things, which is why this +# is a function and not a constant. +# +# 33-36: must NOT be SwiftShader GLES (host-side SELinux/execheap segfault) -- `host` is right. +# 37.x: must NOT be the host GL translator. With `-gpu host` the guest's Gralloc5 mapper +# aborts surfaceflinger in a loop and the device never boots; under ANGLE the same +# assertion fires a handful of times and the boot survives it. Measured, not guessed -- +# docs/api-37-emulator-crash.md has the matrix. +# +# `swangle_indirect` rather than `angle_indirect` for 37: both boot, and swangle names its +# renderer outright instead of resolving through `auto`'s path. +gpu_for_api() { + if [ -n "$GPU_MODE" ]; then + echo "$GPU_MODE" + return + fi + case "$1" in + 37 | 37.*) echo "swangle_indirect" ;; + *) echo "host" ;; + esac +} + +# `lmc_e2e_api37.0` would be a legal AVD name but an awkward one to type and to grep for. +# `37` and `37.0` therefore give two AVD names (`lmc_e2e_api37`, `lmc_e2e_api37_0`) for the one +# image. Harmless -- two AVDs off the same system image cost only disk -- and deliberately not +# normalised, so that `run-e2e.sh 37 37.0` does not have both levels fight over one AVD. +avd_for_api() { echo "lmc_e2e_api${1//./_}"; } + +# Where avdmanager actually put the AVD. `$HOME/.android/avd` is only the default: +# ANDROID_AVD_HOME, ANDROID_USER_HOME and ANDROID_SDK_HOME each move it, and hardcoding the +# default meant a machine that sets any of them silently ran every level at stock RAM and +# userdata size. Rather than encode a precedence that cannot be verified from here, look in +# every location avdmanager honours and let the existence check pick. +avd_config_path() { + local avd="$1" base cfg + for base in "${ANDROID_AVD_HOME:-}" \ + "${ANDROID_USER_HOME:+$ANDROID_USER_HOME/avd}" \ + "${ANDROID_SDK_HOME:+$ANDROID_SDK_HOME/.android/avd}" \ + "$HOME/.android/avd"; do + [ -n "$base" ] || continue + cfg="$base/${avd}.avd/config.ini" + if [ -f "$cfg" ]; then + echo "$cfg" + return 0 + fi + done + return 1 +} + ensure_avd() { local api="$1" avd="$2" - local pkg="system-images;android-${api};google_apis;x86_64" + local pkg + pkg="$(image_pkg_for_api "$api")" + local img_dir="$ANDROID_HOME/system-images/${pkg#system-images;}" + img_dir="${img_dir//;//}" if avdmanager list avd -c 2> /dev/null | grep -qx "$avd"; then echo " reusing existing AVD $avd" else - if [ ! -d "$ANDROID_HOME/system-images/android-${api}/google_apis/x86_64" ]; then + if [ ! -d "$img_dir" ]; then echo " installing $pkg" yes | sdkmanager --install "$pkg" > /dev/null 2>&1 || { echo " FAILED to install $pkg" @@ -144,18 +248,37 @@ ensure_avd() { fi # Written into config.ini rather than passed on the command line, which is how - # reactivecircus/android-emulator-runner applies the same two settings in CI. - local cfg="$HOME/.android/avd/${avd}.avd/config.ini" - sed -i -e '/^disk\.dataPartition\.size=/d' -e '/^hw\.ramSize=/d' "$cfg" - printf 'disk.dataPartition.size=%s\nhw.ramSize=%s\n' "$DISK_SIZE_BYTES" "$RAM_SIZE_MB" >> "$cfg" + # reactivecircus/android-emulator-runner applies the same two settings in CI. On the reuse + # path too, so an AVD left over from an older run gets today's pins. + # + # A level that cannot be pinned FAILS rather than running at the defaults. Unpinned, it + # dies much later with "not enough space", which reads as a device problem -- CI's own + # history is where that lesson comes from -- and nothing points back to a `sed` that + # edited a path this script guessed wrong. + local cfg + if ! cfg="$(avd_config_path "$avd")"; then + echo " FAILED: no config.ini for $avd in any directory avdmanager uses" + echo " (ANDROID_AVD_HOME=${ANDROID_AVD_HOME:-unset}, ANDROID_USER_HOME=${ANDROID_USER_HOME:-unset}," + echo " ANDROID_SDK_HOME=${ANDROID_SDK_HOME:-unset}, HOME=$HOME)" + return 1 + fi + if ! sed -i -e '/^disk\.dataPartition\.size=/d' -e '/^hw\.ramSize=/d' "$cfg"; then + echo " FAILED to rewrite $cfg" + return 1 + fi + if ! printf 'disk.dataPartition.size=%s\nhw.ramSize=%s\n' \ + "$DISK_SIZE_BYTES" "$RAM_SIZE_MB" >> "$cfg"; then + echo " FAILED to write the RAM/disk pins into $cfg" + return 1 + fi } boot_emulator() { - local avd="$1" api="$2" + local avd="$1" api="$2" gpu="$3" local boot_log="$LOG_DIR/emulator-api${api}.log" emulator -avd "$avd" -port "$EMULATOR_PORT" \ - -no-window -gpu "$GPU_MODE" -noaudio -no-boot-anim -camera-back none -no-snapshot \ + -no-window -gpu "$gpu" -noaudio -no-boot-anim -camera-back none -no-snapshot \ > "$boot_log" 2>&1 & EMU_PID=$! @@ -181,6 +304,92 @@ boot_emulator() { return 1 } +# API 37 only, and the reason API 37 can be run at all. +# +# The abort that breaks these images is reached from SurfaceFlinger's RegionSamplingThread, +# which exists only because SystemUI registers a nav-bar luma-sampling listener. Each abort +# kills surfaceflinger, and init responds by SIGKILLing zygote -- so the whole framework +# restarts underneath the test run, which arrives as `Can't find service: package` and +# `INSTRUMENTATION_ABORTED: System has crashed`. Under the host GL renderer that repeats +# forever; under ANGLE it is roughly one every fifteen seconds, which a five-minute suite does +# not survive either. +# +# Removing the listener removes the whole chain. Measured on android-37.0 under +# swangle_indirect: 10-11 aborts per 150 s idle with SystemUI running, and 0 in 180 s with it +# disabled, framework services up throughout. +# +# THIS IS A DEVIATION, and it is deliberately loud rather than silent. The API 37 leg does not +# run the same device configuration as API 33-36 or as the Pixel. It is defensible only +# because nothing in this suite touches SystemUI -- these are Media3, FFmpeg and WorkManager +# tests -- and because the alternative is no API 37 coverage at all. Anything that ever does +# depend on system UI must not trust this leg. docs/api-37-emulator-crash.md explains why. +# +# The retry loop is not defensive padding: at the moment boot_completed flips, the framework +# may be in one of its restarts and `pm` is simply not published yet. The first attempt at this +# failed exactly that way, with `cmd: Can't find service: package`. +# +# The framework restart at the end is not optional, and finding that out cost a run. By the +# time `sys.boot_completed` flips, SystemUI has already registered its region-sampling listener, +# and `pm disable-user` does not retract a registration that already happened -- it only stops +# the package being started again. So the first attempt disabled SystemUI, reported success, and +# then died exactly as before with `Starting 0 tests` and four more aborts. `stop; start` cycles +# zygote deliberately, and the framework that comes back up does not start SystemUI at all. +disable_region_sampling() { + local api="$1" out i before after ready + case "$api" in 37 | 37.*) ;; *) return 0 ;; esac + + out="" + for i in $(seq 1 20); do + out="$(emu_adb shell pm disable-user --user 0 com.android.systemui 2>&1 | tr -d '\r')" + case "$out" in + *"new state: disabled"*) + echo " SystemUI disabled on attempt $i" + break + ;; + esac + out="" + sleep 5 + done + if [ -z "$out" ]; then + echo " WARNING: could not disable SystemUI after 20 attempts." + echo " Expect INSTRUMENTATION_ABORTED -- docs/api-37-emulator-crash.md" + return 0 + fi + + echo " restarting the framework so the region-sampling listener goes with it" + emu_adb shell stop > /dev/null 2>&1 + emu_adb shell start > /dev/null 2>&1 + # There is no property worth waiting on here, and an earlier version of this only looked + # like it was waiting on one: `stop` does not clear sys.boot_completed, so it still reads + # `1` throughout the restart and any loop over it returns at once. The loop below is the + # wait -- and it polls the better thing anyway, since `Can't find service: package` is the + # failure it exists to prevent. + ready=0 + for i in $(seq 1 30); do + if emu_adb shell service check package 2> /dev/null | grep -q ': found' \ + && emu_adb shell service check activity 2> /dev/null | grep -q ': found'; then + ready=1 + break + fi + sleep 5 + done + if [ "$ready" -ne 1 ]; then + echo " WARNING: package and activity services still absent 150 s after the restart." + echo " Expect INSTRUMENTATION_ABORTED -- docs/api-37-emulator-crash.md" + fi + + # Prove it worked rather than assume it. Zero new aborts over this window is what makes the + # difference between a run that completes and one that reports `Starting 0 tests`. + before="$(emu_adb logcat -d -b crash 2> /dev/null | grep -c 'hasReadColorBufferDma')" + emu_adb shell 'sleep 45' > /dev/null 2>&1 + after="$(emu_adb logcat -d -b crash 2> /dev/null | grep -c 'hasReadColorBufferDma')" + echo " quiet check: $((after - before)) new surfaceflinger aborts in 45 s (want 0)" + if [ "$((after - before))" -ne 0 ]; then + echo " WARNING: region sampling is still live; the run may not survive." + fi + return 0 +} + # CI gets this from the action's `disable-animations: true`. disable_animations() { local s @@ -189,17 +398,77 @@ disable_animations() { done } +# `${EMU_PID:-0}` used to guard these three calls, and it guarded the wrong thing: EMU_PID +# is *empty*, not unset, if the background launch never produced a job, and `kill` reads pid +# 0 as "the sender's whole process group" -- this script and, on a terminal, everything else +# in the foreground group with it. The `kill -0` wait loop had the same shape and would have +# spent its full grace period testing the group. Nothing to stop is now a return, never a +# guess. (boot_emulator's own `kill -0 "$EMU_PID"` is unguarded and cannot reach that form: +# it runs only after the assignment.) +# +# max_wait is a parameter so the interrupt path need not sit through the full grace period. stop_emulator() { + local max_wait="${1:-30}" waited=0 + [ -n "${EMU_PID:-}" ] || return 0 emu_adb emu kill > /dev/null 2>&1 - local waited=0 - while kill -0 "${EMU_PID:-0}" 2> /dev/null && [ "$waited" -lt 30 ]; do + while kill -0 "$EMU_PID" 2> /dev/null && [ "$waited" -lt "$max_wait" ]; do sleep 2 waited=$((waited + 2)) done - kill -9 "${EMU_PID:-0}" 2> /dev/null - wait "${EMU_PID:-0}" 2> /dev/null + kill -9 "$EMU_PID" 2> /dev/null + wait "$EMU_PID" 2> /dev/null + EMU_PID="" } +delete_created_avds() { + local avd + [ "${KEEP_AVD:-0}" = "1" ] && return 0 + for avd in ${CREATED_AVDS[@]+"${CREATED_AVDS[@]}"}; do + # A SIGKILLed emulator does not get to remove its own lock files, and avdmanager can + # refuse over them. Staying silent there would leak the very thing this exists to clean. + avdmanager delete avd -n "$avd" > /dev/null 2>&1 \ + || echo " WARNING: could not delete AVD $avd -- 'avdmanager delete avd -n $avd' by hand" + done + CREATED_AVDS=() +} + +# What an interrupted sweep used to leave behind: a headless emulator holding console port +# $EMULATOR_PORT, and an lmc_e2e_apiNN AVD. The next run's `emulator -port` then collides +# with the orphan, and `emu_adb` can resolve to it -- on a workstation that also has the +# Pixel plugged in, exactly the ambiguity the ANDROID_SERIAL pinning exists to prevent. A +# sweep is up to five boots long, so the window for one Ctrl-C is not small. +# +# Idempotent, and called explicitly on the normal path so its output cannot land after the +# summary; the EXIT trap then finds nothing left to do. The emulator logs are deliberately +# NOT removed -- they live in $LOG_DIR and are the only evidence a failed boot leaves. +CLEANED=0 +cleanup() { + [ "$CLEANED" = "1" ] && return 0 + CLEANED=1 + stop_emulator "${1:-30}" + delete_created_avds +} + +# 6 s, not 30: Ctrl-C has already reached the emulator through the foreground process group, +# so this is only waiting for it to finish writing, and `kill -9` follows regardless. The +# EXIT trap is disarmed before exiting so the status below is the one that survives. +# +# bash runs a trap only between commands, so this starts when whatever was in the foreground +# returns -- which for Ctrl-C is immediately, because the same interrupt reached that command +# too. `kill -INT` aimed at this script alone waits for the foreground command to finish. +on_signal() { + echo + echo "interrupted (SIG$1) -- stopping the emulator and removing the AVDs this run created" + echo " emulator logs kept in $LOG_DIR" + cleanup 6 + trap - EXIT + exit "$2" +} + +trap 'on_signal INT 130' INT +trap 'on_signal TERM 143' TERM +trap cleanup EXIT + # The XML is authoritative. The console counter double-counts skips, so a run that reports # "42 tests" on stdout can be 40 in the report. # @@ -236,37 +505,44 @@ PY } # ------------------------------------------------------------------------------ main --- -CREATED_AVDS=() SUMMARY=() overall=0 -for api in "${APIS[@]}"; do - if [ "$api" = "37" ] || [ "$api" = "37.0" ]; then - echo "SKIPPING API $api: the android-37.0 image crash-loops surfaceflinger." - echo " See docs/api-37-emulator-crash.md. Test API 37 on the physical Pixel." - continue - fi +# Whether the red exit is the expected one depends on which level produced it, and only the +# loop knows that -- so it is recorded where `overall` is set rather than guessed from the +# summary afterwards. A note at the end claiming a genuine API 34 failure was "by design" +# would be the same defect it is there to prevent, one layer up. +NON37_RED=0 +mark_red() { + overall=1 + case "$1" in 37 | 37.*) ;; *) NON37_RED=1 ;; esac +} - avd="lmc_e2e_api${api}" +for api in "${APIS[@]}"; do + avd="$(avd_for_api "$api")" + gpu="$(gpu_for_api "$api")" started="$(date '+%Y-%m-%d %H:%M:%S')" echo "==============================================================" - echo "API $api (avd=$avd gpu=$GPU_MODE serial=$SERIAL)" + echo "API $api (avd=$avd gpu=$gpu serial=$SERIAL)" + echo " image: $(image_pkg_for_api "$api")" echo "==============================================================" if ! ensure_avd "$api" "$avd"; then SUMMARY+=("API $api: AVD SETUP FAILED") - overall=1 + mark_red "$api" continue fi - if ! boot_emulator "$avd" "$api"; then + if ! boot_emulator "$avd" "$api" "$gpu"; then host_forensics "$started" + guest_forensics "$api" SUMMARY+=("API $api: BOOT FAILED") - overall=1 + mark_red "$api" stop_emulator continue fi + disable_region_sampling "$api" disable_animations rm -rf "$RESULTS_DIR" @@ -281,23 +557,43 @@ for api in "${APIS[@]}"; do unset ANDROID_SERIAL E2E_EXTRA_GRADLE_ARGS line="$(summarise_results "$api")" + guest_forensics "$api" + # API 37 is in the default list on purpose, and it is expected to be red. Leaving it out would + # put the level back where this whole exercise found it -- untested and unlooked-at -- but a + # summary that just says "2 failures" with no explanation trains people to ignore the exit + # code. So the row says which two, and a THIRD failure is then obviously new. + case "$api" in + 37 | 37.*) + line="$line + expected here: 2 failures, both Media3EngineTest, on c2.goldfish.h264.decoder. + A third is new -- docs/api-37-emulator-crash.md" + ;; + esac if [ "$rc" -ne 0 ]; then line="$line [gradle exit $rc]" - overall=1 + mark_red "$api" host_forensics "$started" fi SUMMARY+=("$line") stop_emulator done -if [ "${KEEP_AVD:-0}" != "1" ]; then - for avd in ${CREATED_AVDS[@]+"${CREATED_AVDS[@]}"}; do - avdmanager delete avd -n "$avd" > /dev/null 2>&1 - done -fi +cleanup echo echo "===================== LOCAL E2E SUMMARY ======================" printf '%s\n' ${SUMMARY[@]+"${SUMMARY[@]}"} echo "==============================================================" + +# An unexplained red exit trains people to stop reading exit codes, and this one is expected +# whenever API 37 is in the sweep -- which the default list makes the common case. Said here +# rather than only in the docs, because this is where it is actually read. Only when 37.x is +# the ONLY thing that went red: a note calling a real failure elsewhere "by design" would be +# worse than no note at all. +if [ "$overall" -ne 0 ] && [ "$NON37_RED" -eq 0 ]; then + echo "note: the only level that went red is API 37, which exits non-zero by design -- it is" + echo " permanently 2 failures short of green. Confirm its row above shows exactly those" + echo " two and nothing else; docs/api-37-emulator-crash.md says why they are the image." +fi + exit "$overall"