Compare commits

..
Author SHA1 Message Date
JMR-dev f8e6bfa2a3 Merge remote-tracking branch 'origin/main' into tools/api-37-emulator 2026-08-22 23:52:58 -05:00
Jason Ross 5a1b8832d3 Merge pull request #48 from JMR-dev/fix/review-app-gaps
Close eleven review findings in app code
2026-08-22 23:52:18 -05:00
JMR-dev 7ae660ee37 Merge remote-tracking branch 'origin/main' into tools/api-37-emulator 2026-08-22 23:21:47 -05:00
JMR-devandClaude Opus 5 775a44753b Say that the default sweep is red on purpose, and narrow two claims
R16 / #25 -- the branch put API 37 into the default APIS list, where it is permanently two
failures short of green, so a bare `run-e2e.sh` exits 1 by design and nothing said so.
Somebody running it from habit, a wrapper or a hook gets a red exit forever and either
stops reading exit codes or debugs a normal state.

Documented rather than suppressed. The script's own comment already argued that an
expected-red level belongs in the exit code -- reversing that is the branch owner's call,
not a correction -- and the review's alternative needs an exact-set comparison of the
failing test names before it can subtract 37's contribution, which is a new mechanism that
cannot be validated without a device. So:

- the header now states the exit code (0 all green / 1 any level red / 2 refused to
  start), says a bare run is 1 by design and why, and gives `run-e2e.sh 33 34 35 36` as
  the sweep that can be green;
- a red sweep prints one note after the summary saying the same thing, because the exit
  code is read in the terminal and not in the docs -- but ONLY when 37.x is the only level
  that went red. `overall` is set by any red level, so a note keyed on "37 was in the
  list" would have called a genuine API 34 failure "by design", which is the defect this
  is meant to prevent, one layer up. mark_red records which level it was, where the loop
  already knows;
- docs/local-emulator.md says it where the default is documented.

R27 / #36 -- 792286a appended the caveat that the harness path reproduces a rate collapse
rather than a clean zero, but left "that is the confirmation that region sampling is the
sole trigger" standing three lines above it, which the caveat contradicts. Now "the
strongest evidence that region sampling is the dominant trigger", with the residue named:
no measurement here separates a second caller of the readback path from a disable that did
not fully take, and the file says so rather than picking one.

R28 / #37 -- "the capability is negotiated regardless of renderer" leaned on the string
search, which shows only that `ANDROID_EMU_read_color_buffer_dma` is implemented in one
shared component, not that it is negotiated on every path. The aborts are the actual
evidence -- the assertion that fires is `!hasReadColorBufferDma` and it fires under ANGLE
too -- and they suffice alone; the string search is demoted to a supporting note. Worth
getting right because the doc says the upstream report should lead with this model.

Two follow-ons that belong with R17 / #26 and land here rather than in their own commit:
bash runs a trap only between commands, so the handler starts when the foreground command
returns -- immediate under Ctrl-C, which reaches that command too, but not under a `kill
-INT` aimed at the script alone; that is now written next to the handler. And
delete_created_avds no longer discards avdmanager's status: an emulator that was SIGKILLed
did not get to remove its own lock files, avdmanager can refuse over them, and silence
there would leak exactly what the trap exists to clean up.

`bash -n` clean; the stub smoke harness (real script, fake SDK binaries, boot-failure path,
no Gradle and no emulator) now also checks that a 37-only red prints the note after the
summary, that a red API 34 alongside it suppresses the note, that a 34-only sweep says
nothing, and that a refused AVD deletion is reported. shellcheck is not installed here.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-22 23:17:19 -05:00
JMR-devandClaude Opus 5 961cfa72a2 Derive the suite size instead of writing it down in two documents
R4 / #13 and R20 / #29 are one defect: an absolute test total in an unregenerated
document, written the same day it went stale. This branch was cut at 22c7914, where
app/src/androidTest held 49 @Test methods; main is 57 (ReattachOnLaunchTest added eight
in ec969c4). So the release instruction "expect 49 / 0 / 0 / 2, and if you get 40 you are
on an old checkout" becomes false the moment this branch merges -- on the one check that
has no CI backstop -- and docs/local-emulator.md's headline promises a 49-test local
baseline main no longer produces.

Re-derived rather than renumbered, because a third total would go stale the same way:

- The total is the size of app/src/androidTest on the checkout that ran, and the reported
  total has equalled that checkout's @Test count everywhere it has been checked: 40 at
  edd6385 (the Pixel run), 49 at 22c7914 (the four local levels and API 37), 57 at
  18c53a3 (counted, not run). The new "Reading these totals" section states that, gives
  the one-line grep, and makes the *mismatch* the signal: a total that disagrees with
  your own checkout's count means an old checkout, a stale build or tests that never ran.
  The pre-release Pixel instruction now reads "that many tests, 0 failures, 0 errors, 2
  skipped" -- the invariant, not the total.
- Measurements are kept verbatim and anchored to 22c7914 (the sweep table, the API 35
  control, the tests="49" XML quote, the 51-on-screen console block). Only the claims
  built on top of them were rewritten.

Two claims went with the number, both of which a rebase would have preserved:

- "47 of 49" is not a defensible ratio when two of the 49 are skips. 49 = 45 passed + 2
  failed + 2 skipped, and that is what it now says.
- "against the Pixel's 49 of 49" and "matches the physical Pixel 10 Pro XL baseline of
  49 / 0 / 0 / 2 exactly" describe a run that never happened: the Pixel measured
  40 / 0 / 0 / 2 at edd6385, nine tests earlier, as the same file says a hundred lines
  further down. Both documents projected the local total onto the phone and called it a
  match. What compares between them is 0 failures and the same two skips.

Also re-derived in the CLAUDE.md wording docs/local-emulator.md proposes, since that text
is meant to be pasted out of the branch and would have carried "49 tests / 2 failures /
2 skipped" with it. CLAUDE.md itself is still untouched.

Counts re-checked with git grep at each of the three commits; nothing here needed a
device, and none was used.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-22 23:09:30 -05:00
JMR-devandClaude Opus 5 da6f2807e9 Stop an interrupted sweep leaking the emulator, the AVD and the port
Four corrections to the harness, none of which changes what a successful sweep does.

R17 / #26 -- no trap. Ctrl-C during a sweep (now up to five boots long) left headless
qemu on console port 5560 and an lmc_e2e_apiNN AVD behind. The next run's `emulator
-port` then collides with the orphan and `emu_adb` can resolve to it -- on a workstation
with the Pixel plugged in, exactly the ambiguity the ANDROID_SERIAL pinning exists to
prevent. `cleanup` (stop_emulator + delete_created_avds, KEEP_AVD honoured) is now on
EXIT, INT and TERM. It is idempotent and the normal path calls it explicitly before the
summary, so cleanup output cannot land after the summary and the EXIT trap finds nothing
to redo. The interrupt path passes a 6-second grace rather than 30: Ctrl-C has already
reached the emulator through the foreground process group, so that wait is only for it to
finish writing, and `kill -9` follows regardless. `exit "$overall"` stays the last line,
so the exit code an EXIT trap could have swallowed is still the one that escapes. The
emulator logs in $LOG_DIR are deliberately kept -- they are the only evidence a failed
boot leaves.

R31 / #40 -- `kill -9 "${EMU_PID:-0}"`. EMU_PID is empty, not unset, if the background
launch never produced a job, so `:-0` converted "nothing to kill" into pid 0, which POSIX
reads as the sender's whole process group. The `kill -0` wait loop had the same shape and
would have spent its full grace period testing the group. All three sites now take a bare
`$EMU_PID` behind one `[ -n ... ] || return 0` guard. boot_emulator's own `kill -0` is
left alone: it runs only after the assignment and cannot reach the group form.

R33 / #42 -- ensure_avd wrote CI's RAM and disk pins to a hardcoded
$HOME/.android/avd/... path and checked nothing. With ANDROID_AVD_HOME (or
ANDROID_USER_HOME, or ANDROID_SDK_HOME) set, the sed failed and the level ran on at
default RAM and userdata, which surfaces much later as "not enough space" and reads as a
device problem. `avd_config_path` now looks in every directory avdmanager honours -- no
precedence is asserted, the existence check decides -- and a level that cannot be found
or written fails instead of running unpinned.

R34 / #43 -- disable_region_sampling's "one blocking wait on the device" did not wait:
`adb shell stop` does not clear sys.boot_completed, so the property still read 1 and the
loop returned at once. Deleted, and the comment now names the service-check loop below it
as the actual wait -- which polls the better thing anyway, since `Can't find service:
package` is the failure it exists to prevent. That loop also says so when it gives up
after 150 s instead of proceeding silently. Deliberately not doing the `setprop
sys.boot_completed 0` variant: the loop tested for an empty value, so a 0 would not have
made it wait either, and the `!= 1` form it would need is an unbounded loop inside `adb
shell` with no timeout.

Checked with `bash -n` and with two stub harnesses in place of a device (shellcheck is
not installed here): one drives the extracted lifecycle functions against fake binaries
and asserts pid 0 really does hit the sender's process group, that an empty EMU_PID now
signals nothing and returns at once, that SIGINT cleans up once and exits 130 within
seconds, that KEEP_AVD survives the trap path, and that an explicit exit status survives
the EXIT trap; the other runs the real script end to end on the boot-failure path, which
stops short of e2e-run.sh, and checks the pins land in config.ini, the created AVD is
removed, a misplaced config.ini fails the level, and `set -u` is not tripped anywhere.
No emulator was booted.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-22 23:06:24 -05:00
JMR-devandClaude Opus 5 792286a2d7 Stop three claims in the API 37 doc outrunning their evidence
Three corrections, all narrowing:

- angle_indirect and swangle_indirect are not two independent renderers here. Both
  logged gles_mode_selected:swangle with the same adapter, differing only in the
  Vulkan backend underneath -- unlike at API 33-36, where angle_indirect resolves to
  ANGLE on llvmpipe. What is 7-for-7 is the host-GLES-versus-not split, not "two
  renderers agree".

- "Disabling SystemUI stops the crashes entirely" was one 180-second measurement on a
  device that had been up twelve minutes. The harness path reproduces a rate collapse,
  not a zero: its own quiet check printed 1 abort in 45 s and 4 across the run. A
  47-second Gradle run survives that; a five-minute one might not.

- "Reproduced twice" conflated two routes. The 49/2/0/2 came back from a hand-driven
  sequence and from the harness, which corroborates the numbers, but the harness path
  itself has one green measurement.

Also records what the doc never said: from 37.1 onward Google ships only 16 KB-page
x86_64 images, so page-size alignment is a prerequisite for that path rather than a
detail. All 20 libraries in the committed FFmpeg AAR are 0x4000-aligned, checked
before the first ps16k boot -- which is why 37.1 reproducing the abort means the
gralloc bug and not a page-size mismatch.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-22 22:10:50 -05:00
JMR-devandClaude Opus 5 739bffa5a0 Re-derive the API 37 emulator failure: it is the renderer, not the image
docs/api-37-emulator-crash.md claimed "Both swiftshader_indirect and host crash...
The crash is in the gralloc mapper, below the renderer." Re-measured, seven runs,
one variable each: that is wrong. The mapper is below the renderer, but whether its
bad path is reached is not.

  -gpu host              gles_mode_selected:host    never boots  (57-71 aborts, looping)
  -gpu swangle_indirect  gles_mode_selected:swangle boots, 85 s  (1 abort)
  -gpu angle_indirect    gles_mode_selected:swangle boots, 112 s (2 aborts)

The old claim rested on two samples of two different things, neither of them ANGLE:
the local swiftshader_indirect sample was void, because on this host every
SwiftShader-GLES launch segfaults the emulator before the guest matters (the
execheap bug in docs/local-emulator.md, not understood when that file was written),
and the CI sample was a single swiftshader_indirect run.

Also re-derived, and null: android-37.1 rev 8 -- a stable REL image the doc's own
"new image revision" trigger was too narrow to catch -- fails identically;
-feature -GLDMA,-GLDMA2,-GLDirectMem is accepted and changes nothing; the image's
advancedFeatures.ini is byte-identical to API 36's but for one camera line; and
there is still no ATD image above API 36.

The mechanism, end to end: SystemUI registers a nav-bar luma-sampling listener,
SurfaceFlinger's RegionSamplingThread locks a GraphicBuffer, Gralloc5 routes into
GoldfishMapper::readFromHost, which asserts, and init SIGKILLs zygote in response --
so the framework restarts under the test run. Disabling SystemUI removes the
listener and the aborts stop dead: 0 in 180 s, against 10-11 per 150 s.

So run-e2e.sh now covers API 37: renderer chosen per level (33-36 need host, 37
must not have it), dotted image labels, SystemUI disabled followed by a deliberate
stop/start, and an abort count printed on every 37 row. The result is 49 tests, 2
failures, 0 errors, 2 skipped, reproduced twice. The two failures are
Media3EngineTest on c2.goldfish.h264.decoder; API 35 under the identical renderer is
49/0/0/2 green, so they are the image and not the renderer.

CI's matrix should still stop at 36, for reasons now written down rather than
assumed. CLAUDE.md is left alone; a replacement bullet is proposed in the doc.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-22 22:09:19 -05:00
3 changed files with 833 additions and 197 deletions
+437 -132
View File
@@ -1,127 +1,381 @@
# API 37 is not tested in CI: a crash in Google's `android-37.0` emulator image
# API 37 on the emulator: a guest gralloc bug that only the host GL renderer triggers
**Status:** open upstream, worked around by removing API 37 from the E2E matrix.
The app itself is verified good on real API 37 hardware — this is an emulator bug only.
**Last verified:** 2026-08-21, against emulator `37.1.11.0` and system image revision 6
**Status:** the bug is real and still open upstream, but the previous diagnosis in this file was
wrong about its most important detail. **The renderer decides whether API 37 boots**, and once it
boots, disabling SystemUI collapses the crash rate far enough to run a suite —
`tools/local-emulator/run-e2e.sh 37` gets through the whole instrumented suite and comes back with
**2 failures, 0 errors and the two by-design skips** (measured 49 / 2 / 0 / 2 at `22c7914`, where
the suite was 49 tests — [Reading these totals](#reading-these-totals) before comparing any total
with another). The crashes do not stop outright, and the two failures are real; both are quantified
below. CI's matrix should still stop at 36 — see
[So should CI take API 37?](#so-should-ci-take-api-37).
**Last verified:** 2026-08-22, emulator `37.1.11.0` (build 15917651), Fedora 44,
against system images `android-37.0` rev 6 **and** `android-37.1` rev 8.
`minSdk` is 33 and `targetSdk` is 37, and the E2E matrix in
[`status_check.yml`](../.github/workflows/status_check.yml) runs API 33 through 36.
API 37 is deliberately absent. This is why.
## The correction
## Summary
This file previously said, under "What was ruled out":
The `android-37.0` emulator system image crashes `surfaceflinger` in a loop. The app
under test never gets a working framework, so every instrumented test fails regardless
of what the app does. The bug is in the emulator image, not in this project.
> **GPU mode.** Both `swiftshader_indirect` and `host` crash, with the same assertion and
> the same frames. The crash is in the gralloc mapper, below the renderer.
The crash is an assertion inside the emulator's own gralloc implementation:
**That is wrong.** The mapper is below the renderer, but *whether the mapper's bad path is
reached* is not. Re-measured on 2026-08-22, seven runs, one variable at a time:
| # | system image | `-gpu` | GLES the emulator chose | booted? | surfaceflinger aborts |
|---|---|---|---|---|---|
| r01 | `android-37.0` rev 6 | `host` | host (Mesa Iris Xe) | **no**, 422 s | 71, looping |
| r02 | `android-37.1` rev 8 | `host` | host (Mesa Iris Xe) | **no**, 362 s | 65, looping |
| r03 | `android-37.0` rev 6 | `swangle_indirect` | ANGLE | **yes, 85 s** | 1 |
| r04 | `android-37.0` rev 6 | `host` + `-feature -GLDMA,-GLDMA2,-GLDirectMem` | host | **no**, 363 s | 57, looping |
| r05 | `android-37.0` rev 6 | `angle_indirect` | ANGLE | **yes, 112 s** | 2 |
| r06 | `android-37.1` rev 8 | `swangle_indirect` | ANGLE | **yes, 285 s** | 23 |
| r07 | `android-37.0` rev 6 | `host` + `-feature -HostComposition` | host | **no**, wedged adb at 208 s | not readable |
The discriminator is exact across all seven: **a run boots if and only if the emulator log says
something other than `gles_mode_selected:host`.**
One caveat about how independent those rows are, because the table flatters itself. `-gpu
angle_indirect` (r05) and `-gpu swangle_indirect` (r03) both logged `gles_mode_selected:swangle`
and both reported the same adapter, differing only in the Vulkan backend beneath
(`vulkan_mode_selected:lavapipe` against `swiftshader`). So they are closer to one GLES path
reached two ways than to two renderers agreeing — note that at API 33–36
[`docs/local-emulator.md`](local-emulator.md) records `angle_indirect` resolving to ANGLE on
*llvmpipe*, a genuinely different adapter, which it did not do here. What is 7-for-7 is the
host-GLES-versus-not split, not "two independent renderers both work".
```
# r01, r02, r04, r07 -- never boots
INFO | emuglConfig_init: vulkan_mode_selected:host gles_mode_selected:host
INFO | Graphics Adapter Android Emulator OpenGL ES Translator (Mesa Intel(R) Iris(R) Xe Graphics (TGL GT2))
# r03, r05, r06 -- boots
INFO | emuglConfig_init: vulkan_mode_selected:swiftshader gles_mode_selected:swangle
INFO | Graphics Adapter Android Emulator OpenGL ES Translator (ANGLE (Google, Vulkan 1.2.0
| (SwiftShader Device (Subzero) (0x0000C0DE)), SwiftShader driver-5.0.0))
```
### Why the wrong claim looked right
It rested on two samples of two different things, and neither of them was ANGLE.
- The **local** `swiftshader_indirect` sample was void. On this workstation *every*
SwiftShader-GLES launch segfaults the host emulator before the guest matters at all —
SELinux denies `execheap` to SwiftShader's Reactor JIT. That is
[`docs/local-emulator.md`](local-emulator.md), and it was not yet understood when this file
was written. So "`swiftshader_indirect` crashes" was true, for an entirely unrelated reason,
and told you nothing about the gralloc assertion.
- The **CI** sample was one `swiftshader_indirect` run on a GPU-less `ubuntu-latest`, and the
**local** sample was one `-gpu host` run. Two renderers, one measurement each, and the pair
written up as "both GPU modes".
`angle_indirect` and `swangle_indirect` — the two modes that work — had never been tried on
API 37. Neither had a second system image.
The lesson is the same one `docs/local-emulator.md` ends on, which makes it worth repeating:
"both backends fail" is a claim about a matrix, and a matrix needs cells, not inference. Two
observations of two different configurations do not establish anything about a third.
## What the bug actually is
`surfaceflinger` aborts inside the emulator's own gralloc mapper:
```
Executable: /system/bin/surfaceflinger
signal 6 (SIGABRT), code -1 (SI_QUEUE), tid: RegionSampling
Abort message: 'Assertion failed: !rcEnc->featureInfo()->hasReadColorBufferDma'
#03 mapper.ranchu.so GoldfishMapper::readFromHost(cb_handle_t const&) const
#04 mapper.ranchu.so GoldfishMapper::GoldfishMapper()::'lambda'(...)::__invoke
#05 libui.so android::Gralloc5Mapper::lock(...)
#06 libui.so android::GraphicBufferMapper::lock(...)
#07 libui.so android::GraphicBuffer::lockAsync(...)
#08 libui.so android::GraphicBuffer::lock(...)
#09 surfaceflinger android::RegionSamplingThread::threadMain()
#03 /vendor/lib64/hw/mapper.ranchu.so GoldfishMapper::readFromHost(cb_handle_t const&) const+543
#04 /vendor/lib64/hw/mapper.ranchu.so GoldfishMapper::GoldfishMapper()::'lambda'(...)::__invoke+704
#05 /system/lib64/libui.so android::Gralloc5Mapper::lock(...)+63
#06 /system/lib64/libui.so android::GraphicBufferMapper::lock(...)+198
#07 /system/lib64/libui.so android::GraphicBuffer::lockAsync(...)+545
#08 /system/lib64/libui.so android::GraphicBuffer::lock(...)+67
#09 /system/bin/surfaceflinger android::RegionSamplingThread::threadMain()+2571
```
`RegionSamplingThread` is SystemUI's navigation-bar luma sampling. It calls
`GraphicBuffer::lock`, which routes into `GoldfishMapper::readFromHost`, which asserts
that the host has *not* negotiated the `ReadColorBufferDma` capability. On this image
the host has, so the assertion fails and `surfaceflinger` aborts. It restarts and
aborts again.
`RegionSamplingThread` is SystemUI's nav-bar luma sampling. It locks a `GraphicBuffer` for CPU
read; that routes through the Gralloc5 mapper into `GoldfishMapper::readFromHost`, which is the
*non-DMA* readback path and asserts that the host has not negotiated `ReadColorBufferDma`. The
host always has, so the assert fires whenever that path is taken.
## Impact
Two facts pin down what "always" means:
The failure surfaces in two different ways depending on how far the job gets, which is
why it took several rounds to identify:
- **The capability is negotiated regardless of renderer.** The evidence is the aborts
themselves: the assertion that fires is `!hasReadColorBufferDma`, and it fires under ANGLE
(r03/r05/r06) as well as under the host translator — just far less often. That is a direct
observation of the guest having negotiated DMA readback under both, and it stands alone.
(Supporting only, and weaker than it first looks: `ANDROID_EMU_read_color_buffer_dma` appears
in exactly one file in the SDK, `emulator/lib64/libgfxstream_backend.so`, which every `-gpu`
mode goes through. A string search establishes where the extension is implemented, not that
it is negotiated on every path.)
- **It is not gated by any feature flag the emulator exposes.** See the ruled-out list below.
| Guest RAM | Where it dies | What CI reports |
|---|---|---|
| 1536 MB | during APK install | `Unknown failure: cmd: Can't find service: package` |
| 2560 MB | during the test run | `There were failing tests` — all of them |
So the renderer does not decide whether the guest *believes* DMA readback exists. It decides how
often `RegionSamplingThread` ends up in `readFromHost` — which under the host GL translator is
constantly, and under ANGLE is occasionally.
At 2560 MB the install succeeds and the tests actually execute, then fail wholesale.
The first failure in the report is misleading:
### Why one abort takes down the whole device
`surfaceflinger` is a critical service. When it dies, `init` kills the framework with it:
```
kotlin.UninitializedPropertyAccessException: lateinit property output has not
been initialized
at Media3EngineTest.tearDown(Media3EngineTest.kt:53)
java.lang.IllegalStateException: WorkManager is not initialized properly.
You have explicitly disabled WorkManagerInitializer in your manifest, ...
08-22 21:40:28.253 I/init: Sending SIGKILL to service 'zygote' (pid 470) process group...
08-22 21:40:28.260 I/init: Service 'zygote' (pid 470) received SIGKILL
```
Neither is a real defect in this project. `tearDown` throws because `setUp` never got
far enough to assign `output`, and WorkManager's `InitializationProvider` never runs
because content-provider installation fails on a framework whose `surfaceflinger` is
crash-looping. The same tests pass at API 33, 34, 35, and 36 in the same CI run, and
the first `surfaceflinger` abort is timestamped *before* the test results are reported.
This is not inference. The full suite was run against a physical API 37 device and
passed — see [Verified on real API 37 hardware](#verified-on-real-api-37-hardware)
below. `ConversionWorkerTest` and `ConcatWorkerTest`, which drive a real WorkManager
round trip and are among the tests that failed this way in CI, both pass there.
Everything above zygote goes with it, which is why the symptoms look nothing like a graphics
bug. Under `-gpu host` the cycle repeats every five to seven seconds forever and
`sys.boot_completed` is never set. Under ANGLE the aborts are sparse enough that the boot
usually completes between them — but they do not stop, and each one is a framework restart.
That is the difference between "boots" and "is usable", and it is the reason this is not simply
fixed by changing the renderer. See [Can the suite run on it?](#can-the-suite-run-on-it) below.
## Environment
Reproduced identically in two unrelated environments, so it is not specific to a host
GPU, driver, or CI runner.
| | GitHub Actions | Local workstation |
|---|---|---|
| Host | `ubuntu-latest`, no GPU | Fedora, Intel Iris Xe (TGL GT2) |
| Host | `ubuntu-latest`, no GPU | Fedora 44, Intel Iris Xe (TGL GT2), kernel `7.1.8-200.fc44` |
| Emulator | `37.1.11.0` (build 15917651) | `37.1.11.0` (build 15917651) |
| GPU mode | `swiftshader_indirect` | `host` |
| Result | boots, aborts during tests | aborts before boot completes |
| GPU mode measured | `swiftshader_indirect` | `host`, `angle_indirect`, `swangle_indirect` |
System image: `system-images;android-37.0;google_apis;x86_64`, `Pkg.Revision=6`,
`AndroidVersion.ApiLevel=37.0`, `AndroidVersion.ExtensionLevel=22`
Images, both reproducing it:
```
Build fingerprint: google/sdk_gphone64_x86_64/emu64xa:17/CE2A.260420.019/15611780:userdebug/dev-keys
Kernel Release: 6.12.58-android16-6-gccafb60de224-ab14828483
system-images;android-37.0;google_apis;x86_64 Pkg.Revision=6 ApiLevel=37.0 ExtensionLevel=22
fingerprint google/sdk_gphone64_x86_64/emu64xa:17/CE2A.260420.019/15611780:userdebug/dev-keys
system-images;android-37.1;google_apis_ps16k;x86_64 Pkg.Revision=8 ApiLevel=37.1 ExtensionLevel=23
ro.build.version.codename=REL (a release image, not a preview)
```
**Note the `ps16k` in the second one — it is not optional, and it is why the 37.1 result is
interpretable.** From API 37.1 onward Google ships *only* 16 KB-page x86_64 images; there is no
plain `google_apis` variant to pick. `sdkmanager --list` for 37.1 and 37.2-beta* offers nothing
but `google_apis_ps16k` and `google_apis_playstore_ps16k`. That makes page-size alignment a
prerequisite rather than a detail: a `.so` that is not 16 KB aligned will not load on such a
guest, and the resulting failure looks like an app bug. Checked before the first `ps16k` boot,
using the same test `build.yml` applies to release APKs — all 20 libraries in the committed
`bin/ffmpeg-kit-next-8.1.1.aar`, both ABIs, report `0x4000`:
```
$ for f in jni/*/*.so; do readelf -lW "$f" | awk '$1=="LOAD"{print $NF}' | sort -u; done
0x4000 (x20: libavcodec, libavdevice, libavfilter, libavformat, libavutil,
libc++_shared, libffmpegkit, libffmpegkit_abidetect, libswresample, libswscale
-- arm64-v8a and x86_64)
```
So when `android-37.1` reproduced the abort, that was the gralloc bug and not a page-size
mismatch. `image_pkg_for_api` in `tools/local-emulator/run-e2e.sh` encodes the `ps16k` tag for
37.1; if this ever fails after an FFmpeg rebuild, re-run the alignment check first.
## What was ruled out, and how
Each of these was tested rather than reasoned about, because the first three attempts
at this bug were plausible fixes that turned out to address earlier, unrelated failures.
**A newer system image.** This file's own revisit trigger was "a new `android-37.0` system image
revision ships (this was revision 6)". That trigger was written too narrowly and would never have
fired: `android-37.0` is *still* revision 6, but Google shipped a whole new minor level.
`android-37.1` `google_apis_ps16k` revision 8 — a `REL` build, not a beta — was installed and
tested (r02, r06) and **behaves identically**: same assertion, same frames, never boots under
`-gpu host`, and *worse* under ANGLE (23 aborts to `37.0`'s 1). `android-37.2-beta3` exists too
but was not needed; two independent images agreeing settles it, and a beta could not be used by
CI anyway.
**Guest memory.** The emulator raises an undersized guest to a minimum on its own, but
only for API levels it recognises, and it does not recognise `"37.0"`. API 33 bumps to
2048 MB and 34/35/36 to 2560 MB, while API 37 logged no bump at all and ran at the
`pixel_6` default of 1536 MB. Setting `ram-size: 2560M` explicitly fixed that asymmetry
and did change the outcome — the job got past install and into the test run — but it is
not the underlying bug. At the moment of failure the guest reported `MemTotal 2527392
kB` with `MemAvailable 1507104 kB`: 1.5 GB free, and no OOM kills.
**An ATD image.** Still does not exist for API 37. `sdkmanager --list` offers `aosp_atd` and
`google_atd` for API 30 through 36 and nothing above:
**GPU mode.** Both `swiftshader_indirect` and `host` crash, with the same assertion and
the same frames. The crash is in the gralloc mapper, below the renderer.
```
system-images;android-36;google_atd;x86_64 | 1 | Google APIs ATD Intel x86_64 Atom System Image
(no android-37 ATD of any kind)
```
**Disabling the DMA feature.** `GLDMA` is the host feature that most plausibly backs the
guest's `hasReadColorBufferDma`. Launching with `-feature -GLDMA` was accepted by the
emulator — the log confirms `Feature 'GLDMA' (51) is overridden to 'disabled'` — and
`surfaceflinger` still aborted 13 times and the device never finished booting. Whatever
sets that guest capability, it is not this flag.
For API 37 the only x86_64 images are `google_apis`, `google_apis_playstore`, their `ps16k`
16 KB-page variants, and Wear OS. Check again when revisiting.
**An ATD image.** `google_atd` / `aosp_atd` images are built for automated testing and
ship without the SystemUI package set, which is what drives `RegionSamplingThread` in
the first place. That would likely sidestep the bug class entirely, but **no ATD image
exists for `android-37.0`** — only `google_apis`, `google_apis_playstore`, the `ps16k`
16 KB-page variants, and Wear OS. Check again when revisiting; if an ATD image appears,
try it before anything else here.
**The DMA feature flags.** `GLDMA` alone was ruled out previously; `GLDMA2` and `GLDirectMem`
were not, and the per-image `advancedFeatures.ini` turns all three on. Disabling all three
together (r04) is accepted by the emulator and changes nothing:
```
INFO | Feature 'GLDMA' (51) is overridden to 'disabled'
INFO | Feature 'GLDMA2' (52) is overridden to 'disabled'
INFO | Feature 'GLDirectMem' (53) is overridden to 'disabled'
... 57 surfaceflinger aborts, device never boots
```
**Host composition.** `-feature -HostComposition` (r07) was the best remaining guess at what
forces the readback. It did not help; it made things worse, wedging adb entirely at 208 s so the
crash buffer could not even be read. Recorded as inconclusive rather than ruled out, because no
evidence came back from it.
**Guest feature negotiation differing from API 36.** It does not. The image-level
`advancedFeatures.ini` for `android-37.0` is byte-identical to `android-36`'s except for one
unrelated line:
```
$ diff android-36/google_apis/x86_64/advancedFeatures.ini android-37.0/google_apis/x86_64/advancedFeatures.ini
+QemuCameraSensorOrientation = on
```
`GLDMA`, `GLDMA2`, `GLDirectMem`, `GrallocSync`, `HostComposition` and `YUVCache` are on in
both. API 36 boots and passes. So nothing about the host/guest feature handshake changed — the
regression is in the guest's Gralloc5 mapper or in what API 37's `RegionSamplingThread` asks of
it, not in what the emulator advertises.
**Guest memory.** Ruled out previously and not revisited; every run above used
`hw.ramSize=2560`, the same value the E2E matrix pins, and none of them OOMed.
**A host-side crash.** Not this bug, and worth stating because the other emulator failure on this
workstation *is* host-side. Every run above left `coredumpctl` empty and produced zero
`avc: denied` lines, and the qemu process was still alive at the end of the ones that never
booted (`emulator_alive=yes`). The host emulator is fine; the guest is not.
## Can the suite run on it?
**Almost.** `tools/local-emulator/run-e2e.sh 37` now runs the whole suite locally, and all of it
passes except two tests. Measured at `22c7914`: **49 tests, 2 failures, 0 errors, 2 skipped** — 45
passed, the two `Media3EngineTest` failures dissected below, and the two `assumeTrue` skips every
level has. It costs two deviations from how every other level is run, and both are worth
understanding before trusting the leg.
Two things about that total before it is compared with anything. It is the size of the suite on
the checkout that ran, not a property of API 37 — `app/src/androidTest` held 49 `@Test` methods at
`22c7914`, and a newer checkout reports its own count; see
[Reading these totals](#reading-these-totals). And **the Pixel has never run 49**: its green run
was 40 / 0 / 0 / 2 at `edd6385`, the same suite nine tests earlier. What compares across the two
is two failures against none, and the same two skips — not the totals.
The same numbers and the same two test names came back twice, which is real corroboration — but
by two different routes, and only one of them is the harness. The first was driven by hand
(`pm disable-user`, then several minutes of incidental framework restarts, then `e2e-run.sh`
directly); the second went through `disable_region_sampling`'s `stop; start`. **The harness path
itself has one green measurement.** What would make this routine is a second consecutive
`run-e2e.sh 37` whose only failures are the same two.
### Booting is not the same as being usable
Changing the renderer gets the device to `sys.boot_completed=1`, and that is all it gets you. The
aborts do not stop, and each one is a framework restart. A five-minute test run does not survive
that. What it looks like from Gradle:
```
Shell command failed (1): rm -rf "/sdcard/Android/media/org.libremediaconverter/..."
rm: ...: Transport endpoint is not connected
Starting 0 tests on lmc_e2e_api37(AVD) - 17
Shell command failed (20): am get-current-user
cmd: Can't find service: activity
Device emulator-5572 failed to uninstall test APK org.libremediaconverter.
[cmd: Can't find service: package]
Test run failed to complete. No test results.
onError: commandError=false message=INSTRUMENTATION_ABORTED: System has crashed.
```
Measured idle rate on `android-37.0` under `swangle_indirect`: **10 aborts in 150 s, then 11 more
in the next 150 s**. Steady, not a start-up transient.
### The fix is to remove the region-sampling listener, not to survive it
`RegionSamplingThread` exists only because SystemUI registers a nav-bar luma-sampling listener.
Take SystemUI away and the thread is never started, so the mapper's bad path is never called:
```
$ adb shell pm disable-user --user 0 com.android.systemui
Package com.android.systemui new state: disabled-user
=== aborts at start of measurement: 36
=== idle 180s with SystemUI disabled ===
=== aborts after: 36 NEW IN WINDOW: 0
--- services still up? ---
activity Service activity: found
package Service package: found
window Service window: found
```
**Zero in 180 s, against 10–11 per 150 s.** That is the strongest evidence that region sampling
is the dominant trigger, and it is worth recording even by someone who never wants the workaround.
It does not establish it as the *only* trigger: the paragraph below has an abort surviving the
disable, and nothing measured here says whether that residue is a second caller of the readback
path or a disable that did not fully take.
Do not read that as "the crashes stop", though, because the harness path does not reproduce a
clean zero. Its own post-disable check on the run recorded below printed
```
quiet check: 1 new surfaceflinger aborts in 45 s (want 0)
surfaceflinger hasReadColorBufferDma aborts: 4 (whole run)
```
So what is reliably achieved is a **rate collapse** — from roughly one abort every fourteen
seconds to one every forty-five — which a 47-second Gradle run survives and a five-minute one
might not. The 180-second zero above is one measurement on a device that had been up for twelve
minutes and had already cycled its framework several times. The harness prints the quiet-check
delta on every run precisely so this is visible rather than assumed.
One ordering detail cost a whole run and is now encoded in `disable_region_sampling`: by the time
`sys.boot_completed` flips, SystemUI has **already registered**, and `pm disable-user` does not
retract an existing registration — it only stops the package being started again. Disabling it
and proceeding straight to the tests fails exactly as before. The harness therefore does
`stop; start` afterwards, so the framework that comes back never starts SystemUI at all.
### The two deviations, stated plainly
1. **The renderer is ANGLE, not the host GPU.** Shared with nothing else in the matrix — API
33–36 run `-gpu host` locally, and CI runs `swiftshader_indirect`.
2. **SystemUI is disabled.** The API 37 leg does not run the same device configuration as any
other leg or as the Pixel. It is defensible here only because nothing in this suite touches
system UI — these are Media3, FFmpeg and WorkManager tests — and because the alternative is no
local API 37 coverage at all. **Anything that ever does depend on system UI must not trust
this leg.**
### The two remaining failures are the same bug, one layer down
```
org.libremediaconverter.convert.Media3EngineTest > runsFromAThreadWithNoLooper FAILED
org.libremediaconverter.convert.Media3EngineTest > transcodesH264ToH265AndReportsProgress FAILED
androidx.media3.transformer.ExportException: Codec exception:
CodecInfo{type=VideoDecoder, ..., mime=video/avc, name=c2.goldfish.h264.decoder}
at androidx.media3.transformer.DefaultCodec.maybeDequeueOutputBuffer(DefaultCodec.java:398)
Caused by: android.media.MediaCodec$CodecException:
at android.media.MediaCodec.native_dequeueOutputBuffer(Native Method)
```
Three measurements say this is the emulator image and not this app, and not the software
renderer:
- **Control at API 35 under the identical renderer.** `GPU_MODE=swangle_indirect
tools/local-emulator/run-e2e.sh 35` → **49 / 0 / 0 / 2** at `22c7914`, green.
`c2.goldfish.h264.decoder` is perfectly happy under ANGLE one API level down, so the renderer is
not what breaks it.
- **Real API 37 hardware passes**, see below. There is no `c2.goldfish.*` codec on a Pixel.
- The failing call is `dequeueOutputBuffer` on the *goldfish* decoder — the emulator's own codec,
which like `RegionSamplingThread` gets its frames out of a host-side colour buffer. Same
readback machinery, one layer down. This is inference rather than a measurement, and is flagged
as such; what is measured is the first two bullets.
**Do not try `-feature -HardwareDecoder`.** It is the obvious next idea and it is much worse:
forcing the guest onto software decoders took the run from 2 failures to **46**, across
`RemuxTest`, `ForcedFailureTest`, `HardwareFallbackTest` and `UnopenableUriTest` as well. The
suite depends on those decoders existing.
### So should CI take API 37?
**No, and the matrix should still stop at 36.** Three reasons, in order of weight:
1. CI runs `swiftshader_indirect` on a GPU-less runner. The aborts happen under ANGLE too — they
are merely sparser — so nothing here says a runner would be stable.
2. The working configuration needs SystemUI disabled and a framework restart mid-job. That is a
lot of bespoke device surgery to put behind a merge gate, and it silently weakens what the leg
proves.
3. Even at its best two tests fail, so the leg would be permanently red or permanently
allow-listed. Neither is a gate worth having.
What has changed is the *local* story: API 37 is no longer a level nobody can look at. A
regression that shows up at 37 and not at 36 can now be reproduced on this workstation in about
four minutes, which is what the missing matrix row was really costing.
## Verified on real API 37 hardware
The bug is confined to the emulator image. On 2026-08-21 the whole instrumented suite
was run against a physical device and passed:
Unchanged and still true. On 2026-08-21 the whole instrumented suite ran green on a physical
device:
```
Device: Pixel 10 Pro XL (mustang), arm64-v8a
@@ -132,65 +386,94 @@ API: 37 (Android 17, codename REL -- a release build, not a preview)
40 tests, 0 failures, 0 errors, 2 skipped BUILD SUCCESSFUL
```
**That "40" is a measurement of the tree it ran on, not a baseline for today**, and it is not a
contradiction of the totals in [`docs/local-emulator.md`](local-emulator.md) either.
### Reading these totals
Every total in this file and in [`docs/local-emulator.md`](local-emulator.md) is the size of
`app/src/androidTest` on the checkout that produced it, and nothing else. The reported total has
equalled that checkout's `@Test` count everywhere it has been checked:
| checkout | `@Test` methods | total the run reported |
|---|---|---|
| `edd6385` | 40 | 40 — the Pixel run above |
| `22c7914` | 49 | 49 — the four local levels, and API 37 |
| `18c53a3` | 57 | not run |
So the number to expect is not written down here. It is derived from the checkout in front of
you, which is the only thing that cannot go stale:
```bash
grep -rho '@Test' app/src/androidTest | wc -l
```
**Before each release, run the suite on the Pixel 10 Pro XL and expect that many tests, 0
failures, 0 errors, 2 skipped.** The failure, error and skip counts are the invariant; the total
is not. A total that disagrees with your own checkout's count is the signal — an old checkout, a
stale build, or tests that never ran — and it is worth stopping on either way.
The two skips are `RealMediaBenchmark.hardwareVersusSoftwareOnRealVideo` and
`av1InputRoutesAccordingToDeviceDecodeSupport`, which `assumeTrue` their sample files
are present and skip when they are not. That is by design and unrelated to API level.
`av1InputRoutesAccordingToDeviceDecodeSupport`, which `assumeTrue` their sample files are present
and skip when they are not. That is by design and unrelated to API level.
One harmless warning appears during the run and can be ignored:
`No UID for androidx.test.services in user 0`, from an `appops` call the test services
package makes before it is fully registered.
So the app is correct on Android 17. What is missing is only *automated* coverage in
CI. Until the image is fixed, run the suite on a physical API 37 device before release;
that is the substitute for the missing matrix row.
`No UID for androidx.test.services in user 0`, from an `appops` call the test services package
makes before it is fully registered.
## Reproducing it
Locally, with `-gpu host` so the emulator itself does not segfault on Intel graphics:
Both halves, so the renderer claim can be checked rather than taken on trust:
```bash
export ANDROID_HOME="$HOME/Android/Sdk"
export PATH="$ANDROID_HOME/platform-tools:$ANDROID_HOME/emulator:$ANDROID_HOME/cmdline-tools/latest/bin:$PATH"
sdkmanager --install "system-images;android-37.0;google_apis;x86_64"
echo no | avdmanager create avd -n api37_repro \
-k "system-images;android-37.0;google_apis;x86_64" -d pixel_6 --force
$ANDROID_HOME/emulator/emulator -avd api37_repro \
-no-window -gpu host -noaudio -no-boot-anim -camera-back none -no-snapshot &
# never boots -- surfaceflinger aborts every ~6 s, forever
emulator -avd api37_repro -no-window -gpu host \
-noaudio -no-boot-anim -camera-back none -no-snapshot &
# Boot never completes. Count the aborts:
adb logcat -d -b crash | grep -c hasReadColorBufferDma
# boots in ~85 s, having aborted once or twice on the way
emulator -avd api37_repro -no-window -gpu swangle_indirect \
-noaudio -no-boot-anim -camera-back none -no-snapshot &
```
`sys.boot_completed` never reaches `1`, `pgrep -f system_server` stays empty, and
`keystore2`'s watchdog logs `await_boot_completed ... Overdue` indefinitely.
Count the aborts either way:
To see the CI-side form instead, restore the API 37 row in the E2E matrix of
`status_check.yml` (`api-level: "37.0"` — a bare `37` fails earlier still, during SDK
setup, because there is no `platforms;android-37`).
```bash
adb -s emulator-5554 logcat -d -b crash | grep -c hasReadColorBufferDma
```
Under `-gpu host`, `sys.boot_completed` never reaches `1`, `pgrep -f system_server` stays empty,
and `keystore2`'s watchdog logs `await_boot_completed ... Overdue` indefinitely.
`tools/local-emulator/run-e2e.sh 37` does all of this, with the working renderer picked
automatically — see `gpu_for_api` in that file.
## Filing this upstream
Not yet filed. To file it:
Not yet filed. The report is stronger than it was, because the renderer dependency narrows it:
1. Go to <https://issuetracker.google.com/> and sign in with a Google account.
2. Choose **Report an issue**, then pick the component for the Android emulator — search
the component picker for "Emulator"; it sits under the Android Studio component tree.
If the picker is unclear, Android Studio's **Help → Submit Feedback** opens the same
tracker with the component preselected, and the emulator's own **Extended controls →
Help → File a bug** does likewise.
3. Title it for the mechanism, not the symptom, so it is searchable — for example:
`surfaceflinger aborts in GoldfishMapper::readFromHost (hasReadColorBufferDma) on
android-37.0 google_apis x86_64`.
4. Paste the assertion and backtrace from the top of this document, the environment
table, and the reproduction steps above. State explicitly that it reproduces on two
unrelated hosts under both GPU modes — that is the detail that stops it being closed
as a local graphics problem.
5. List what was ruled out. Bugs that arrive with `-feature -GLDMA` already eliminated
tend not to bounce back asking for it.
6. Attach:
- the guest tombstone, via `adb pull /data/tombstones` (or the `pbtombstone` output
the crash log names)
1. Go to <https://issuetracker.google.com/>, **Report an issue**, and pick the Android emulator
component (search the component picker for "Emulator"; Android Studio's **Help → Submit
Feedback** opens the same tracker with it preselected).
2. Title it for the mechanism: `surfaceflinger aborts in GoldfishMapper::readFromHost
(hasReadColorBufferDma) on android-37.0 and android-37.1 x86_64 -- fatal under -gpu host,
intermittent under ANGLE`.
3. Paste the assertion and backtrace, the environment block, and the seven-row matrix. The
matrix is the valuable part: it shows the abort is not renderer-specific but its *frequency*
is, which points at the readback path rather than at any one GL implementation.
4. State that it reproduces on two independent system images (`37.0` rev 6 and `37.1` rev 8) and
on two unrelated hosts, and that `-feature -GLDMA,-GLDMA2,-GLDirectMem` does not suppress it.
5. Attach:
- the guest tombstone, via `adb pull /data/tombstones` (or the `pbtombstone` output the crash
log names)
- `adb logcat -d -b crash > crash.txt`
- the emulator's own stdout log, captured by redirecting the launch command
- the emulator's own stdout log (`-verbose -debug all`, redirected)
- the AVD's `config.ini`
- a link to a failing CI job, which shows it on hardware you do not control:
<https://github.com/JMR-dev/LibreMediaConverter/actions/runs/32545625459/job/96963461184>
@@ -199,16 +482,38 @@ Record the issue number here once filed.
## When to revisit
Re-add the API 37 row when any of these happens:
The old trigger list named "a new `android-37.0` revision", which is why nothing ever fired even
though a new API level shipped. Watch for these instead:
- a new `android-37.0` system image revision ships (this was revision 6)
- an ATD image appears for API 37
- the upstream issue is marked fixed
- **any new API 37.x system image**, not just a new revision of `37.0` — `37.1` rev 8 and
`37.2-beta*` already exist, and more will. Test with `-gpu host`: if it boots, the guest mapper
is fixed.
- **an ATD image for API 37.** Still none as of 2026-08-22. ATD images ship without SystemUI,
which is what drives `RegionSamplingThread`, so one would very likely sidestep the bug
entirely. Try it before anything else here.
- **the upstream issue being marked fixed.**
Until then the gap is narrower than the missing row suggests. `targetSdk` is 37, so the
app is compiled and unit-tested against it; the API-dependent behaviour this matrix
exists to exercise — the foreground service type, absent below 34, `dataSync` at 34,
`mediaProcessing` from 35 — is covered at 35 and 36; and the full instrumented suite has
been run green on real API 37 hardware. What is missing is *automated* API 37 coverage,
so a regression there would not be caught by a pull request. Run the suite on a physical
API 37 device before each release for as long as this row is absent.
## Correction owed to `CLAUDE.md`
`CLAUDE.md` currently says:
> - **The API 37 image is broken.** `android-37.0` crash-loops surfaceflinger inside its own
> gralloc mapper, so every test fails there regardless of this app.
> `docs/api-37-emulator-crash.md` records the evidence and the ruled-out fixes; CI's matrix
> therefore stops at API 36 even though targetSdk is 37.
The first sentence is right, and now under-specified in one direction and over-specified in the
other: it is not only `android-37.0` (it is `37.1` too), and it does not crash-loop under every
renderer. Proposed replacement, offered for review rather than applied here — `CLAUDE.md` is left
alone deliberately, because several branches touch it:
> - **The API 37 images crash-loop surfaceflinger under the host GL renderer.** Both
> `android-37.0` and `android-37.1` abort inside their own gralloc mapper
> (`RegionSamplingThread` → `GoldfishMapper::readFromHost`), and when surfaceflinger dies init
> SIGKILLs zygote, so the framework restarts under the test run. Under `-gpu host` it never
> boots at all; under `-gpu swangle_indirect` it boots and the aborts merely become
> intermittent. `docs/api-37-emulator-crash.md` has the seven-run matrix and the ruled-out
> list, and `tools/local-emulator/run-e2e.sh` picks the working renderer per API level.
> CI's matrix stops at 36 because CI runs `swiftshader_indirect` on a GPU-less runner and the
> aborts continue there too. **API 37 still needs a manual check on the Pixel 10 Pro XL before
> each release.**
+52 -17
View File
@@ -1,8 +1,10 @@
# Emulators do run on this host: the segfault is SwiftShader's JIT against SELinux
**Status:** solved. Local instrumented runs work with `-gpu host`, and the suite is green
on API 33–36 — 49 tests, 0 failures, 0 errors, 2 skipped on every level. See
[The sweep, run](#the-sweep-run).
**Status:** solved. Local instrumented runs work with `-gpu host`, and the whole suite is green
on API 33–36 — 0 failures, 0 errors and the two by-design skips on every level, measured as
49 / 0 / 0 / 2 at `22c7914`, where the suite was 49 tests. See [The sweep, run](#the-sweep-run),
and [Reading these totals](api-37-emulator-crash.md#reading-these-totals) before comparing any
total with another checkout's.
**Last verified:** 2026-08-22, emulator `37.1.11.0` (build 15917651), Fedora 44,
kernel `7.1.8-200.fc44`, `selinux-policy-44.6-1.fc44`
@@ -190,11 +192,28 @@ emulator -avd <name> -no-window -gpu host \
`.github/scripts/e2e-run.sh` for the run itself. Use it rather than the raw command:
```bash
tools/local-emulator/run-e2e.sh # API 33 34 35 36
tools/local-emulator/run-e2e.sh # API 33 34 35 36 37
tools/local-emulator/run-e2e.sh 35 # one level
tools/local-emulator/run-e2e.sh 37 37.1 # both API 37 images
GPU_MODE=swangle_indirect tools/local-emulator/run-e2e.sh 35
```
Levels are the labels above, not SDK ints: API 37's SDK directories are dotted
(`android-37.0`, `android-37.1`) and there is no `android-37`, so `37` is accepted as a
spelling of `37.0`. Setting `GPU_MODE` forces one renderer on every level, which is what
you want when measuring a mode; leaving it unset lets `gpu_for_api` pick, which is what
you want when running the suite — 33–36 need `host` and 37 must not have it.
**A bare `run-e2e.sh` exits 1, and that is the design.** API 37 is in the default list
deliberately — leaving it out is what left the level unlooked-at for as long as it was — and
it is permanently two failures short of green: `Media3EngineTest` cannot drive the emulator's
`c2.goldfish.h264.decoder` on those images, which
[`api-37-emulator-crash.md`](api-37-emulator-crash.md) pins on the image and not on this app
(API 35 under the same renderer is green). The summary row names the two expected failures so
that a third is visibly new, and the script repeats the point on the way out. Anything that
treats a non-zero exit as breakage — a wrapper, a hook, a habit — should name the levels it
wants: `run-e2e.sh 33 34 35 36` is the sweep that can be green.
`swangle_indirect` is the fallback worth knowing about. It is entirely software, so it
does not depend on reaching the session's GPU — useful over plain SSH, where `-gpu host`
has not been tested and may not find a device. It is also the closest local analogue to
@@ -236,8 +255,9 @@ after an AGP upgrade.
## The sweep, run
`tools/local-emulator/run-e2e.sh`, one invocation per level so each got a freshly created
AVD, `-gpu host` throughout, 2026-08-22 19:42–19:56. Every level matches the physical
Pixel 10 Pro XL (API 37) baseline of 49 / 0 / 0 / 2 exactly:
AVD, `-gpu host` throughout, 2026-08-22 19:42–19:56, on `22c7914`. All four levels agree exactly,
and 49 is that checkout's whole suite — every `@Test` in `app/src/androidTest`, two of which skip
by design everywhere:
| API | Android | AVD | Boot | `connectedDebugAndroidTest` | Tests | Failures | Errors | Skipped |
|---|---|---|---|---|---|---|---|---|
@@ -246,6 +266,12 @@ Pixel 10 Pro XL (API 37) baseline of 49 / 0 / 0 / 2 exactly:
| 35 | 15 | `lmc_e2e_api35` | 40 s | 3 m 46 s | 49 | 0 | 0 | 2 |
| 36 | 16 | `lmc_e2e_api36` | 90 s | 2 m 18 s | 49 | 0 | 0 | 2 |
The physical Pixel has never reported 49, and an earlier version of this paragraph said the
sweep matched it exactly. Its green API 37 run was 40 / 0 / 0 / 2, at `edd6385` — the same suite
nine tests earlier. What matches is 0 failures, 0 errors and the same two skips; totals only ever
match between runs of one checkout, which
[`api-37-emulator-crash.md`](api-37-emulator-crash.md#reading-these-totals) sets out.
Thirteen and a half minutes for the four levels, AVD creation and cold boots included;
fifteen with the pre-warm build in front of them. Nothing needed a retry, and no level
produced a `diagnostics-api*.txt` — `e2e-run.sh` writes that only on the failure path, so
@@ -352,9 +378,11 @@ and the same binaries against the same kernel boot fine under `-gpu host`.
before `sys.boot_completed` is ever set. Nothing in the guest — system image variant,
RAM, disk size, ATD versus `google_apis` — can influence a host-side `mprotect` denial,
so none of those axes was varied. (The API 37 failure in
[`api-37-emulator-crash.md`](api-37-emulator-crash.md) is genuinely guest-side and
genuinely unrelated: there the host emulator survives and the guest's `surfaceflinger`
aborts.)
[`api-37-emulator-crash.md`](api-37-emulator-crash.md) is genuinely guest-side — there the host
emulator survives and the guest's `surfaceflinger` aborts — but it is **not** unrelated, as this
paragraph originally claimed. Both are decided by the renderer, in opposite directions: below 37
you must avoid SwiftShader GLES and `-gpu host` is the answer; at 37 you must avoid the *host* GL
translator and `-gpu host` is the thing that never boots.)
**Turning the SELinux boolean on** — deliberately *not* done, though it would almost
certainly work:
@@ -401,8 +429,13 @@ are easy to forget to look at.
"SwiftShader 4.0.0.1" as reported by the GLES translator.
- **If `-gpu host` regresses** after a Mesa or kernel update, fall back to
`GPU_MODE=swangle_indirect`, which needs no GPU at all.
- **This changes nothing about API 37.** That image is broken for a different reason and
still must be checked on the physical Pixel 10 Pro XL before each release.
- **API 37 needs the opposite renderer, and this file used to say it needed nothing.** The
original bullet here read "This changes nothing about API 37"; that turned out to be wrong.
The API 37 images abort `surfaceflinger` under the *host* GL translator and boot under ANGLE —
the exact mirror of the rule above — and `run-e2e.sh` therefore picks the renderer per API
level. See [`api-37-emulator-crash.md`](api-37-emulator-crash.md), which was rewritten on
2026-08-22 with the seven-run matrix. API 37 still must be checked on the physical Pixel 10 Pro
XL before each release.
## Correction owed to `CLAUDE.md`
@@ -432,12 +465,14 @@ Proposed replacement for the section, offered for review rather than applied her
> `auto` (the default), `off` and `guest` do when headless. `-gpu host` works, and the
> harness both picks it and refuses the others. `docs/local-emulator.md` has the
> backtrace and the mode matrix.
> - **The API 37 image is broken.** `android-37.0` crash-loops surfaceflinger inside its
> own gralloc mapper, so every test fails there regardless of this app —
> `docs/api-37-emulator-crash.md` records the evidence and the ruled-out fixes. This is
> unrelated to the renderer above: it is a guest-side bug that CI hits too, which is why
> the matrix stops at API 36 even though targetSdk is 37. **API 37 needs a manual check
> on the Pixel 10 Pro XL before each release.**
> - **API 37 needs the opposite renderer, and SystemUI turned off.** Both `android-37.0` and
> `android-37.1` abort surfaceflinger inside their own gralloc mapper, and init SIGKILLs
> zygote each time. Under `-gpu host` they never boot; under `-gpu swangle_indirect` they
> boot, and disabling SystemUI removes the trigger. `run-e2e.sh` does all of that per level,
> and the local API 37 result is two failures and the two usual skips, not a clean run. CI's
> matrix still stops at 36.
> `docs/api-37-emulator-crash.md` has the matrix and the reasoning. **API 37 needs a manual
> check on the Pixel 10 Pro XL before each release.**
The wording is worth getting right rather than merely correcting, because the original was
not a careless sentence — it was a reasonable inference from three crashes, written down
+344 -48
View File
@@ -3,13 +3,24 @@
# Runs the instrumented suite on a local emulator, on this workstation, for one or more
# API levels.
#
# Usage: tools/local-emulator/run-e2e.sh [API ...] # default: 33 34 35 36
# Usage: tools/local-emulator/run-e2e.sh [API ...] # default: 33 34 35 36 37
#
# GPU_MODE=host renderer to use; see the refusal list below
# API levels are the labels below, not SDK ints: 33-36, plus `37` (= `37.0`) and `37.1`.
#
# GPU_MODE= force one renderer on every level; unset means per-API (gpu_for_api)
# EMULATOR_PORT=5560 console port, so the serial is deterministic
# BOOT_TIMEOUT=300 seconds to wait for sys.boot_completed
# KEEP_AVD=1 do not delete an AVD this script created
#
# EXIT CODE: 0 only if every level was green; 1 if any level failed, wedged or could not be
# set up; 2 if it refused to start at all. **A bare `run-e2e.sh` therefore exits 1 by design.**
# API 37 is in the default list on purpose -- leaving it out is what left the level unlooked-at
# for as long as it was -- and it is permanently two failures short of green, on the emulator's
# own c2.goldfish.h264.decoder rather than on anything this app does. The summary names the two,
# so a third is visibly new, and the last line printed says the same thing. Anything that reads a
# non-zero exit as breakage should name the levels it wants: `run-e2e.sh 33 34 35 36` is the
# sweep that can be green. docs/api-37-emulator-crash.md has the measurements.
#
# WHY THIS EXISTS, AND WHAT IT DELIBERATELY DOES NOT DO
#
# It is a *launcher*, not a second test harness. The diagnostics -- the FAILED-vs-WEDGED
@@ -34,6 +45,14 @@
# those modes was measured crashing. docs/local-emulator.md has the backtrace, the faulting
# page's RW-without-E segment flags, and the full mode matrix.
#
# AND THE ONE THING API 37 NEEDS THAT 33-36 DO NOT: the opposite renderer. On the API 37
# images the guest's Gralloc5 mapper aborts surfaceflinger from RegionSamplingThread
# (`Assertion failed: !rcEnc->featureInfo()->hasReadColorBufferDma`). Under `-gpu host` that
# repeats every few seconds and the device never boots; under ANGLE it fires a handful of
# times and the boot survives. So `host` is required below 37 and forbidden at 37, which is
# why the renderer is chosen per level in gpu_for_api rather than set once.
# docs/api-37-emulator-crash.md has that matrix.
#
# THE OTHER LOCAL-ONLY HAZARD: a physical Pixel is usually plugged into this machine, so
# `adb` is ambiguous in a way it never is on a runner, and an unpinned run would install
# and execute this suite on the phone. Every path below pins the emulator serial.
@@ -65,12 +84,14 @@ export ANDROID_HOME="${ANDROID_HOME:-$HOME/Android/Sdk}"
export ANDROID_SDK_ROOT="$ANDROID_HOME"
export PATH="$ANDROID_HOME/platform-tools:$ANDROID_HOME/emulator:$ANDROID_HOME/cmdline-tools/latest/bin:$PATH"
GPU_MODE="${GPU_MODE:-host}"
# Empty means "let each level pick" -- see gpu_for_api. Setting GPU_MODE forces one renderer
# on every level, which is what you want when measuring a mode, not when running the suite.
GPU_MODE="${GPU_MODE:-}"
EMULATOR_PORT="${EMULATOR_PORT:-5560}"
BOOT_TIMEOUT="${BOOT_TIMEOUT:-300}"
SERIAL="emulator-${EMULATOR_PORT}"
APIS=("$@")
[ "${#APIS[@]}" -eq 0 ] && APIS=(33 34 35 36)
[ "${#APIS[@]}" -eq 0 ] && APIS=(33 34 35 36 37)
# Matches CI. `disk-size: 8G` because the FFmpeg libraries do not fit the default userdata
# partition; `ram-size: 2560M` because the emulator's own floor varies by API level and
@@ -83,21 +104,28 @@ RESULTS_DIR="app/build/outputs/androidTest-results"
LOG_DIR="${TMPDIR:-/tmp}/lmc-local-e2e"
mkdir -p "$LOG_DIR"
# The two things that outlive a level, declared here rather than where they are first
# assigned, because the cleanup trap below can fire before either has been reached.
EMU_PID=""
CREATED_AVDS=()
# ---------------------------------------------------------------- renderer preflight ---
case "$GPU_MODE" in
swiftshader_indirect | auto | off | guest)
echo "REFUSING to launch with -gpu $GPU_MODE."
echo "On this host that resolves to SwiftShader's GLES, whose JIT is denied execheap by"
echo "SELinux; the emulator segfaults (exit 139) before boot. See docs/local-emulator.md."
echo "Working modes: host (default), angle_indirect, swangle_indirect."
exit 2
;;
host | angle_indirect | swangle_indirect) ;;
*)
echo "Unrecognised GPU_MODE '$GPU_MODE'. Known-good: host, angle_indirect, swangle_indirect."
exit 2
;;
esac
if [ -n "$GPU_MODE" ]; then
case "$GPU_MODE" in
swiftshader_indirect | auto | off | guest)
echo "REFUSING to launch with -gpu $GPU_MODE."
echo "On this host that resolves to SwiftShader's GLES, whose JIT is denied execheap by"
echo "SELinux; the emulator segfaults (exit 139) before boot. See docs/local-emulator.md."
echo "Working modes: host, angle_indirect, swangle_indirect."
exit 2
;;
host | angle_indirect | swangle_indirect) ;;
*)
echo "Unrecognised GPU_MODE '$GPU_MODE'. Known-good: host, angle_indirect, swangle_indirect."
exit 2
;;
esac
fi
# A courtesy, not a gate: the boolean being on means SwiftShader would work too, and the
# refusal list above could be relaxed. It is off on a stock Fedora.
@@ -121,14 +149,90 @@ host_forensics() {
journalctl --since "$since" --no-pager 2> /dev/null | grep -E 'avc: .*denied' | tail -10 || echo " (none)"
}
# The API 37 counterpart of host_forensics. `-gpu host` there aborts surfaceflinger in a loop
# and the device never boots; the working renderers abort it a few times and survive. Either way
# the count is the number to look at, and the crash buffer is where it lives -- so print it on
# every 37 level, not only on the failure path, because a level that passed with 40 aborts is
# telling you something a level that passed with 1 is not.
guest_forensics() {
local api="$1" n
case "$api" in 37 | 37.*) ;; *) return 0 ;; esac
n="$(emu_adb logcat -d -b crash 2> /dev/null | grep -c 'hasReadColorBufferDma')"
echo " surfaceflinger hasReadColorBufferDma aborts: ${n:-?} (docs/api-37-emulator-crash.md)"
}
# API label -> system image. API 33-36 are plain integers with a `google_apis` image. API 37
# is not: its SDK directories are dotted minor versions (`android-37.0`, `android-37.1`), there
# is no `android-37`, and from 37.1 onwards Google ships only 16 KB-page (`ps16k`) images for
# x86_64. `37` is accepted as a spelling of `37.0` because that is what people type.
image_pkg_for_api() {
case "$1" in
37 | 37.0) echo "system-images;android-37.0;google_apis;x86_64" ;;
37.1) echo "system-images;android-37.1;google_apis_ps16k;x86_64" ;;
*) echo "system-images;android-$1;google_apis;x86_64" ;;
esac
}
# The renderer requirement is per-API and the two levels want OPPOSITE things, which is why this
# is a function and not a constant.
#
# 33-36: must NOT be SwiftShader GLES (host-side SELinux/execheap segfault) -- `host` is right.
# 37.x: must NOT be the host GL translator. With `-gpu host` the guest's Gralloc5 mapper
# aborts surfaceflinger in a loop and the device never boots; under ANGLE the same
# assertion fires a handful of times and the boot survives it. Measured, not guessed --
# docs/api-37-emulator-crash.md has the matrix.
#
# `swangle_indirect` rather than `angle_indirect` for 37: both boot, and swangle names its
# renderer outright instead of resolving through `auto`'s path.
gpu_for_api() {
if [ -n "$GPU_MODE" ]; then
echo "$GPU_MODE"
return
fi
case "$1" in
37 | 37.*) echo "swangle_indirect" ;;
*) echo "host" ;;
esac
}
# `lmc_e2e_api37.0` would be a legal AVD name but an awkward one to type and to grep for.
# `37` and `37.0` therefore give two AVD names (`lmc_e2e_api37`, `lmc_e2e_api37_0`) for the one
# image. Harmless -- two AVDs off the same system image cost only disk -- and deliberately not
# normalised, so that `run-e2e.sh 37 37.0` does not have both levels fight over one AVD.
avd_for_api() { echo "lmc_e2e_api${1//./_}"; }
# Where avdmanager actually put the AVD. `$HOME/.android/avd` is only the default:
# ANDROID_AVD_HOME, ANDROID_USER_HOME and ANDROID_SDK_HOME each move it, and hardcoding the
# default meant a machine that sets any of them silently ran every level at stock RAM and
# userdata size. Rather than encode a precedence that cannot be verified from here, look in
# every location avdmanager honours and let the existence check pick.
avd_config_path() {
local avd="$1" base cfg
for base in "${ANDROID_AVD_HOME:-}" \
"${ANDROID_USER_HOME:+$ANDROID_USER_HOME/avd}" \
"${ANDROID_SDK_HOME:+$ANDROID_SDK_HOME/.android/avd}" \
"$HOME/.android/avd"; do
[ -n "$base" ] || continue
cfg="$base/${avd}.avd/config.ini"
if [ -f "$cfg" ]; then
echo "$cfg"
return 0
fi
done
return 1
}
ensure_avd() {
local api="$1" avd="$2"
local pkg="system-images;android-${api};google_apis;x86_64"
local pkg
pkg="$(image_pkg_for_api "$api")"
local img_dir="$ANDROID_HOME/system-images/${pkg#system-images;}"
img_dir="${img_dir//;//}"
if avdmanager list avd -c 2> /dev/null | grep -qx "$avd"; then
echo " reusing existing AVD $avd"
else
if [ ! -d "$ANDROID_HOME/system-images/android-${api}/google_apis/x86_64" ]; then
if [ ! -d "$img_dir" ]; then
echo " installing $pkg"
yes | sdkmanager --install "$pkg" > /dev/null 2>&1 || {
echo " FAILED to install $pkg"
@@ -144,18 +248,37 @@ ensure_avd() {
fi
# Written into config.ini rather than passed on the command line, which is how
# reactivecircus/android-emulator-runner applies the same two settings in CI.
local cfg="$HOME/.android/avd/${avd}.avd/config.ini"
sed -i -e '/^disk\.dataPartition\.size=/d' -e '/^hw\.ramSize=/d' "$cfg"
printf 'disk.dataPartition.size=%s\nhw.ramSize=%s\n' "$DISK_SIZE_BYTES" "$RAM_SIZE_MB" >> "$cfg"
# reactivecircus/android-emulator-runner applies the same two settings in CI. On the reuse
# path too, so an AVD left over from an older run gets today's pins.
#
# A level that cannot be pinned FAILS rather than running at the defaults. Unpinned, it
# dies much later with "not enough space", which reads as a device problem -- CI's own
# history is where that lesson comes from -- and nothing points back to a `sed` that
# edited a path this script guessed wrong.
local cfg
if ! cfg="$(avd_config_path "$avd")"; then
echo " FAILED: no config.ini for $avd in any directory avdmanager uses"
echo " (ANDROID_AVD_HOME=${ANDROID_AVD_HOME:-unset}, ANDROID_USER_HOME=${ANDROID_USER_HOME:-unset},"
echo " ANDROID_SDK_HOME=${ANDROID_SDK_HOME:-unset}, HOME=$HOME)"
return 1
fi
if ! sed -i -e '/^disk\.dataPartition\.size=/d' -e '/^hw\.ramSize=/d' "$cfg"; then
echo " FAILED to rewrite $cfg"
return 1
fi
if ! printf 'disk.dataPartition.size=%s\nhw.ramSize=%s\n' \
"$DISK_SIZE_BYTES" "$RAM_SIZE_MB" >> "$cfg"; then
echo " FAILED to write the RAM/disk pins into $cfg"
return 1
fi
}
boot_emulator() {
local avd="$1" api="$2"
local avd="$1" api="$2" gpu="$3"
local boot_log="$LOG_DIR/emulator-api${api}.log"
emulator -avd "$avd" -port "$EMULATOR_PORT" \
-no-window -gpu "$GPU_MODE" -noaudio -no-boot-anim -camera-back none -no-snapshot \
-no-window -gpu "$gpu" -noaudio -no-boot-anim -camera-back none -no-snapshot \
> "$boot_log" 2>&1 &
EMU_PID=$!
@@ -181,6 +304,92 @@ boot_emulator() {
return 1
}
# API 37 only, and the reason API 37 can be run at all.
#
# The abort that breaks these images is reached from SurfaceFlinger's RegionSamplingThread,
# which exists only because SystemUI registers a nav-bar luma-sampling listener. Each abort
# kills surfaceflinger, and init responds by SIGKILLing zygote -- so the whole framework
# restarts underneath the test run, which arrives as `Can't find service: package` and
# `INSTRUMENTATION_ABORTED: System has crashed`. Under the host GL renderer that repeats
# forever; under ANGLE it is roughly one every fifteen seconds, which a five-minute suite does
# not survive either.
#
# Removing the listener removes the whole chain. Measured on android-37.0 under
# swangle_indirect: 10-11 aborts per 150 s idle with SystemUI running, and 0 in 180 s with it
# disabled, framework services up throughout.
#
# THIS IS A DEVIATION, and it is deliberately loud rather than silent. The API 37 leg does not
# run the same device configuration as API 33-36 or as the Pixel. It is defensible only
# because nothing in this suite touches SystemUI -- these are Media3, FFmpeg and WorkManager
# tests -- and because the alternative is no API 37 coverage at all. Anything that ever does
# depend on system UI must not trust this leg. docs/api-37-emulator-crash.md explains why.
#
# The retry loop is not defensive padding: at the moment boot_completed flips, the framework
# may be in one of its restarts and `pm` is simply not published yet. The first attempt at this
# failed exactly that way, with `cmd: Can't find service: package`.
#
# The framework restart at the end is not optional, and finding that out cost a run. By the
# time `sys.boot_completed` flips, SystemUI has already registered its region-sampling listener,
# and `pm disable-user` does not retract a registration that already happened -- it only stops
# the package being started again. So the first attempt disabled SystemUI, reported success, and
# then died exactly as before with `Starting 0 tests` and four more aborts. `stop; start` cycles
# zygote deliberately, and the framework that comes back up does not start SystemUI at all.
disable_region_sampling() {
local api="$1" out i before after ready
case "$api" in 37 | 37.*) ;; *) return 0 ;; esac
out=""
for i in $(seq 1 20); do
out="$(emu_adb shell pm disable-user --user 0 com.android.systemui 2>&1 | tr -d '\r')"
case "$out" in
*"new state: disabled"*)
echo " SystemUI disabled on attempt $i"
break
;;
esac
out=""
sleep 5
done
if [ -z "$out" ]; then
echo " WARNING: could not disable SystemUI after 20 attempts."
echo " Expect INSTRUMENTATION_ABORTED -- docs/api-37-emulator-crash.md"
return 0
fi
echo " restarting the framework so the region-sampling listener goes with it"
emu_adb shell stop > /dev/null 2>&1
emu_adb shell start > /dev/null 2>&1
# There is no property worth waiting on here, and an earlier version of this only looked
# like it was waiting on one: `stop` does not clear sys.boot_completed, so it still reads
# `1` throughout the restart and any loop over it returns at once. The loop below is the
# wait -- and it polls the better thing anyway, since `Can't find service: package` is the
# failure it exists to prevent.
ready=0
for i in $(seq 1 30); do
if emu_adb shell service check package 2> /dev/null | grep -q ': found' \
&& emu_adb shell service check activity 2> /dev/null | grep -q ': found'; then
ready=1
break
fi
sleep 5
done
if [ "$ready" -ne 1 ]; then
echo " WARNING: package and activity services still absent 150 s after the restart."
echo " Expect INSTRUMENTATION_ABORTED -- docs/api-37-emulator-crash.md"
fi
# Prove it worked rather than assume it. Zero new aborts over this window is what makes the
# difference between a run that completes and one that reports `Starting 0 tests`.
before="$(emu_adb logcat -d -b crash 2> /dev/null | grep -c 'hasReadColorBufferDma')"
emu_adb shell 'sleep 45' > /dev/null 2>&1
after="$(emu_adb logcat -d -b crash 2> /dev/null | grep -c 'hasReadColorBufferDma')"
echo " quiet check: $((after - before)) new surfaceflinger aborts in 45 s (want 0)"
if [ "$((after - before))" -ne 0 ]; then
echo " WARNING: region sampling is still live; the run may not survive."
fi
return 0
}
# CI gets this from the action's `disable-animations: true`.
disable_animations() {
local s
@@ -189,17 +398,77 @@ disable_animations() {
done
}
# `${EMU_PID:-0}` used to guard these three calls, and it guarded the wrong thing: EMU_PID
# is *empty*, not unset, if the background launch never produced a job, and `kill` reads pid
# 0 as "the sender's whole process group" -- this script and, on a terminal, everything else
# in the foreground group with it. The `kill -0` wait loop had the same shape and would have
# spent its full grace period testing the group. Nothing to stop is now a return, never a
# guess. (boot_emulator's own `kill -0 "$EMU_PID"` is unguarded and cannot reach that form:
# it runs only after the assignment.)
#
# max_wait is a parameter so the interrupt path need not sit through the full grace period.
stop_emulator() {
local max_wait="${1:-30}" waited=0
[ -n "${EMU_PID:-}" ] || return 0
emu_adb emu kill > /dev/null 2>&1
local waited=0
while kill -0 "${EMU_PID:-0}" 2> /dev/null && [ "$waited" -lt 30 ]; do
while kill -0 "$EMU_PID" 2> /dev/null && [ "$waited" -lt "$max_wait" ]; do
sleep 2
waited=$((waited + 2))
done
kill -9 "${EMU_PID:-0}" 2> /dev/null
wait "${EMU_PID:-0}" 2> /dev/null
kill -9 "$EMU_PID" 2> /dev/null
wait "$EMU_PID" 2> /dev/null
EMU_PID=""
}
delete_created_avds() {
local avd
[ "${KEEP_AVD:-0}" = "1" ] && return 0
for avd in ${CREATED_AVDS[@]+"${CREATED_AVDS[@]}"}; do
# A SIGKILLed emulator does not get to remove its own lock files, and avdmanager can
# refuse over them. Staying silent there would leak the very thing this exists to clean.
avdmanager delete avd -n "$avd" > /dev/null 2>&1 \
|| echo " WARNING: could not delete AVD $avd -- 'avdmanager delete avd -n $avd' by hand"
done
CREATED_AVDS=()
}
# What an interrupted sweep used to leave behind: a headless emulator holding console port
# $EMULATOR_PORT, and an lmc_e2e_apiNN AVD. The next run's `emulator -port` then collides
# with the orphan, and `emu_adb` can resolve to it -- on a workstation that also has the
# Pixel plugged in, exactly the ambiguity the ANDROID_SERIAL pinning exists to prevent. A
# sweep is up to five boots long, so the window for one Ctrl-C is not small.
#
# Idempotent, and called explicitly on the normal path so its output cannot land after the
# summary; the EXIT trap then finds nothing left to do. The emulator logs are deliberately
# NOT removed -- they live in $LOG_DIR and are the only evidence a failed boot leaves.
CLEANED=0
cleanup() {
[ "$CLEANED" = "1" ] && return 0
CLEANED=1
stop_emulator "${1:-30}"
delete_created_avds
}
# 6 s, not 30: Ctrl-C has already reached the emulator through the foreground process group,
# so this is only waiting for it to finish writing, and `kill -9` follows regardless. The
# EXIT trap is disarmed before exiting so the status below is the one that survives.
#
# bash runs a trap only between commands, so this starts when whatever was in the foreground
# returns -- which for Ctrl-C is immediately, because the same interrupt reached that command
# too. `kill -INT` aimed at this script alone waits for the foreground command to finish.
on_signal() {
echo
echo "interrupted (SIG$1) -- stopping the emulator and removing the AVDs this run created"
echo " emulator logs kept in $LOG_DIR"
cleanup 6
trap - EXIT
exit "$2"
}
trap 'on_signal INT 130' INT
trap 'on_signal TERM 143' TERM
trap cleanup EXIT
# The XML is authoritative. The console counter double-counts skips, so a run that reports
# "42 tests" on stdout can be 40 in the report.
#
@@ -236,37 +505,44 @@ PY
}
# ------------------------------------------------------------------------------ main ---
CREATED_AVDS=()
SUMMARY=()
overall=0
for api in "${APIS[@]}"; do
if [ "$api" = "37" ] || [ "$api" = "37.0" ]; then
echo "SKIPPING API $api: the android-37.0 image crash-loops surfaceflinger."
echo " See docs/api-37-emulator-crash.md. Test API 37 on the physical Pixel."
continue
fi
# Whether the red exit is the expected one depends on which level produced it, and only the
# loop knows that -- so it is recorded where `overall` is set rather than guessed from the
# summary afterwards. A note at the end claiming a genuine API 34 failure was "by design"
# would be the same defect it is there to prevent, one layer up.
NON37_RED=0
mark_red() {
overall=1
case "$1" in 37 | 37.*) ;; *) NON37_RED=1 ;; esac
}
avd="lmc_e2e_api${api}"
for api in "${APIS[@]}"; do
avd="$(avd_for_api "$api")"
gpu="$(gpu_for_api "$api")"
started="$(date '+%Y-%m-%d %H:%M:%S')"
echo "=============================================================="
echo "API $api (avd=$avd gpu=$GPU_MODE serial=$SERIAL)"
echo "API $api (avd=$avd gpu=$gpu serial=$SERIAL)"
echo " image: $(image_pkg_for_api "$api")"
echo "=============================================================="
if ! ensure_avd "$api" "$avd"; then
SUMMARY+=("API $api: AVD SETUP FAILED")
overall=1
mark_red "$api"
continue
fi
if ! boot_emulator "$avd" "$api"; then
if ! boot_emulator "$avd" "$api" "$gpu"; then
host_forensics "$started"
guest_forensics "$api"
SUMMARY+=("API $api: BOOT FAILED")
overall=1
mark_red "$api"
stop_emulator
continue
fi
disable_region_sampling "$api"
disable_animations
rm -rf "$RESULTS_DIR"
@@ -281,23 +557,43 @@ for api in "${APIS[@]}"; do
unset ANDROID_SERIAL E2E_EXTRA_GRADLE_ARGS
line="$(summarise_results "$api")"
guest_forensics "$api"
# API 37 is in the default list on purpose, and it is expected to be red. Leaving it out would
# put the level back where this whole exercise found it -- untested and unlooked-at -- but a
# summary that just says "2 failures" with no explanation trains people to ignore the exit
# code. So the row says which two, and a THIRD failure is then obviously new.
case "$api" in
37 | 37.*)
line="$line
expected here: 2 failures, both Media3EngineTest, on c2.goldfish.h264.decoder.
A third is new -- docs/api-37-emulator-crash.md"
;;
esac
if [ "$rc" -ne 0 ]; then
line="$line [gradle exit $rc]"
overall=1
mark_red "$api"
host_forensics "$started"
fi
SUMMARY+=("$line")
stop_emulator
done
if [ "${KEEP_AVD:-0}" != "1" ]; then
for avd in ${CREATED_AVDS[@]+"${CREATED_AVDS[@]}"}; do
avdmanager delete avd -n "$avd" > /dev/null 2>&1
done
fi
cleanup
echo
echo "===================== LOCAL E2E SUMMARY ======================"
printf '%s\n' ${SUMMARY[@]+"${SUMMARY[@]}"}
echo "=============================================================="
# An unexplained red exit trains people to stop reading exit codes, and this one is expected
# whenever API 37 is in the sweep -- which the default list makes the common case. Said here
# rather than only in the docs, because this is where it is actually read. Only when 37.x is
# the ONLY thing that went red: a note calling a real failure elsewhere "by design" would be
# worse than no note at all.
if [ "$overall" -ne 0 ] && [ "$NON37_RED" -eq 0 ]; then
echo "note: the only level that went red is API 37, which exits non-zero by design -- it is"
echo " permanently 2 failures short of green. Confirm its row above shows exactly those"
echo " two and nothing else; docs/api-37-emulator-crash.md says why they are the image."
fi
exit "$overall"