main landed 25 commits while this branch was open, including a third @FailsOnEmulatorApi37 on Media3EngineTest.cancellingARunningExportStopsIt and a batch of new instrumented tests. Every number this branch touches moved with them. Re-derived rather than adjusted, and cross-checked against run 34020234606: the API 34 leg (no filter) reports 68 tests and the API 37 gating leg 64, which is 68 minus main's four markers. With the picker test marked that is five markers, baseline 5, and 63 on the gating leg. The conflict in FailsOnEmulatorApi37.kt is resolved main's way: it had replaced the hardcoded "grows by two" with a reference to the constant, which is the same drift this file exists to prevent and a better fix than the number I put there. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
54 KiB
API 37 on the emulator: a guest gralloc bug that only the host GL renderer triggers
Status: the bug is real and still open upstream, but the previous diagnosis in this file was
wrong about its most important detail. The renderer decides whether API 37 boots, and once it
boots, disabling SystemUI collapses the crash rate far enough to run a suite —
tools/local-emulator/run-e2e.sh 37 gets through the whole instrumented suite and comes back with
2 failures, 0 errors and the two by-design skips (measured 49 / 2 / 0 / 2 at 22c7914, where
the suite was 49 tests — Reading these totals before comparing any total
with another). The crashes do not stop outright, and the two failures are real; both are quantified
below. CI now takes API 37 as two jobs — a gating leg and an advisory one for those two
failures — see So should CI take API 37?.
Last verified: 2026-08-22, emulator 37.1.11.0 (build 15917651), Fedora 44,
against system images android-37.0 rev 6 and android-37.1 rev 8.
The correction
This file previously said, under "What was ruled out":
GPU mode. Both
swiftshader_indirectandhostcrash, with the same assertion and the same frames. The crash is in the gralloc mapper, below the renderer.
That is wrong. The mapper is below the renderer, but whether the mapper's bad path is reached is not. Re-measured on 2026-08-22, seven runs, one variable at a time:
| # | system image | -gpu |
GLES the emulator chose | booted? | surfaceflinger aborts |
|---|---|---|---|---|---|
| r01 | android-37.0 rev 6 |
host |
host (Mesa Iris Xe) | no, 422 s | 71, looping |
| r02 | android-37.1 rev 8 |
host |
host (Mesa Iris Xe) | no, 362 s | 65, looping |
| r03 | android-37.0 rev 6 |
swangle_indirect |
ANGLE | yes, 85 s | 1 |
| r04 | android-37.0 rev 6 |
host + -feature -GLDMA,-GLDMA2,-GLDirectMem |
host | no, 363 s | 57, looping |
| r05 | android-37.0 rev 6 |
angle_indirect |
ANGLE | yes, 112 s | 2 |
| r06 | android-37.1 rev 8 |
swangle_indirect |
ANGLE | yes, 285 s | 23 |
| r07 | android-37.0 rev 6 |
host + -feature -HostComposition |
host | no, wedged adb at 208 s | not readable |
The discriminator is exact across the six runs that reported: a run boots if and only if the
emulator log says something other than gles_mode_selected:host. r07 is excluded on purpose — it
wedged adb at 208 s and is recorded below as inconclusive rather than ruled out, and a row this
page calls inconclusive cannot also be counted as evidence. Excluding it costs nothing: r07 is a
host row, so the discriminator predicts it would not boot, and confirming a prediction with the
one run whose evidence did not come back would add no information either way.
One caveat about how independent those rows are, because the table flatters itself. -gpu angle_indirect (r05) and -gpu swangle_indirect (r03) both logged gles_mode_selected:swangle
and both reported the same adapter, differing only in the Vulkan backend beneath
(vulkan_mode_selected:lavapipe against swiftshader). So they are closer to one GLES path
reached two ways than to two renderers agreeing — note that at API 33–36
docs/local-emulator.md records angle_indirect resolving to ANGLE on
llvmpipe, a genuinely different adapter, which it did not do here. What is 7-for-7 is the
host-GLES-versus-not split, not "two independent renderers both work".
# r01, r02, r04, r07 -- never boots
INFO | emuglConfig_init: vulkan_mode_selected:host gles_mode_selected:host
INFO | Graphics Adapter Android Emulator OpenGL ES Translator (Mesa Intel(R) Iris(R) Xe Graphics (TGL GT2))
# r03, r05, r06 -- boots
INFO | emuglConfig_init: vulkan_mode_selected:swiftshader gles_mode_selected:swangle
INFO | Graphics Adapter Android Emulator OpenGL ES Translator (ANGLE (Google, Vulkan 1.2.0
| (SwiftShader Device (Subzero) (0x0000C0DE)), SwiftShader driver-5.0.0))
Why the wrong claim looked right
It rested on two samples of two different things, and neither of them was ANGLE.
- The local
swiftshader_indirectsample was void. On this workstation every SwiftShader-GLES launch segfaults the host emulator before the guest matters at all — SELinux deniesexecheapto SwiftShader's Reactor JIT. That isdocs/local-emulator.md, and it was not yet understood when this file was written. So "swiftshader_indirectcrashes" was true, for an entirely unrelated reason, and told you nothing about the gralloc assertion. - The CI sample was one
swiftshader_indirectrun on a GPU-lessubuntu-latest, and the local sample was one-gpu hostrun. Two renderers, one measurement each, and the pair written up as "both GPU modes".
angle_indirect and swangle_indirect — the two modes that work — had never been tried on
API 37. Neither had a second system image.
The lesson is the same one docs/local-emulator.md ends on, which makes it worth repeating:
"both backends fail" is a claim about a matrix, and a matrix needs cells, not inference. Two
observations of two different configurations do not establish anything about a third.
What the bug actually is
surfaceflinger aborts inside the emulator's own gralloc mapper:
Executable: /system/bin/surfaceflinger
signal 6 (SIGABRT), code -1 (SI_QUEUE), tid: RegionSampling
Abort message: 'Assertion failed: !rcEnc->featureInfo()->hasReadColorBufferDma'
#03 /vendor/lib64/hw/mapper.ranchu.so GoldfishMapper::readFromHost(cb_handle_t const&) const+543
#04 /vendor/lib64/hw/mapper.ranchu.so GoldfishMapper::GoldfishMapper()::'lambda'(...)::__invoke+704
#05 /system/lib64/libui.so android::Gralloc5Mapper::lock(...)+63
#06 /system/lib64/libui.so android::GraphicBufferMapper::lock(...)+198
#07 /system/lib64/libui.so android::GraphicBuffer::lockAsync(...)+545
#08 /system/lib64/libui.so android::GraphicBuffer::lock(...)+67
#09 /system/bin/surfaceflinger android::RegionSamplingThread::threadMain()+2571
RegionSamplingThread is SystemUI's nav-bar luma sampling. It locks a GraphicBuffer for CPU
read; that routes through the Gralloc5 mapper into GoldfishMapper::readFromHost, which is the
non-DMA readback path and asserts that the host has not negotiated ReadColorBufferDma. The
host always has, so the assert fires whenever that path is taken.
Two facts pin down what "always" means:
- The capability is negotiated regardless of renderer. The evidence is the aborts
themselves: the assertion that fires is
!hasReadColorBufferDma, and it fires under ANGLE (r03/r05/r06) as well as under the host translator — just far less often. That is a direct observation of the guest having negotiated DMA readback under both, and it stands alone. (Supporting only, and weaker than it first looks:ANDROID_EMU_read_color_buffer_dmaappears in exactly one file in the SDK,emulator/lib64/libgfxstream_backend.so, which every-gpumode goes through. A string search establishes where the extension is implemented, not that it is negotiated on every path.) - It is not gated by any feature flag the emulator exposes. See the ruled-out list below.
So the renderer does not decide whether the guest believes DMA readback exists. It decides how
often RegionSamplingThread ends up in readFromHost — which under the host GL translator is
constantly, and under ANGLE is occasionally.
Why one abort takes down the whole device
surfaceflinger is a critical service. When it dies, init kills the framework with it:
08-22 21:40:28.253 I/init: Sending SIGKILL to service 'zygote' (pid 470) process group...
08-22 21:40:28.260 I/init: Service 'zygote' (pid 470) received SIGKILL
Everything above zygote goes with it, which is why the symptoms look nothing like a graphics
bug. Under -gpu host the cycle repeats every five to seven seconds forever and
sys.boot_completed is never set. Under ANGLE the aborts are sparse enough that the boot
usually completes between them — but they do not stop, and each one is a framework restart.
That is the difference between "boots" and "is usable", and it is the reason this is not simply
fixed by changing the renderer. See Can the suite run on it? below.
Environment
| GitHub Actions | Local workstation | |
|---|---|---|
| Host | ubuntu-latest, no GPU |
Fedora 44, Intel Iris Xe (TGL GT2), kernel 7.1.8-200.fc44 |
| Emulator | 37.1.11.0 (build 15917651) |
37.1.11.0 (build 15917651) |
| GPU mode measured | swiftshader_indirect |
host, angle_indirect, swangle_indirect |
Images, both reproducing it:
system-images;android-37.0;google_apis;x86_64 Pkg.Revision=6 ApiLevel=37.0 ExtensionLevel=22
fingerprint google/sdk_gphone64_x86_64/emu64xa:17/CE2A.260420.019/15611780:userdebug/dev-keys
system-images;android-37.1;google_apis_ps16k;x86_64 Pkg.Revision=8 ApiLevel=37.1 ExtensionLevel=23
ro.build.version.codename=REL (a release image, not a preview)
Note the ps16k in the second one — it is not optional, and it is why the 37.1 result is
interpretable. From API 37.1 onward Google ships only 16 KB-page x86_64 images; there is no
plain google_apis variant to pick. sdkmanager --list for 37.1 and 37.2-beta* offers nothing
but google_apis_ps16k and google_apis_playstore_ps16k. That makes page-size alignment a
prerequisite rather than a detail: a .so that is not 16 KB aligned will not load on such a
guest, and the resulting failure looks like an app bug. Checked before the first ps16k boot,
using the same test build.yml applies to release APKs — all 20 libraries in the committed
bin/ffmpeg-kit-next-8.1.1.aar, both ABIs, report 0x4000:
$ for f in jni/*/*.so; do readelf -lW "$f" | awk '$1=="LOAD"{print $NF}' | sort -u; done
0x4000 (x20: libavcodec, libavdevice, libavfilter, libavformat, libavutil,
libc++_shared, libffmpegkit, libffmpegkit_abidetect, libswresample, libswscale
-- arm64-v8a and x86_64)
So when android-37.1 reproduced the abort, that was the gralloc bug and not a page-size
mismatch. image_pkg_for_api in tools/local-emulator/run-e2e.sh encodes the ps16k tag for
37.1; if this ever fails after an FFmpeg rebuild, re-run the alignment check first.
What was ruled out, and how
A newer system image. This file's own revisit trigger was "a new android-37.0 system image
revision ships (this was revision 6)". That trigger was written too narrowly and would never have
fired: android-37.0 is still revision 6, but Google shipped a whole new minor level.
android-37.1 google_apis_ps16k revision 8 — a REL build, not a beta — was installed and
tested (r02, r06) and behaves identically: same assertion, same frames, never boots under
-gpu host, and worse under ANGLE (23 aborts to 37.0's 1). android-37.2-beta3 exists too
but was not needed; two independent images agreeing settles it, and a beta could not be used by
CI anyway.
An ATD image. Still does not exist for API 37. sdkmanager --list offers aosp_atd and
google_atd for API 30 through 36 and nothing above:
system-images;android-36;google_atd;x86_64 | 1 | Google APIs ATD Intel x86_64 Atom System Image
(no android-37 ATD of any kind)
For API 37 the only x86_64 images are google_apis, google_apis_playstore, their ps16k
16 KB-page variants, and Wear OS. Check again when revisiting.
The DMA feature flags. GLDMA alone was ruled out previously; GLDMA2 and GLDirectMem
were not, and the per-image advancedFeatures.ini turns all three on. Disabling all three
together (r04) is accepted by the emulator and changes nothing:
INFO | Feature 'GLDMA' (51) is overridden to 'disabled'
INFO | Feature 'GLDMA2' (52) is overridden to 'disabled'
INFO | Feature 'GLDirectMem' (53) is overridden to 'disabled'
... 57 surfaceflinger aborts, device never boots
Host composition. -feature -HostComposition (r07) was the best remaining guess at what
forces the readback. It did not help; it made things worse, wedging adb entirely at 208 s so the
crash buffer could not even be read. Recorded as inconclusive rather than ruled out, because no
evidence came back from it.
Guest feature negotiation differing from API 36. It does not. The image-level
advancedFeatures.ini for android-37.0 is byte-identical to android-36's except for one
unrelated line:
$ diff android-36/google_apis/x86_64/advancedFeatures.ini android-37.0/google_apis/x86_64/advancedFeatures.ini
+QemuCameraSensorOrientation = on
GLDMA, GLDMA2, GLDirectMem, GrallocSync, HostComposition and YUVCache are on in
both. API 36 boots and passes. So nothing about the host/guest feature handshake changed — the
regression is in the guest's Gralloc5 mapper or in what API 37's RegionSamplingThread asks of
it, not in what the emulator advertises.
Guest memory. Ruled out previously and not revisited; every run above used
hw.ramSize=2560, the same value the E2E matrix pins, and none of them OOMed.
A host-side crash. Not this bug, and worth stating because the other emulator failure on this
workstation is host-side. Every run above left coredumpctl empty and produced zero
avc: denied lines, and the qemu process was still alive at the end of the ones that never
booted (emulator_alive=yes). The host emulator is fine; the guest is not.
Can the suite run on it?
Almost, and less so than it was. tools/local-emulator/run-e2e.sh 37 runs the whole suite
locally. Measured at 22c7914: 49 tests, 2 failures, 0 errors, 2 skipped — 45 passed, the two
Media3EngineTest failures dissected below, and the two assumeTrue skips every level has. It
costs two deviations from how every other level is run, and both are worth understanding before
trusting the leg.
That was the high-water mark. On 2026-08-24 a test that touches system UI joined the suite,
and the level stopped finishing rather than merely failing two —
see below.
Two Media3EngineTest failures is what CI's gating leg expects, because it filters on
notAnnotation; a local run-e2e.sh 37 does not filter and sees more.
Two things about that total before it is compared with anything. It is the size of the suite on
the checkout that ran, not a property of API 37 — app/src/androidTest held 49 @Test methods at
22c7914, and a newer checkout reports its own count; see
Reading these totals. And the Pixel has never run 49: its green run
was 40 / 0 / 0 / 2 at edd6385, the same suite nine tests earlier. What compares across the two
is two failures against none, and the same two skips — not the totals.
The same numbers and the same two test names came back twice, which is real corroboration — but
by two different routes, and only one of them is the harness. The first was driven by hand
(pm disable-user, then several minutes of incidental framework restarts, then e2e-run.sh
directly); the second went through disable_region_sampling's stop; start. The harness path
itself has one green measurement. What would make this routine is a second consecutive
run-e2e.sh 37 whose only failures are the same two.
Booting is not the same as being usable
Changing the renderer gets the device to sys.boot_completed=1, and that is all it gets you. The
aborts do not stop, and each one is a framework restart. A five-minute test run does not survive
that. What it looks like from Gradle:
Shell command failed (1): rm -rf "/sdcard/Android/media/org.libremediaconverter/..."
rm: ...: Transport endpoint is not connected
Starting 0 tests on lmc_e2e_api37(AVD) - 17
Shell command failed (20): am get-current-user
cmd: Can't find service: activity
Device emulator-5572 failed to uninstall test APK org.libremediaconverter.
[cmd: Can't find service: package]
Test run failed to complete. No test results.
onError: commandError=false message=INSTRUMENTATION_ABORTED: System has crashed.
Measured idle rate on android-37.0 under swangle_indirect: 10 aborts in 150 s, then 11 more
in the next 150 s. Steady, not a start-up transient.
The fix is to remove the region-sampling listener, not to survive it
RegionSamplingThread exists only because SystemUI registers a nav-bar luma-sampling listener.
Take SystemUI away and the thread is never started, so the mapper's bad path is never called:
$ adb shell pm disable-user --user 0 com.android.systemui
Package com.android.systemui new state: disabled-user
=== aborts at start of measurement: 36
=== idle 180s with SystemUI disabled ===
=== aborts after: 36 NEW IN WINDOW: 0
--- services still up? ---
activity Service activity: found
package Service package: found
window Service window: found
Zero in 180 s, against 10–11 per 150 s. That is the strongest evidence that region sampling is the dominant trigger, and it is worth recording even by someone who never wants the workaround. It does not establish it as the only trigger: the paragraph below has an abort surviving the disable, and nothing measured here says whether that residue is a second caller of the readback path or a disable that did not fully take.
Do not read that as "the crashes stop", though, because the harness path does not reproduce a clean zero. Its own post-disable check on the run recorded below printed
quiet check: 1 new surfaceflinger aborts in 45 s (want 0)
surfaceflinger hasReadColorBufferDma aborts: 4 (whole run)
So what is reliably achieved is a rate collapse — from roughly one abort every fourteen seconds to one every forty-five — which a 47-second Gradle run survives and a five-minute one might not.
And that restart has never happened — which is how the disable turned out not to work either. Corrected 2026-09-05; this replaces the two paragraphs above rather than qualifying them.
adb shell stop and start are root-only, adbd is not root on a booted emulator, and all three
copies of this logic called them without adb root. On CI both printed Must be root, between
lines that read as if the restart had happened; run-e2e.sh sent them to /dev/null, so its
Must be root was never even visible. Neither number in those logs was an observation either —
the pidof loop breaks when the process is gone and otherwise falls out at its last iteration,
and the old code printed the iteration count either way, so system_server down after ~40 s is
what a stop that did nothing looks like.
Adding adb root made the restart real, and that is what proved the disable ineffective.
api37-debug run 34010167885, disable_system_ui=true:
--- disable round 1 ---
pm attempt 1: Package com.android.systemui new state: disabled-user
restarting the framework
adbd is running as root
system_server down after 2 s
services back after 10 s
NOT DISABLED after the restart -- the package state did not survive
Three rounds of that, then final state: SystemUI STILL ENABLED, and the leg reported
expected: 0, received: 0 — Starting 0 tests, the exact failure this function exists to
prevent.
Bisected locally on android-37.0, which explains the lost state and nothing else:
| arm | sequence | disabled after the restart? |
|---|---|---|
| A | pm disable-user, then stop at once |
no |
| B | pm disable-user, wait 15 s, then stop |
yes |
That is PackageManager's delayed write of package restrictions: the stop kills system_server
before the settings are flushed, and arm A is what CI did. Arm B does not help either, which
is the measurement that matters. With the package verified disabled-user before and after a
further clean restart:
package still disabled? YES
processes:
9275 00:17 system_server
9695 00:14 com.android.systemui <- started 3 s after system_server
CI's own logcat says the same without any restart at all. In the gating leg of run 34006456986,
pm disable-user is accepted at 02:28:37.9 and the package really is in pm list packages -d at
02:29:33 — and SystemUI is started at 02:28:39.5 and again at 02:28:52.3, the second of which
(pid 4275) is alive for the whole instrumentation run.
So pm disable-user --user 0 com.android.systemui does not stop SystemUI starting on this
image, with or without a framework restart, on CI or locally. The premise this section was
built on — "the framework that comes back never starts SystemUI at all" — is false.
Two things follow, pointing in opposite directions.
- The restart is removed rather than repaired, in all three copies. It cost a leg every test
it had and there is nothing for it to buy. What is kept is the 45-second window with zero new
aborts, which was always the part doing the work: in that same run the boot aborts land at
02:28:18 and 02:28:43, and the wait is what puts instrumentation at 02:32:42 — after them
rather than inside one. The
pm disable-usercall is kept too, for a narrower reason than it was written for: every green leg and every number quoted about this row was measured with it applied, and changing the configuration while fixing a flake is not a trade worth making. - The rate collapse recorded above is not evidence of what it says. Both arms of that comparison had SystemUI running. What it measured is a device twelve minutes into its uptime against one that had just booted — a real difference, and a different claim. The quiet gate is still worth having on exactly that reading.
The two deviations, stated plainly
- The renderer is ANGLE, not the host GPU. Shared with nothing else in the matrix — API
33–36 run
-gpu hostlocally, and CI runsswiftshader_indirect. - SystemUI is asked to be disabled, and runs anyway. This was written as the deviation that
mattered — "anything that ever does depend on system UI must not trust this leg" — and the
measurements above say the deviation does not exist: the package is marked
disabled-userandcom.android.systemuiis up for the whole leg regardless. The correction is good news rather than bad. This row is more comparable to API 33–36 and to the Pixel than it has been claiming, not less, and the test that depends on system UI (see the section below) was never running in the exotic configuration this bullet describes. Whatpm disable-userleaves behind is a package-manager flag nothing acts on.
Something does depend on system UI now, and half of it is excluded
Added 2026-08-24, and the first entry on this page that is not a codec.
SafPickerRoundTripTest drives the real system file picker and rotates the display. Both reach
the gralloc mapper — DocumentsUI is another app's windows, and a rotation rebuilds every surface
on screen — and disabling SystemUI does not help. Two reasons now, and only the first was
known when this was written: it removes the idle trigger (RegionSamplingThread's nav-bar luma
sampling) and not this one, and — see the section above — it does not remove SystemUI either.
Measured one method per fresh emulator, android-37.0, swangle_indirect, with the disable
applied and verified quiet — separately, because inferring the second from the first is the
mistake this page's opening correction is about:
| test | result on android-37.0 | hasReadColorBufferDma aborts in the window |
|---|---|---|
thePickedInputSurvivesARealRotation |
fails: INSTRUMENTATION_ABORTED: System has crashed., Expected 1 tests, received 0. The framework dies during it, so the JUnit XML carries a failure with no text at all. |
3 |
pickingAFileThroughTheSystemPickerFillsInTheFileCard |
passes | 4 |
So a rotation, which rebuilds every surface at once, is what the mapper does not survive. Merely
starting DocumentsUI is not. Only the rotation test carries @FailsOnEmulatorApi37; the picker
test runs on the gating leg like anything else.
That last sentence was wrong for twelve days, and the aborts in the table said so
Corrected 2026-09-05. Read the second row again: the picker test passes and takes four
hasReadColorBufferDma aborts with it. This section counted them, put them in the table, and then
drew the conclusion from the pass/fail column alone. The right question is not "does the test
pass" but "does the image survive it", and the answer had been printed in the right-hand column
from the day it was written.
Four gating API 37 runs read logcat-first — 34006456986, 34001744574, 34001377499, and the green
34002313300 — say it without ambiguity. Each carries exactly two aborts before the suite starts
(both surfaceflinger, during boot and the SystemUI disable) and then exactly one during it:
| run | picker test window | the run's only in-suite abort | leg |
|---|---|---|---|
| 34006456986 | 02:33:04.2 → 02:34:46.9, failed | 02:34:46.845 | red, failed: 1 |
| 34001744574 | 00:55:41.4 → 00:57:23.9, failed | 00:57:23.794 | red, failed: 1 |
| 34001377499 | 00:35:53.3 → 00:36:00.6, passed | 00:35:59.662 | red, failed: 0 |
| 34002313300 | 00:58:12.7 → 00:58:19.8, passed | 00:58:19.218 | green |
Every one is system_server, thread TaskSnapshotPer, and every one lands inside that test's
window. Nothing else in the gating set reached the mapper at all. So the picker test is
deterministic in what it does to the image and a coin flip in what the leg reports: 34001377499
passed it and lost the leg from teardown with no failing test to name, and 34002313300 passed it
0.6 s after the abort and went green.
That is #108, which had been filed against this behaviour in August and left open because the
trigger was unknown. The trigger is this test. It now carries @FailsOnEmulatorApi37 too, and the
marker's KDoc had to widen from "does not pass on this image" to "cannot be run on this image" to
say so honestly.
The stack, for the record, is a different caller from either of the two above:
Cmdline: system_server name: TaskSnapshotPer
Abort message: 'Assertion failed: !rcEnc->featureInfo()->hasReadColorBufferDma'
#04 mapper.ranchu.so GoldfishMapper::readFromHost(cb_handle_t const&) const+543
#06 libui.so android::Gralloc5Mapper::lock(...)+63
#10 libandroid_runtime.so android::lockImageFromBuffer(...)+374
#15 framework.jar android.media.ImageReader$SurfaceImage.getPlanes+50
#17 services.jar com.android.server.wm.TaskSnapshotConvertUtil.copyToSwBitmapDirect+56
#28 services.jar com.android.server.wm.SnapshotPersistQueue$StoreWriteQueueItem.writeBuffer+66
#32 services.jar com.android.server.wm.SnapshotPersistQueue$1.run+186
WindowManager writing a task snapshot to disk, which needs the buffer as a software bitmap, which
is the non-DMA readback path. PickActivity is started into the app's own task (Task #11 A=10234:org.libremediaconverter in the logcat), so the snapshot being persisted is that task's,
and the churn at the end of the pick is what schedules it.
There is no shell knob for task snapshots, and that was checked rather than assumed
#108 asks whether TaskSnapshotPersister is suppressible the way the region-sampling listener was.
Probed on a local android-37.0 google_apis x86_64 AVD, 2026-09-05:
getprop | grep -i snapshot # nothing but apexd-snapshotde
settings list global | grep -iE 'snapshot|recents' # empty
device_config list window_manager | grep -i snapshot # empty
cmd window help # no snapshot or screenshot command
dumpsys window | grep -i snapshot # mSnapshotEnabled=true, for Task and Activity
mSnapshotEnabled is real state and there is nothing that sets it from outside. The only
device_config hits anywhere in the tree are aconfig flags — e.g.
windowing_frontend/com.android.window.flags.respect_requested_task_snapshot_resolution — which
tune the snapshot rather than disable it. So the marker is the available answer, not the lazy one.
When the picker test does fail, the abort is the coda and not the cause
Worth separating, because the failure message points the wrong way. In both runs where the test itself went red, it had been broken for 98 seconds before the abort landed. The discriminator is one line, present in both reds and absent from the green:
I/InputDispatcher: No new touched window at (539.0, 525.0) in display 0
(539, 525) is the centre of the fixture's root row — the same coordinates the green run clicks.
The touch reaches no window and is discarded; UiObject2.click() cannot see that and returns
normally. DocumentsUI then logs nothing at all, where the green run logs DocumentStack and
Creating new directory loader 40 ms after its click. The walk waits out its timeout twice for a
fixture it never navigated to, and by the time the back presses start, WindowManager is still
saying no window has focus but ...PickActivity may eventually add a window when it finishes starting up — for another 63 s. All four presses are dropped, DocumentsUI ANRs on
Input dispatching timed out, and only then does the abort fire and make the failure message
read no windows at all.
SafPickerRoundTripTest.forceStopThePicker is the answer to that half: am force-stop goes around
input entirely, so the picker's process can be removed from a task no key press can reach and
pickTheFixture's whole-picker retry — which exists for exactly this — becomes reachable again.
That is a fix to the test on every level, not to API 37.
It was made to bite before it was believed. On a local API 36 emulator, with the walk cut short
so the picker is left open and in front and with device.pressBack() removed, so that nothing but
the force-stop can close it:
| result | |
|---|---|
with forceStopThePicker() |
passes — ActivityManager: Force stopping com.google.android.documentsui ... from pid 5334, Killing 5269:com.google.android.documentsui (adj 0), a second PickActivity opens, the retry completes the pick |
| with the one call removed | fails — the system picker would not close: after 4 back presses ... com.google.android.documentsui is in front, which is the API 37 failure verbatim |
The unmutated class passes on that emulator either way, which is the point of running the mutation at all: the recovery path is unreachable on a healthy device, so a green suite says nothing about it.
The correction that produced that table
The first version of this section said both tests failed, and put the marker on the class. The
picker test had indeed failed at API 37 — with androidx.test.uiautomator.StaleObjectException,
which looked like a framework restart invalidating an accessibility node, because that is exactly
what it looks like.
It was the test's own bug. UiObject2 caches the AccessibilityNodeInfo it was found with, and
DocumentsUI is still settling when a node first appears; the handle went stale before click().
CI then reproduced it deterministically at API 33, 34 and 35 — every cold runner emulator, not
intermittently — which is what made it obviously not an API 37 property. It had passed locally
only because the emulator was warm.
The lesson is worth more than the measurement: an annotation is a claim about an image, and a broken test makes every image look broken. Re-measure after fixing a test before deciding what the platform did. Both the abort and the stale node produce "the run fell over", and only one of them was the image.
Two consequences worth stating rather than discovering
run-e2e.sh 37applies no annotation filter, unlike CI, so a local API 37 run includes the rotation test and therefore does not finish: its totals come back short and which later tests ran is arbitrary. The summary row says so.- The advisory job is still named
E2E API 37 Media3 hardware transcode (advisory)and now carries a test that is neither Media3 nor a transcode. Renaming a check is a branch-protection change and was deliberately not made in the same PR; the name is stale, the behaviour is correct.
The two remaining failures are the same bug, one layer down
org.libremediaconverter.convert.Media3EngineTest > runsFromAThreadWithNoLooper FAILED
org.libremediaconverter.convert.Media3EngineTest > transcodesH264ToH265AndReportsProgress FAILED
androidx.media3.transformer.ExportException: Codec exception:
CodecInfo{type=VideoDecoder, ..., mime=video/avc, name=c2.goldfish.h264.decoder}
at androidx.media3.transformer.DefaultCodec.maybeDequeueOutputBuffer(DefaultCodec.java:398)
Caused by: android.media.MediaCodec$CodecException:
at android.media.MediaCodec.native_dequeueOutputBuffer(Native Method)
Three measurements say this is the emulator image and not this app, and not the software renderer. A fourth bullet offers a mechanism, and is inference rather than measurement:
-
Control at API 35 under the identical renderer.
GPU_MODE=swangle_indirect tools/local-emulator/run-e2e.sh 35→ 49 / 0 / 0 / 2 at22c7914, green.c2.goldfish.h264.decoderis perfectly happy under ANGLE one API level down, so the renderer is not what breaks it. -
Real API 37 hardware passes, see below. There is no
c2.goldfish.*codec on a Pixel. -
API 36 against API 37 on CI, back to back, everything else held. Same two tests, same
-gpu swiftshader_indirect, same SystemUI-disable path —pm disable-user,stop, wait forsystem_serverto actually be gone,start, then verify againstpm list packages -d. Both runs were narrowed to the two failing tests:-Pandroid.testInstrumentationRunnerArguments.class=\ org.libremediaconverter.convert.Media3EngineTest#transcodesH264ToH265AndReportsProgress,\ org.libremediaconverter.convert.Media3EngineTest#runsFromAThreadWithNoLooperand the filter is confirmed three independent ways:
tests="2"in the XML,Expected 2 testsin the abort message, andrun started: 2 testsin the guest logcat.run api result XML 32660148155 37.0 tests="2" failures="2" errors="0" skipped="0"32660152961 36 tests="2" failures="0" errors="0" skipped="0" time="4.603"API 37 fails with the signature above —
name=c2.goldfish.h264.decoder,MediaCodec$CodecExceptionatdequeueOutputBuffer(MediaCodec.java:4274). API 36 passes both in 4.603 s, andc2.goldfish.h264.decoderis in its logcat too (44 mentions), so the two runs are not being served by different decoder names. What this falsifies is "the stripped configuration is what breaks these tests" — a reading none of the other measurements addresses, because they all compare against a device that still had SystemUI. Here SystemUI is absent and the framework has been restarted on both sides, and the healthy image is green anyway.Two things it does not control, which is why it narrows the claim rather than closing it:
- The restarts were not performed under equal conditions. API 36 did its
stop/startwithdma_aborts=0; API 37's did the same restart with two aborts already logged. "A framework restart performed while the abort loop is running" therefore remains uncontrolled. - The images differ on the encoder side. These tests transcode H.264 → H.265. The API 37
logcat carries
c2.goldfish.hevc.decoder(16 mentions in the control run) where API 36 carriesc2.android.hevc.encoder(32). The pipeline is not identical end to end, which is a second reason "the image ships a broken h264 decoder" is the wrong shape of claim: what is measured is that these two tests fail on the API 37 image, pass at API 36 under the same renderer and the same disable path, and pass at 33–36 without needing that path at all — because nothing below 37 has the bug it works around.
- The restarts were not performed under equal conditions. API 36 did its
-
The failing call is
dequeueOutputBufferon the goldfish decoder — the emulator's own codec, which likeRegionSamplingThreadgets its frames out of a host-side colour buffer. Same readback machinery, one layer down. This is inference rather than a measurement, and is flagged as such; what is measured is the first three bullets.
Do not try -feature -HardwareDecoder. It is the obvious next idea and it is much worse:
forcing the guest onto software decoders took the run from 2 failures to 46, across
RemuxTest, ForcedFailureTest, HardwareFallbackTest and UnopenableUriTest as well. The
suite depends on those decoders existing.
The intact-SystemUI counterfactual cannot be measured on CI
The control the block above still lacks is the obvious one: run those same two tests at API 37 with SystemUI left running. Passing would put the failure on the disable rather than on the image; failing on the decoder would make the decoder attribution direct instead of inferred.
Seven dispatches of api37-debug.yml, zero verdicts. Not bad luck — a mechanism, which is why
this is written down rather than left as a gap for the next person to spend seven runs on:
| arm | run | result XML | what actually happened |
|---|---|---|---|
| E1 | 32660528355 | tests="1" failures="1", <failure> body empty |
Expected 2 tests, received 0. INSTRUMENTATION_ABORTED: System has crashed. |
| E2 | 32660533845 | tests="0" |
never installed: Failed to commit install session ... Failure calling service package: Broken pipe (32) |
| E3 | 32660539259 | tests="0" |
Test run failed to complete. No test results. |
| E4 | 32661117237 | tests="2" failures="2" |
both failed in @Before, never reached MediaCodec |
| E5 | 32661121972 | tests="2" failures="2" |
same |
| S1 | 32661127224 | tests="1" failures="1" |
same, single-test arm |
| S2 | 32661132024 | tests="1" failures="1" |
same |
While the framework is crash-looping, the guest cannot reliably create per-user private
directories. An app installed during the loop has no cache directory — and Media3EngineTest
copies its H.264 fixture into context.cacheDir in @Before, so it dies there, before any
MediaCodec exists:
W/ContextImpl( 8216): Failed to ensure /data/user/0/org.libremediaconverter/cache
I/TestRunner( 8216): run started: 1 tests
E/TestRunner( 8216): failed: transcodesH264ToH265AndReportsProgress(...)
E/TestRunner( 8216): java.io.FileNotFoundException:
/data/user/0/org.libremediaconverter/cache/sample_h264.mp4: open failed: ENOENT
at org.libremediaconverter.convert.Media3EngineTest.setUp(Media3EngineTest.kt:47)
Not app-specific: com.google.android.googlesdksetup and com.google.android.apps.nexuslauncher
hit the same Failed to ensure /data/user/0/<pkg>/cache in the same logcats.
The result XML masks this, and reading only the report gets you the wrong bug. What E4, E5, S1 and S2 report is
<failure>kotlin.UninitializedPropertyAccessException: lateinit property output has not been initialized
at org.libremediaconverter.convert.Media3EngineTest.tearDown(Media3EngineTest.kt:56)
— tearDown failing because setUp threw before it assigned output. That looks like a
teardown defect in this repository and is not one; the cause is only in the guest logcat.
So the obstacle is structural: install, data-directory creation and instrumentation start-up do
not fit between framework kills, and four of the seven runs show the directory creation itself is
broken during the loop. More dispatches of this shape would repeat these outcomes. The
counterfactual is still open on the Pixel 10 Pro XL, the one API 37 device here that is not an
emulator — but a Pixel has no c2.goldfish.* codec at all, so it answers "does the app work at
API 37", not "is that codec broken".
Abort cadence, corrected
.github/workflows/api37-debug.yml carried "roughly every 20 s" for the kill cycle in its own
comments. That number was the watchdog's sampling interval, not the cadence, and the two got
conflated. Measured across the six runs whose crash buffer could be read — r07 wedged adb
before one could be taken, so it contributes no gaps — successive hasReadColorBufferDma aborts
run 20 s to 90 s, median 60–70 s — three to five aborts in a four-minute window.
Slower than assumed, and still not slow enough: install, data-directory creation and
instrumentation start-up do not fit inside one gap.
sys.boot_completed held at 1 throughout every one of those test windows. The device reports
itself booted while zygote is being killed under it, which is why no boot-state check catches
this and why stop/start waits must poll pidof system_server and service check instead
(see disable_region_sampling in tools/local-emulator/run-e2e.sh).
So should CI take API 37?
Yes, as two jobs: a gating E2E API 37 and an advisory leg carrying the two tests that do not
pass. That reverses the answer this section gave, and the reversal is measured rather than
argued — two of its three reasons were claims about CI, and CI had never been measured. The
instrument that measured it is .github/workflows/api37-debug.yml,
dispatch-only, a copy of the E2E job with the matrix replaced by inputs.
Every run below is ubuntu-latest, KVM on, pixel_6, x86_64, disk 8G, RAM 2560M, emulator
37.1.11.0 build 15917651 — the same emulator build the local investigation used. Every API
37 row is system-images;android-37.0;google_apis;x86_64; c2 is the API 36 control and runs
that level's own image, which is the whole point of it.
| # | run | api | -gpu |
SystemUI | suite | verdict |
|---|---|---|---|---|---|---|
| c1 | 32644947334 | 37.0 | swiftshader_indirect | running | Starting 0 tests |
FAIL |
| c2 | 32644965828 | 36 | swiftshader_indirect | running | 57 tests, BUILD SUCCESSFUL | green control |
| c3 | 32644970240 | 37.0 | swangle_indirect | running | Starting 0 tests |
FAIL |
| c5 | 32645543238 | 37.0 | swangle_indirect | disabled | 57 / 2 / 0 / 2 | suite ran |
| c6 | 32646029143 | 37.0 | swiftshader_indirect | one-shot disable, did not hold | Starting 0 tests |
FAIL |
| c8 | 32646611485 | 37.0 | swiftshader_indirect | disabled, verified | 57 / 2 / 0 / 2 | suite ran |
| c9 | 32646615706 | 37.0 | swiftshader_indirect | disabled, verified | 57 / 2 / 0 / 2 | suite ran |
| c10 | 32646619472 | 37.0 | swiftshader_indirect | disabled, verified | 57 / 2 / 0 / 2 | suite ran |
| c11 | 32647138060 | 37.0 | swiftshader_indirect | disabled, verified | 57 / 2 / 0 / 2 | suite ran |
57 is that checkout's own @Test count at acc71bc, so those are whole-suite runs and not
truncated ones — see Reading these totals. Taking the three old reasons
in turn:
- "Nothing says a runner would be stable" — measured, and it is.
-gpu swiftshader_indirecton a GPU-less runner resolves togles_mode_selected:swiftshader, a third renderer that locally never survives to say anything (Fedora's SELinux deniesexecheapto SwiftShader's Reactor JIT — seelocal-emulator.md). It bootsandroid-37.0in about 60 s. The local discriminator — fatal iffgles_mode_selected:host— holds, and a runner with no GPU can never select host, so CI was never in the fatal class. Switching CI's-gpuchanges nothing either way: c1 and c3 both fail with SystemUI up, under swiftshader and swangle respectively, and c5 and c8–c11 show the suite running under either once SystemUI is gone. - "Bespoke device surgery" — still true, and now a written caveat rather than a reason to skip
the level. It is one env flag,
E2E_DISABLE_SYSTEM_UI, read by.github/scripts/e2e-run.shand unset on every other leg. What it costs is stated where it can be read from the failing check: the API 37 row runs a device configuration no other leg and no Pixel run uses. What makes it dependable is verification, not repetition — c6 is the counter-case, a one-shotpm disable-userthat reportednew state: disabled-userand then started SystemUI eight more times. The verified form is 4/4; the unverified form was 3/4. - "Permanently red or permanently allow-listed" — this was the real objection, and it is the
one the split answers. The two failures are marked
@FailsOnEmulatorApi37inapp/src/androidTest. The gating job runsnotAnnotationon that marker and must be green; the advisory job runsannotationon the same marker, reports, and never blocks. One marker rather than two lists, so a test cannot silently end up in neither job — which would read as green.
The cost is about three minutes on the API 37 leg — the stop/start plus a 45 s quiet window,
and another round when the first does not verify. Measured wall clock for the whole job, boot
included: 6–7 minutes at API 37 against ~6 at API 36.
Two things this does not buy. The advisory job is expected red, so a third failure there is
the signal and the run's logcat is the only thing that distinguishes it — which is why that job
uploads diagnostics unconditionally. And a green E2E API 37 still does not replace the release
check on the Pixel: the emulator leg runs without SystemUI, and the Pixel does not.
The local story changed at the same time and independently: API 37 is no longer a level nobody can look at. A regression that shows up at 37 and not at 36 can be reproduced on this workstation in about four minutes.
Verified on real API 37 hardware
Unchanged and still true. On 2026-08-21 the whole instrumented suite ran green on a physical device:
Device: Pixel 10 Pro XL (mustang), arm64-v8a
Build: google/mustang/mustang:17/CP2A.260805.005/15828068:user/release-keys
API: 37 (Android 17, codename REL -- a release build, not a preview)
./gradlew :app:connectedDebugAndroidTest -PabiFilters=arm64-v8a
40 tests, 0 failures, 0 errors, 2 skipped BUILD SUCCESSFUL
That "40" is a measurement of the tree it ran on, not a baseline for today, and it is not a
contradiction of the totals in docs/local-emulator.md either.
Reading these totals
Every total in this file and in docs/local-emulator.md is the size of
app/src/androidTest on the checkout that produced it, and nothing else. The reported total has
equalled that checkout's @Test count everywhere it has been checked:
| checkout | @Test methods |
total the run reported |
|---|---|---|
edd6385 |
40 | 40 — the Pixel run above |
22c7914 |
49 | 49 — the four local levels, and API 37 |
18c53a3 |
57 | not run |
So the number to expect is not written down here. It is derived from the checkout in front of you, which is the only thing that cannot go stale:
grep -rho '@Test' app/src/androidTest | wc -l
Before each release, run the suite on the Pixel 10 Pro XL and expect that many tests, 0 failures, 0 errors, 2 skipped. The failure, error and skip counts are the invariant; the total is not. A total that disagrees with your own checkout's count is the signal — an old checkout, a stale build, or tests that never ran — and it is worth stopping on either way.
The two skips are RealMediaBenchmark.hardwareVersusSoftwareOnRealVideo and
av1InputRoutesAccordingToDeviceDecodeSupport, which assumeTrue their sample files are present
and skip when they are not. That is by design and unrelated to API level.
One harmless warning appears during the run and can be ignored:
No UID for androidx.test.services in user 0, from an appops call the test services package
makes before it is fully registered.
Reproducing it
Both halves, so the renderer claim can be checked rather than taken on trust:
export ANDROID_HOME="$HOME/Android/Sdk"
export PATH="$ANDROID_HOME/platform-tools:$ANDROID_HOME/emulator:$ANDROID_HOME/cmdline-tools/latest/bin:$PATH"
sdkmanager --install "system-images;android-37.0;google_apis;x86_64"
echo no | avdmanager create avd -n api37_repro \
-k "system-images;android-37.0;google_apis;x86_64" -d pixel_6 --force
# never boots -- surfaceflinger aborts every ~6 s, forever
emulator -avd api37_repro -no-window -gpu host \
-noaudio -no-boot-anim -camera-back none -no-snapshot &
# boots in ~85 s, having aborted once or twice on the way
emulator -avd api37_repro -no-window -gpu swangle_indirect \
-noaudio -no-boot-anim -camera-back none -no-snapshot &
Count the aborts either way:
adb -s emulator-5554 logcat -d -b crash | grep -c hasReadColorBufferDma
Under -gpu host, sys.boot_completed never reaches 1, pgrep -f system_server stays empty,
and keystore2's watchdog logs await_boot_completed ... Overdue indefinitely.
tools/local-emulator/run-e2e.sh 37 does all of this, with the working renderer picked
automatically — see gpu_for_api in that file.
Filing this upstream
Not yet filed. The report is stronger than it was, because the renderer dependency narrows it:
- Go to https://issuetracker.google.com/, Report an issue, and pick the Android emulator component (search the component picker for "Emulator"; Android Studio's Help → Submit Feedback opens the same tracker with it preselected).
- Title it for the mechanism:
surfaceflinger aborts in GoldfishMapper::readFromHost (hasReadColorBufferDma) on android-37.0 and android-37.1 x86_64 -- fatal under -gpu host, intermittent under ANGLE. - Paste the assertion and backtrace, the environment block, and the seven-row matrix. The matrix is the valuable part: it shows the abort is not renderer-specific but its frequency is, which points at the readback path rather than at any one GL implementation.
- State that it reproduces on two independent system images (
37.0rev 6 and37.1rev 8) and on two unrelated hosts, and that-feature -GLDMA,-GLDMA2,-GLDirectMemdoes not suppress it. - Attach:
- the guest tombstone, via
adb pull /data/tombstones(or thepbtombstoneoutput the crash log names) adb logcat -d -b crash > crash.txt- the emulator's own stdout log (
-verbose -debug all, redirected) - the AVD's
config.ini - a link to a failing CI job, which shows it on hardware you do not control: https://github.com/JMR-dev/LibreMediaConverter/actions/runs/32545625459/job/96963461184
- the guest tombstone, via
Record the issue number here once filed.
When to revisit
The old trigger list named "a new android-37.0 revision", which is why nothing ever fired even
though a new API level shipped. Watch for these instead:
-
any new API 37.x system image, not just a new revision of
37.0—37.1rev 8 and37.2-beta*already exist, and more will. Test with-gpu host: if it boots, the guest mapper is fixed. -
an ATD image for API 37. Still none as of 2026-08-22. ATD images ship without SystemUI, which is what drives
RegionSamplingThread, so one would very likely sidestep the bug entirely. Try it before anything else here. -
the upstream issue being marked fixed.
-
E2E API 37 Media3 hardware transcode (advisory)going green. Nothing announces this: the job iscontinue-on-error, so it fixing itself looks exactly like a check nobody reads quietly ceasing to be red. It is listed here because that makes it the least likely of these triggers to be noticed, not the most. When it happens, delete@FailsOnEmulatorApi37from everything carrying it rather than deleting the job — the gating leg picks them back up on its own, and the advisory job then runs nothing and can go.It is not two tests any more. As of 2026-08-24 the marker is on
Media3EngineTest's two methods and onSafPickerRoundTripTestas a class, and the two groups fail for unrelated reasons — a codec and the gralloc mapper. They can go green independently, so check both before concluding the marker is done; and the job's name still says "Media3 hardware transcode", which half of what it runs is not.
Correction owed to CLAUDE.md
CLAUDE.md currently says:
- The API 37 image is broken.
android-37.0crash-loops surfaceflinger inside its own gralloc mapper, so every test fails there regardless of this app.docs/api-37-emulator-crash.mdrecords the evidence and the ruled-out fixes; CI's matrix therefore stops at API 36 even though targetSdk is 37.
The first sentence is right, and now under-specified in one direction and over-specified in the
other: it is not only android-37.0 (it is 37.1 too), and it does not crash-loop under every
renderer. The last clause is now simply false: CI's matrix does not stop at API 36 any more.
Proposed replacement, offered for review rather than applied here — CLAUDE.md is left alone
deliberately, because several branches touch it:
- The API 37 images crash-loop surfaceflinger under the host GL renderer. Both
android-37.0andandroid-37.1abort inside their own gralloc mapper (RegionSamplingThread→GoldfishMapper::readFromHost), and when surfaceflinger dies init SIGKILLs zygote, so the framework restarts under the test run. Under-gpu hostit never boots at all; under-gpu swangle_indirectit boots and the aborts merely become intermittent.docs/api-37-emulator-crash.mdhas the seven-run matrix and the ruled-out list, andtools/local-emulator/run-e2e.shpicks the working renderer per API level. CI takes API 37 as two jobs: a gatingE2E API 37that disables SystemUI first, and an advisory leg carrying the two@FailsOnEmulatorApi37tests. The gating leg therefore runs a device configuration nothing else does. API 37 still needs a manual check on the Pixel 10 Pro XL before each release — it is the only API 37 run with SystemUI intact.