Gating API 37 leg aborts in system_server TaskSnapshotPersister after every test passes #108
Closed
opened 2026-08-25 20:10:42 +00:00 by JMR-dev
·
7 comments
No Branch/Tag Specified
main
fix/102-picker-back-press-overshoot
fix/268-saf-picker-determinism
feat/ogg-vorbis-libvorbis
feat/expedited-conversion-work
fix/gate-cache-in-worktrees
chore/gate-runs-shellcheck
test/publish-delete-arm-real-provider
docs/e8-instrumented-coverage
docs/api37-carrier-count-drift
docs/e7-second-constraint
test/publish-to-a-real-saf-destination
fix/api37-report-match-line
fix/api37-task-snapshot-crash
test/join-failure-message-on-device
docs/e2e-read-findings-e7
test/cancelling-a-running-export
test/reattach-to-a-running-job
test/content-uri-reaches-ffmpeg
fix/launcher-wiring-waits-for-the-pick
fix/cancel-tests-need-a-slower-encode
test/cancelling-a-running-join
test/cancelling-a-running-session
test/notification-cancel-action
test/ffmpeg-progress-is-observed
test/fallback-asserts-the-path
test/flac-and-opus-assert-their-format
docs/e2e-read-findings
docs/wave4-coverage-numbers
fix/injectable-startup-sweep-scope
test/session-outcome-seam
test/launcher-callback-identity
test/theme-follows-system-dark
test/audio-drop-arm
fix/rotation-waits-for-recreation
fix/convert-guards-on-ready
test/retry-save-mime
test/hardware-progress-reaches-workmanager
test/ffprobe-mapping-seam
test/device-codec-enumeration-seam
test/unknown-container-row
test/null-message-fallbacks
test/cancel-reaches-workmanager
docs/coverage-wave3-recovery
test/concat-engine-seam
docs/coverage-wave3
test/mediaprobe-merge-seam
test/adaptive-shell-wiring
test/aac-audio-args
test/notification-progress-text
test/media3-muxer-guard
test/hardware-fallback-and-cancellation
test/unprobeable-join-clip
test/one-branch-outcomes
test/foreground-type-regimes
fix/bound-wedge-diagnostics
docs/coverage-wave2
test/screen-wiring
test/viewmodel-setters
test/join-state-mapping
test/conversion-state-mapping
test/dedupe-user-messages
fix/restore-stack-merges
test/refused-jobs
test/concatworker-failure-arms
test/container-capabilities-audio
test/readspec-enum-fallbacks
test/outputpublisher-seams
test/mediaprobe-track-seam
test/outputpublisher-partial-branches
test/fake-provider-scaffolding
docs/coverage-read-findings
chore/gitignore-kotlin
test/bound-the-hangs
docs/coverage-remeasure
ci/baseline-counter-precision
fix/invalid-suggestion-chip
ci/wedged-leg-report
fix/reattachment-overwrites-pick
test/theme-live-branches
fix/failed-save-retry
fix/empty-composition-crash
ci/advisory-failure-report
docs/seven-run-counts
test/release-permission-guard
ci/build-workflow-permissions
docs/api37-point-release
docs/benchmark-populate-path
fix/dead-assertion-probe-test
ci/actionlint
test/device-codecs-encode-consequence
fix/sdkmanager-pipefail
docs/readme-restart-claim
fix/saf-picker-root-discovery
fix/probe-dispatcher-seam
test/media3engine-mime-tables
test/mediaprobe-pure-helpers
fix/codec-vocabulary-drift
docs/robolectric-rationale-correction
docs/api37-advisory-counts
test/r38-8-saf-e2e
test/r38-7-join-states
test/r38-6-conversion-states
test/r38-5-state-seam
fix/jacoco-robolectric-coverage
docs/instrumented-tests-correction
test/r38-2-filecard
test/r38-4-advanced-picker
test/r38-3-pickers
tools/file-issue-script
tools/api-37-emulator
fix/review-app-gaps
docs/review-corrections
No results found.
Labels
Clear labels
above-cut
accessibility
backlog
bug
confirmed
documentation
duplicate
enhancement
good first issue
help wanted
invalid
plausible
question
sev:high
sev:low
sev:medium
wontfix
Worked autonomously overnight: local, JVM-verifiable, no product decision
Barrier affecting people with disabilities
Held for manual review: product/UX call, CI/workflow, hardware, or unverifiable here
Something isn't working
Reviewer demonstrated the defect
Improvements or additions to documentation
This issue or pull request already exists
New feature or request
Good for newcomers
Extra attention is needed
This doesn't seem right
Reviewer could not fully demonstrate it; treat as unproven
Further information is requested
High severity
Low severity
Medium severity
This will not be worked on
Milestone
No items
No Milestone
Projects
Clear projects
No projects
Notifications
Due Date
No due date set.
Dependencies
No dependencies set.
Reference: JMR-dev/LibreMediaConverter#108
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
The gating API 37 leg has failed three consecutive PRs today with zero test failures. This is a third caller of the abort
docs/api-37-emulator-crash.mdalready documents, and the SystemUI mitigation does not cover it.What the leg reports
E2E API 37(gating, not the advisory job), on #99, #104 and #106 — diffs of a KDoc comment, aworkflow step and a
permissions:block, none of which can reach the E2E matrix:Every test passes. The job fails anyway —
am get-current-userreturning 20 meanssystem_serverhas stopped answering.The crash, from the job's own native-crash tail
Why this is new rather than already covered
docs/api-37-emulator-crash.mddocuments this assertion twice, and neither is this:surfaceflingerRegionSamplingE2E_DISABLE_SYSTEM_UI=1removes the listener@FailsOnEmulatorApi37on the two testssystem_serverTaskSnapshotPersisterThe doc's own words for the second entry apply exactly here: "disabling SystemUI does not help,
because it removes the idle trigger (RegionSamplingThread's nav-bar luma sampling) and not this
one."
TaskSnapshotPersisteris a third one — WindowManager writing task snapshots, insystem_serverrather thansurfaceflinger, on a path the disable never touched.The gating row does set
disable-system-ui: "1", and it is working as documented. It simplydoes not reach this caller.
Why it matters more than the other two
This one aborts after the suite has passed, so it converts a green run into a red check with no
failing test to point at. A reader seeing
E2E API 37red will look for a broken test and find 56passes — which is the most confusing possible failure, and it has now happened three times in a row
on unrelated PRs.
Frequency
Three consecutive gating-leg failures (#99, #104, #106) in roughly ninety minutes. That leg was
stable at 56/0 for the whole of the preceding session. This is a change in behaviour, not a
long-standing flake — worth saying, because the honest response to a rare flake and to a
regression are different.
What to establish first
these are load-dependent and invisible on this workstation.
TaskSnapshotPersistersuppressible the way the region-sampling listener was — awmsetting, or a device config — without disabling more than the leg already disables?record it the way the other two are recorded. Neither should be taken quietly — this row
gates, and #56 added it deliberately after measuring that 55 of 57 tests do pass there.
Correcting the frequency claim in this ticket. The crash evidence stands; the "regression" framing does not.
I wrote that the gating leg "was stable at 56/0 for the whole of the preceding session" and that this is "a change in behaviour, not a long-standing flake." Measured, that is wrong.
Across the last 24
status_check.ymlruns, theE2E API 37gating job:Successes are interleaved throughout, including at 15:13 and 15:14 — between the failures I cited as consecutive. What I actually saw was three PRs holding a failed run at one moment, which is not the same as three consecutive runs, and I generalised from the first to the second.
This is the same error I corrected on #49 this morning — reading a cluster as a rate. Twice in one day, so it is worth naming the shape rather than just the instance: when several PRs are queued behind each other, they sample the same window, and a moderate flake will show up in all of them at once and look like a step change.
What stands, unchanged
The crash evidence is from an actual native-crash dump and is not affected:
TaskSnapshotPersisterinsystem_serveris a third caller of an assertion this repo hasdocumented twice, and
E2E_DISABLE_SYSTEM_UI=1genuinely does not cover it — the doc's ownsentence about the idle trigger applies. That part needed no frequency argument.
So does the reason it is confusing: it aborts after the suite passes, turning a green run into a
red check with 56 passes and nothing to point at.
What changes about the ticket
annoyance; one that presents as "red with zero failing tests" costs a reader real time. The
argument for fixing it is legibility.
question is whether
TaskSnapshotPersisteris suppressible the way the region-sampling listenerwas, and if not, whether a teardown-only abort should fail the leg at all.
have introduced the activity churn that triggers task snapshots. The failure rate does not step at
#80's merge, so that is not supported.
Re-measured per leg-attempt: 3.75%, not ~15%
The earlier figure on this ticket counted whole runs. A leg that is re-run to green disappears from
that count, so the denominator was wrong in one direction and the numerator in the other. Counting
leg-attempts for
E2E API 37since 2026-08-24:TaskSnapshotPersister/system_serverSIGABRT trace."aborts after every test passes". That is 3/80 ≈ 3.75%.
pickingAFileThroughTheSystemPickerFillsInTheFileCardpickingAFileThroughTheSystemPickerFillsInTheFileCardroutesAFastMp4JobByDeviceCapabilitypickingAFileThroughTheSystemPickerFillsInTheFileCardpickingAFileThroughTheSystemPickerFillsInTheFileCardpickingAFileThroughTheSystemPickerFillsInTheFileCardIn 5 of the 8, a test failed as well. So the abort is not purely a post-run tidy-up crash — when
it happens it can take a test with it. That is the strongest single piece of evidence for #102's
shared-cause reading, and it is why the two tickets should not be closed independently.
Three further failures show neither an abort trace nor a test failure, so the trace is not always
captured; treat 3.75% as a floor for the pure signature and 10% as the rate at which the trace is
present at all.
Method note: historical attempts must be read via
gh api --allow-escape-sequences /repos/{owner}/{repo}/actions/jobs/{job_id}/logs.gh run view --job <id> --logresolves by run and serves the latest attempt, so it hands back agreen log for a red attempt.
This is not a second emulator bug. It is the same one, reached by a different caller.
I compared the tombstones from all 8 API 37 legs that carry this abort against the crash the
E2E_DISABLE_SYSTEM_UImitigation in.github/scripts/e2e-run.shalready handles. They are thesame assertion in the same mapper:
hasReadColorBufferDmaGoldfishMapper::readFromHostRegionSamplingThreadTaskSnapshotPersystem_serveritselfAll 8:
hasReadColorBufferDma=1,RegionSamplingThread=0, threadTaskSnapshotPer, frames throughTaskSnapshotConvertUtil.copyToSwBitmapDirect→copySWBitmap.That is why the existing mitigation cannot help. It works by removing the caller —
pm disable-user com.android.systemuideletes the region-sampling registration, and the script's owncomment says so: "RegionSamplingThread exists only because SystemUI registers a nav-bar luma-sampling
listener, so removing the package removes the whole chain." Task snapshots are captured by
WindowManager inside
system_server. There is no package to disable, so nothing about the SystemUIwork touches this path.
So the shape of a fix is "remove the second caller" — stop the emulator taking task snapshots —
rather than anything about this app or its tests. Recents thumbnails have no bearing on what the
suite asserts.
One thing I checked and am reporting as a non-finding
count_aborts()grepshasReadColorBufferDma, which both callers produce, and it drives themitigation's "45 s with zero new aborts" exit test. So in principle a
TaskSnapshotPerabort landinginside that window would be misread as "SystemUI disable failed" and burn another round.
Measured across 12 recorded runs: it has never happened. Every one reported
abort rate, SystemUI disabled: 0 new in 45 s, and the 4 runs that reached round 2 got there via theother exit — the
pmstate not surviving the framework restart — not via the abort count. TheTaskSnapshotPeraborts land later, during the test run, well outside the verification window.Recording it because the reasoning is sound and someone will notice the shared grep later: it is a
latent imprecision, not a live defect, and narrowing that grep on current evidence would be a
speculative change to the one function keeping API 37 alive. Leave it until a run actually shows a
non-zero delta.
Note on the numbers above it
The
total 2in those log lines is cumulative crash-buffer content, so aborts had already occurredbefore the disable finished. That is consistent with the two callers being independent.
Four consecutive API 37 attempts, same abort, three different outcomes
From #117, where the agent re-ran the API 37 leg twice (four attempts total). Every attempt carried
TaskSnapshotPer >>> system_serverandhasReadColorBufferDma, and the visible outcome differedeach time:
expected: 57 / received: 57 / failed: 0)SafPickerRoundTripTest.pickingAFileThroughTheSystemPickerFillsInTheFileCardConversionWorkerTest.routesAFastMp4JobByDeviceCapabilityNo test failed twice, and nothing in that PR's area ever failed — the change was to
ContainerCapabilities.validateVideo, nowhere near workers or the picker.This is the cleanest evidence yet for what this ticket describes: the abort is the constant, and
which test it takes down is arbitrary. It also strengthens the reading on #102 that API 37 picker
failures are not necessarily a #93 regression — here the picker was collateral in one attempt out of
four, with the same abort present in all four.
Combined with the earlier census (8 aborting legs, 5 of which also failed a test), the pattern is
consistent: when
system_servergoes down mid-run it sometimes lands on a test and sometimes doesnot, and the test it lands on tells you nothing about the code.
The abort clusters at one point in the run — it is triggered, not ambient load
Last progress line before the abort, across every API 37 log I hold that carries
TaskSnapshotPer:50/56 completed. (2 skipped) (0 failed)51/57 completed. (2 skipped) (0 failed)53/5654/5658/5659/5750/56and51/57are the same position — the suite grew by one test when #113 landed. So6 of 11 aborts happen at the same index, and the rest are spread thin.
That is hard to reconcile with "the runner was busy". An abort caused by ambient load should land
uniformly across a five-minute run; this one has a mode.
What sits at that index. The agent working #49 observed directly that on both #124 and #117 the
framework goes down under
SafPickerRoundTripTest, having reported0 failedthrough test 51. Ihave verified the
51/57boundary on #117 (job98051840467) independently. I have not verifiedthe test identity from the gradle log — it does not name tests — so that half rests on the agent's
reading of the run, and instrumented test order, while deterministic in practice, is not a guarantee.
Why this matters beyond this ticket.
SafPickerRoundTripTestis the only test in the suite thattouches system UI, and it is now implicated in two separate CI failure modes on two disjoint sets of
API levels:
system_server.thePickedInputSurvivesARealRotationhangs indefinitely.Its own KDoc already says why it is special: "This class is the first thing in the suite that
touches system UI, and the android-37.x images are where that stops being free." That was written
about 37. The 33/34 hang says the sentence may be broader than its author knew.
A cheap next measurement, for whoever picks this up: the
e2e-diagnostics-apiNNartifacts on anyaborting run should contain the TestRunner logcat, which names every test start and finish. That
would settle the test-identity question directly rather than by inference — the same way reading the
wedge artifact settled #122.
I read the diagnostics artifact, and it corrects my previous comment
I said the clustering meant the abort was "triggered, not ambient load", and relayed an attribution to
SafPickerRoundTripTest. The first half is wrong and the attribution does not survive.From
e2e-diagnostics-api37on run32926994200(job98051840467), timestamps from the logcat:Three of the five aborts fire before the first test starts. The instrumentation had not begun; there
was nothing to trigger them. So the abort is a property of the emulator image running at all, not of
anything the suite does — which is what this ticket originally said and what I talked myself out of.
The SAF picker test was hit by an abort at 03:42:07 and still finished 1.1 s later. It is a victim
in this run, not a cause.
A better explanation for the clustering, which is still real
6 of 11 aborts landing at the same index is a genuine observation and I am not withdrawing it. But
exposure explains it without causation: aborts recur every 20-90 s for the emulator's whole life
(the cadence
e2e-run.sh's own comment records for the surfaceflinger variant), and the run dieswhen one lands during instrumentation. The tests around that index include the slowest in the suite
—
SafPickerRoundTripTesttook 8.9 s here and 9.6 s on API 33, against a median well under a second— so they present by far the largest target for a randomly-timed abort.
That fits every observation: aborts before the run, aborts that a test survives, arbitrary victims
across four attempts on #117, and a mode at the index where the long tests sit. "Which test it lands
on tells you nothing about the code" (my earlier comment) was right; "it is triggered" was not.
Consequence for #122
This weakens the link I drew to #122.
SafPickerRoundTripTestis implicated in both, but forpossibly unrelated reasons: on 33/34 its rotation test hangs with the framework fully alive and no
abort anywhere, and on 37 it is one of several tests an ongoing abort can interrupt. Being the
slowest, most system-UI-dependent test in the suite is enough to put it at the scene of both without
the two sharing a cause. Treat them as separate until something links them.
Method note
The logcat artifact concatenates buffers, so line order is not time order — my first pass at this
read the tail of the file as "the last events" and got a picture four minutes out of date. Sort by
timestamp before drawing any conclusion from it.
This is now hitting the gating API 37 leg repeatedly, and it lands on a different test each time — which is the strongest evidence yet that it is the environment rather than any test.
Four sightings on 2026-09-06, across four different PRs, all on
E2E API 37(gating,disable-system-ui: "1"), all with the same abort:34015541233androidTestfilesSafPickerRoundTripTest.pickingAFileThroughTheSystemPickerFillsInTheFileCard34017895958androidTestfileSafPickerRoundTripTest.pickingAFileThroughTheSystemPickerFillsInTheFileCard34019238667androidTestfile + a constNotificationCancelActionTest.theNotificationsCancelActionCancelsThatJob34020234606NotificationCancelActionTest.theNotificationsCancelActionCancelsThatJobThe last row is the useful one: a docs-only diff cannot reach any instrumented test, and it failed the gating API 37 leg twice in a row before passing on a third attempt.
Why the varying victim matters
This ticket's title says the abort happens "after every test passes". These four show it also aborts mid-run —
received: 63of 64 — and the test it interrupts is whichever one happened to be executing. Two different tests, in two different packages, one of which drives system UI and one of which does not.NotificationCancelActionTestfires aPendingIntentthroughWorkManager.createCancelPendingIntent, so it does touchsystem_server;SafPickerRoundTripTesttouches DocumentsUI. Neither is a common cause with the other beyond "was running when system_server went down".Deliberately not proposed: marking either with
@FailsOnEmulatorApi37. That marker means "cannot pass on this image", and both pass on the same leg most of the time —NotificationCancelActionTestpassed all five legs when it landed in #235. Marking an intermittent failure would move a working test out of the gating set and make the advisory job expect a failure that may not happen.Cost, since that is what #190 says is the thing worth measuring
Every one of these needed a human to read the log, decide it was environmental, and re-run. Four times in one session, three of them on PRs whose diffs could not reach the E2E matrix at all.