Gating API 37 leg aborts in system_server TaskSnapshotPersister after every test passes #108

Closed
opened 2026-08-25 20:10:42 +00:00 by JMR-dev · 7 comments
JMR-dev commented 2026-08-25 20:10:42 +00:00 (Migrated from github.com)

The gating API 37 leg has failed three consecutive PRs today with zero test failures. This is a third caller of the abort docs/api-37-emulator-crash.md already documents, and the SystemUI mitigation does not cover it.

What the leg reports

E2E API 37 (gating, not the advisory job), on #99, #104 and #106 — diffs of a KDoc comment, a
workflow step and a permissions: block, none of which can reach the E2E matrix:

Tests 53/56 completed. (2 skipped) (0 failed)
...
Shell command failed (20): am get-current-user
Device emulator-5554 failed to uninstall test APK org.libremediaconverter.
> Task :app:connectedDebugAndroidTest FAILED

Every test passes. The job fails anyway — am get-current-user returning 20 means
system_server has stopped answering.

The crash, from the job's own native-crash tail

Executable: /system/bin/app_process64
Cmdline: system_server
pid: 3909, tid: 4200, name: TaskSnapshotPer  >>> system_server <<<
signal 6 (SIGABRT)
Abort message: 'Assertion failed: !rcEnc->featureInfo()->hasReadColorBufferDma'

Why this is new rather than already covered

docs/api-37-emulator-crash.md documents this assertion twice, and neither is this:

caller process thread covered by
documented, original surfaceflinger RegionSampling E2E_DISABLE_SYSTEM_UI=1 removes the listener
documented 2026-08-24 app under test picker / rotation @FailsOnEmulatorApi37 on the two tests
this system_server TaskSnapshotPersister nothing

The doc's own words for the second entry apply exactly here: "disabling SystemUI does not help,
because it removes the idle trigger (RegionSamplingThread's nav-bar luma sampling) and not this
one."
TaskSnapshotPersister is a third one — WindowManager writing task snapshots, in
system_server rather than surfaceflinger, on a path the disable never touched.

The gating row does set disable-system-ui: "1", and it is working as documented. It simply
does not reach this caller.

Why it matters more than the other two

This one aborts after the suite has passed, so it converts a green run into a red check with no
failing test to point at. A reader seeing E2E API 37 red will look for a broken test and find 56
passes — which is the most confusing possible failure, and it has now happened three times in a row
on unrelated PRs.

Frequency

Three consecutive gating-leg failures (#99, #104, #106) in roughly ninety minutes. That leg was
stable at 56/0 for the whole of the preceding session. This is a change in behaviour, not a
long-standing flake
— worth saying, because the honest response to a rare flake and to a
regression are different.

What to establish first

  1. Does it reproduce on a local API 37 emulator, or only on CI runners? #96 established that some of
    these are load-dependent and invisible on this workstation.
  2. Is TaskSnapshotPersister suppressible the way the region-sampling listener was — a
    wm setting, or a device config — without disabling more than the leg already disables?
  3. If not: the honest options are to accept the leg as advisory-only for teardown failures, or to
    record it the way the other two are recorded. Neither should be taken quietly — this row
    gates, and #56 added it deliberately after measuring that 55 of 57 tests do pass there.
_The **gating** API 37 leg has failed three consecutive PRs today with zero test failures. This is a third caller of the abort `docs/api-37-emulator-crash.md` already documents, and the SystemUI mitigation does not cover it._ ### What the leg reports `E2E API 37` (gating, not the advisory job), on #99, #104 and #106 — diffs of a KDoc comment, a workflow step and a `permissions:` block, none of which can reach the E2E matrix: ``` Tests 53/56 completed. (2 skipped) (0 failed) ... Shell command failed (20): am get-current-user Device emulator-5554 failed to uninstall test APK org.libremediaconverter. > Task :app:connectedDebugAndroidTest FAILED ``` **Every test passes. The job fails anyway** — `am get-current-user` returning 20 means `system_server` has stopped answering. ### The crash, from the job's own native-crash tail ``` Executable: /system/bin/app_process64 Cmdline: system_server pid: 3909, tid: 4200, name: TaskSnapshotPer >>> system_server <<< signal 6 (SIGABRT) Abort message: 'Assertion failed: !rcEnc->featureInfo()->hasReadColorBufferDma' ``` ### Why this is new rather than already covered `docs/api-37-emulator-crash.md` documents this assertion twice, and neither is this: | caller | process | thread | covered by | |---|---|---|---| | documented, original | `surfaceflinger` | `RegionSampling` | `E2E_DISABLE_SYSTEM_UI=1` removes the listener | | documented 2026-08-24 | app under test | picker / rotation | `@FailsOnEmulatorApi37` on the two tests | | **this** | **`system_server`** | **`TaskSnapshotPersister`** | **nothing** | The doc's own words for the second entry apply exactly here: *"disabling SystemUI does not help, because it removes the idle trigger (RegionSamplingThread's nav-bar luma sampling) and not this one."* `TaskSnapshotPersister` is a third one — WindowManager writing task snapshots, in `system_server` rather than `surfaceflinger`, on a path the disable never touched. The gating row **does** set `disable-system-ui: "1"`, and it is working as documented. It simply does not reach this caller. ### Why it matters more than the other two This one aborts **after the suite has passed**, so it converts a green run into a red check with no failing test to point at. A reader seeing `E2E API 37` red will look for a broken test and find 56 passes — which is the most confusing possible failure, and it has now happened three times in a row on unrelated PRs. ### Frequency Three consecutive gating-leg failures (#99, #104, #106) in roughly ninety minutes. That leg was stable at 56/0 for the whole of the preceding session. **This is a change in behaviour, not a long-standing flake** — worth saying, because the honest response to a rare flake and to a regression are different. ### What to establish first 1. Does it reproduce on a local API 37 emulator, or only on CI runners? #96 established that some of these are load-dependent and invisible on this workstation. 2. Is `TaskSnapshotPersister` suppressible the way the region-sampling listener was — a `wm` setting, or a device config — without disabling more than the leg already disables? 3. If not: the honest options are to accept the leg as advisory-only for teardown failures, or to record it the way the other two are recorded. **Neither should be taken quietly** — this row gates, and #56 added it deliberately after measuring that 55 of 57 tests do pass there.
JMR-dev commented 2026-08-25 20:15:59 +00:00 (Migrated from github.com)

Correcting the frequency claim in this ticket. The crash evidence stands; the "regression" framing does not.

I wrote that the gating leg "was stable at 56/0 for the whole of the preceding session" and that this is "a change in behaviour, not a long-standing flake." Measured, that is wrong.

Across the last 24 status_check.yml runs, the E2E API 37 gating job:

success 17   failure 3   still running 4
failure rate among concluded: 3/20 = 15%

Successes are interleaved throughout, including at 15:13 and 15:14 — between the failures I cited as consecutive. What I actually saw was three PRs holding a failed run at one moment, which is not the same as three consecutive runs, and I generalised from the first to the second.

This is the same error I corrected on #49 this morning — reading a cluster as a rate. Twice in one day, so it is worth naming the shape rather than just the instance: when several PRs are queued behind each other, they sample the same window, and a moderate flake will show up in all of them at once and look like a step change.

What stands, unchanged

The crash evidence is from an actual native-crash dump and is not affected:

Cmdline: system_server
name: TaskSnapshotPer  >>> system_server <<<
Abort message: 'Assertion failed: !rcEnc->featureInfo()->hasReadColorBufferDma'

TaskSnapshotPersister in system_server is a third caller of an assertion this repo has
documented twice, and E2E_DISABLE_SYSTEM_UI=1 genuinely does not cover it — the doc's own
sentence about the idle trigger applies. That part needed no frequency argument.

So does the reason it is confusing: it aborts after the suite passes, turning a green run into a
red check with 56 passes and nothing to point at.

What changes about the ticket

  • Severity is about confusion, not breakage. A 15% leg failure that clears on rerun is an
    annoyance; one that presents as "red with zero failing tests" costs a reader real time. The
    argument for fixing it is legibility.
  • Do not investigate this as a regression. There is no "what changed today" to find. The
    question is whether TaskSnapshotPersister is suppressible the way the region-sampling listener
    was, and if not, whether a teardown-only abort should fail the leg at all.
  • One hypothesis I raised and then falsified, so nobody re-runs it: I thought #80's SAF tests might
    have introduced the activity churn that triggers task snapshots. The failure rate does not step at
    #80's merge, so that is not supported.
**Correcting the frequency claim in this ticket. The crash evidence stands; the "regression" framing does not.** I wrote that the gating leg *"was stable at 56/0 for the whole of the preceding session"* and that this is *"a change in behaviour, not a long-standing flake."* **Measured, that is wrong.** Across the last 24 `status_check.yml` runs, the `E2E API 37` gating job: ``` success 17 failure 3 still running 4 failure rate among concluded: 3/20 = 15% ``` Successes are interleaved throughout, including at 15:13 and 15:14 — between the failures I cited as consecutive. What I actually saw was **three PRs holding a failed run at one moment**, which is not the same as three consecutive runs, and I generalised from the first to the second. **This is the same error I corrected on #49 this morning** — reading a cluster as a rate. Twice in one day, so it is worth naming the shape rather than just the instance: when several PRs are queued behind each other, they sample the same window, and a moderate flake will show up in all of them at once and look like a step change. ### What stands, unchanged The crash evidence is from an actual native-crash dump and is not affected: ``` Cmdline: system_server name: TaskSnapshotPer >>> system_server <<< Abort message: 'Assertion failed: !rcEnc->featureInfo()->hasReadColorBufferDma' ``` `TaskSnapshotPersister` in `system_server` **is** a third caller of an assertion this repo has documented twice, and `E2E_DISABLE_SYSTEM_UI=1` genuinely does not cover it — the doc's own sentence about the idle trigger applies. That part needed no frequency argument. So does the reason it is confusing: it aborts **after** the suite passes, turning a green run into a red check with 56 passes and nothing to point at. ### What changes about the ticket - **Severity is about confusion, not breakage.** A 15% leg failure that clears on rerun is an annoyance; one that presents as "red with zero failing tests" costs a reader real time. The argument for fixing it is legibility. - **Do not investigate this as a regression.** There is no "what changed today" to find. The question is whether `TaskSnapshotPersister` is suppressible the way the region-sampling listener was, and if not, whether a teardown-only abort should fail the leg at all. - One hypothesis I raised and then falsified, so nobody re-runs it: I thought #80's SAF tests might have introduced the activity churn that triggers task snapshots. The failure rate does not step at #80's merge, so that is not supported.
JMR-dev commented 2026-08-26 01:50:50 +00:00 (Migrated from github.com)

Re-measured per leg-attempt: 3.75%, not ~15%

The earlier figure on this ticket counted whole runs. A leg that is re-run to green disappears from
that count, so the denominator was wrong in one direction and the numerator in the other. Counting
leg-attempts for E2E API 37 since 2026-08-24:

  • 80 leg-attempts, 12 failures.
  • 8 of the 12 carry the TaskSnapshotPersister / system_server SIGABRT trace.
  • 3 of those 8 have no named test failure at all — which is this ticket's exact signature,
    "aborts after every test passes". That is 3/80 ≈ 3.75%.
run / attempt abort trace named test failures
32808298995 a2 yes none
32811493607 a1 yes none
32864742077 a1 yes none
32806342548 a1 yes pickingAFileThroughTheSystemPickerFillsInTheFileCard
32808298995 a1 yes pickingAFileThroughTheSystemPickerFillsInTheFileCard
32810469166 a1 yes routesAFastMp4JobByDeviceCapability
32812892103 a1 yes pickingAFileThroughTheSystemPickerFillsInTheFileCard
32899061992 a1 yes pickingAFileThroughTheSystemPickerFillsInTheFileCard
32813885120 a1 no none
32862120199 a1 no none
32865281555 a1 no none
32865281555 a2 no pickingAFileThroughTheSystemPickerFillsInTheFileCard

In 5 of the 8, a test failed as well. So the abort is not purely a post-run tidy-up crash — when
it happens it can take a test with it. That is the strongest single piece of evidence for #102's
shared-cause reading, and it is why the two tickets should not be closed independently.

Three further failures show neither an abort trace nor a test failure, so the trace is not always
captured; treat 3.75% as a floor for the pure signature and 10% as the rate at which the trace is
present at all.

Method note: historical attempts must be read via
gh api --allow-escape-sequences /repos/{owner}/{repo}/actions/jobs/{job_id}/logs.
gh run view --job <id> --log resolves by run and serves the latest attempt, so it hands back a
green log for a red attempt.

## Re-measured per leg-attempt: 3.75%, not ~15% The earlier figure on this ticket counted whole runs. A leg that is re-run to green disappears from that count, so the denominator was wrong in one direction and the numerator in the other. Counting **leg-attempts** for `E2E API 37` since 2026-08-24: - **80 leg-attempts, 12 failures.** - **8** of the 12 carry the `TaskSnapshotPersister` / `system_server` SIGABRT trace. - **3** of those 8 have *no named test failure at all* — which is this ticket's exact signature, "aborts after every test passes". That is **3/80 ≈ 3.75%**. | run / attempt | abort trace | named test failures | |---|---|---| | 32808298995 a2 | yes | none | | 32811493607 a1 | yes | none | | 32864742077 a1 | yes | none | | 32806342548 a1 | yes | `pickingAFileThroughTheSystemPickerFillsInTheFileCard` | | 32808298995 a1 | yes | `pickingAFileThroughTheSystemPickerFillsInTheFileCard` | | 32810469166 a1 | yes | `routesAFastMp4JobByDeviceCapability` | | 32812892103 a1 | yes | `pickingAFileThroughTheSystemPickerFillsInTheFileCard` | | 32899061992 a1 | yes | `pickingAFileThroughTheSystemPickerFillsInTheFileCard` | | 32813885120 a1 | no | none | | 32862120199 a1 | no | none | | 32865281555 a1 | no | none | | 32865281555 a2 | no | `pickingAFileThroughTheSystemPickerFillsInTheFileCard` | **In 5 of the 8, a test failed as well.** So the abort is not purely a post-run tidy-up crash — when it happens it can take a test with it. That is the strongest single piece of evidence for #102's shared-cause reading, and it is why the two tickets should not be closed independently. Three further failures show neither an abort trace nor a test failure, so the trace is not always captured; treat 3.75% as a floor for the pure signature and 10% as the rate at which the trace is present at all. Method note: historical attempts must be read via `gh api --allow-escape-sequences /repos/{owner}/{repo}/actions/jobs/{job_id}/logs`. `gh run view --job <id> --log` resolves by run and serves the latest attempt, so it hands back a green log for a red attempt.
JMR-dev commented 2026-08-26 02:16:08 +00:00 (Migrated from github.com)

This is not a second emulator bug. It is the same one, reached by a different caller.

I compared the tombstones from all 8 API 37 legs that carry this abort against the crash the
E2E_DISABLE_SYSTEM_UI mitigation in .github/scripts/e2e-run.sh already handles. They are the
same assertion in the same mapper:

Abort message: 'Assertion failed: !rcEnc->featureInfo()->hasReadColorBufferDma'
  ... GoldfishMapper::readFromHost
already mitigated this ticket
assertion hasReadColorBufferDma same
mapper frame GoldfishMapper::readFromHost same
aborting thread RegionSamplingThread TaskSnapshotPer
who calls it SystemUI's nav-bar luma sampler WindowManager, in system_server itself

All 8: hasReadColorBufferDma=1, RegionSamplingThread=0, thread TaskSnapshotPer, frames through
TaskSnapshotConvertUtil.copyToSwBitmapDirect → copySWBitmap.

That is why the existing mitigation cannot help. It works by removing the caller — pm disable-user com.android.systemui deletes the region-sampling registration, and the script's own
comment says so: "RegionSamplingThread exists only because SystemUI registers a nav-bar luma-sampling
listener, so removing the package removes the whole chain." Task snapshots are captured by
WindowManager inside system_server. There is no package to disable, so nothing about the SystemUI
work touches this path.

So the shape of a fix is "remove the second caller" — stop the emulator taking task snapshots —
rather than anything about this app or its tests. Recents thumbnails have no bearing on what the
suite asserts.

One thing I checked and am reporting as a non-finding

count_aborts() greps hasReadColorBufferDma, which both callers produce, and it drives the
mitigation's "45 s with zero new aborts" exit test. So in principle a TaskSnapshotPer abort landing
inside that window would be misread as "SystemUI disable failed" and burn another round.

Measured across 12 recorded runs: it has never happened. Every one reported
abort rate, SystemUI disabled: 0 new in 45 s, and the 4 runs that reached round 2 got there via the
other exit — the pm state not surviving the framework restart — not via the abort count. The
TaskSnapshotPer aborts land later, during the test run, well outside the verification window.

Recording it because the reasoning is sound and someone will notice the shared grep later: it is a
latent imprecision, not a live defect, and narrowing that grep on current evidence would be a
speculative change to the one function keeping API 37 alive. Leave it until a run actually shows a
non-zero delta.

Note on the numbers above it

The total 2 in those log lines is cumulative crash-buffer content, so aborts had already occurred
before the disable finished. That is consistent with the two callers being independent.

## This is not a second emulator bug. It is the same one, reached by a different caller. I compared the tombstones from all 8 API 37 legs that carry this abort against the crash the `E2E_DISABLE_SYSTEM_UI` mitigation in `.github/scripts/e2e-run.sh` already handles. They are the **same assertion in the same mapper**: ``` Abort message: 'Assertion failed: !rcEnc->featureInfo()->hasReadColorBufferDma' ... GoldfishMapper::readFromHost ``` | | already mitigated | this ticket | |---|---|---| | assertion | `hasReadColorBufferDma` | **same** | | mapper frame | `GoldfishMapper::readFromHost` | **same** | | aborting thread | `RegionSamplingThread` | **`TaskSnapshotPer`** | | who calls it | SystemUI's nav-bar luma sampler | WindowManager, in `system_server` itself | All 8: `hasReadColorBufferDma=1`, `RegionSamplingThread=0`, thread `TaskSnapshotPer`, frames through `TaskSnapshotConvertUtil.copyToSwBitmapDirect` → `copySWBitmap`. **That is why the existing mitigation cannot help.** It works by removing the *caller* — `pm disable-user com.android.systemui` deletes the region-sampling registration, and the script's own comment says so: "RegionSamplingThread exists only because SystemUI registers a nav-bar luma-sampling listener, so removing the package removes the whole chain." Task snapshots are captured by WindowManager inside `system_server`. There is no package to disable, so nothing about the SystemUI work touches this path. So the shape of a fix is **"remove the second caller"** — stop the emulator taking task snapshots — rather than anything about this app or its tests. Recents thumbnails have no bearing on what the suite asserts. ## One thing I checked and am reporting as a *non*-finding `count_aborts()` greps `hasReadColorBufferDma`, which both callers produce, and it drives the mitigation's "45 s with zero new aborts" exit test. So in principle a `TaskSnapshotPer` abort landing inside that window would be misread as "SystemUI disable failed" and burn another round. **Measured across 12 recorded runs: it has never happened.** Every one reported `abort rate, SystemUI disabled: 0 new in 45 s`, and the 4 runs that reached round 2 got there via the other exit — the `pm` state not surviving the framework restart — not via the abort count. The `TaskSnapshotPer` aborts land later, during the test run, well outside the verification window. Recording it because the reasoning is sound and someone will notice the shared grep later: it is a latent imprecision, **not a live defect**, and narrowing that grep on current evidence would be a speculative change to the one function keeping API 37 alive. Leave it until a run actually shows a non-zero delta. ## Note on the numbers above it The `total 2` in those log lines is cumulative crash-buffer content, so aborts had already occurred before the disable finished. That is consistent with the two callers being independent.
JMR-dev commented 2026-08-26 03:35:34 +00:00 (Migrated from github.com)

Four consecutive API 37 attempts, same abort, three different outcomes

From #117, where the agent re-ran the API 37 leg twice (four attempts total). Every attempt carried
TaskSnapshotPer >>> system_server and hasReadColorBufferDma
, and the visible outcome differed
each time:

outcome count
no test failed at all (expected: 57 / received: 57 / failed: 0) 2
took down SafPickerRoundTripTest.pickingAFileThroughTheSystemPickerFillsInTheFileCard 1
took down ConversionWorkerTest.routesAFastMp4JobByDeviceCapability 1

No test failed twice, and nothing in that PR's area ever failed — the change was to
ContainerCapabilities.validateVideo, nowhere near workers or the picker.

This is the cleanest evidence yet for what this ticket describes: the abort is the constant, and
which test it takes down is arbitrary. It also strengthens the reading on #102 that API 37 picker
failures are not necessarily a #93 regression — here the picker was collateral in one attempt out of
four, with the same abort present in all four.

Combined with the earlier census (8 aborting legs, 5 of which also failed a test), the pattern is
consistent: when system_server goes down mid-run it sometimes lands on a test and sometimes does
not, and the test it lands on tells you nothing about the code.

## Four consecutive API 37 attempts, same abort, three different outcomes From #117, where the agent re-ran the API 37 leg twice (four attempts total). **Every attempt carried `TaskSnapshotPer >>> system_server` and `hasReadColorBufferDma`**, and the visible outcome differed each time: | outcome | count | |---|---| | no test failed at all (`expected: 57 / received: 57 / failed: 0`) | 2 | | took down `SafPickerRoundTripTest.pickingAFileThroughTheSystemPickerFillsInTheFileCard` | 1 | | took down `ConversionWorkerTest.routesAFastMp4JobByDeviceCapability` | 1 | **No test failed twice, and nothing in that PR's area ever failed** — the change was to `ContainerCapabilities.validateVideo`, nowhere near workers or the picker. This is the cleanest evidence yet for what this ticket describes: the abort is the constant, and *which* test it takes down is arbitrary. It also strengthens the reading on #102 that API 37 picker failures are not necessarily a #93 regression — here the picker was collateral in one attempt out of four, with the same abort present in all four. Combined with the earlier census (8 aborting legs, 5 of which also failed a test), the pattern is consistent: when `system_server` goes down mid-run it sometimes lands on a test and sometimes does not, and the test it lands on tells you nothing about the code.
JMR-dev commented 2026-08-26 03:52:20 +00:00 (Migrated from github.com)

The abort clusters at one point in the run — it is triggered, not ambient load

Last progress line before the abort, across every API 37 log I hold that carries TaskSnapshotPer:

stopping point count
50/56 completed. (2 skipped) (0 failed) 5
51/57 completed. (2 skipped) (0 failed) 1
53/56 1
54/56 1
58/56 1
59/57 2

50/56 and 51/57 are the same position — the suite grew by one test when #113 landed. So
6 of 11 aborts happen at the same index, and the rest are spread thin.

That is hard to reconcile with "the runner was busy". An abort caused by ambient load should land
uniformly across a five-minute run; this one has a mode.

What sits at that index. The agent working #49 observed directly that on both #124 and #117 the
framework goes down under SafPickerRoundTripTest, having reported 0 failed through test 51. I
have verified the 51/57 boundary on #117 (job 98051840467) independently. I have not verified
the test identity from the gradle log — it does not name tests — so that half rests on the agent's
reading of the run, and instrumented test order, while deterministic in practice, is not a guarantee.

Why this matters beyond this ticket. SafPickerRoundTripTest is the only test in the suite that
touches system UI, and it is now implicated in two separate CI failure modes on two disjoint sets of
API levels:

  • API 37 — this ticket: the gralloc assertion, aborting system_server.
  • API 33/34 — #122: thePickedInputSurvivesARealRotation hangs indefinitely.

Its own KDoc already says why it is special: "This class is the first thing in the suite that
touches system UI, and the android-37.x images are where that stops being free."
That was written
about 37. The 33/34 hang says the sentence may be broader than its author knew.

A cheap next measurement, for whoever picks this up: the e2e-diagnostics-apiNN artifacts on any
aborting run should contain the TestRunner logcat, which names every test start and finish. That
would settle the test-identity question directly rather than by inference — the same way reading the
wedge artifact settled #122.

## The abort clusters at one point in the run — it is triggered, not ambient load Last progress line before the abort, across every API 37 log I hold that carries `TaskSnapshotPer`: | stopping point | count | |---|---| | `50/56 completed. (2 skipped) (0 failed)` | 5 | | `51/57 completed. (2 skipped) (0 failed)` | 1 | | `53/56` | 1 | | `54/56` | 1 | | `58/56` | 1 | | `59/57` | 2 | `50/56` and `51/57` are **the same position** — the suite grew by one test when #113 landed. So **6 of 11 aborts happen at the same index**, and the rest are spread thin. That is hard to reconcile with "the runner was busy". An abort caused by ambient load should land uniformly across a five-minute run; this one has a mode. **What sits at that index.** The agent working #49 observed directly that on both #124 and #117 the framework goes down *under* `SafPickerRoundTripTest`, having reported `0 failed` through test 51. I have verified the `51/57` boundary on #117 (job `98051840467`) independently. I have **not** verified the test identity from the gradle log — it does not name tests — so that half rests on the agent's reading of the run, and instrumented test order, while deterministic in practice, is not a guarantee. **Why this matters beyond this ticket.** `SafPickerRoundTripTest` is the only test in the suite that touches system UI, and it is now implicated in two separate CI failure modes on two disjoint sets of API levels: - **API 37** — this ticket: the gralloc assertion, aborting `system_server`. - **API 33/34** — #122: `thePickedInputSurvivesARealRotation` hangs indefinitely. Its own KDoc already says why it is special: *"This class is the first thing in the suite that touches system UI, and the android-37.x images are where that stops being free."* That was written about 37. The 33/34 hang says the sentence may be broader than its author knew. **A cheap next measurement, for whoever picks this up:** the `e2e-diagnostics-apiNN` artifacts on any aborting run should contain the TestRunner logcat, which names every test start and finish. That would settle the test-identity question directly rather than by inference — the same way reading the wedge artifact settled #122.
JMR-dev commented 2026-08-26 03:55:18 +00:00 (Migrated from github.com)

I read the diagnostics artifact, and it corrects my previous comment

I said the clustering meant the abort was "triggered, not ambient load", and relayed an attribution to
SafPickerRoundTripTest. The first half is wrong and the attribution does not survive.

From e2e-diagnostics-api37 on run 32926994200 (job 98051840467), timestamps from the logcat:

abort  03:37:38.733     <- no test running
abort  03:38:07.092     <- no test running
abort  03:38:07.907     <- no test running
   ...
TEST RUN BEGINS  03:41:36.818
abort  03:42:07.504     <- during SafPickerRoundTripTest.pickingAFile... (which then FINISHED, 03:42:08.563)
abort  03:42:10.162     <- during ConversionWorkerTest.routesAFastMp4JobByDeviceCapability
TEST RUN ENDS    03:42:10.657

Three of the five aborts fire before the first test starts. The instrumentation had not begun; there
was nothing to trigger them. So the abort is a property of the emulator image running at all, not of
anything the suite does — which is what this ticket originally said and what I talked myself out of.

The SAF picker test was hit by an abort at 03:42:07 and still finished 1.1 s later. It is a victim
in this run, not a cause.

A better explanation for the clustering, which is still real

6 of 11 aborts landing at the same index is a genuine observation and I am not withdrawing it. But
exposure explains it without causation: aborts recur every 20-90 s for the emulator's whole life
(the cadence e2e-run.sh's own comment records for the surfaceflinger variant), and the run dies
when one lands during instrumentation
. The tests around that index include the slowest in the suite
— SafPickerRoundTripTest took 8.9 s here and 9.6 s on API 33, against a median well under a second
— so they present by far the largest target for a randomly-timed abort.

That fits every observation: aborts before the run, aborts that a test survives, arbitrary victims
across four attempts on #117, and a mode at the index where the long tests sit. "Which test it lands
on tells you nothing about the code" (my earlier comment) was right; "it is triggered" was not.

Consequence for #122

This weakens the link I drew to #122. SafPickerRoundTripTest is implicated in both, but for
possibly unrelated reasons: on 33/34 its rotation test hangs with the framework fully alive and no
abort anywhere
, and on 37 it is one of several tests an ongoing abort can interrupt. Being the
slowest, most system-UI-dependent test in the suite is enough to put it at the scene of both without
the two sharing a cause. Treat them as separate until something links them.

Method note

The logcat artifact concatenates buffers, so line order is not time order — my first pass at this
read the tail of the file as "the last events" and got a picture four minutes out of date. Sort by
timestamp before drawing any conclusion from it.

## I read the diagnostics artifact, and it corrects my previous comment I said the clustering meant the abort was "triggered, not ambient load", and relayed an attribution to `SafPickerRoundTripTest`. **The first half is wrong and the attribution does not survive.** From `e2e-diagnostics-api37` on run `32926994200` (job `98051840467`), timestamps from the logcat: ``` abort 03:37:38.733 <- no test running abort 03:38:07.092 <- no test running abort 03:38:07.907 <- no test running ... TEST RUN BEGINS 03:41:36.818 abort 03:42:07.504 <- during SafPickerRoundTripTest.pickingAFile... (which then FINISHED, 03:42:08.563) abort 03:42:10.162 <- during ConversionWorkerTest.routesAFastMp4JobByDeviceCapability TEST RUN ENDS 03:42:10.657 ``` **Three of the five aborts fire before the first test starts.** The instrumentation had not begun; there was nothing to trigger them. So the abort is a property of the emulator image running at all, not of anything the suite does — which is what this ticket originally said and what I talked myself out of. The SAF picker test was *hit* by an abort at 03:42:07 and **still finished** 1.1 s later. It is a victim in this run, not a cause. ## A better explanation for the clustering, which is still real 6 of 11 aborts landing at the same index is a genuine observation and I am not withdrawing it. But exposure explains it without causation: aborts recur every 20-90 s for the emulator's whole life (the cadence `e2e-run.sh`'s own comment records for the surfaceflinger variant), and **the run dies when one lands during instrumentation**. The tests around that index include the slowest in the suite — `SafPickerRoundTripTest` took 8.9 s here and 9.6 s on API 33, against a median well under a second — so they present by far the largest target for a randomly-timed abort. That fits every observation: aborts before the run, aborts that a test survives, arbitrary victims across four attempts on #117, and a mode at the index where the long tests sit. "Which test it lands on tells you nothing about the code" (my earlier comment) was right; "it is triggered" was not. ## Consequence for #122 This weakens the link I drew to #122. `SafPickerRoundTripTest` is implicated in both, but for possibly unrelated reasons: on 33/34 its rotation test hangs with **the framework fully alive and no abort anywhere**, and on 37 it is one of several tests an ongoing abort can interrupt. Being the slowest, most system-UI-dependent test in the suite is enough to put it at the scene of both without the two sharing a cause. Treat them as separate until something links them. ## Method note The logcat artifact concatenates buffers, so **line order is not time order** — my first pass at this read the tail of the file as "the last events" and got a picture four minutes out of date. Sort by timestamp before drawing any conclusion from it.
JMR-dev commented 2026-09-06 08:22:26 +00:00 (Migrated from github.com)

This is now hitting the gating API 37 leg repeatedly, and it lands on a different test each time — which is the strongest evidence yet that it is the environment rather than any test.

Four sightings on 2026-09-06, across four different PRs, all on E2E API 37 (gating, disable-system-ui: "1"), all with the same abort:

Test run failed to complete. Expected 64 tests, received 63.
  onError: commandError=false message=INSTRUMENTATION_ABORTED: System has crashed.
run PR diff test it took down
34015541233 #240 two androidTest files SafPickerRoundTripTest.pickingAFileThroughTheSystemPickerFillsInTheFileCard
34017895958 #242 one androidTest file SafPickerRoundTripTest.pickingAFileThroughTheSystemPickerFillsInTheFileCard
34019238667 #243 one androidTest file + a const NotificationCancelActionTest.theNotificationsCancelActionCancelsThatJob
34020234606 #244 documentation only NotificationCancelActionTest.theNotificationsCancelActionCancelsThatJob

The last row is the useful one: a docs-only diff cannot reach any instrumented test, and it failed the gating API 37 leg twice in a row before passing on a third attempt.

Why the varying victim matters

This ticket's title says the abort happens "after every test passes". These four show it also aborts mid-run — received: 63 of 64 — and the test it interrupts is whichever one happened to be executing. Two different tests, in two different packages, one of which drives system UI and one of which does not.

NotificationCancelActionTest fires a PendingIntent through WorkManager.createCancelPendingIntent, so it does touch system_server; SafPickerRoundTripTest touches DocumentsUI. Neither is a common cause with the other beyond "was running when system_server went down".

Deliberately not proposed: marking either with @FailsOnEmulatorApi37. That marker means "cannot pass on this image", and both pass on the same leg most of the time — NotificationCancelActionTest passed all five legs when it landed in #235. Marking an intermittent failure would move a working test out of the gating set and make the advisory job expect a failure that may not happen.

Cost, since that is what #190 says is the thing worth measuring

Every one of these needed a human to read the log, decide it was environmental, and re-run. Four times in one session, three of them on PRs whose diffs could not reach the E2E matrix at all.

**This is now hitting the gating API 37 leg repeatedly, and it lands on a different test each time** — which is the strongest evidence yet that it is the environment rather than any test. Four sightings on 2026-09-06, across four different PRs, all on `E2E API 37` (gating, `disable-system-ui: "1"`), all with the same abort: ``` Test run failed to complete. Expected 64 tests, received 63. onError: commandError=false message=INSTRUMENTATION_ABORTED: System has crashed. ``` | run | PR | diff | test it took down | |---|---|---|---| | [`34015541233`](https://github.com/JMR-dev/LibreMediaConverter/actions/runs/34015541233) | #240 | two `androidTest` files | `SafPickerRoundTripTest.pickingAFileThroughTheSystemPickerFillsInTheFileCard` | | [`34017895958`](https://github.com/JMR-dev/LibreMediaConverter/actions/runs/34017895958) | #242 | one `androidTest` file | `SafPickerRoundTripTest.pickingAFileThroughTheSystemPickerFillsInTheFileCard` | | [`34019238667`](https://github.com/JMR-dev/LibreMediaConverter/actions/runs/34019238667) | #243 | one `androidTest` file + a const | `NotificationCancelActionTest.theNotificationsCancelActionCancelsThatJob` | | [`34020234606`](https://github.com/JMR-dev/LibreMediaConverter/actions/runs/34020234606) | #244 | **documentation only** | `NotificationCancelActionTest.theNotificationsCancelActionCancelsThatJob` | The last row is the useful one: a docs-only diff cannot reach any instrumented test, and it failed the gating API 37 leg twice in a row before passing on a third attempt. ### Why the varying victim matters This ticket's title says the abort happens "after every test passes". These four show it also aborts *mid-run* — `received: 63` of 64 — and the test it interrupts is whichever one happened to be executing. Two different tests, in two different packages, one of which drives system UI and one of which does not. `NotificationCancelActionTest` fires a `PendingIntent` through `WorkManager.createCancelPendingIntent`, so it does touch `system_server`; `SafPickerRoundTripTest` touches DocumentsUI. Neither is a common cause with the other beyond "was running when system_server went down". **Deliberately not proposed:** marking either with `@FailsOnEmulatorApi37`. That marker means "cannot pass on this image", and both pass on the same leg most of the time — `NotificationCancelActionTest` passed all five legs when it landed in #235. Marking an intermittent failure would move a working test out of the gating set and make the advisory job expect a failure that may not happen. ### Cost, since that is what #190 says is the thing worth measuring Every one of these needed a human to read the log, decide it was environmental, and re-run. Four times in one session, three of them on PRs whose diffs could not reach the E2E matrix at all.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: JMR-dev/LibreMediaConverter#108