E2E_DISABLE_SYSTEM_UI does not disable SystemUI: pm disable-user is set, survives, and is not acted on #246

Open
opened 2026-09-06 13:52:33 +00:00 by JMR-dev · 0 comments
JMR-dev commented 2026-09-06 13:52:33 +00:00 (Migrated from github.com)

Split out of #245, which stops the harness claiming otherwise but does not fix this.

E2E_DISABLE_SYSTEM_UI exists to remove RegionSamplingThread, the nav-bar luma sampler that
reaches this image's gralloc mapper and aborts surfaceflinger
(docs/api-37-emulator-crash.md). It has never done that, for two independent reasons, and
both were measured on 2026-09-05.

1. The framework restart never ran

adb shell stop and adb shell start are root-only and adbd is not root on a booted emulator.
All three copies of the logic called them without adb root, so every API 37 leg ever run
printed Must be root twice — green legs and red alike — between lines that read as if the
restart had happened:

  restarting the framework
Must be root
  system_server down after ~40 s
Must be root
  services back after ~5 s

Neither number was an observation: the pidof loop breaks when the process is gone and
otherwise falls out at its last iteration, and the old code printed the iteration count either
way. tools/local-emulator/run-e2e.sh was worse — it sent both to /dev/null, so its
Must be root was never even visible.

2. And pm disable-user does not keep SystemUI down anyway

This is the part that matters, because it is the one a restart would not have fixed.

On CI, gating leg of run 34006456986. pm disable-user is accepted at 02:28:37.9 and the
package really is in pm list packages -d at 02:29:33 — and SystemUI is started at 02:28:39.5
and again at 02:28:52.3, the second of which (pid 4275) is alive for the whole instrumentation
run, logging WindowManagerShell … app=com.android.systemui minutes after the harness prints
final state: SystemUI disabled. The two restarts are the image's own gralloc aborts killing
the framework; SystemUI comes back through both, disabled or not.

Locally on android-37.0, with the package verified disabled-user before and after a
deliberate, working stop; start:

  package still disabled? YES
  processes:
   9275  00:17 system_server
   9695  00:14 com.android.systemui     <- started 3 s after system_server

So the flag is set, survives, and is not acted on.

What #245 already did, and what it deliberately did not

It removed the restart from all three copies rather than repairing it, because making it real
is worse than leaving it broken: api37-debug run 34010167885 shows a stop landing ~2 s
after the pm call killing system_server before PackageManager flushes its delayed write of
package restrictions, so the state is lost (NOT DISABLED after the restart, three rounds) and
the leg reported expected: 0, received: 0 — Starting 0 tests. A 15 s pause before the stop
does make the state survive (bisected locally) and still does not help, per §2.

It kept the pm disable-user call, because every green leg and every number quoted about this
row was measured with it applied, and kept the 45-second no-new-aborts window, which is the part
that was always doing the work: in 34006456986 the boot aborts land at 02:28:18 and 02:28:43 and
that wait is what puts instrumentation at 02:32:42, after them rather than inside one. The log
now prints the true thing beside the misleading one, verified on run 34011072884:

  pm list packages -d: com.android.systemui is in it
  com.android.systemui pid: 1987
final state: com.android.systemui is disabled in pm (it still runs)

Why this is still worth a ticket

Two reasons, and the first is not about API 37.

  • Nobody knows what would keep SystemUI down on this image, and nothing was tried beyond
    pm disable-user. Candidates not evaluated: pm disable rather than disable-user,
    am force-stop after boot completes, killing the process and letting the disabled state stop
    the restart, or leaving SystemUI alone and disabling only the nav bar. If one works, the
    region-sampling trigger really can be removed and the abort rate on this leg falls — which is
    the thing the 45-second wait is currently working around rather than fixing.
  • A misnamed knob is a trap for the next reader. E2E_DISABLE_SYSTEM_UI, the
    disable_region_sampling function and api37-debug.yml's disable_system_ui input all name
    something that does not happen. #245 keeps the names and documents them, because a rename
    touches the matrix row, both workflows and two documents; that trade is worth revisiting
    separately rather than inside a flake fix.

Done means

Either a mechanism that demonstrably keeps com.android.systemui out of ps across a framework
restart on android-37.0 — with the abort rate measured before and after, since that is the only
reason to want it — or a recorded decision that the 45-second quiet window is the answer and the
three knobs get renamed to say so.

Do not accept pm list packages -d as evidence on its own. That is exactly what made this
look like it worked for two weeks; pidof com.android.systemui is the check that bites.

_Split out of #245, which stops the harness claiming otherwise but does not fix this._ `E2E_DISABLE_SYSTEM_UI` exists to remove `RegionSamplingThread`, the nav-bar luma sampler that reaches this image's gralloc mapper and aborts `surfaceflinger` (`docs/api-37-emulator-crash.md`). It has never done that, for two independent reasons, and both were measured on 2026-09-05. ### 1. The framework restart never ran `adb shell stop` and `adb shell start` are root-only and adbd is not root on a booted emulator. All three copies of the logic called them without `adb root`, so every API 37 leg ever run printed `Must be root` twice — green legs and red alike — between lines that read as if the restart had happened: ``` restarting the framework Must be root system_server down after ~40 s Must be root services back after ~5 s ``` Neither number was an observation: the `pidof` loop breaks when the process is gone and otherwise falls out at its last iteration, and the old code printed the iteration count either way. `tools/local-emulator/run-e2e.sh` was worse — it sent both to `/dev/null`, so its `Must be root` was never even visible. ### 2. And `pm disable-user` does not keep SystemUI down anyway This is the part that matters, because it is the one a restart would not have fixed. **On CI, gating leg of run 34006456986.** `pm disable-user` is accepted at 02:28:37.9 and the package really is in `pm list packages -d` at 02:29:33 — and SystemUI is started at 02:28:39.5 and again at 02:28:52.3, the second of which (pid 4275) is alive for the whole instrumentation run, logging `WindowManagerShell … app=com.android.systemui` minutes after the harness prints `final state: SystemUI disabled`. The two restarts are the image's own gralloc aborts killing the framework; SystemUI comes back through both, disabled or not. **Locally on `android-37.0`**, with the package verified `disabled-user` before *and* after a deliberate, working `stop; start`: ``` package still disabled? YES processes: 9275 00:17 system_server 9695 00:14 com.android.systemui <- started 3 s after system_server ``` So the flag is set, survives, and is not acted on. ### What #245 already did, and what it deliberately did not It removed the restart from all three copies rather than repairing it, because making it real is **worse than leaving it broken**: `api37-debug` run 34010167885 shows a `stop` landing ~2 s after the `pm` call killing `system_server` before PackageManager flushes its delayed write of package restrictions, so the state is lost (`NOT DISABLED after the restart`, three rounds) and the leg reported `expected: 0, received: 0` — `Starting 0 tests`. A 15 s pause before the stop does make the state survive (bisected locally) and still does not help, per §2. It kept the `pm disable-user` call, because every green leg and every number quoted about this row was measured with it applied, and kept the 45-second no-new-aborts window, which is the part that was always doing the work: in 34006456986 the boot aborts land at 02:28:18 and 02:28:43 and that wait is what puts instrumentation at 02:32:42, after them rather than inside one. The log now prints the true thing beside the misleading one, verified on run 34011072884: ``` pm list packages -d: com.android.systemui is in it com.android.systemui pid: 1987 final state: com.android.systemui is disabled in pm (it still runs) ``` ### Why this is still worth a ticket Two reasons, and the first is not about API 37. - **Nobody knows what would keep SystemUI down on this image**, and nothing was tried beyond `pm disable-user`. Candidates not evaluated: `pm disable` rather than `disable-user`, `am force-stop` after boot completes, killing the process and letting the disabled state stop the restart, or leaving SystemUI alone and disabling only the nav bar. If one works, the region-sampling trigger really can be removed and the abort rate on this leg falls — which is the thing the 45-second wait is currently working around rather than fixing. - **A misnamed knob is a trap for the next reader.** `E2E_DISABLE_SYSTEM_UI`, the `disable_region_sampling` function and `api37-debug.yml`'s `disable_system_ui` input all name something that does not happen. #245 keeps the names and documents them, because a rename touches the matrix row, both workflows and two documents; that trade is worth revisiting separately rather than inside a flake fix. ### Done means Either a mechanism that demonstrably keeps `com.android.systemui` out of `ps` across a framework restart on `android-37.0` — with the abort rate measured before and after, since that is the only reason to want it — or a recorded decision that the 45-second quiet window is the answer and the three knobs get renamed to say so. **Do not accept `pm list packages -d` as evidence on its own.** That is exactly what made this look like it worked for two weeks; `pidof com.android.systemui` is the check that bites.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: JMR-dev/LibreMediaConverter#246