test/readspec-enum-fallbacks
5
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
225ecdd7e6 |
Split the API 37 leg so the part that works can gate
CI has never run the API level this app targets. The reason it did not was never "API 37 is untestable" -- it was that two tests fail on the emulator image, so one row would be permanently red or permanently allow-listed. This splits that row instead of choosing between those two. E2E API 37 gates. It runs 55 of the suite's 57 instrumented tests and must be green. E2E API 37 Media3 hardware transcode runs the other two, reports, and never blocks (continue-on-error). Both are driven off ONE marker, @FailsOnEmulatorApi37: the gating job passes notAnnotation, the advisory job passes annotation. Two lists would drift, and drift is silent in both directions -- a test that ends up in neither job reads as green. Excluding by class was not an option either: Media3EngineTest has four tests and two of them pass here, so notClass would have thrown away real coverage. The advisory job is named for what it runs, not for what we think is wrong. Both its tests drive a full H.264 -> H.265 hardware transcode, which is what distinguishes them from the two Media3EngineTest cases that pass -- those never decode video. The goldfish-decoder theory sits in a comment inside the job, where it can be corrected without renaming a check people have learned to look for; docs/api-37-emulator-crash.md keeps measurement and inference apart. The SystemUI disable moves into .github/scripts/e2e-run.sh behind E2E_DISABLE_SYSTEM_UI, unset everywhere but the two API 37 jobs, so the other four legs run byte-identical commands -- the same shape as E2E_EXTRA_GRADLE_ARGS. It runs BEFORE the streamed logcat starts, deliberately: `adb shell stop` would end that logcat and nothing restarts it, so a disable placed after it would cost the leg its diagnostics for the part of the run that matters. The body is probe v2 from api37-debug.yml -- the version measured 4/4 -- not the older one-round form: three rounds, waits for system_server to actually be gone, verifies against `pm list packages -d`, and requires a 45 s window with zero new aborts. The weaker probe reported success on a run that then started SystemUI eight more times. The caveat is written next to the row rather than left implicit: this leg runs with SystemUI disabled and the framework restarted under it, a device configuration no other leg and no Pixel run uses. Anything that touches system UI must not trust it, and the Pixel check before each release is still the only API 37 run with SystemUI intact. docs/api-37-emulator-crash.md's "So should CI take API 37?" said no on three reasons. Two were claims about CI that had never been measured; the section now carries the eight runs that measured them, and the third reason is what the split answers. docs/local-emulator.md and api37-debug.yml's header carried the same "the matrix stops at 36" claim and are corrected with it. CLAUDE.md is left alone deliberately -- its "CI's matrix therefore stops at API 36" clause is now false, and that correction is parked in the doc's existing "Correction owed to CLAUDE.md" section, where two others are already waiting. Making E2E API 37 an actually-required check is a repository-settings change and must come after this is on main: adding a required context that does not exist on the default branch blocks every PR. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
b3a705e3da |
Measure the API 36 control and record what CI cannot measure
Three additions to docs/api-37-emulator-crash.md, all from a CI investigation run through .github/workflows/api37-debug.yml. A third measured bullet: API 36 against API 37, back to back, same two tests, same renderer, same SystemUI-disable path. 37.0 fails both on c2.goldfish.h264.decoder (32660148155); 36 passes both in 4.603 s with the same decoder in its logcat (32660152961). That falsifies "the stripped configuration is what breaks these tests" -- a reading the other measurements never addressed, because they all compare against a device that still had SystemUI. It carries its two uncontrolled variables rather than dropping them: API 36's framework restart happened with zero aborts logged where API 37's had two, so a restart under an active abort loop is still uncontrolled; and the images differ on the encoder side, which is a second reason "broken h264 decoder" is the wrong shape of claim. The decoder-mechanism bullet is unchanged and still labelled inference. This adds a measurement next to it; it does not retract anything. The intact-SystemUI counterfactual is unmeasurable on a GitHub runner, and now says why. Seven dispatches, zero verdicts, with a mechanism rather than bad luck: while the framework crash-loops the guest cannot reliably create per-user private directories, so an app installed during the loop has no cache dir and the fixture copy dies in @Before before any codec exists. googlesdksetup and nexuslauncher hit the same thing. The result XML masks it behind an UninitializedPropertyAccessException in tearDown, which reads as a defect in this repository and is not one. Abort cadence corrected. "Roughly every 20 s" was the watchdog's sampling interval, not the cadence: measured gaps are 20-90 s, median 60-70 s, three to five per run, with sys.boot_completed held at 1 throughout. The wrong figure lived in api37-debug.yml's own comments, so that line is corrected too. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
acc71bcaee |
Verify the SystemUI disable instead of trusting what pm reported
Four dispatches of one configuration -- API 37.0, swiftshader_indirect,
SystemUI disabled -- came back three green and one not, and the odd one out
was not a different failure so much as the same run without the fix applied.
In 32646029143 `pm disable-user` reported `new state: disabled-user` and
SystemUI then started eight more times:
14:41:24 ActivityManager: Start proc 6412:com.android.systemui ... GradientColorWallpaper
14:45:05 ActivityManager: Start proc 17299:com.android.systemui ... GradientColorWallpaper
with ten more RegionSampling aborts and a surfaceflinger pid that never sat
still (489, 1570, 3524, 4396, 6038, 7987, 9732, 11520, 13208, 15048). The
framework is being SIGKILLed every twenty seconds while this runs, so a
package-state change can go down with the system_server that accepted it.
Two things were wrong, and the second is why the first went unnoticed:
- one disable attempt was treated as sufficient
- the wait after `adb shell stop` was not a wait. It asked `service check`
0.3 s later and got `found` from the system_server that was still on its
way out, so it never waited for anything. Both the good and the bad run
printed `services back after 5 s`, which is how a broken fix looked
identical to a working one.
Now: up to three rounds of disable -> take the framework down and confirm
system_server is actually gone -> bring it back -> verify the package is in
`pm list packages -d` -> require a 45 s window with zero new aborts. Nothing
is believed because a command said so.
Also adds measure_baseline, default true. The 45 s pre-measurement is what
makes the rate comparable with the local figures, but it is 45 s of
crash-looping before the disable has to land, which is a worse starting
point than a real leg would have.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
93398c4616 |
Resolve adb by path in the API 37 watchdog
The first three dispatches came back with every watchdog sample reading `boot=? surfaceflinger=none zygote64=none dma_aborts=0`, on runs where the device demonstrably booted and the action's own adb was working two steps away. The watchdog was not measuring anything. The emulator action puts platform-tools on PATH with core.addPath, which writes GITHUB_PATH and therefore only affects LATER steps. The watchdog is started before the action -- that is the whole point of it -- so it inherits the runner's own PATH, where a bare `adb` is not necessarily anything. Every call failed into `2>/dev/null` and the sampler dutifully recorded the silence as zero. It now resolves adb by path, preferring ANDROID_HOME, re-resolving on every iteration in case platform-tools arrives later, and echoing the path it settled on. The launch step prints ANDROID_HOME and `command -v adb` for the same reason: a repeat of this failure should be one line to spot, not three runs of quiet zeros. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
8fdad6e20b |
Add a dispatch-only workflow for the API 37 CI question
status_check.yml stops its E2E matrix at 36 and says the android-37.0 image is why. That is established locally under -gpu host and under ANGLE, and it is not established for CI: runners use -gpu swiftshader_indirect, and the one local measurement of that mode was void for a local reason -- Fedora denies execheap to SwiftShader's JIT, so the emulator died before the guest mattered. What CI does at API 37 has therefore never actually been measured. This is that E2E job with the matrix replaced by workflow_dispatch inputs, so a hypothesis costs a dispatch rather than a commit: renderer, API level, image target, channel, SystemUI disable, boot timeout, whether the suite runs at all, and free-form emulator and Gradle arguments. It triggers on nothing else and gates nothing. It calls .github/scripts/e2e-run.sh rather than forking it, and pins the same disk-size, ram-size, action SHAs and KVM setup as the job it copies, so a run here measures the renderer and not a different device. The watchdog is load-bearing rather than decorative. The emulator action calls killEmulator() from its own catch block, so a run whose emulator never boots is torn down before any script: line executes and leaves nothing behind -- which is the exact failure shape API 37 is suspected of. It starts before the action, samples sys.boot_completed, the surfaceflinger and zygote pids and the hasReadColorBufferDma abort count every 20 s, and keeps a rolling copy of the crash buffer so the last read survives the teardown. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |