Commit Graph
13 Commits
Author SHA1 Message Date
JMR-dev 7ae660ee37 Merge remote-tracking branch 'origin/main' into tools/api-37-emulator 2026-08-22 23:21:47 -05:00
JMR-devandClaude Opus 5 775a44753b Say that the default sweep is red on purpose, and narrow two claims
R16 / #25 -- the branch put API 37 into the default APIS list, where it is permanently two
failures short of green, so a bare `run-e2e.sh` exits 1 by design and nothing said so.
Somebody running it from habit, a wrapper or a hook gets a red exit forever and either
stops reading exit codes or debugs a normal state.

Documented rather than suppressed. The script's own comment already argued that an
expected-red level belongs in the exit code -- reversing that is the branch owner's call,
not a correction -- and the review's alternative needs an exact-set comparison of the
failing test names before it can subtract 37's contribution, which is a new mechanism that
cannot be validated without a device. So:

- the header now states the exit code (0 all green / 1 any level red / 2 refused to
  start), says a bare run is 1 by design and why, and gives `run-e2e.sh 33 34 35 36` as
  the sweep that can be green;
- a red sweep prints one note after the summary saying the same thing, because the exit
  code is read in the terminal and not in the docs -- but ONLY when 37.x is the only level
  that went red. `overall` is set by any red level, so a note keyed on "37 was in the
  list" would have called a genuine API 34 failure "by design", which is the defect this
  is meant to prevent, one layer up. mark_red records which level it was, where the loop
  already knows;
- docs/local-emulator.md says it where the default is documented.

R27 / #36 -- 792286a appended the caveat that the harness path reproduces a rate collapse
rather than a clean zero, but left "that is the confirmation that region sampling is the
sole trigger" standing three lines above it, which the caveat contradicts. Now "the
strongest evidence that region sampling is the dominant trigger", with the residue named:
no measurement here separates a second caller of the readback path from a disable that did
not fully take, and the file says so rather than picking one.

R28 / #37 -- "the capability is negotiated regardless of renderer" leaned on the string
search, which shows only that `ANDROID_EMU_read_color_buffer_dma` is implemented in one
shared component, not that it is negotiated on every path. The aborts are the actual
evidence -- the assertion that fires is `!hasReadColorBufferDma` and it fires under ANGLE
too -- and they suffice alone; the string search is demoted to a supporting note. Worth
getting right because the doc says the upstream report should lead with this model.

Two follow-ons that belong with R17 / #26 and land here rather than in their own commit:
bash runs a trap only between commands, so the handler starts when the foreground command
returns -- immediate under Ctrl-C, which reaches that command too, but not under a `kill
-INT` aimed at the script alone; that is now written next to the handler. And
delete_created_avds no longer discards avdmanager's status: an emulator that was SIGKILLed
did not get to remove its own lock files, avdmanager can refuse over them, and silence
there would leak exactly what the trap exists to clean up.

`bash -n` clean; the stub smoke harness (real script, fake SDK binaries, boot-failure path,
no Gradle and no emulator) now also checks that a 37-only red prints the note after the
summary, that a red API 34 alongside it suppresses the note, that a 34-only sweep says
nothing, and that a refused AVD deletion is reported. shellcheck is not installed here.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-22 23:17:19 -05:00
JMR-devandClaude Opus 5 961cfa72a2 Derive the suite size instead of writing it down in two documents
R4 / #13 and R20 / #29 are one defect: an absolute test total in an unregenerated
document, written the same day it went stale. This branch was cut at 22c7914, where
app/src/androidTest held 49 @Test methods; main is 57 (ReattachOnLaunchTest added eight
in ec969c4). So the release instruction "expect 49 / 0 / 0 / 2, and if you get 40 you are
on an old checkout" becomes false the moment this branch merges -- on the one check that
has no CI backstop -- and docs/local-emulator.md's headline promises a 49-test local
baseline main no longer produces.

Re-derived rather than renumbered, because a third total would go stale the same way:

- The total is the size of app/src/androidTest on the checkout that ran, and the reported
  total has equalled that checkout's @Test count everywhere it has been checked: 40 at
  edd6385 (the Pixel run), 49 at 22c7914 (the four local levels and API 37), 57 at
  18c53a3 (counted, not run). The new "Reading these totals" section states that, gives
  the one-line grep, and makes the *mismatch* the signal: a total that disagrees with
  your own checkout's count means an old checkout, a stale build or tests that never ran.
  The pre-release Pixel instruction now reads "that many tests, 0 failures, 0 errors, 2
  skipped" -- the invariant, not the total.
- Measurements are kept verbatim and anchored to 22c7914 (the sweep table, the API 35
  control, the tests="49" XML quote, the 51-on-screen console block). Only the claims
  built on top of them were rewritten.

Two claims went with the number, both of which a rebase would have preserved:

- "47 of 49" is not a defensible ratio when two of the 49 are skips. 49 = 45 passed + 2
  failed + 2 skipped, and that is what it now says.
- "against the Pixel's 49 of 49" and "matches the physical Pixel 10 Pro XL baseline of
  49 / 0 / 0 / 2 exactly" describe a run that never happened: the Pixel measured
  40 / 0 / 0 / 2 at edd6385, nine tests earlier, as the same file says a hundred lines
  further down. Both documents projected the local total onto the phone and called it a
  match. What compares between them is 0 failures and the same two skips.

Also re-derived in the CLAUDE.md wording docs/local-emulator.md proposes, since that text
is meant to be pasted out of the branch and would have carried "49 tests / 2 failures /
2 skipped" with it. CLAUDE.md itself is still untouched.

Counts re-checked with git grep at each of the three commits; nothing here needed a
device, and none was used.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-22 23:09:30 -05:00
JMR-devandClaude Opus 5 614af35647 Hold the corrections themselves to the standard they impose
Three defects in the three preceding commits, found on review. A commit set
whose subject is stale dates and inferred status cannot carry either.

Dates. Both correction blocks were stamped 2026-08-23. The commits are dated
2026-08-22, as is every other date in these two files and the review that
produced them -- a day in the future, in the one place a reader checks to see
how fresh a correction is. Corrected to the commit date, and the D6 note now
carries one too.

Coherence. The Status line was changed to say fix status "tracks main,
re-checked at 18c53a3" while "Last verified: 2026-08-22, against main at
903b43c" stood two lines below it, unchanged. A reader would take the whole
document as anchored to 903b43c -- exactly the failure being corrected. The two
anchors now say what each covers and that they move independently: the as-found
bodies are frozen at 903b43c, the status is not.

Inferred count. "detekt 0, lint clean" was carried over from the old text on the
strength of a BUILD SUCCESSFUL, which means "nothing above threshold", not
"nothing found" -- and this project deliberately keeps a real lint finding
visible (`informational += "UsableSpace"`), so "clean" was wrong as well as
unmeasured. Read off the reports instead: detekt.xml has zero <error> elements,
lint-results-debug.txt says 0 errors, 0 warnings, 1 hint. Stated that way, which
is also the convention the audit's own opening table already uses.

README's correction note is tightened from nine lines to six. It sits on the
first screen and nothing load-bearing is dropped.

Gate green: ktlintCheck, detekt, lintDebug.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-22 23:01:43 -05:00
JMR-devandClaude Opus 5 7e7f1301a7 Re-date the audit's testing section, which its own follow-up falsified
R14 / #23, R15 / #24. "On testing these" proposed a plan; the twelve fixes then
executed it, so the section describes a state that no longer exists. It is
re-dated rather than deleted -- the reasoning is why the test stack looks the
way it does -- with each stale claim marked where it stands.

R15 / #24, four statements, each checked against this checkout rather than
against another document:

- "exactly one dependency, testImplementation(libs.junit)". There are five:
  junit, robolectric, androidx.work.testing, the Compose BOM platform and
  compose-ui-test-junit4.
- "a testOptions { unitTests.isIncludeAndroidResources = true } block, which
  this module does not currently have at all". app/build.gradle.kts:105-110.
- "work-testing, compose-ui-test-junit4 and espresso-core ... have zero users."
  By import, work-testing has seven files under app/src/test and
  androidx.compose.ui.test has one. espresso-core really is still at zero, so
  that third is kept as the only part still standing.
- The preamble's "OutputPublisher, both ViewModels, both Workers and
  MainActivity have no JVM unit tests at all". 25 JVM test files were added over
  that set, 180 tests to 257.

That last one is also the derivation of CLAUDE.md's ~31% coverage figure, so the
reasoning is kept verbatim and only its tense and scope are fixed: the ~31% is
what those ~1,200 untested lines produced at 903b43c, it predates the new tests,
and jacoco has not been re-run. The figure itself is deliberately not touched
here -- re-measuring it and updating all three sites together is its own change.

R14 / #23. The audit asserted in two places that instrumented tests cannot run
on this host, and recorded the opposite in a third. 22c7914 is merged and
tools/local-emulator/run-e2e.sh runs API 33-36 locally, so the section's
impossibility argument for "pure seams plus Robolectric" is restated on the
grounds that survive -- speed and determinism, which is the weaker claim and
worth making honestly. D6's "it is an instrumented test, so it runs on CI, not
locally" gets the same treatment, and went further the other way: the
StateRestorationTester test that entry asked for is AppRootRestorationTest,
which runs under Robolectric on :app:testDebugUnitTest. The API 37 rule is
explicitly left standing -- that image is broken and the Pixel check before each
release is unaffected.

No as-found body was rewritten; D6's note is appended to its "Fix direction and
test" guidance, not to its description of the defect. Gate green: ktlintCheck,
detekt, lintDebug.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-22 22:57:01 -05:00
JMR-devandClaude Opus 5 9f0bc9d19b Correct the defect audit's status metadata, which went stale in hours
R3 / #12, R12 / #21, R13 / #22. Three status claims in the audit were false
against `main` at 18c53a3. The audit is read as the work queue for the ticket
phase, so each correction is stated in the document rather than made quietly --
a reader who believed the old claim needs to see that it changed.

R3 / #12. D5 and D7 were marked `open` and "in progress on
fix/space-proxy-and-notification". Both merged hours before: b86df47 (D5) and
c2e6344 (D7) are ancestors of 18c53a3 (`git merge-base --is-ancestor`), and so
is the branch. Anyone working the table would have re-implemented merged work.
The header count was also wrong in its own arithmetic -- "ten fixed, two in
progress, one parked" covers thirteen of sixteen entries, dropping D12, D15 and
D16. It now states four numbers that sum, and says the Fix column tracks `main`
at a named commit so the next reader knows what it is relative to.

R12 / #21. "Where the fixes live" sent the reader to `feat/defect-fixes-base`,
for which `git show-ref` finds nothing -- no local ref, no remote, deleted when
it merged -- and quoted 242 JVM tests. A real run on this branch gives 257 / 0 /
0 / 0 across 34 classes; the 15 missing are UnknownInputSizeTest, SpaceCheckTest
and ProgressNotificationTest. The replacement names `main`, anchors the total to
this commit and gives the command to re-derive it, since the number is only as
fresh as the document.

The same paragraph's "API 33-36 now run locally, 49 tests each" is anchored to
22c7914, the commit that measured it: androidTest is 57 `@Test` on `main`, and
an unanchored total invites a reader to mistake drift for breakage. The
surviving invariant -- 0 failures, 0 errors, 2 skipped, same total at every
level and on the Pixel -- is stated instead.

R13 / #22. D11 was marked `merged`, but 7db3200's own body says one of its four
rows was deliberately skipped: "Not touched: OutputPublisher's hasSpaceFor
KDoc". Still true -- nothing in OutputPublisher.kt mentions D1 -- and that row
is the one place a reader of the code would learn the parked defect exists. The
summary row now says "less the OutputPublisher KDoc row -- held with D1", with
the reasoning at the entry.

The as-found bodies are untouched throughout, including D11's own item table:
corrections are carried as marked editorial notes beside them, the pattern D1's
entry already uses. Gate green: ktlintCheck, detekt, lintDebug.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-22 22:54:33 -05:00
JMR-devandClaude Opus 5 792286a2d7 Stop three claims in the API 37 doc outrunning their evidence
Three corrections, all narrowing:

- angle_indirect and swangle_indirect are not two independent renderers here. Both
  logged gles_mode_selected:swangle with the same adapter, differing only in the
  Vulkan backend underneath -- unlike at API 33-36, where angle_indirect resolves to
  ANGLE on llvmpipe. What is 7-for-7 is the host-GLES-versus-not split, not "two
  renderers agree".

- "Disabling SystemUI stops the crashes entirely" was one 180-second measurement on a
  device that had been up twelve minutes. The harness path reproduces a rate collapse,
  not a zero: its own quiet check printed 1 abort in 45 s and 4 across the run. A
  47-second Gradle run survives that; a five-minute one might not.

- "Reproduced twice" conflated two routes. The 49/2/0/2 came back from a hand-driven
  sequence and from the harness, which corroborates the numbers, but the harness path
  itself has one green measurement.

Also records what the doc never said: from 37.1 onward Google ships only 16 KB-page
x86_64 images, so page-size alignment is a prerequisite for that path rather than a
detail. All 20 libraries in the committed FFmpeg AAR are 0x4000-aligned, checked
before the first ps16k boot -- which is why 37.1 reproducing the abort means the
gralloc bug and not a page-size mismatch.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-22 22:10:50 -05:00
JMR-devandClaude Opus 5 739bffa5a0 Re-derive the API 37 emulator failure: it is the renderer, not the image
docs/api-37-emulator-crash.md claimed "Both swiftshader_indirect and host crash...
The crash is in the gralloc mapper, below the renderer." Re-measured, seven runs,
one variable each: that is wrong. The mapper is below the renderer, but whether its
bad path is reached is not.

  -gpu host              gles_mode_selected:host    never boots  (57-71 aborts, looping)
  -gpu swangle_indirect  gles_mode_selected:swangle boots, 85 s  (1 abort)
  -gpu angle_indirect    gles_mode_selected:swangle boots, 112 s (2 aborts)

The old claim rested on two samples of two different things, neither of them ANGLE:
the local swiftshader_indirect sample was void, because on this host every
SwiftShader-GLES launch segfaults the emulator before the guest matters (the
execheap bug in docs/local-emulator.md, not understood when that file was written),
and the CI sample was a single swiftshader_indirect run.

Also re-derived, and null: android-37.1 rev 8 -- a stable REL image the doc's own
"new image revision" trigger was too narrow to catch -- fails identically;
-feature -GLDMA,-GLDMA2,-GLDirectMem is accepted and changes nothing; the image's
advancedFeatures.ini is byte-identical to API 36's but for one camera line; and
there is still no ATD image above API 36.

The mechanism, end to end: SystemUI registers a nav-bar luma-sampling listener,
SurfaceFlinger's RegionSamplingThread locks a GraphicBuffer, Gralloc5 routes into
GoldfishMapper::readFromHost, which asserts, and init SIGKILLs zygote in response --
so the framework restarts under the test run. Disabling SystemUI removes the
listener and the aborts stop dead: 0 in 180 s, against 10-11 per 150 s.

So run-e2e.sh now covers API 37: renderer chosen per level (33-36 need host, 37
must not have it), dotted image labels, SystemUI disabled followed by a deliberate
stop/start, and an abort count printed on every 37 row. The result is 49 tests, 2
failures, 0 errors, 2 skipped, reproduced twice. The two failures are
Media3EngineTest on c2.goldfish.h264.decoder; API 35 under the identical renderer is
49/0/0/2 green, so they are the image and not the renderer.

CI's matrix should still stop at 36, for reasons now written down rather than
assumed. CLAUDE.md is left alone; a replacement bullet is proposed in the doc.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-22 22:09:19 -05:00
Jason Ross fc2afff32e Merge branch 'main' into tools/local-emulator 2026-08-22 21:08:42 -05:00
JMR-devandClaude Opus 5 ef3d87e12d Write down what is actually wrong with this app, and how we know
detekt reports zero findings and there is no baseline, no @Suppress and no
tools:ignore anywhere -- so the static-analysis gate is green and honest, and
it is not where the defects are. They are in the Android-framework edge the
linters cannot see into: OutputPublisher, both ViewModels, both Workers and
MainActivity, which between them have no JVM unit tests at all and account for
most of the ~31% coverage figure.

Sixteen entries. Each records what is wrong, how confident we are that it is
wrong, how to provoke it, and what a fix would have to decide. The confidence
labels are load-bearing: four entries were driven on a physical Pixel 10 Pro XL
running API 37, and they are marked differently from the ones that are still
inspection only.

The device pass earned its keep by contradicting us. D1 -- the one defect that
was already known and deferred, the UsableSpace lint finding -- did not
reproduce. getAllocatableBytes measured 500 MiB SMALLER than usableSpace, and
writing 3 GB into the app's own cache moved both numbers identically, so no
cache counted as reclaimable at 66% free. The entry keeps the falsified
prediction next to the measurement that killed it, because that is the useful
part.

Two entries, D15 and D16, were found while fixing others and are recorded
rather than folded in silently. D16 is the one worth reading: two individually
correct fixes compose into a gap neither of them owns.

Entry bodies describe each defect as found and are deliberately not rewritten
as fixes land. This is the record of what was wrong, not a changelog; the
summary table carries the fix status.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-22 20:34:42 -05:00
JMR-devandClaude Opus 5 22c7914395 Find out why the emulators segfault, and make them run
CLAUDE.md has said "Emulators segfault on this host -- qemu dies on every AVD"
since the E2E matrix landed, and the PR that introduced it called the failure
"exit 139 across three AVDs and both GPU backends, environmental". That is
accurate about the symptom and wrong about the cause, and the cost of being
wrong was the whole instrumented suite being unrunnable here.

SwiftShader's Reactor JIT writes generated GLES shader code onto the heap and
mprotects it executable. Fedora's SELinux policy denies that -- execheap is not
granted to unconfined_t and selinuxuser_execheap is off -- so the mprotect
fails and the emulator takes SIGSEGV the moment it calls the routine it just
generated. The AVC denial and the core are the same event, one second apart.

The predictor is mechanical and held 7 for 7 across every -gpu mode: a run
crashes if and only if it dlopens gles_swiftshader/libGLESv2.so. host,
angle_indirect and swangle_indirect boot. auto, off, guest and
swiftshader_indirect crash -- and auto is the default, which is why the failure
looked universal rather than renderer-specific.

tools/local-emulator/run-e2e.sh picks a renderer that works and refuses the
ones that do not. It reuses .github/scripts/e2e-run.sh rather than forking it,
so the local and CI diagnostics cannot drift; the one change there adds an
optional E2E_EXTRA_GRADLE_ARGS that is unset in CI, so CI runs byte-identical
commands.

The API 33-36 sweep has now been run and is written down. All four levels are green
on a local emulator and match the physical Pixel 10 Pro XL baseline exactly: 49 tests,
0 failures, 0 errors, 2 skipped, every level. Those counts come from the result XML,
not the UTP console counter, which double-counts skips and reported "Finished 51 tests"
on all four. No boot log dlopens SwiftShader GLES and the sweep window holds no AVC
denial and no qemu core -- which is confirmation of the mode matrix's first row rather
than new coverage, since every one of these runs is -gpu host. The table is still seven
modes measured once each.

Two things the sweep surfaced that the doc now records: pre-build before sweeping, or a
fresh checkout spends API 33's 20-minute wrapper budget compiling and wedges before a
test runs; and the device pinning is untested by this run, because the Pixel dropped off
USB five seconds before it started.

Still offered for review rather than applied: the CLAUDE.md correction the doc drafts.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-22 20:02:46 -05:00
JMR-devandClaude Opus 5 edd6385bf7 Record that the suite passes on real API 37 hardware
The doc reasoned that the WorkManager and lateinit failures in CI were
downstream of the broken framework rather than real defects, but said so as
inference and flagged that only a healthy API 37 device could settle it.

One was available. The full instrumented suite runs green on a Pixel 10 Pro XL
on Android 17 -- a release build, not a preview -- with 40 tests, 0 failures,
2 skipped, both skips being benchmarks that assume sample files present.
ConversionWorkerTest and ConcatWorkerTest drive a real WorkManager round trip
and are among the tests that failed that way in CI; they pass on hardware.

So the bug is confined to the emulator image, and the gap left by the missing
matrix row is automated coverage rather than confidence in the app. Noted that
the suite should be run on a physical API 37 device before each release while
the row is absent.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-21 21:37:49 -05:00
JMR-devandClaude Opus 5 4e6fe6b75a Drop API 37 from the E2E matrix and write down why
The android-37.0 emulator image crash-loops surfaceflinger inside its own
gralloc mapper: RegionSamplingThread calls GraphicBuffer::lock, which reaches
GoldfishMapper::readFromHost, which asserts that the host has not negotiated
ReadColorBufferDma. It has, so surfaceflinger aborts, restarts, and aborts
again. Nothing this app does can survive that, and it reproduces on a GitHub
runner under swiftshader_indirect and on a workstation under -gpu host alike.

There is no ATD image at android-37.0 to fall back to, and -feature -GLDMA is
accepted by the emulator but does not prevent the assertion.

Correcting the previous commit, which is already pushed so its message stands:
ram-size was not the cause of that failure. Setting it did move the job from
failing at install to failing during the test run, which is how the real
crash became visible, but at 2560M the guest had 1.5 GB free when it died.
The setting is kept because the emulator's own floor varies by API level --
2048M at 33, 2560M at 34 to 36 -- and pinning it makes the matrix uniform.

Also corrected: a comment claiming this could not be reproduced locally. It
can, and the local crash was the same one all along.

Dropped the dmesg probe. adb shell is not root, so klogctl is denied and it
only ever printed a permission error -- which a later reader would reasonably
misread as "no OOM kills".

docs/api-37-emulator-crash.md carries the evidence, the ruled-out fixes, the
reproduction, and how to file it upstream, so re-adding the row later starts
from what is already known rather than from scratch.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-21 21:35:04 -05:00