Commit Graph
7 Commits
Author SHA1 Message Date
JMR-devandClaude Opus 5 7e7f1301a7 Re-date the audit's testing section, which its own follow-up falsified
R14 / #23, R15 / #24. "On testing these" proposed a plan; the twelve fixes then
executed it, so the section describes a state that no longer exists. It is
re-dated rather than deleted -- the reasoning is why the test stack looks the
way it does -- with each stale claim marked where it stands.

R15 / #24, four statements, each checked against this checkout rather than
against another document:

- "exactly one dependency, testImplementation(libs.junit)". There are five:
  junit, robolectric, androidx.work.testing, the Compose BOM platform and
  compose-ui-test-junit4.
- "a testOptions { unitTests.isIncludeAndroidResources = true } block, which
  this module does not currently have at all". app/build.gradle.kts:105-110.
- "work-testing, compose-ui-test-junit4 and espresso-core ... have zero users."
  By import, work-testing has seven files under app/src/test and
  androidx.compose.ui.test has one. espresso-core really is still at zero, so
  that third is kept as the only part still standing.
- The preamble's "OutputPublisher, both ViewModels, both Workers and
  MainActivity have no JVM unit tests at all". 25 JVM test files were added over
  that set, 180 tests to 257.

That last one is also the derivation of CLAUDE.md's ~31% coverage figure, so the
reasoning is kept verbatim and only its tense and scope are fixed: the ~31% is
what those ~1,200 untested lines produced at 903b43c, it predates the new tests,
and jacoco has not been re-run. The figure itself is deliberately not touched
here -- re-measuring it and updating all three sites together is its own change.

R14 / #23. The audit asserted in two places that instrumented tests cannot run
on this host, and recorded the opposite in a third. 22c7914 is merged and
tools/local-emulator/run-e2e.sh runs API 33-36 locally, so the section's
impossibility argument for "pure seams plus Robolectric" is restated on the
grounds that survive -- speed and determinism, which is the weaker claim and
worth making honestly. D6's "it is an instrumented test, so it runs on CI, not
locally" gets the same treatment, and went further the other way: the
StateRestorationTester test that entry asked for is AppRootRestorationTest,
which runs under Robolectric on :app:testDebugUnitTest. The API 37 rule is
explicitly left standing -- that image is broken and the Pixel check before each
release is unaffected.

No as-found body was rewritten; D6's note is appended to its "Fix direction and
test" guidance, not to its description of the defect. Gate green: ktlintCheck,
detekt, lintDebug.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-22 22:57:01 -05:00
JMR-devandClaude Opus 5 9f0bc9d19b Correct the defect audit's status metadata, which went stale in hours
R3 / #12, R12 / #21, R13 / #22. Three status claims in the audit were false
against `main` at 18c53a3. The audit is read as the work queue for the ticket
phase, so each correction is stated in the document rather than made quietly --
a reader who believed the old claim needs to see that it changed.

R3 / #12. D5 and D7 were marked `open` and "in progress on
fix/space-proxy-and-notification". Both merged hours before: b86df47 (D5) and
c2e6344 (D7) are ancestors of 18c53a3 (`git merge-base --is-ancestor`), and so
is the branch. Anyone working the table would have re-implemented merged work.
The header count was also wrong in its own arithmetic -- "ten fixed, two in
progress, one parked" covers thirteen of sixteen entries, dropping D12, D15 and
D16. It now states four numbers that sum, and says the Fix column tracks `main`
at a named commit so the next reader knows what it is relative to.

R12 / #21. "Where the fixes live" sent the reader to `feat/defect-fixes-base`,
for which `git show-ref` finds nothing -- no local ref, no remote, deleted when
it merged -- and quoted 242 JVM tests. A real run on this branch gives 257 / 0 /
0 / 0 across 34 classes; the 15 missing are UnknownInputSizeTest, SpaceCheckTest
and ProgressNotificationTest. The replacement names `main`, anchors the total to
this commit and gives the command to re-derive it, since the number is only as
fresh as the document.

The same paragraph's "API 33-36 now run locally, 49 tests each" is anchored to
22c7914, the commit that measured it: androidTest is 57 `@Test` on `main`, and
an unanchored total invites a reader to mistake drift for breakage. The
surviving invariant -- 0 failures, 0 errors, 2 skipped, same total at every
level and on the Pixel -- is stated instead.

R13 / #22. D11 was marked `merged`, but 7db3200's own body says one of its four
rows was deliberately skipped: "Not touched: OutputPublisher's hasSpaceFor
KDoc". Still true -- nothing in OutputPublisher.kt mentions D1 -- and that row
is the one place a reader of the code would learn the parked defect exists. The
summary row now says "less the OutputPublisher KDoc row -- held with D1", with
the reasoning at the entry.

The as-found bodies are untouched throughout, including D11's own item table:
corrections are carried as marked editorial notes beside them, the pattern D1's
entry already uses. Gate green: ktlintCheck, detekt, lintDebug.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-22 22:54:33 -05:00
Jason Ross fc2afff32e Merge branch 'main' into tools/local-emulator 2026-08-22 21:08:42 -05:00
JMR-devandClaude Opus 5 ef3d87e12d Write down what is actually wrong with this app, and how we know
detekt reports zero findings and there is no baseline, no @Suppress and no
tools:ignore anywhere -- so the static-analysis gate is green and honest, and
it is not where the defects are. They are in the Android-framework edge the
linters cannot see into: OutputPublisher, both ViewModels, both Workers and
MainActivity, which between them have no JVM unit tests at all and account for
most of the ~31% coverage figure.

Sixteen entries. Each records what is wrong, how confident we are that it is
wrong, how to provoke it, and what a fix would have to decide. The confidence
labels are load-bearing: four entries were driven on a physical Pixel 10 Pro XL
running API 37, and they are marked differently from the ones that are still
inspection only.

The device pass earned its keep by contradicting us. D1 -- the one defect that
was already known and deferred, the UsableSpace lint finding -- did not
reproduce. getAllocatableBytes measured 500 MiB SMALLER than usableSpace, and
writing 3 GB into the app's own cache moved both numbers identically, so no
cache counted as reclaimable at 66% free. The entry keeps the falsified
prediction next to the measurement that killed it, because that is the useful
part.

Two entries, D15 and D16, were found while fixing others and are recorded
rather than folded in silently. D16 is the one worth reading: two individually
correct fixes compose into a gap neither of them owns.

Entry bodies describe each defect as found and are deliberately not rewritten
as fixes land. This is the record of what was wrong, not a changelog; the
summary table carries the fix status.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-22 20:34:42 -05:00
JMR-devandClaude Opus 5 22c7914395 Find out why the emulators segfault, and make them run
CLAUDE.md has said "Emulators segfault on this host -- qemu dies on every AVD"
since the E2E matrix landed, and the PR that introduced it called the failure
"exit 139 across three AVDs and both GPU backends, environmental". That is
accurate about the symptom and wrong about the cause, and the cost of being
wrong was the whole instrumented suite being unrunnable here.

SwiftShader's Reactor JIT writes generated GLES shader code onto the heap and
mprotects it executable. Fedora's SELinux policy denies that -- execheap is not
granted to unconfined_t and selinuxuser_execheap is off -- so the mprotect
fails and the emulator takes SIGSEGV the moment it calls the routine it just
generated. The AVC denial and the core are the same event, one second apart.

The predictor is mechanical and held 7 for 7 across every -gpu mode: a run
crashes if and only if it dlopens gles_swiftshader/libGLESv2.so. host,
angle_indirect and swangle_indirect boot. auto, off, guest and
swiftshader_indirect crash -- and auto is the default, which is why the failure
looked universal rather than renderer-specific.

tools/local-emulator/run-e2e.sh picks a renderer that works and refuses the
ones that do not. It reuses .github/scripts/e2e-run.sh rather than forking it,
so the local and CI diagnostics cannot drift; the one change there adds an
optional E2E_EXTRA_GRADLE_ARGS that is unset in CI, so CI runs byte-identical
commands.

The API 33-36 sweep has now been run and is written down. All four levels are green
on a local emulator and match the physical Pixel 10 Pro XL baseline exactly: 49 tests,
0 failures, 0 errors, 2 skipped, every level. Those counts come from the result XML,
not the UTP console counter, which double-counts skips and reported "Finished 51 tests"
on all four. No boot log dlopens SwiftShader GLES and the sweep window holds no AVC
denial and no qemu core -- which is confirmation of the mode matrix's first row rather
than new coverage, since every one of these runs is -gpu host. The table is still seven
modes measured once each.

Two things the sweep surfaced that the doc now records: pre-build before sweeping, or a
fresh checkout spends API 33's 20-minute wrapper budget compiling and wedges before a
test runs; and the device pinning is untested by this run, because the Pixel dropped off
USB five seconds before it started.

Still offered for review rather than applied: the CLAUDE.md correction the doc drafts.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-22 20:02:46 -05:00
JMR-devandClaude Opus 5 edd6385bf7 Record that the suite passes on real API 37 hardware
The doc reasoned that the WorkManager and lateinit failures in CI were
downstream of the broken framework rather than real defects, but said so as
inference and flagged that only a healthy API 37 device could settle it.

One was available. The full instrumented suite runs green on a Pixel 10 Pro XL
on Android 17 -- a release build, not a preview -- with 40 tests, 0 failures,
2 skipped, both skips being benchmarks that assume sample files present.
ConversionWorkerTest and ConcatWorkerTest drive a real WorkManager round trip
and are among the tests that failed that way in CI; they pass on hardware.

So the bug is confined to the emulator image, and the gap left by the missing
matrix row is automated coverage rather than confidence in the app. Noted that
the suite should be run on a physical API 37 device before each release while
the row is absent.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-21 21:37:49 -05:00
JMR-devandClaude Opus 5 4e6fe6b75a Drop API 37 from the E2E matrix and write down why
The android-37.0 emulator image crash-loops surfaceflinger inside its own
gralloc mapper: RegionSamplingThread calls GraphicBuffer::lock, which reaches
GoldfishMapper::readFromHost, which asserts that the host has not negotiated
ReadColorBufferDma. It has, so surfaceflinger aborts, restarts, and aborts
again. Nothing this app does can survive that, and it reproduces on a GitHub
runner under swiftshader_indirect and on a workstation under -gpu host alike.

There is no ATD image at android-37.0 to fall back to, and -feature -GLDMA is
accepted by the emulator but does not prevent the assertion.

Correcting the previous commit, which is already pushed so its message stands:
ram-size was not the cause of that failure. Setting it did move the job from
failing at install to failing during the test run, which is how the real
crash became visible, but at 2560M the guest had 1.5 GB free when it died.
The setting is kept because the emulator's own floor varies by API level --
2048M at 33, 2560M at 34 to 36 -- and pinning it makes the matrix uniform.

Also corrected: a comment claiming this could not be reproduced locally. It
can, and the local crash was the same one all along.

Dropped the dmesg probe. adb shell is not root, so klogctl is denied and it
only ever printed a permission error -- which a later reader would reasonably
misread as "no OOM kills".

docs/api-37-emulator-crash.md carries the evidence, the ruled-out fixes, the
reproduction, and how to file it upstream, so re-adding the row later starts
from what is already known rather than from scratch.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-21 21:35:04 -05:00