Commit Graph
402 Commits
Author SHA1 Message Date
JMR-devandClaude Opus 5 7d1d3191a9 Resolve the gate's cache dir with --git-common-dir, and say when it cannot (#258)
CACHE_DIR was the literal ".git/lmc-verify". In a linked worktree `.git` is a
FILE containing `gitdir: ...`, so `mkdir -p .git/lmc-verify` fails with "Not a
directory" -- and because the write is the last thing the script does, it failed
while the gate still printed green and exited 0. Every commit and push from a
worktree then re-swept API 33-36 for nothing, silently. That is the worst shape a
cache can fail in: invisible and expensive, and it was found by an agent paying
for it four times over rather than by the tool saying anything.

Measured both ways: in a worktree the old expression gives
`mkdir: cannot create directory '.git': Not a directory`, exit 1; `git rev-parse
--git-common-dir` gives the real path and exit 0.

--git-common-dir rather than --git-dir so the cache is SHARED between worktrees.
The key is the app/src tree hash, and identical content is identical content
whichever worktree produced it -- a sweep run in one is evidence for all of them.

The write also stops being silent. record_sweep() prints when it cannot record,
because a cache that never fills looks exactly like one that is working.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-06 17:54:12 -05:00
Jason Ross 68b863fbdb Merge pull request #257 from JMR-dev/chore/gate-runs-shellcheck
Run shellcheck in the local gate, at CI's exact pin
2026-09-06 16:28:29 -05:00
JMR-devandClaude Opus 5 f4174e5b06 Run actionlint in the gate too, at CI's exact pin
The other half of the hole the previous commit closed. `git ls-files '*.sh'` does
not match workflow `run:` blocks, and a good deal of this repo's bash lives
there -- so a workflow edit was still the case where the gate passed and CI's
Static analysis leg went red.

Pinned by digest, read out of status_check.yml rather than copied, for the reason
the shellcheck section gives and for actionlint's own: its documented install is
`curl | bash` off a moving branch, which does not belong in a repo that pins every
action by SHA.

The container runtime detection and the SELinux `:z` mount option are hoisted out
of the shellcheck branch so both checks share one answer rather than deciding it
twice and drifting.

Verified that it bites rather than assumed: status_check.yml was given a
`needs: [a-job-that-does-not-exist]`, and the real pre-commit hook blocked with
actionlint's own message -- `job "static-analysis" needs job
"a-job-that-does-not-exist" which does not exist in this workflow [job-needs]`.
Workflow restored; nothing but the hook is in this diff.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-06 16:19:49 -05:00
JMR-devandClaude Opus 5 dd01f9f27c Run shellcheck in the local gate, at CI's exact pin
The gate checked ktlint, detekt and Android lint but not shellcheck, so a new or
edited .sh file was precisely the case where the hook passed and CI's Static
analysis leg still went red. The first file it could not check was itself, and it
was caught by hand twice before it was caught here.

THE DIGEST IS READ OUT OF status_check.yml RATHER THAN COPIED. shellcheck 0.9.0
and 0.11.0 disagree about how to report a trap handler -- SC2317 on seven body
lines against SC2329 once on the declaration, same script, same directive, one
red and one green. That is why CI pins by digest, and it is also why a second
copy of the digest in this file would be worse than none: when it drifts, the
symptom is the gate passing and CI failing, which is the exact failure this
section prevents.

Runs over `git ls-files '*.sh'` -- all tracked files, not the diff -- because
that is what CI does, and the job here is to predict that leg rather than audit
the change. podman is preferred over docker for the mount's SELinux relabel;
neither present, or the digest unreadable, reports the check as NOT COVERED
rather than skipping it quietly.

Verified that it bites rather than assumed: a probe script whose only fault was
an unquoted `ls $foo` was staged, and the real pre-commit hook blocked on SC2086
before it reached the JVM gate. Probe removed; all tracked .sh are clean under
the pinned digest.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-06 16:16:23 -05:00
Jason Ross dd76229e90 Merge pull request #256 from JMR-dev/test/publish-delete-arm-real-provider
Delete the document a failed save could not write (#250)
2026-09-06 16:11:13 -05:00
JMR-devandClaude Opus 5 8105291f6a Make the gate name the levels it ran instead of claiming all of them
The closing line was `green at every supported API level`, printed on both
paths -- including the one that had just said `NOT COVERED LOCALLY: API 37` two
lines above. A false claim, printed by the tool whose entire purpose is to stop
false claims reaching CI, on its first run.

It now names them: `green on API 33, 34, 35, 36` when the Pixel is absent, and
`green on API 33, 34, 35, 36, 37` when it is attached and passed.

Nothing else changes. The app/src subtree is untouched, so this exercises the
cache scoping from the previous commit: the sweep is skipped as already green and
only the JVM gate runs -- which is the whole reason that key was moved off the
repo tree.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-06 16:02:47 -05:00
JMR-devandClaude Opus 5 68bd24e54a Gate commits and pushes on a local sweep at every supported API level
New rule, and a hook rather than a habit. Source work needs the unit tests and
the instrumented tests green at every supported API level before it is committed
or pushed; test work needs the whole suite green at every level.

tools/git-hooks/local-gate.sh is wired in as pre-commit and pre-push (symlinks,
so shellcheck sees one file), enabled with
`git config core.hooksPath tools/git-hooks`.

WHAT "EVERY LEVEL" CAN MEAN HERE, measured rather than assumed. 33-36 run the
whole suite on emulators. API 37 CANNOT be run on an emulator on this host at
all -- not "is red", cannot run: the image logs `3 new surfaceflinger aborts in
45 s (want 0)` and the APK install then fails with `Can't find service: package`,
because the framework is gone before Gradle installs anything. Starting 0 tests.
So 37 runs on the attached Pixel 10 Pro XL when it is there, and the hook says
plainly that the level is uncovered when it is not, rather than claiming five
levels having run four.

The first cut passed a notAnnotation filter through E2E_EXTRA_GRADLE_ARGS, which
run-e2e.sh:587 overwrites with --rerun -- so that argument was discarded and
would have been discarded silently.

The sweep is cached under the app/src SUBTREE hash, not the whole repo tree. The
first cut used the whole tree and that was wrong in a way that would teach people
to resent this hook: editing a comment in CLAUDE.md discarded a sweep of
byte-identical application code and re-ran forty minutes of emulators to prove
nothing. Any change under app/src still invalidates it; the JVM gate always runs.

There is deliberately no skip variable -- that would be --no-verify wearing a
different hat.

Why it is worth the time: #256 spent several gating legs learning one leg at a
time what a sweep answers in one pass, and the failing leg MOVED between runs
(API 35 red then green, API 34 green then red). One leg at a time reads as
someone else's flake; as a sweep it is one signal.

Also here, and the reason the rule arrived now: awaitNode treated "the app has no
composition right now" as a failure rather than as not-yet. fetchSemanticsNodes
throws IllegalStateException when nothing is attached and waitUntil propagates it
on the first poll instead of waiting out the deadline. This class spends much of
its time behind the picker, the save dialog and the permission dialog, so there
is always a window where the app is coming back with no composition -- and on run
34057706195's API 34 leg both SAF tests died in it. Now it is not-yet, with the
last composition error carried into the timeout message so a genuinely dead app
stays diagnosable.

Verified: this commit's own hook swept API 33, 34, 35 and 36 at 71/71 failed=0,
API 34 included.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-06 15:52:21 -05:00
JMR-devandClaude Opus 5 19e35394e1 Bound the conversion against API 35's encode, not API 34's
The API 35 leg of #256 went red on aSaveWritesToTheDocumentTheSystemPickerCreated
with a 120 s ComposeTimeoutException on action.saveFile. It was not a cancelled
job and not the new teardown: run 34056545386's logcat has

  20:05:26.897 FFmpegEngine: ffmpeg ... -c:v libx265 -crf 24 -preset veryfast
  20:07:41.693 ConversionWorker: Routing worker_sample.mp4 ...

134.8 s between the encode starting and the next job in the suite, with no cancel
between them. The conversion was healthy and still running when the bound fired.

CONVERSION_TIMEOUT_MS was 120_000, and its KDoc justified that with "the whole
test takes 11.8 s on the API 34 CI leg" -- a real measurement generalised to an
API level it was never taken on. Adding a second picker test made this class
encode twice, so the second one runs on a more contended emulator and crossed a
line that was already marginal. Now 300_000, justified against the 134.8 s, with
a note not to re-tighten it from a fast leg's timing.

cancelAllWork was SUSPECTED of causing this and did not. A local API 35 run with
it passed, which is what sent me to the logcat. pruneWork is kept because it is
the narrower call -- only finished records need to go, and cancelling live work
is a wider blast radius than teardown in a shared process needs -- and its KDoc
now says it fixed nothing rather than claiming a cause it does not have.

That correction is the point: the first version of that KDoc asserted a cause
from one red CI leg and one green local run on a different machine. A test
carrying a confident wrong explanation is the failure mode this whole read has
been about.

Verified: API 35 at 71/71 failed=0 with the raised bound, and the full gate green.
Production is untouched -- git diff origin/main -- app/src/main is empty.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-06 15:20:51 -05:00
JMR-devandClaude Opus 5 cbbaf74285 Delete the document a failed save could not write (#250)
#226 proved D4's premise -- SAF hands back a document reporting exactly zero
bytes, so destinationIsKnownEmpty can answer true -- and then drove the success
path, where publish's catch is never entered. So deletePartialOutput had still
never run against a real DocumentsProvider; its only assertions were
OutputPublisherPublishTest's, against FakeSafProvider under Robolectric. That is
the same "asserted only against a fake built to match it" shape #226 was filed to
break, one layer down.

RecordingPublisher.failOpen makes openDestination return null, which publish
turns into error("Could not open destination for writing") AFTER its size probe
has run -- so the catch is reached with destinationWasEmpty true on a document
DocumentsUI created seconds earlier. Null rather than a throw because
openDestination's KDoc says a provider that is present and declines is the half
no fake can produce on demand, so that arm is also taken for the first time.

Mutation, measured: delete the deletePartialOutput call and this test fails with
"publish did not delete the document it could not write". Nothing anywhere went
red for that line before.

TWO DEAD ACCESSORS #226 LEFT, and the reason is the same one:

FixtureDocumentsProvider is declared by the test APK and runs in
org.libremediaconverter.test; instrumentation runs in the app's process. A static
in the provider is a different object from the one a test can see, so
deletedDocumentIds() would have read empty forever, and reset(File) deletes under
a filesDir that is not the provider's. Both are removed rather than worked
around. That is E7's process wall from a third side, after ACTION_OPEN_DOCUMENT
and ActivityScenario.

The oracle is the document instead, which crosses the boundary because the app
holds a URI grant for it. Still the path rather than the artefact: the size query
proves the document existed and was empty moments earlier, and one that no longer
answers a query is one something deleted.

CLEANUP IS IN TEARDOWN, and the mutation run is why. A failed save keeps its
staged file deliberately, so this test ends with a finished job for the next
launch to reattach to; its sibling then opened on Converted with no "Choose file"
to tap. The first fix tapped Start over at the end of the test body, which does
not run when the test fails -- so the mutation run turned one real failure into
two, the second looking like an unrelated flake. One cause must produce one red
test.

Baseline 6 -> 7, with the derived counts in CLAUDE.md, the marker KDoc and
status_check.yml moved in the same diff. 71 - 7 is 64, the same gating figure for
the third consecutive time, which is how that paragraph goes stale unnoticed.

Verified: three API 34 runs at 71/71 failed=0, the mutation red on the right
assertion, and the full gate plus pinned actionlint green.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-06 14:58:40 -05:00
Jason Ross fac8e67db2 Merge pull request #255 from JMR-dev/docs/e8-instrumented-coverage
Record the first instrumented coverage measurement, and classify the 32 it found
2026-09-06 14:54:10 -05:00
JMR-devandClaude Opus 5 4de169c99b Record the first instrumented coverage measurement, and classify the 32 it found
E8. The instrumented suite had never been measured: enableAndroidTestCoverage was
unset, so a connected run emitted no .ec at all, and jacocoTestReport reads only
testDebugUnitTest. Measured on API 34 by setting the flag temporarily.

  JVM     2236/2374 line 94.2%   1171/1338 branch 87.5%
  E2E     1711/2374 line 72.1%    669/1354 branch 49.4%
  UNION   2342/2374 line 98.7%   1212/1354 branch 89.5%

The JVM row reproduced the committed figure exactly, which is the control that
says both exec sets match the current class files.

The device suite closes 106 lines the JVM suite misses, and the first four are
the 81 wave 4 wrote off as native or device edges -- FFmpegEngine 32,
Media3Engine 24, ConcatEngine 15, MainActivity 10. The union leaves one. That
confirms the read's own hypothesis rather than overturning it; nobody had
measured past the boundary it named.

All 32 lines reached by neither suite were read, and none is an e2e test gap:
nine are compiler-generated, ten are getForegroundInfo() for expedited work this
app never enqueues (#252), three are F5, three are uncalled members (#253), one
is the Vorbis encode arm (#254), and four are F4-shaped error guards.

MediaProbe:210-212 gained a measurement rather than an assumption. F7 ruled
probeWithExtractor's catch unreachable because Robolectric's MediaExtractor never
throws; probeWithFFprobe calls native ffmpeg-kit, so that reasoning does not
transfer. But probe() calls both, and RemuxTest drives it with garbage bytes on a
device -- so the ffprobe path has had malformed input on real hardware and did
not throw. Same conclusion as F7, different mechanism, now on record.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-06 14:22:24 -05:00
Jason Ross e27d7601b1 Merge pull request #251 from JMR-dev/docs/api37-carrier-count-drift
Re-derive the API 37 carrier counts, and correct what #226 left behind
2026-09-06 14:13:56 -05:00
JMR-devandClaude Opus 5 e81403c5f3 Measure the fourth claim rather than asserting it, and fix three slips
Review of the previous commit found three things of exactly the kind it corrects.

"Both of the advisory runs that exist" asserted exhaustiveness that had not been
checked -- two jobs were read, and the #248 branch had three status_check runs.
All four advisory runs at baseline 6 are now read: 34041156680, 34041593697,
34042397320 and 34043502322 each report expected: 6, received: 4, failed: 4,
and the only SAF test reporting in any of them is the rotation one. The claim
was right; the wording claimed more than the evidence.

CONVERSION_TIMEOUT_MS's KDoc said the bound is "two orders of magnitude" clear
of the real cost. 120 s against a measured 11.8 s is one.

status_check.yml dated the save test's marker to 2026-09-05. fa10d94 is dated
2026-09-06.

Also reworded that comment's summary: "two abort the framework and two never
report" double-counts the picker test, so a reader summing gets seven.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-06 11:20:05 -05:00
JMR-devandClaude Opus 5 54167c052f Re-derive the API 37 carrier counts, and correct what #226 left behind
The 2026-09-06 re-check of the instrumented suite. Every drifted line it found
came from #226, the last PR of the e2e read's own wave.

The suite is 70 tests in 14 classes, 6 carrying @FailsOnEmulatorApi37, gating
leg 64. The committed baseline says 6 and the advisory job agrees. Four places
still said five carriers of 69:

  - CLAUDE.md, three sites
  - FailsOnEmulatorApi37.kt's KDoc
  - two comments in status_check.yml

The gating figure is what hid it. 69 - 5 and 70 - 6 are both 64, so the one
number a reader checks against a run had not moved -- which is exactly why
CLAUDE.md says to derive these rather than remember them.

Two KDoc claims in SafPickerRoundTripTest described a draft rather than the
code. The save test says MP3 was chosen so the setup could not depend on device
codecs; the code converts at the default MP4_H265/FAST, which routes on
canEncode(H265). The negation of the stated reason was true. That is E1 and E3's
failure mode committed by the wave that found it, so it is written down as such
rather than quietly corrected.

Neither picker test has ever reported on the advisory leg. The marker's KDoc
said the picker test fails there behind the rotation test; with six carriers the
rotation test truncates the run first, and both advisory runs since #226 --
34042397320 and 34043502322 -- report expected: 6, received: 4, the four being
the three Media3 tests plus the rotation. The save test is therefore marked by
inheritance, not measurement, and both KDocs now say so.

FixtureDocumentsProvider.deletedDocumentIds() has no callers: #226 proved D4's
premise and drove only the success path, so deletePartialOutput against a real
DocumentsProvider is still asserted nowhere. Filed as #250 with the forcing
condition and the mutation; the accessor is kept with a KDoc naming that ticket
rather than removed and re-added.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-06 11:16:15 -05:00
Jason Ross 1f21557b8d Merge pull request #249 from JMR-dev/docs/e7-second-constraint
Record E7's second constraint, and that the premise held
2026-09-06 10:57:35 -05:00
JMR-devandClaude Opus 5 3b0c262030 Record E7's second constraint, and that the premise held (#226)
Doing #226 turned up a second obstacle underneath E7's, with the same
cause. The obvious way to avoid driving the app was a host Activity in
androidTest owning its own CreateDocument launcher; it cannot be started at
all, because instrumentation runs in the target app's process and the
component is in the instrumentation one. That is the same fact as E7's
second bullet arriving from the other side, and it leaves the app's own
Save button as the only launcher available to drive.

And the answer #226 was filed for: on API 34, stock DocumentsUI hands back
a document URI reporting a size of exactly zero, so destinationIsKnownEmpty
can return true and D4's fix is live rather than inert.

A "no defect found", and not one that could have been reached by reading --
which is the argument for having done it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-06 10:49:09 -05:00
Jason Ross 18aff51c98 Merge pull request #248 from JMR-dev/test/publish-to-a-real-saf-destination
Save to a document stock DocumentsUI created
2026-09-06 10:48:21 -05:00
JMR-devandClaude Opus 5 b23ff0f082 Wait for the app to come back before asking Compose about it
The save test failed an API 35 leg with "No compose hierarchies found in
the app". Dismissing the POST_NOTIFICATIONS dialog presses back and waits
for the permission UI to be gone, but going away and the app being in front
again are not the same moment, and the next Compose query landed in the gap.

Asked of UiAutomator rather than through awaitAppFocus, which is the
opposite of what this class argues for elsewhere and is right here:
awaitAppFocus goes through composeRule.waitUntil, so it would raise the very
error it is being used to avoid.

Two more local API 34 runs at 70/0/0/3.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-06 10:28:02 -05:00
JMR-devandClaude Opus 5 fa10d94192 Save to a document stock DocumentsUI created (#226)
publish deletes a destination it could not write to -- D4's fix, so a
failed save does not leave a truncated file at the name the user chose --
but only when that destination was positively zero bytes first.
destinationIsKnownEmpty is careful that "I could not tell" never authorises
a delete, which makes the precondition load-bearing.

Until now that precondition was asserted only against a fake built to match
it: OutputPublisherPublishTest writes ByteArray(0) into FakeSafProvider
before each case, under a comment stating this is how CreateDocument
behaves. If it were false in production, D4's fix would be inert and every
existing test would still pass.

It is not false. Measured on an API 34 emulator against the real dialog:
the document SAF hands back is a document URI and reports a size of exactly
zero before anything writes to it. RecordingPublisher reads both at the
moment publish sees them, through the ConversionDependencies seam, then
lets the real copy proceed so the bytes are checked too.

This has to go through the picker, and through the app, and both are
platform constraints rather than choices. E7 in docs/e2e-read-findings.md
records the first: a DocumentsProvider is reachable only through a
picker-issued grant. The second was measured here -- a host Activity in
this source set owning its own CreateDocument launcher cannot be started at
all, because instrumentation runs in the target app's process and
ActivityScenario refuses with "Intent in process org.libremediaconverter
resolved to different process org.libremediaconverter.test". So #226 has no
cheap half, which is what its comment now says.

Three things the flow needed, each measured rather than guessed:

Both taps scroll first. On Ready the screen carries a file card, five
pickers and then the button, so Convert is below the fold; performClick on
an off-screen node dispatches where nothing is and throws nothing, while
assertIsEnabled passes either way. The first version sat waiting for a
Converted that could never come.

The format stays at its default. FixtureDocumentsProvider advertises
video/mp4 so the picker's MIME filter has a mutation with a shape, and
DocumentsUI honours that on the save side too: choosing MP3 makes the
destination audio/mpeg and the fixture root is filtered out of the save
dialog entirely.

The notification dialog is dismissed rather than pre-granted. Convert
converts from the permission callback whichever way the answer goes, so
denying is a real user's path and enough. Granting programmatically did not
take -- GrantPermissionsActivity appeared anyway and swallowed the tap.

The provider gains create, write and delete support, which it needs to be a
save target at all. It carries @FailsOnEmulatorApi37 because anything that
puts DocumentsUI on screen aborts system_server on that image, as #245
established for the other two; baseline 5 -> 6.

Verified on a local API 34 emulator: three full-suite runs at 70/0/0/3, and
a publish that writes no bytes fails it with "array lengths differed,
expected.length=58677 actual.length=0".

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-06 10:12:37 -05:00
Jason Ross 69d5392227 Merge pull request #247 from JMR-dev/fix/api37-report-match-line
Stop the advisory match line claiming a failure count nobody measured
2026-09-06 09:11:24 -05:00
JMR-devandClaude Opus 5 e7c3e5688f Stop the match line claiming a failure count nobody measured
This PR's own advisory leg caught it. With five markers and a truncated run it
printed

  failed:            4
  ...
  baseline: matches (5 expected, 5 failed)

three lines apart. The match line has always printed the baseline twice, which
was true while `failed` had to equal it to get there -- and the previous commit
removed that requirement for truncated runs without noticing this line depended
on it.

So the truncated spelling says what happened: `matches (5 expected; 4 of 5
failed, on a run the abort truncated -- not compared)`. Pinned by a third case
beside the two from that commit.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-06 09:03:21 -05:00
Jason Ross b3d4318273 Merge pull request #245 from JMR-dev/fix/api37-task-snapshot-crash
Take the picker test off the API 37 gating leg, and stop the SystemUI disable pretending
2026-09-06 09:01:56 -05:00
JMR-devandClaude Opus 5 07f7ed4259 Merge #221, and make the two counts derived rather than remembered
#221 landed while this was in review, adding one instrumented test: 68 -> 69, and
64 on the gating leg. Its own commit was "Move CLAUDE.md's instrumented counts
with the test that changes them", and it still arrived stale -- it says three
markers and 61 tests, both true of the main it was branched from and neither true
of the main it merged into.

That is the third time these two numbers have gone stale in a day, so the
paragraph now says where they come from: a grep for @Test over app/src/androidTest
minus the marker count, cross-checkable against any run's shape, since a leg below
37 reports the first as `expected` and the API 37 gating leg reports the second.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-06 08:53:59 -05:00
Jason Ross dc6ee3dc9b Merge pull request #221 from JMR-dev/test/join-failure-message-on-device
Assert the join failure message against a real FFmpeg session
2026-09-06 08:50:59 -05:00
JMR-devandClaude Opus 5 7fd95ddede Merge main, and re-derive every count it moved
main landed 25 commits while this branch was open, including a third
@FailsOnEmulatorApi37 on Media3EngineTest.cancellingARunningExportStopsIt and a
batch of new instrumented tests. Every number this branch touches moved with them.

Re-derived rather than adjusted, and cross-checked against run 34020234606: the
API 34 leg (no filter) reports 68 tests and the API 37 gating leg 64, which is
68 minus main's four markers. With the picker test marked that is five markers,
baseline 5, and 63 on the gating leg.

The conflict in FailsOnEmulatorApi37.kt is resolved main's way: it had replaced
the hardcoded "grows by two" with a reference to the constant, which is the same
drift this file exists to prevent and a better fix than the number I put there.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-06 08:49:59 -05:00
Jason Ross 350b179c9e Merge branch 'main' into test/join-failure-message-on-device 2026-09-06 08:43:20 -05:00
Jason Ross 0cc4c4f3a3 Merge pull request #244 from JMR-dev/docs/e2e-read-findings-e7
Record what working the e2e tickets found
2026-09-06 03:32:19 -05:00
JMR-devandClaude Opus 5 c5b2dc0f55 Record what working the e2e tickets found (E7, and #238)
Two results from #223-#230 that belong with the read rather than only in
their own tickets.

E7 re-scoped its own ticket. #226 split into a cheap headless half and an
expensive picker-driven one, on the premise that a real DocumentsProvider
can be reached without DocumentsUI. It cannot: an unprotected one is
refused at install, instrumentation runs in the app's uid so the test APK's
own identity is no help, and adopting shell identity is denied too -- each
denial naming ACTION_OPEN_DOCUMENT as the only way in. Measured three ways.
So #226 is one item at the picker's cost, not two.

The useful half of that distinction is that the input bridge needs no
documents provider at all. getSafParameterForRead opens a descriptor
through the resolver, so any readable content:// URI exercises it, which is
what kept #225 headless.

And that is how the read's one production defect surfaced. #238: joining
files picked through the system picker failed outright on the stream-copy
path, because the concat demuxer whitelists protocols separately from
-safe 0 and ffkitsaf was not on the list. Only STREAM_COPY feeds the
demuxer a list file, and every existing join test passed Uri.fromFile, so
the one broken combination was the only one a user could reach.

Worth stating plainly next to the coverage entry: it was not a missed line
and not an unasserted value, but two covered things no test put together --
the gap shape a coverage number is worst at, and the reason the read
happened.

E4 is marked fixed; #243 made that KDoc name the constant rather than
restate it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-06 02:50:37 -05:00
Jason Ross 8db9a6f52f Merge pull request #243 from JMR-dev/test/cancelling-a-running-export
Cancel a running Media3 export, completing the third engine
2026-09-06 02:48:22 -05:00
JMR-devandClaude Opus 5 31f249ae04 Cancel a running Media3 export, completing #224's third engine
The two FFmpeg engines were done in ad2a75d and d293646. This is
Media3Engine.transcode's invokeOnCancellation, which posts
transformer.cancel() onto the engine's own HandlerThread because cancel()
has the same single-thread requirement as start().

The assertion is the output file here, where it could not be for FFmpeg.
That side deletes the partial on cancellation, and on POSIX ffmpeg keeps
writing to the unlinked inode, so the path stays gone whether or not the
cancel landed -- it asserts the session's return code instead. Media3Engine
deletes nothing, the partial being ConversionWorker's to clean up, so the
file is the evidence.

A cancelled export reports itself two ways and both mean interrupted: no
video track, or MediaExtractor refusing the file outright with "Failed to
instantiate extractor" because there is no moov atom. The first version
treated only the null as success and the exception failed the test, which
is how that was measured. Only a playable file counts as a miss.

The wait before reading is several times the export's own length, so a
cancel that did not land has certainly finished by then: the failure
direction is "the file became playable", never "we did not wait long
enough". The attempt is retried for the reason the other two engines
measured -- a 3 s 320x240 export outruns a naive cancel on a loaded runner
-- and an export that never wrote a file at all is recorded as
inconclusive rather than allowed to pass as a cancellation.

It carries @FailsOnEmulatorApi37, so FAILS_ON_EMULATOR_API37_BASELINE moves
3 -> 4 in this diff. That file also said removing the marker would grow the
gating leg "by two", which has been wrong since the third marker landed; it
now names the constant instead of restating it.

Verified on a local API 34 emulator: 68 tests, 0 failures, 3 skipped; and
with transformer.cancel() removed all five attempts produce a playable
video/hevc and the test fails, naming each one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-06 02:28:18 -05:00
Jason Ross d6e1e3bf86 Merge pull request #242 from JMR-dev/test/reattach-to-a-running-job
Reattach to a conversion that is still running
2026-09-06 02:17:06 -05:00
JMR-devandClaude Opus 5 6992f0e783 Reattach to a conversion that is still running (#230)
Reattachment.rank gives RUNNING the highest rank of all -- "live work
outranks a finished result because a running job is holding a foreground
service" -- and no test on either source set had ever produced one.
ReattachOnLaunchTest covers a job that finished, one whose staged file is
gone, an ambiguous pair, one still queued, and one the user cancelled.
ReattachmentTest exercises the ranking as a pure function over fabricated
snapshots. What was missing is a ViewModel meeting a real running job,
which is also the likeliest reattachment there is: the user starts a
conversion, leaves, and comes back while it is still going.

The engine is a fake, deliberately. The job has to still be running when
the ViewModel is built, and every real conversion in this suite finishes in
about a second -- racing that is what made the cancellation tests flaky
enough to need retries. A SoftwareTranscoder that blocks until released
removes the race outright. Nothing about reattachment depends on which
engine is transcoding: the tag query, Reattachment.choose over live
WorkManager state, and observe's mapping to Converting all run identically
whatever is doing the work.

This is what #230 can actually deliver, and the ticket asked for the answer
either way. Process death itself stays device-manual. D3/D13 already record
that am kill refuses a process holding a foreground service, and there is a
more basic obstacle underneath it: instrumentation runs in the app's own
process, so any route that really killed it would take the test runner with
it and leave nothing to assert with. Observing a relaunch needs two
instrumentation runs, which the runner does not provide. So the closest
observable analogue is a fresh ViewModel, with no memory of the work,
meeting a job that is genuinely mid-flight.

The teardown now resets ConversionDependencies. The suite runs without
Android Test Orchestrator, so a BlockingTranscoder left in place would hang
the next class that converts anything.

Verified on a local API 34 emulator: 67 tests, 0 failures, 3 skipped; and
making RUNNING unreattachable in Reattachment.rank fails this test and
nothing else -- which is also the evidence that the JVM ranking test was
not already covering it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-06 01:58:24 -05:00
Jason Ross bb920b5bd0 Merge pull request #239 from JMR-dev/test/content-uri-reaches-ffmpeg
Let a join read the files the user actually picked
2026-09-06 01:51:10 -05:00
JMR-devandClaude Opus 5 802997439d Let a join read the files the user actually picked (#238, #225)
Joining files picked through the system picker failed outright whenever
the strategy was stream copy -- the matched-files case the UI advertises
as "joined without re-encoding, no quality loss".

  [ffkitsaf @ ...] Protocol 'ffkitsaf' not on whitelist 'file,crypto,data'!
  Error opening input file .../joined_from_content.concat_list.txt

JoinScreen picks with OpenMultipleDocuments, so real inputs are always
content://. ConcatEngine maps each through getSafParameterForRead and
FFmpegConcatCommand writes the resulting ffkitsaf: paths into the concat
list file. The demuxer applies its own protocol whitelist, defaulting to
file,crypto,data, and -safe 0 does not touch it: that permits absolute
paths, this permits the scheme they carry. Two separate gates, and only
one was open.

Nothing caught it because the two halves of the bug never met. Only
STREAM_COPY feeds the demuxer a list file -- REENCODE passes each input
with its own -i, where the whitelist does not apply -- so joining over SAF
worked for mismatched clips. And every join test passed Uri.fromFile,
which takes ConcatEngine's uri.path arm instead of the bridge, so
matchingClipsAreJoinedByStreamCopy exercised stream copy with a file:
path and passed. The one broken combination was the one no test produced
and the only one a user can reach.

That is #225's gap: FFmpegKitConfig.getSafParameterForRead is on every
real conversion and join, and was on no passing test -- only on
UnopenableUriTest's failure side, which proves the error message rather
than the bridge. ContentUriInputTest now drives both the convert and join
paths from a real content:// URI.

It uses a plain ContentProvider, because the documents provider cannot be
reached. Measured three ways: a DOCUMENTS_PROVIDER without MANAGE_DOCUMENTS
is refused at install, instrumentation runs in the target app's process so
Instrumentation.getContext() still carries the app's uid and is denied, and
adoptShellPermissionIdentity(MANAGE_DOCUMENTS) is denied identically -- the
denial naming ACTION_OPEN_DOCUMENT as the only way in. The bridge needs no
documents provider: it opens a descriptor through the resolver, so any
readable content:// URI exercises it, and an ordinary provider may be
exported unprotected. The whole class stays headless. Recorded on #226,
which that also settles: its cheap half does not exist.

Verified on a local API 34 emulator: 66 tests, 0 failures, 3 skipped; and
removing the -protocol_whitelist pair reproduces the production failure
verbatim in the join test and nothing else. FFmpegConcatCommandTest pins
the flag on the JVM.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-06 01:42:03 -05:00
Jason Ross 495eaa4ab8 Merge pull request #241 from JMR-dev/fix/launcher-wiring-waits-for-the-pick
Wait for the pick the launcher test is about
2026-09-06 01:41:55 -05:00
JMR-devandClaude Opus 5 bc8e67888e Wait for the pick the launcher test is about (#220)
onInputPicked does not reach Ready on the calling thread. It hops twice --
withContext(pickDispatcher) { InputQuery.describe(...) } and then the probe
-- and pickDispatcher defaults to Dispatchers.IO, a real background thread
Compose's idling knows nothing about. So deliver() returned with the state
still Idle, and asserting immediately was a race the test usually won.

It lost five times on CI in one day, on PRs whose diffs were instrumented
tests and documentation and could not reach it. Two of those failures came
alongside #125's deadlock and could be argued as fallout; three did not.

waitUntil polls through waitForIdle, draining the main looper each time, so
it sees the recomposition the IO hop eventually posts back.

Injecting the dispatcher would be better and is not available here.
pickDispatcher is a constructor parameter precisely so a test can pin it,
but this test composes the real ConverterScreen, which resolves its own
ViewModel through viewModel() -- the seam is one layer below the launcher
edge this class exists to cover, and reaching for it would mean not testing
that edge.

The evidence is the mutation rather than the repetition count, per the note
#218 left: transposing the two launcher callbacks at ConverterScreen.kt:70
and :83 -- the exact defect this test guards -- still fails it, so the wait
did not make it vacuous. A transposed callback leaves the screen in Idle
forever and it fails on the timeout with the meaning it had before.
Supporting evidence, eight consecutive green runs of the class.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-06 01:34:17 -05:00
Jason Ross 706eea8709 Merge pull request #240 from JMR-dev/fix/cancel-tests-need-a-slower-encode
Stop the cancel tests losing their race on a loaded runner
2026-09-06 01:20:57 -05:00
JMR-devandClaude Opus 5 163ce54b77 Stop the cancel tests losing their race on a loaded runner
Both cancellation tests I added in ad2a75d and d293646 wait for
SessionState.RUNNING and then cancel. That is not enough. The conversion
one passed four consecutive local runs and all five CI legs, then failed
the API 34 and 35 legs of the next PR with state=COMPLETED rc=0, on a diff
that could not reach it. On a loaded runner the thread that observed
RUNNING can be descheduled long enough for a short encode to finish before
it calls cancel. A longer timeout does not help: the wait already
succeeded.

Two changes, because neither is sufficient alone.

A slower encode. The conversion test now targets WEBM_VP9 at BEST, the
slowest thing FFmpegCommandBuilder emits -- libvpx-vp9 -crf 31 -b:v 0,
with -deadline realtime added only on FAST. Probed on an API 34 emulator:
that session is still RUNNING at 1 s and finished by 2 s, against well
under a second for x265 -preset medium.

A bounded retry. An attempt whose session finished before the cancel
landed has not tested anything, so it is a miss rather than a failure and
is retried; only exhausting five attempts fails, and the message reports
every attempt's state and return code so a real breakage is distinguishable
from a slow machine.

The retry does not soften the test. With FFmpegKit.cancel removed from
both engines, every attempt ends COMPLETED, so both still fail -- verified,
each listing five [state=COMPLETED rc=0] outcomes. Two clean runs
beforehand at 64/0/0/3.

The join test gets the same treatment. It has not flaked yet, but it is
the same mechanism and the same fragility, and finding out on CI again is
not worth the round trip.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-06 01:04:03 -05:00
Jason Ross d293646f69 Merge pull request #237 from JMR-dev/test/cancelling-a-running-join
Cancel a running join session too
2026-09-06 00:11:33 -05:00
JMR-devandClaude Opus 5 5416788274 Cancel a running join session too (#224)
The FFmpegEngine half landed in ad2a75d; this is the same gap in
ConcatEngine. Between them, a real native session being asked to stop is
now covered on both FFmpeg paths.

The assertion is the session's return code again, and here that is not
merely the better choice but close to the only one: ConcatEngine does not
delete its output on cancellation at all. Its invokeOnCancellation is
FFmpegKit.cancel and nothing else, where FFmpegEngine's also deletes the
partial. Whether that asymmetry is deliberate is a separate question, so
this asserts what is true of both engines rather than depending on it.

The cancel triggers on SessionState.RUNNING rather than on progress.
ConcatWorker publishes no progress at all, so there is no callback to hang
it on even in principle -- and the conversion side already measured the
deeper reason, that the committed clips outrun a callback-triggered cancel.

The inputs are the mismatched pair on purpose, so ConcatPlanner chooses
REENCODE. A stream copy of two 2 s clips is close to instantaneous and
would leave nothing to interrupt; re-encoding is also the case where a user
would actually reach for Cancel.

Verified on a local API 34 emulator: three runs green at 64/0/0/3, and
replacing invokeOnCancellation's body with an empty block fails this test
and nothing else, with state=COMPLETED rc=0.

Media3Engine's transformer.cancel() is still uncovered and #224 stays open
for it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-06 00:02:52 -05:00
Jason Ross ad2a75d9a0 Merge pull request #236 from JMR-dev/test/cancelling-a-running-session
Cancel a running FFmpeg session, which nothing had ever done
2026-09-05 23:53:13 -05:00
Jason Ross 2b921fafe4 Merge pull request #235 from JMR-dev/test/notification-cancel-action
Press the Cancel button in the notification
2026-09-05 23:45:53 -05:00
JMR-devandClaude Opus 5 98c0e4dba2 Cancel a running FFmpeg session, which nothing had ever done (#224)
Every cancel in app/src/androidTest is WorkManager.cancelWorkById against
work that is queued or already finished: ReattachOnLaunchTest cancels a job
carrying a one-hour initial delay, and another immediately after enqueue.
On the JVM, WorkerCancellationTest and HardwareFallbackTest's cancellation
case drive a SoftwareTranscoder double that records the call. No test on
any source set had asked a real native session to stop. That is
docs/defect-audit.md D10's forcing condition.

It is the one path where cancelling wrong is silently expensive rather than
loudly broken: a missed FFmpegKit.cancel leaves the native process encoding
to completion while the UI says the job is cancelled.

Two things were measured rather than assumed, and both changed the test.

The output file cannot be the assertion. invokeOnCancellation deletes the
path, and on POSIX unlinking a file ffmpeg still holds open leaves ffmpeg
writing to the unlinked inode -- so the path stays gone whether or not the
cancel reached the session, and removing FFmpegKit.cancel passes that check
every time. The session's own verdict is what separates them: a cancelled
session ends with the cancel return code, a completed one does not.

Cancelling from the first progress callback loses the race. It was tried
first and failed with state=COMPLETED rc=0: every committed fixture is
2-3 s at 320x240, and the encode finishes before the first statistics
callback is delivered and acted on. FFmpegKit.listSessions shows the
session RUNNING far earlier, so that is what the test waits for.
QualityTier.BEST is deliberate for the same reason -- preset medium leaves
more of the encode ahead of the cancel.

Verified on a local API 34 emulator. Four consecutive runs green at
62/0/0/3, and removing FFmpegKit.cancel while keeping output.delete fails
this test and nothing else, with state=COMPLETED rc=1.

ConcatEngine and Media3Engine carry the same shape and are not covered
here; #224 stays open for them.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-05 23:45:34 -05:00
JMR-devandClaude Opus 5 557b3edab4 Compare the advisory failure count only on a run that finished
The verification dispatch of the reworked harness (34011072884) came back
4 expected / 3 received / 3 failed, where the one before it (34008889182) had
been 4/4/4 on the identical configuration. Nothing about the test list changed
between them: the abort landed one test earlier and the picker test never
started.

The baseline check would have called that "one now passes", which is the wrong
reading and the kind of notice #120 is about -- a deviation that is wrong often
enough to teach everyone to skim past deviation notices. So `failed` is compared
only when `completed cleanly` is yes, and `expected` is compared always, because
`Starting N tests` is printed before anything can abort and is what actually
answers "is the marked set the size the baseline says".

Two cases in e2e-report-shape-test.sh, as a pair: a truncated run short by one is
not a deviation, and a CLEAN run short by one still is -- so the first cannot have
bought its quiet by disabling the check.

This is a consequence of adding the fourth marker rather than a pre-existing bug
worth its own ticket: with three, the advisory leg had been completing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-05 23:26:56 -05:00
Jason Ross cf540f1ecc Merge pull request #234 from JMR-dev/test/ffmpeg-progress-is-observed
Read the progress percentage FFmpeg has always been computing
2026-09-05 23:21:54 -05:00
JMR-devandClaude Opus 5 e0412329ff Press the Cancel button in the notification (#227)
ConversionNotifications.build attaches one action, wired to
WorkManager.createCancelPendingIntent(id). Until now
createCancelPendingIntent had no references anywhere outside its own
declaration -- no JVM test, no instrumented test.

That is worth more than an ordinary uncovered line. A conversion runs in a
foreground service and the user is invited to leave the app; once they do,
this action is the only way to stop it. If the PendingIntent carries the
wrong id the button does nothing, the notification stays, and the job runs
to completion, with no error, no log and no screen to look at.

The obvious version of this test reads NotificationManager's active
notifications for id 1001 and taps what it finds. Rejected: the
instrumented suite grants no runtime permissions, so POST_NOTIFICATIONS is
denied throughout, and whether a suppressed foreground-service notification
is returned there is a platform detail that varies. The test would be
asserting something about notification visibility rather than about
cancellation. The PendingIntent is the subject and where it is read from is
incidental, so this builds the notification for a real live work id and
fires its action -- a real dispatch reaching real WorkManager, the same way
on every API level.

The job carries an initial delay so it stays ENQUEUED. A conversion of the
3 s fixture finishes in well under a second on an emulator, so racing a
cancel against a running job would be flaky in the direction that fails,
and cancelWorkById acts on ENQUEUED identically. What is under test is
whether firing the action reaches WorkManager with the right id.

Verified both ways on a local API 34 emulator: 61 tests, 0 failures, 3
skipped as written; and building the PendingIntent from a random UUID
instead of the request's id fails the new test and nothing else.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-05 23:20:49 -05:00
JMR-devandClaude Opus 5 97558c259f The root fix was wrong, and this is what it found: the disable does nothing
Three commits back I gave `disable_region_sampling` the `adb root` it needed, on
the strength of `Must be root` appearing in every API 37 leg's log. That part was
right and the conclusion drawn from it was not. api37-debug run 34010167885, with
the restart finally real:

  pm attempt 1: Package com.android.systemui new state: disabled-user
  restarting the framework
  adbd is running as root
  system_server down after 2 s
  NOT DISABLED after the restart -- the package state did not survive

three rounds of it, `final state: SystemUI STILL ENABLED`, and the leg reported
`expected: 0, received: 0`. Making the restart work cost the leg every test it had.

Bisected locally on android-37.0: a `stop` 2 s after `pm disable-user` kills
system_server before PackageManager flushes its delayed write, and a 15 s pause
makes the state survive. That repairs the wrong thing. With the package verified
disabled before AND after a clean restart, `com.android.systemui` comes up 3 s
after `system_server` regardless -- and CI's own logcat says the same with no
restart at all: run 34006456986 verifies the package disabled at 02:29:33 and has
SystemUI pid 4275 alive from 02:28:52 for the whole run.

So `pm disable-user` does not stop SystemUI starting on this image, with or without
a restart, and the restart is removed from all three copies rather than repaired.
What is kept is the 45-second window with no new aborts, which is what was always
doing the work: the boot aborts land at 02:28:18 and 02:28:43 and the wait is what
puts instrumentation at 02:32:42, after them rather than inside one. `pm
disable-user` is kept too, because every green leg and every number quoted about
this row was measured with it applied.

The prose the earlier commits got wrong is corrected in place, and one of the
corrections is good news: status_check.yml's caveat that this row runs a
configuration no other leg or Pixel run uses, so nothing depending on system UI
may trust it, describes a state that has never existed. The row is more comparable
to API 33-36 than it has been claiming, not less.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-05 23:15:57 -05:00
JMR-devandClaude Opus 5 948d53b67e Read the progress percentage FFmpeg has always been computing (#229)
FFmpegEngine derives progress as stats.time / durationMs * 100, and the
statistics callback runs on every conversion in FFmpegEngineTest -- they
all pass durationMs = 3_000. But every call site omits onProgress, so
nothing on any source set had ever looked at the number. Replacing percent
with a constant reddened nothing.

What already existed covers the plumbing downstream and not this: #196
covered the worker's progress lambda with a fake engine that reports
whatever the test tells it to, and ProgressNotificationTest covers the
throttling the same way. The arithmetic was the one part with no reader.

The new test passes 30 s as the duration for a fixture that is exactly
3.000 s, so the conversion still encodes the whole clip and the reported
percentage tops out around 10 rather than 100.

That is what makes it bite. A range check alone is worthless: a constant 0
satisfies both "every value is in 0..100" and "the values never go
backwards", and so does a list of [0, 100]. Pinning the band rejects every
constant, and because the band sits a tenth of the way up it also rejects
an implementation that ignores durationMs, which would report ~100 for the
same run. The bound is loose -- 5..25 for an expected 10 -- because the
last statistics callback can land slightly before the final frame.

Verified on a local API 34 emulator: 61 tests, 0 failures, 3 skipped.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-05 23:13:36 -05:00
Jason Ross 54932e97c6 Merge pull request #232 from JMR-dev/test/fallback-asserts-the-path
Make the fallback test say when it cannot test the fallback
2026-09-05 23:06:41 -05:00
JMR-devandClaude Opus 5 745c4f62ce The local runner had the same missing root, and hid it in /dev/null
Third copy of the same defect. tools/local-emulator/run-e2e.sh sends `adb shell
stop` and `start` to /dev/null, so its `Must be root` was never printed and the
framework restart it credits has never happened either.

That matters for what docs/api-37-emulator-crash.md's abort numbers are evidence
of, so the caveat goes next to them rather than in a commit message: on API 37
the image restarts its own framework every minute or so, and a restart landing
after a successful `pm disable-user` brings back a SystemUI-less zygote on its
own. That produces the recorded rate collapse by accident, and it is why the same
code bought nothing on CI's much quieter swiftshader legs, where the logcat shows
SystemUI alive for the whole run.

Also verified, because the previous commit asserted it: the advisory leg really
does run thePickedInputSurvivesARealRotation before the picker test -- run
34008889182 logs the four in the order Media3, Media3, rotation, picker.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-05 22:58:10 -05:00