Compare commits

...
Author SHA1 Message Date
JMR-dev 7f2a6e1376 Merge remote-tracking branch 'origin/main' into m-129-tmp 2026-08-26 00:34:02 -05:00
Jason Ross dc8b7c3944 Merge pull request #127 from JMR-dev/test/bound-the-hangs
Bound the JVM suite's hangs so a deadlock ends in minutes with a stack
2026-08-26 00:21:48 -05:00
JMR-devandClaude Opus 5 1d80e88f9b Ignore Gradle's .kotlin/, which every local build leaves in the repo root
It has never been committed, so nothing is wrong today -- but nothing stops it
either, and `git add -A` would stage Kotlin build-session state into history.

It belongs beside /build, .gradle and .cxx, which are the same category and are
already here. Placed with them rather than in a section of its own, and left
without a comment: unlike tools/ffmpeg/out/ and .claude/, there is no non-obvious
choice here to explain.

Verified rather than assumed:

  $ git check-ignore -v .kotlin
  .gitignore:16:.kotlin	.kotlin

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 00:14:31 -05:00
JMR-dev 27d7cc0a86 Merge remote-tracking branch 'origin/main' into m-127b-tmp 2026-08-26 00:14:07 -05:00
Jason Ross 6df1bdf31a Merge pull request #115 from JMR-dev/docs/coverage-remeasure
Re-measure coverage, because the figure here predates the test push
2026-08-26 00:04:40 -05:00
JMR-dev c6b581ab5a Merge remote-tracking branch 'origin/main' into m-115b-tmp 2026-08-25 23:57:17 -05:00
JMR-dev c27881ab64 Merge remote-tracking branch 'origin/main' into m-127-tmp 2026-08-25 23:57:00 -05:00
Jason Ross 34641df19d Merge pull request #126 from JMR-dev/ci/baseline-counter-precision
Count the annotation, not the comment saying a test does not carry it
2026-08-25 23:49:01 -05:00
JMR-devandClaude Opus 5 0f39964193 Say which misfire the hang watchdog actually has
The comment described adopting a later build's worker as an edge case. It is the
ordinary CI shape: the worker is found by scanning this daemon's descendants for
GradleWorkerMain, which cannot tell one invocation from the next, and the Unit
tests job runs testDebugUnitTest and jacocoTestReport back to back against one
daemon. Still harmless -- the watchdog only reads and writes -- but a reader
should not have to rediscover that.

Refs #125.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-25 23:47:16 -05:00
JMR-dev 71141b5734 Merge remote-tracking branch 'origin/main' into m-126-tmp 2026-08-25 23:40:37 -05:00
JMR-devandClaude Opus 5 81ad102f2a Stop a deadlocked unit-test run, and make it say what it deadlocked on
The JVM suite had no timeout of any kind, so #125's Room/WorkManager lock-order
inversion ran until something outside it gave up: 47 minutes locally, and on CI
it would burn the Unit tests job's 30-minute cap and report as a job timeout
with no cause. The deadlock is monitor contention, which no interrupt breaks, so
nothing inside the JVM could have ended it either.

The obvious fix does not work here. A JUnit `Timeout` -- as a rule or as
`@Test(timeout = ...)` -- runs the test body on a separate thread, and every Compose test in this
source set goes through Robolectric's paused main looper. Both forms fail with
"main looper can only be controlled from main thread"; the same tests with the
timeout removed pass, so it is the mechanism and not the probe.

So the bound comes from outside the test JVM, where it moves no threads:
`timeout` on the Test tasks kills the forked worker, and a watchdog jstacks that
worker two minutes earlier. The jstack is the point. Gradle's timeout on its own
kills silently, a timed-out run writes no XML for the class that hung, and the
JVM's own "Found one Java-level deadlock" section naming both monitors is the
only reason #125 could be described at all -- so it goes to stdout as well as to
a file, because the Unit tests job uploads only reports/tests/.

Ten minutes is against the slowest observed passing run, not the typical one:
eight CI samples of the whole invocation ranged 62-90s, so this is ~6.7x that
and a third of the job cap. A timeout that fires on a healthy slow runner turns
a real signal into noise.

Both numbers live in a build script that nothing compiles, so HangBoundTest
reads them back and the build script joins build.yml as a declared input --
without that the guard would go stale on exactly the edit it exists to catch.

Refs #125.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-25 23:35:13 -05:00
JMR-devandClaude Opus 5 0f41bc3f6b Count the annotation, not the comment saying a test does not carry it
The advisory baseline check has announced a deviation on every PR since #113:
the tree "carries 4 tests marked @FailsOnEmulatorApi37" where it carries three
and FAILS_ON_EMULATOR_API37_BASELINE says three. The fourth is a KDoc in
Media3EngineTest saying the opposite -- "Deliberately not
`@FailsOnEmulatorApi37`: nothing here decodes or encodes" -- which the old
matcher counted because it looked for the string anywhere on any line.

Neither ingredient was wrong on its own, and the number is not the real damage.
#83 added this check so that a new failure joining the known ones could not be
invisible; a notice that is wrong every single time teaches everyone to skim
past deviation notices, which is precisely the signal it was built to create.
Editing the baseline to 4 would have silenced it by breaking it -- the check
would then have been wrong the moment someone added or removed a real marker.

Anchor the pattern at line start and require whitespace or end-of-line after the
name. The second half is the part that is easy to get wrong: "only the
annotation on a line of its own" also stops counting `@FailsOnEmulatorApi37
@Test`, which is legal Kotlin, and undercounting is the dangerous direction --
it hides a genuine new marker, the one thing this exists to catch. Measured
against a fixture carrying every shape at once: the old matcher 5, own-line-only
2, this one 3; on the real tree 4 / 3 / 3, so the baseline is untouched.
`grep -v import` goes too, since `^[[:space:]]*@` cannot match an import.

The check is a pure function of the working tree, so the fixture is committed
and e2e-report-shape-test.sh runs the real report against it -- inside a
throwaway repo root, which the script finds from BASH_SOURCE, so no knob had to
be added that could point the live count somewhere else. The fixture sits under
.github/, where Gradle does not compile it and :app's ktlint and detekt do not
see it; running the report against the real root with it committed still
reports 3.

Every other path through the report is byte-identical to the previous version on
both stdout and the job summary -- passing, failing, wedged, no-run, and
advisory-with-an-unreadable-baseline all diff empty -- and the two advisory legs
differ only by the false line disappearing. No job's status or pass/fail rules
change; the advisory leg stays continue-on-error and stays red by design.

The test is deliberately not wired into CI: adding a step to Static analysis
would add a new way for a gating job to go red, which #120 ruled out. shellcheck
still covers the file, since that step reads `git ls-files '*.sh'`.

Closes #120

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-25 23:30:55 -05:00
JMR-devandClaude Opus 5 4f1a007b58 Quote the settled figure, now that the work it was waiting on has landed
This PR was opened quoting 81.4%, measured on b53f326. Holding it until #116,
#117, #119, #121 and #124 merged was the point: by the time it was ready the
number had moved three points, which is the same staleness the entry is about.

Measured on 93ebfa6 with ./gradlew :app:jacocoTestReport:

  LINE    1971/2321   84.9%   (81.4% four hours earlier, 69.2% on 2026-08-24)
  BRANCH   900/1410   63.8%   (60.2%, then 53.2%)

454 JVM tests in 67 classes, all green.

The note now says the entry went stale while it was open, because that is a
better argument for the rule than the rule restating itself.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-25 23:24:54 -05:00
JMR-dev bf78e06969 Merge remote-tracking branch 'origin/main' into m-115-tmp 2026-08-25 23:24:29 -05:00
Jason Ross 93ebfa6b4a Merge pull request #117 from JMR-dev/fix/invalid-suggestion-chip
Offer a fix that works when the file has no video to copy
2026-08-25 23:21:09 -05:00
JMR-dev d94906ef42 Merge remote-tracking branch 'origin/main' into m-117b-tmp 2026-08-25 23:13:02 -05:00
Jason Ross 959bd8be13 Merge pull request #121 from JMR-dev/ci/wedged-leg-report
Say in the run-shape table when the wedge timeout was what killed the leg
2026-08-25 23:11:39 -05:00
JMR-dev b93ef79931 Merge remote-tracking branch 'origin/main' into m-117-tmp 2026-08-25 23:04:11 -05:00
JMR-dev d739b425c0 Merge remote-tracking branch 'origin/main' into m-121-tmp 2026-08-25 23:04:09 -05:00
Jason Ross cd77aceff4 Merge pull request #124 from JMR-dev/fix/reattachment-overwrites-pick
Let the user's pick keep the screen a reattachment was about to take
2026-08-25 22:58:57 -05:00
JMR-dev 3d8b89bfab Merge remote-tracking branch 'origin/main' into merge-121-tmp 2026-08-25 22:47:44 -05:00
JMR-dev 3f731d8ea7 Merge remote-tracking branch 'origin/main' into merge-117-tmp 2026-08-25 22:35:01 -05:00
JMR-devandClaude Opus 5 25aac95db9 Say in the run-shape table when the wedge timeout was what killed the leg
The report added by #111 runs on every path out of e2e-run.sh, including the
wedge, and until now it answered a question it had not been asked. On job
98035980326 -- API 34, a docs-only PR -- it printed `received: 59` and
`completed cleanly: yes` six seconds before `##[warning] ... WEDGED`, for a leg
the WEDGE_TIMEOUT had killed 22 minutes in. `completed cleanly` means only
"instrumentation was not aborted", which was true; a reader scanning the table
had to notice a separate warning line to learn the leg had died.

The wedge cannot be read out of the log, which is why it is passed in: a wedge
is gradle never returning, so gradle printed no verdict, no truncation line and
no INSTRUMENTATION_ABORTED, and the log it leaves is the log of a run that just
stops. Only e2e-run.sh saw `timeout` exit 124. It now derives that fact once and
tells the report as E2E_WEDGED_AFTER, and reuses the same variable for
capture_wedge so the two cannot drift.

The table gains a `wedged:` row above `completed cleanly`, and `completed
cleanly` flips to no -- but only where it would have said yes. An abort already
says no and names the abort, which the wedge row does not, and a run that left
no evidence still says unknown; a wedge on top of either prints both facts.

`received`'s source line told the same lie in the same table -- "the run was not
truncated, so every expected test reported" is only "gradle never got as far as
saying so" when the leg was killed -- so it is qualified on that path. The
number itself is unchanged, and so is `failed: unknown`: gradle printed no
summary line, so that count genuinely is not knowable.

Nothing here decides anything. No exit status, no pass/fail rule, no baseline
comparison and no `::notice::` behaviour changes; the leg already failed
correctly and still does.

Verified against captured CI output rather than a live emulator, as #111 was and
for the same reason -- this host cannot run API 37 and cannot wedge on demand.
Four real logs (the wedged leg, a green API 34 leg, a failing gating leg, and an
advisory leg with its baseline deviation) through both versions of the script,
in both env states, comparing stdout and the job summary: only the wedged run
with the signal set differs, byte for byte.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-25 21:59:54 -05:00
JMR-devandClaude Opus 5 dce516224c Offer a fix that works when the file has no video to copy
Refusing "copy the video" for a file that has none built its one suggestion by
hand — drop the video track and leave everything else alone. That is valid only
when the audio axis already happened to be fine. For any audio the target cannot
carry (Vorbis or PCM into MP4, MP3 into WebM) the offer is refused in the next
breath, so the Advanced picker showed a one-tap fix leading straight to a second
error. Nothing unsafe shipped — ConversionWorker re-validates — but it is a dead
end, and it contradicted the promise Validation.Invalid makes in its own KDoc.

Route it through the shared repair-and-filter path instead, as every other branch
does. Excluding what the *user* asked for rather than the already-repaired spec
is what keeps the case that worked working: an MP3 into MP4 still gets its copy
offered, because the repair of a copyable track is that same copy.

Only a branch that builds its own list can break that promise at all, since
suggestions() ends by filtering on validate().isValid. The property test now
covers both of them — this one and the image output — rather than reaching them
by luck, which is how a dead-end chip survived two earlier widenings of it. Its
failures name the probe too: three rows share a spec and differ only in the input.

Closes #114

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-25 21:59:37 -05:00
JMR-devandClaude Opus 5 68f841e84f Re-measure coverage, because the figure here was quoted from before the test push
CLAUDE.md's own rule is "re-measure before quoting", and the figure it carried
was measured on 2026-08-24 -- before the #52 children, the MediaProbe and codec
tests, and the guards from #100/#107 landed. Quoting it now would understate the
suite by twelve points, which is the same failure the bullet directly below it
was written to describe.

Measured on b53f326 with ./gradlew :app:jacocoTestReport:

  LINE    1847/2268   81.4%   (was 1519/2194, 69.2%)
  BRANCH   837/1390   60.2%   (was 53.2%)

against 417 JVM tests in 60 classes, all green.

The denominator moved too, 2194 -> 2268: the same push added production code of
its own, so this is not a pure numerator gain and the note now says so. The
Robolectric/isIncludeNoLocationClasses history is left exactly as it was -- it
explains why every pre-2026-08-24 figure was an artifact, and that is still the
most useful thing in the entry. "That date" is now spelled out, since the
headline date above it has moved and the phrase no longer points at itself.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-25 21:07:28 -05:00
10 changed files with 625 additions and 17 deletions
+193
View File
@@ -0,0 +1,193 @@
#!/usr/bin/env bash
#
# Exercises e2e-report-shape.sh's baseline counter against fixture source, with no emulator and
# no CI run. Run it directly:
#
# .github/scripts/e2e-report-shape-test.sh
#
# WHY THIS CAN EXIST AT ALL: the counter is a pure function of the working tree. It greps
# `app/src/androidTest` for `@FailsOnEmulatorApi37` and compares the total against the number
# committed in FailsOnEmulatorApi37.kt. Nothing about that needs a device, which is the whole
# reason #120 could be measured rather than argued about.
#
# WHY A THROWAWAY REPO ROOT rather than a knob on the script. The report finds its own root from
# `BASH_SOURCE`, so a copy of it placed at `<root>/.github/scripts/` reads `<root>/app/src/...`.
# Building that root is three mkdirs and costs the shipped script nothing:
#
# - the REAL script is what runs, byte for byte, so reverting the matcher reddens this test
# rather than a testing-only code path beside it;
# - no environment variable exists that could point the LIVE count somewhere else, which is
# the failure mode #83 built the baseline check to prevent in the first place;
# - XML_DIR resolves inside the throwaway root, so a stale app/build/outputs left by a real
# run on a developer machine cannot leak into the numbers here.
#
# WHAT IT GUARDS (#120). The old matcher looked for the string anywhere on any line, so a KDoc
# saying `Deliberately not @FailsOnEmulatorApi37` counted as a marked test and every PR got a
# deviation notice that was wrong. The obvious repair -- count only lines that are nothing but
# the annotation -- silently stops counting `@FailsOnEmulatorApi37 @Test`, which is legal Kotlin,
# and undercounting is the direction that hides a genuine new marker. The fixture carries every
# shape at once -- three that count and three that must not, enumerated in its own header -- so
# both mistakes fail here instead of on a PR: against testdata/marker-shapes the old matcher says
# 5, own-line-only says 2, and the shipped one 3.
#
# NOT WIRED INTO CI, deliberately and as a known gap. Adding a step to the Static analysis job
# would add a new way for a gating job to go red, and #120 was explicit that nothing about it may
# change any job's status or the pass/fail rules. shellcheck still covers this file, since that
# step reads `git ls-files '*.sh'` rather than a fixed list.
set -uo pipefail
SCRIPT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)"
REPORT="$SCRIPT_DIR/e2e-report-shape.sh"
FIXTURE_DIR="$SCRIPT_DIR/testdata/marker-shapes"
FIXTURE="$FIXTURE_DIR/MarkerShapes.kt"
TMP="$(mktemp -d)"
# Single quotes: the path is expanded when the trap fires, not when it is set.
trap 'rm -rf -- "$TMP"' EXIT
failures=0
pass() { printf 'ok %s\n' "$1"; }
fail() {
failures=$((failures + 1))
printf 'FAIL %s\n' "$1"
shift
printf ' %s\n' "$@"
}
# assert_contains <name> <haystack> <needle>
# `case` rather than grep: the strings being matched carry backticks and an em dash, and this
# way neither the shell nor a regex engine gets an opinion about them.
assert_contains() {
case "$2" in
*"$3"*) pass "$1" ;;
*) fail "$1" "wanted to find: $3" "in:" "$2" ;;
esac
}
assert_absent() {
case "$2" in
*"$3"*) fail "$1" "did NOT want to find: $3" "in:" "$2" ;;
*) pass "$1" ;;
esac
}
# make_root <marked-tree-dir> <baseline-number>
# Assembles a throwaway repo root around the given tree and prints its path.
make_root() {
local tree="$1" baseline="$2" root testdir
root="$(mktemp -d "$TMP/root.XXXXXX")"
testdir="$root/app/src/androidTest/java/org/libremediaconverter"
mkdir -p "$root/.github/scripts" "$testdir"
cp -- "$REPORT" "$root/.github/scripts/"
cp -- "$tree"/*.kt "$testdir/"
# The synthetic stand-in for the committed baseline. Its KDoc names the marker the way the real
# file does -- in brackets, never with an `@` -- because the real file lives inside the tree
# being counted, so an `@` spelling here would add a phantom to every number below.
cat > "$testdir/FailsOnEmulatorApi37.kt" <<EOF
package org.libremediaconverter
/** Stand-in for the real marker file. Only [FAILS_ON_EMULATOR_API37_BASELINE] is read. */
const val FAILS_ON_EMULATOR_API37_BASELINE = $baseline
EOF
# A clean, untruncated run of exactly <baseline> tests, all failing -- which is what the
# advisory leg looks like when nothing has drifted. It leaves the marked count as the only
# field that can deviate, so every assertion below is about the thing under test.
cat > "$root/gradle.log" <<EOF
> Task :app:connectedDebugAndroidTest
Starting $baseline tests on test(AVD) - 16
There was $baseline failure(s).
EOF
printf '%s\n' "$root"
}
# run_report <root> -- stdout of the real script; its summary lands in <root>/summary.md.
run_report() {
E2E_WEDGED_AFTER='' GITHUB_STEP_SUMMARY="$1/summary.md" \
bash "$1/.github/scripts/e2e-report-shape.sh" 37 "$1/gradle.log" \
"$1/app/src/androidTest/java/org/libremediaconverter/FailsOnEmulatorApi37.kt"
}
# ---------------------------------------------------------------------------
# The fixture still carries every shape.
#
# Three of the checks below are covered twice over -- deleting a real annotation moves the count
# and fails a case further down. The two decoys are not: drop the KDoc mention and the count
# stays 3, so the precision this whole ticket is about would stop being tested and nothing would
# say so. That asymmetry is why the shapes are asserted by name rather than only by their effect
# on the total.
# ---------------------------------------------------------------------------
fixture_text="$(cat -- "$FIXTURE")"
assert_contains "fixture: the import" "$fixture_text" 'import org.libremediaconverter.FailsOnEmulatorApi37'
assert_contains "fixture: annotation own line" "$fixture_text" '
@FailsOnEmulatorApi37
@Test'
assert_contains "fixture: annotation with @Test on one line" "$fixture_text" '@FailsOnEmulatorApi37 @Test'
assert_contains "fixture: annotation nested and indented" "$fixture_text" '
@FailsOnEmulatorApi37'
assert_contains "fixture: KDoc mention (this is #120)" "$fixture_text" "* Deliberately not \`@FailsOnEmulatorApi37\`"
assert_contains "fixture: commented-out annotation" "$fixture_text" '// @FailsOnEmulatorApi37'
# ---------------------------------------------------------------------------
# 1. The fixture's three real annotations against a baseline of 3: no deviation.
#
# This one case fails under both wrong matchers -- the old one counts 5, own-line-only counts 2 --
# which is why it is first.
# ---------------------------------------------------------------------------
root="$(make_root "$FIXTURE_DIR" 3)"
out="$(run_report "$root")"
assert_contains "3 real markers, baseline 3: reports a match" "$out" ' baseline: matches (3 expected, 3 failed)'
assert_absent "3 real markers, baseline 3: says nothing about the tree" "$out" 'the tree carries'
assert_contains "3 real markers, baseline 3: summary agrees" \
"$(cat -- "$root/summary.md")" '**Matches the committed baseline of 3**'
# ---------------------------------------------------------------------------
# 2. A fourth REAL annotation. The count has to move and the deviation has to fire.
#
# The important half of #120: precision was the bug, but a matcher that stopped noticing a new
# marker would have been a worse one, silently.
# ---------------------------------------------------------------------------
plus_one="$(mktemp -d "$TMP/plusone.XXXXXX")"
cp -- "$FIXTURE" "$plus_one/"
cat > "$plus_one/FourthMarker.kt" <<'EOF'
package org.libremediaconverter.fixture
class FourthMarker {
@FailsOnEmulatorApi37
@Test
fun addedToday() = Unit
}
EOF
root="$(make_root "$plus_one" 3)"
out="$(run_report "$root")"
assert_contains "a 4th real marker: the deviation fires, and counts 4" "$out" \
" baseline DEVIATION: the tree carries 4 tests marked \`@FailsOnEmulatorApi37\` but the baseline says 3 — update FAILS_ON_EMULATOR_API37_BASELINE"
assert_contains "a 4th real marker: the summary carries it too" "$(cat -- "$root/summary.md")" \
"- the tree carries 4 tests marked \`@FailsOnEmulatorApi37\` but the baseline says 3"
# ---------------------------------------------------------------------------
# 3. Delete the same-line annotation and the count must drop to 2.
#
# This is the trap, pinned down. `@FailsOnEmulatorApi37 @Test` on one line is what separates the
# shipped matcher from `^[[:space:]]*@NAME[[:space:]]*$`, and without this case the fixture entry
# guarding it could be deleted as decoration -- case 1 would then pass under the wrong matcher.
# Here the same-line entry is worth exactly one, and it is asserted to be.
# ---------------------------------------------------------------------------
minus_same_line="$(mktemp -d "$TMP/minus.XXXXXX")"
sed -e '/@FailsOnEmulatorApi37 @Test/d' -- "$FIXTURE" > "$minus_same_line/MarkerShapes.kt"
root="$(make_root "$minus_same_line" 3)"
out="$(run_report "$root")"
assert_contains "same-line annotation removed: counts 2, so it was worth 1" "$out" \
" baseline DEVIATION: the tree carries 2 tests marked \`@FailsOnEmulatorApi37\` but the baseline says 3 — update FAILS_ON_EMULATOR_API37_BASELINE"
echo
if [ "$failures" -eq 0 ]; then
echo "e2e-report-shape-test.sh: all checks passed"
exit 0
fi
echo "e2e-report-shape-test.sh: $failures check(s) failed"
exit 1
+63 -4
View File
@@ -30,16 +30,31 @@
#
# Usage:
# e2e-report-shape.sh <label> <gradle-log> [<baseline-file>]
# E2E_WEDGED_AFTER=<seconds> the wrapper timeout killed gradle after that many seconds
#
# With a third argument the run is compared against the baseline in that file (advisory mode)
# and a `::notice::` is emitted per deviation. NEVER `::error::`: the advisory job is
# `continue-on-error: true` and stays that way, and an error annotation would be a new way for
# a diagnostic to change a conclusion.
#
# WHY THE WEDGE ARRIVES AS AN ENV VAR (#118) rather than being read out of the log like every
# other field: there is nothing in the log to read. A wedge is gradle never returning, so gradle
# never printed a verdict, never printed a truncation line, and never aborted instrumentation --
# the log of a wedged leg is the log of a run that simply stops. Measured on job 98035980326:
# `expected: 59`, `received: 59`, `completed cleanly: yes`, six seconds before the wedge warning,
# for a leg that the timeout had killed 22 minutes in. Only e2e-run.sh knows, because only it
# saw `timeout` exit 124, so it says so. Guessing it from a log that ends abruptly would call
# every cancelled run a wedge.
#
# It is read as a STRING and only ever interpolated into one. `[ -n ... ]`, never `-gt`: it
# crosses a process boundary from a shell that deliberately sets it EMPTY on every non-wedge
# path, and an arithmetic test on an empty string is the header's rule four paragraphs up.
set -uo pipefail
LABEL="${1:-unknown}"
LOG="${2:-}"
BASELINE_FILE="${3:-}"
WEDGED_AFTER="${E2E_WEDGED_AFTER:-}"
SCRIPT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)"
REPO_ROOT="$(cd -- "$SCRIPT_DIR/../.." && pwd)"
@@ -115,6 +130,13 @@ elif [ -n "$abort_received" ]; then
elif [ -n "$expected" ] && [ -z "$abort_line" ]; then
received="$expected"
received_src="the run was not truncated, so every expected test reported"
# ... unless it was killed, in which case "not truncated" is only "gradle never got as far as
# saying so". This is the branch the wedged leg in #118 took -- with no XML written yet, the
# number is what the runner was TOLD to run, and the source line said the opposite in the same
# table that called the leg clean. The number is deliberately left alone: it is still the best
# available answer, and only the claim about where it came from was wrong.
[ -n "$WEDGED_AFTER" ] \
&& received_src="no test XML was written and gradle never printed a truncation line — but the leg was killed mid-run, so this is what it was told to run, not what reported"
fi
failed="unknown"
@@ -152,6 +174,12 @@ elif [ "$no_run" = "nothing" ]; then
# "cleanly" would be a lie about a run that left no evidence it happened.
completed="unknown"
completed_src="no runner output to read"
elif [ -n "$WEDGED_AFTER" ]; then
# The wedge is checked LAST of the four, so it only ever overrides the `yes`. The two "no"s
# above are already right and name the abort, which the wedge row does not; `unknown` is
# already right too. A wedge on top of an abort is both facts, and both get printed.
completed="**no**"
completed_src="the wrapper timeout killed gradle after ${WEDGED_AFTER}s — instrumentation itself was never aborted, which is why nothing in the log says so"
else
completed="yes"
completed_src="no truncation line and no \`INSTRUMENTATION_ABORTED\`"
@@ -180,11 +208,34 @@ if [ -n "$BASELINE_FILE" ]; then
advisory="yes"
[ -f "$BASELINE_FILE" ] \
&& baseline="$(sed -nE 's/^const val FAILS_ON_EMULATOR_API37_BASELINE = ([0-9]+).*/\1/p' "$BASELINE_FILE" | head -1)"
# The #81 check, verbatim: what the tree actually carries. Reported next to the baseline so a
# stale baseline shows up here rather than only once the emulator disagrees with it.
# What the tree actually carries. Reported next to the baseline so a stale baseline shows up
# here rather than only once the emulator disagrees with it.
#
# ANCHORED AT LINE START, AND WHITESPACE-OR-END-OF-LINE AFTER THE NAME (#120). The #81 check
# this replaces matched the string anywhere on any line, so #113's KDoc reading `Deliberately
# not @FailsOnEmulatorApi37` counted as a fourth marked test and the report announced a
# deviation on every PR. That is worse than a wrong number: #83 built this so a new failure
# could not be invisible, and a notice that is wrong every time teaches everyone to skim past
# deviation notices.
#
# THE OBVIOUS REPAIR IS A TRAP, and the reason for the second half of the pattern.
# `^[[:space:]]*@NAME[[:space:]]*$` -- "the annotation on a line of its own" -- also stops
# counting `@FailsOnEmulatorApi37 @Test`, which is legal Kotlin, and UNDERcounting is the
# dangerous direction: it hides a genuine new marker, which is the one thing this exists to
# catch. Measured against `testdata/marker-shapes`, a fixture carrying every shape at once:
# the old matcher says 5, own-line-only says 2, this one says 3. On the real tree, 4 / 3 / 3.
# e2e-report-shape-test.sh runs that fixture through this whole script.
#
# `^[[:space:]]*@` cannot match an `import` line, so the old `grep -v import` goes with it
# rather than staying to imply a filter is still doing work.
#
# This is a regex over source text and not a parser. An annotation inside a multi-line string,
# or inside a `/* */` block that opened mid-line, would still be counted. Neither exists here;
# if one ever does, this check wants a different tool rather than a longer regex.
if [ -d "$REPO_ROOT/app/src/androidTest" ]; then
marked="$(grep -rn "@FailsOnEmulatorApi37" "$REPO_ROOT/app/src/androidTest" --include='*.kt' \
| grep -v import | grep -c FailsOn || true)"
marked="$(grep -rhcE '^[[:space:]]*@FailsOnEmulatorApi37([[:space:]]|$)' \
"$REPO_ROOT/app/src/androidTest" --include='*.kt' \
| awk '{ total += $1 } END { print total + 0 }' || true)"
fi
fi
@@ -217,6 +268,11 @@ echo "----- RUN SHAPE (api${LABEL}) -----"
echo " expected: $expected"
echo " received: $received"
echo " failed: $failed"
# Above `completed cleanly`, because it is the line that says what happened to the leg and the
# other one only qualifies it. A reader who stops after three rows still sees it.
if [ -n "$WEDGED_AFTER" ]; then
echo " wedged: yes -- gradle was killed after ${WEDGED_AFTER}s and never returned"
fi
echo " completed cleanly: ${completed//\*/}"
if [ -n "$abort_received" ]; then
echo " received before the abort: $abort_received"
@@ -249,6 +305,9 @@ if [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
echo "| expected | $expected | $expected_src |"
echo "| received | $received | $received_src |"
echo "| failed | $failed | $failed_src |"
if [ -n "$WEDGED_AFTER" ]; then
echo "| wedged | **yes** | \`timeout\` fired after ${WEDGED_AFTER}s and killed gradle (exit 124), which is what e2e-run.sh then captured the wedge diagnostics for |"
fi
echo "| completed cleanly | $completed | $completed_src |"
if [ -n "$abort_received" ]; then
echo "| received before the abort | $abort_received | the same line — the XML above counts the truncated test as a failure, this number does not |"
+17 -3
View File
@@ -239,20 +239,34 @@ timeout -k 30s "$WEDGE_TIMEOUT" \
${E2E_EXTRA_GRADLE_ARGS:-} 2>&1 | tee "$GRADLE_LOG" || status=$?
echo "::endgroup::"
# Whether the wrapper timeout fired, decided ONCE. 124 is `timeout` saying it killed the
# command, and two places downstream need that fact: capture_wedge below, and the report, which
# otherwise calls a killed leg `completed cleanly: yes` (#118). Deriving it twice is how those
# two would drift apart -- the report would keep printing after someone changed what a wedge
# means here. It stays a string: empty on every other path, so those legs pass an empty
# E2E_WEDGED_AFTER and the report behaves exactly as before.
wedged=""
[ "$status" -eq 124 ] && wedged="$WEDGE_TIMEOUT"
# The run-shape report: expected/received/failed and whether the run finished, every time,
# green or red. It never changes `status` -- it is a diagnostic, and the header's rule about
# diagnostics applies to it as much as to every probe below.
#
# E2E_WEDGED_AFTER is the wedge, told to the report rather than left for it to infer. It cannot
# be inferred: a wedge is gradle never returning, so gradle printed no verdict at all, and the
# log the report reads looks like a run that simply stopped. Only this script knows the
# difference, because only this script saw the exit status.
#
# The baseline argument, and only it, turns on the comparison, and only the advisory API 37 job
# passes E2E_ADVISORY=1. Comparing on the gating legs would announce a deviation on all five of
# them every run, since they run the whole suite rather than the marked three. They still get
# the report: a truncated run reporting fewer results than it ran is what #108 looks like, and
# `completed cleanly` is the field that shows it.
if [ "${E2E_ADVISORY:-}" = "1" ]; then
bash "$SCRIPT_DIR/e2e-report-shape.sh" "$LABEL" "$GRADLE_LOG" \
E2E_WEDGED_AFTER="$wedged" bash "$SCRIPT_DIR/e2e-report-shape.sh" "$LABEL" "$GRADLE_LOG" \
"$REPO_ROOT/app/src/androidTest/java/org/libremediaconverter/FailsOnEmulatorApi37.kt" || true
else
bash "$SCRIPT_DIR/e2e-report-shape.sh" "$LABEL" "$GRADLE_LOG" || true
E2E_WEDGED_AFTER="$wedged" bash "$SCRIPT_DIR/e2e-report-shape.sh" "$LABEL" "$GRADLE_LOG" || true
fi
if [ "$status" -eq 0 ]; then
@@ -260,7 +274,7 @@ if [ "$status" -eq 0 ]; then
exit 0
fi
if [ "$status" -eq 124 ]; then
if [ -n "$wedged" ]; then
capture_wedge "api${LABEL}"
else
echo "::error::E2E api${LABEL} failed (exit $status)"
@@ -0,0 +1,54 @@
// NOT A TEST, AND NEVER COMPILED. This is fixture data for e2e-report-shape-test.sh, which
// copies it into a throwaway repo root and runs the real report script against that. It lives
// under .github/ deliberately: Gradle only compiles app/src/**, ktlint and detekt are applied to
// :app only, and the report's own count reads app/src/androidTest -- so nothing here can reach
// the build, the linters, or the number the advisory job compares against. Verified by running
// the report against the real repo root with this file committed: still 3.
//
// It carries every shape the counter has to tell apart, in one file, because the bug in #120 was
// exactly that two of them look alike to a substring match. Three count and three must not:
//
// COUNTS the annotation on its own line
// COUNTS the annotation sharing a line with @Test -- legal Kotlin, and the case the
// obvious "own line only" repair silently drops
// COUNTS the annotation indented inside a nested class
// must NOT a KDoc mentioning it -- this is #120 itself, copied from Media3EngineTest
// must NOT a commented-out annotation
// must NOT the import
//
// Three count. That is what the synthetic baseline in the test is set to, so the fixture and the
// baseline agree exactly the way the real tree and FAILS_ON_EMULATOR_API37_BASELINE do.
//
// The `@Test` here is spelled the way a real test spells it so the fixture reads like source
// rather than like a regex exercise. Nothing runs it.
package org.libremediaconverter.fixture
import org.junit.Test
import org.libremediaconverter.FailsOnEmulatorApi37
class MarkerShapes {
@FailsOnEmulatorApi37
@Test
fun ownLine() = Unit
@FailsOnEmulatorApi37 @Test
fun sameLineAsTest() = Unit
/**
* Deliberately not `@FailsOnEmulatorApi37`: nothing here decodes or encodes, so no emulator
* codec is involved and the API 37 image has no quarrel with it.
*/
@Test
fun mentionedInKdoc() = Unit
// @FailsOnEmulatorApi37 -- taken off on 2026-01-01, kept as a note rather than deleted
@Test
fun commentedOut() = Unit
class Nested {
@FailsOnEmulatorApi37
@Test
fun indentedDeeper() = Unit
}
}
+1
View File
@@ -13,6 +13,7 @@
.externalNativeBuild
.cxx
local.properties
.kotlin
# The FFmpeg AAR is committed under bin/ so test runs do not depend on a rebuild.
# Build outputs from tools/ffmpeg are not.
+25 -6
View File
@@ -98,6 +98,13 @@ days. Read it as the current answer, and see the git history if you need the old
is written anyway and says nothing about it — `.github/scripts/e2e-report-shape.sh` is where that
is measured and explained.
Every leg prints that table, advisory or not, and **on the wedge path it also carries a `wedged:`
row** (#118). `completed cleanly` only ever meant "instrumentation was not aborted", which stays
true of a leg the `WEDGE_TIMEOUT` killed 22 minutes in — so without that row the table read
`received: 59, completed cleanly: yes` for a leg that had just died. The wedge is passed to the
report as `E2E_WEDGED_AFTER` by `e2e-run.sh`, which is the only thing that can know it: a wedge
is gradle never returning, so the log it left says nothing about it.
Still true, and the reason the advisory job is not simply deleted: **API 37 needs a manual check on
the Pixel 10 Pro XL before each release.** Those three tests are the one thing CI cannot answer
for.
@@ -123,10 +130,10 @@ install for code that can never run — and on API 37 the full APK does not fit
- The `model` package is excluded from `ReturnCount` and `CyclomaticComplexMethod` only. It is the
decision layer, where one branch is one documented user-visible outcome and the metric counts
answers rather than complexity. Every other rule still applies there.
- **Coverage is reported, not gated** — **69.2% of lines (1519/2194), 53.2% of branches**,
measured 2026-08-24 with `./gradlew :app:jacocoTestReport`.
- **Coverage is reported, not gated** — **84.9% of lines (1971/2321), 63.8% of branches**,
measured 2026-08-26 with `./gradlew :app:jacocoTestReport`, against 454 JVM tests in 67 classes.
**Every figure this file carried before that date was an artifact, roughly half the real one.**
**Every figure this file carried before 2026-08-24 was an artifact, roughly half the real one.**
Robolectric loads classes through its own sandbox classloader with no source location, JaCoCo
skips no-location classes by default, and nothing told it otherwise — so **not one Robolectric
test counted**, and Robolectric is what exercises the framework edge here. The
@@ -140,9 +147,12 @@ install for code that can never run — and on API 37 the full APK does not fit
disproportionately Robolectric, so each one added denominator and no numerator — the measurement
was punishing exactly the tests that were hardest to write.
Two things still hold. A floor needs a baseline that has settled, and this one has now moved by
39 points in a single build change, so it has not. And **re-measure before quoting** — that
instruction is the only reason this was caught.
Two things still hold. A floor needs a baseline that has settled, and this one has not: it moved
39 points in a single build change on 2026-08-24, then another 16 as the #52 test push and the
fixes it turned up landed — 69.2% -> 84.9% line, 53.2% -> 63.8% branch — while the denominator
grew 2194 -> 2321, because that work added production code of its own. And **re-measure before
quoting**: this entry was written quoting 81.4%, measured four hours earlier, and was already
three points stale by the time it was ready to merge.
- **Testable code is not done until it is tested.** If a piece is unit testable, it gets unit
tests before it counts as done. If it is e2e testable, it gets e2e tests. Both clauses apply —
a change that is both needs both.
@@ -236,3 +246,12 @@ Because versions float, a build can change without a commit. `./gradlew :app:dep
`@OptIn`. Android lint's `UnsafeOptInUsageError` catches a missed one.
- **Release builds ship both ABIs.** `-PabiFilters` is a test-run override only; `build.yml`
verifies the released APK carries every ABI and that all native libraries are 16 KB aligned.
- **A JUnit `Timeout` — rule or `@Test(timeout=)` — cannot be used in the JVM suite.** Both run the
test body on a separate thread, and every Compose test here goes through Robolectric's paused
main looper: `UnsupportedOperationException: main looper can only be controlled from main
thread`, from `ShadowPausedLooper` under `RobolectricIdlingStrategy.runUntilIdle`. The identical
tests pass with the timeout removed, so it is the mechanism, not the test. What bounds a hung
run instead is `timeout` on the `Test` tasks plus the jstack watchdog beside it in
`app/build.gradle.kts`, neither of which moves a thread. `HangBoundTest` guards both numbers,
and **a timed-out run writes no XML for the class that hung** — the dump is its only
attribution, so do not delete the watchdog as stray config.
+111
View File
@@ -1,5 +1,7 @@
import org.gradle.api.tasks.PathSensitivity
import org.gradle.testing.jacoco.tasks.JacocoReport
import java.io.File
import java.time.Duration
plugins {
// Applied by id: these two come from the root buildscript classpath, which is what
@@ -217,6 +219,115 @@ tasks.withType<Test>().configureEach {
.withPropertyName("releaseWorkflow")
.withPathSensitivity(PathSensitivity.RELATIVE)
// Same reasoning, same trap: HangBoundTest reads the two numbers below out of this file, and
// they are the one part of the change that does not compile. Without this the task stays
// UP-TO-DATE when the build script changes, so the guard would go stale on exactly the edit
// it exists to catch.
inputs.file(project.file("build.gradle.kts"))
.withPropertyName("moduleBuildScript")
.withPathSensitivity(PathSensitivity.RELATIVE)
// --- Bounding a hung run (#125) -----------------------------------------------------------
//
// This suite had no timeout of any kind, so a hang ran until something outside it gave up.
// #125 is a real Java-level deadlock -- a lock-order inversion between Room's
// TransactionExecutor and WorkManager's SerialExecutorImpl, reached through the WorkInfo flow
// -- and one local run sat in it for 47 minutes. On CI it would burn the Unit tests job's
// 30-minute cap and report as a job timeout with no cause at all.
//
// WHY NOT A JUnit `Timeout` RULE, which is the obvious answer: it runs the test body on a
// separate thread, and this suite is thread-affine. Measured here, `@Rule Timeout` and
// `@Test(timeout = ...)` against a `createComposeRule()` Robolectric test both give:
//
// java.lang.UnsupportedOperationException: main looper can only be controlled from main
// at org.robolectric.shadows.ShadowPausedLooper.executeOnLooper
// at androidx.compose.ui.test.RobolectricIdlingStrategy.runUntilIdle
//
// The same two tests with the timeout removed pass, so that is the mechanism and not the
// probe. Nothing that moves a test off its own thread can be used here.
//
// `Task.timeout` moves nothing -- it stops the forked test JVM from outside. Its weakness is
// that it kills without a thread dump, and the jstack is the only reason #125 could be named
// at all; the watchdog below is what answers that, and it only dumps.
//
// THE NUMBER, against the slowest observed *pass* rather than the typical one. Eight CI runs
// sampled 2026-08-26, whole `./gradlew :app:testDebugUnitTest` invocation with compilation in
// it and this task a subset: 62, 76, 77, 79, 81, 84, 86 and 90 seconds. Locally the task
// itself is ~11 s over 454 tests. Ten minutes is ~6.7x the slowest of those and a third of
// the job's 30-minute cap, so a fired timeout still has room to be reported and uploaded. It
// is deliberately nowhere near the observed duration: a timeout that fires on a healthy slow
// runner turns a real signal into noise and teaches people to re-run reflexively.
timeout.set(Duration.ofMinutes(10))
// The dump, two minutes before the kill. jstack is what turned #125 from "CI timed out" into
// a named lock-order inversion, and `Task.timeout` on its own would have thrown it away.
//
// It is deliberately incapable of failing a build: it reads a live process and writes a file.
// Nothing here kills, interrupts or signals anything, so the worst a misfire can do is leave a
// stack trace nobody needed. It has one, and it is the ordinary CI shape rather than an exotic
// case: the worker is found by scanning this daemon's descendants for GradleWorkerMain, which
// cannot tell one invocation's worker from the next, and the Unit tests job runs
// testDebugUnitTest and jacocoTestReport back to back against the same daemon. If this task's
// own worker lived and died inside a single poll, the watchdog can adopt the following one.
//
// Everything it needs is read here, at configuration time, and captured by value. Reaching
// back through the task or the project from inside the action would not survive the
// configuration cache, which `gradle.properties` turns on for every build.
val threadDump = layout.buildDirectory.file("reports/hang/$name-threads.txt").get().asFile
val taskPath = path
val dumpAfterNanos = Duration.ofMinutes(8).toNanos()
val captureWindowNanos = Duration.ofMinutes(1).toNanos()
val pollMillis = 1_000L
doFirst {
val watchdog = Thread {
val startedAt = System.nanoTime()
var worker: ProcessHandle? = null
while (true) {
Thread.sleep(pollMillis)
val elapsed = System.nanoTime() - startedAt
val watched = worker
if (watched == null) {
// Gradle forks the worker moments after this task starts. If none has shown
// up by the end of the capture window there is nothing to watch, and going on
// polling would only risk adopting some other build's.
if (elapsed > captureWindowNanos) return@Thread
worker = ProcessHandle.current().descendants()
.filter { it.info().commandLine().orElse("").contains("GradleWorkerMain") }
.findFirst().orElse(null)
} else if (!watched.isAlive) {
return@Thread // the run finished; this is the healthy exit
} else if (elapsed >= dumpAfterNanos) {
val jstack = File(File(System.getProperty("java.home"), "bin"), "jstack")
threadDump.parentFile.mkdirs()
if (jstack.canExecute()) {
ProcessBuilder(jstack.absolutePath, "-l", watched.pid().toString())
.redirectErrorStream(true)
.redirectOutput(threadDump)
.start()
.waitFor()
} else {
threadDump.writeText("no jstack at ${jstack.absolutePath}\n")
}
// To stdout as well as to the file, and that is the half that matters on CI:
// the Unit tests job uploads app/build/reports/tests/ and nothing else, so a
// dump that only ever existed under reports/hang/ would be unreachable from a
// red run -- which is the "timed out with no cause" this exists to end. The
// step log always survives, and needs no workflow edit to say so.
println(
"$taskPath is still running after ${Duration.ofNanos(elapsed).toMinutes()} " +
"minutes and is about to be timed out. Thread dump of pid " +
"${watched.pid()}, also written to $threadDump -- look for 'Found one " +
"Java-level deadlock' (that is #125).\n" + threadDump.readText(),
)
return@Thread
}
}
}
watchdog.isDaemon = true
watchdog.name = "hang-watchdog"
watchdog.start()
}
extensions.configure<JacocoTaskExtension> {
isIncludeNoLocationClasses = true
excludes = listOf("jdk.internal.*")
@@ -178,7 +178,12 @@ object ContainerCapabilities {
if (!probe.hasVideo) {
return Validation.Invalid(
"This file has no video track to copy.",
listOf(spec.copy(videoCodec = VideoCodec.NONE)),
// Dropping the video is the right shape of answer, but it is only half of one:
// `spec.copy(videoCodec = NONE)` is valid exactly when the audio axis already
// happened to be fine, and refused otherwise — a Vorbis or PCM source into MP4,
// an MP3 into WebM. Handing it to the shared path repairs both axes and drops
// anything that still fails, so the chip cannot lead to a second error.
suggestions(spec.copy(videoCodec = VideoCodec.NONE), probe, exclude = spec),
)
}
val source = CodecNames.videoFromName(probe.videoCodec)
@@ -0,0 +1,98 @@
package org.libremediaconverter.ci
import org.junit.Assert.assertTrue
import org.junit.Test
import java.io.File
/**
* That a hung unit-test run still ends by itself, and still says why.
*
* `:app:testDebugUnitTest` had no timeout of any kind until #125 was filed. That ticket is a real
* Java-level deadlock between Room's `TransactionExecutor` and WorkManager's `SerialExecutorImpl`,
* reached through the WorkInfo flow the ViewModel collects, and one local run sat in it for 47
* minutes. Nothing inside the suite could break it: the deadlock is monitor contention, which is
* not interruptible, so it runs until something outside the JVM gives up.
*
* Two numbers in `app/build.gradle.kts` are what bound it now, and neither compiles, so nothing
* else would notice their removal:
*
* - `timeout.set(...)` on every `Test` task, which stops the forked test JVM.
* - the watchdog's `dumpAfterNanos`, which jstacks that JVM *before* the timeout kills it.
*
* The second is the one worth guarding hardest, and the one that most looks like stray config.
* Gradle's timeout kills without a thread dump, and the jstack -- with its "Found one Java-level
* deadlock" section naming both monitors -- is the only reason #125 could be described at all.
* The ordering between the two numbers is what makes it work: dump first, kill second. Reverse
* them, or delete the watchdog, and the suite still stops hanging but every hang from then on
* reports as a bare "Timeout has been exceeded" with nothing to read. Measured against a probe
* that hung one test: no test XML was written for the class that hung, so the hanging test itself
* gets no attribution from the report at all.
*
* The range on the timeout is not decoration either, and it is the half a future edit is most
* likely to get wrong. Below it, a healthy-but-slow runner trips the bound and a real signal
* becomes noise people learn to re-run through; above it, CI's 30-minute job cap fires first and
* the bound never gets to say anything.
*
* `ReleasePermissionTest` is the precedent and its caveat applies here too. This asserts the two
* numbers are present, sanely sized and correctly ordered. It cannot assert that the timeout
* fires -- that needs a hang, which is what the whole change exists to prevent. Refs #125.
*/
class HangBoundTest {
@Test
fun `every Test task is bounded, and bounded between the slow runner and the job cap`() {
assertTrue(
"app/build.gradle.kts sets its Test task timeout to ${timeoutMinutes}m, which is " +
"outside $SANE_MINUTES. Under that range a slow CI runner trips a bound meant for " +
"deadlocks -- the slowest observed passing run of the whole invocation was 90s. " +
"Over it, the Unit tests job's own 30-minute cap kills the job first and the " +
"timeout never reports. `null` means the line is gone or the block was rewritten, " +
"and without it #125's deadlock has nothing to stop it: monitor contention breaks " +
"no interrupt, so it runs until CI gives up and reports a timeout with no cause.",
timeoutMinutes in SANE_MINUTES,
)
}
@Test
fun `the thread dump is taken before the timeout kills the JVM it would dump`() {
assertTrue(
"app/build.gradle.kts takes its hang thread dump after ${dumpAfterMinutes}m but times " +
"the task out at ${timeoutMinutes}m, so the JVM is already dead when jstack runs " +
"and every future hang reports as a bare `Timeout has been exceeded`. The dump " +
"has to come first -- it is the only attribution a hanging test gets, since the " +
"test XML never names it.",
(dumpAfterMinutes ?: 0) < (timeoutMinutes ?: 0),
)
}
/** Minutes given to a whole `Test` task before Gradle stops the forked JVM. */
private val timeoutMinutes: Int?
get() = minutesIn("""timeout\.set\(Duration\.ofMinutes\((\d+)\)\)""")
/** Minutes the watchdog waits before jstacking the forked JVM. */
private val dumpAfterMinutes: Int?
get() = minutesIn("""val dumpAfterNanos = Duration\.ofMinutes\((\d+)\)""")
/**
* Read out of the build script rather than from a model: the numbers live in a Kotlin DSL block
* that no unit test can instantiate, and a scan that reports `null` when the shape changes is a
* better trade than not checking them at all.
*/
private fun minutesIn(pattern: String): Int? =
Regex(pattern).find(buildScript.readText())?.groupValues?.get(1)?.toInt()
/**
* Found by walking up rather than by a fixed relative path: Gradle's working directory for the
* unit tests is the module, but that is a default rather than a promise.
*/
private val buildScript: File
get() = generateSequence(File(".").absoluteFile) { it.parentFile }
.map { File(it, "app/build.gradle.kts") }
.firstOrNull { it.isFile }
?: error("could not find app/build.gradle.kts above ${File(".").absolutePath}")
private companion object {
/** Above the slowest observed passing run, below the Unit tests job's `timeout-minutes`. */
val SANE_MINUTES = 3..29
}
}
@@ -11,6 +11,11 @@ import org.junit.Test
* `OutputFormat` used to be twelve hand-picked triples, and its KDoc defended that on the grounds
* that a closed set was what made routing decidable. Opening it up moves that burden here, so this
* is where decidability now has to be proven.
*
* That includes what a refusal offers instead. `Validation.Invalid` promises every suggestion is
* itself valid and names this class as the proof, so a branch that assembles its own suggestion
* list rather than going through `suggestions()` is only checked here if some row happens to reach
* it — which is how a dead-end chip survived two widenings of that table.
*/
class ContainerCapabilitiesTest {
@@ -35,6 +40,29 @@ class ContainerCapabilitiesTest {
container = Container.MP3,
)
/**
* An audio-only input carrying a codec MP4 has no place for at all.
*
* Vorbis lives in Ogg and Matroska; MP4 carries AAC, MP3, Opus and FLAC. That gap is what turns
* a suggestion which merely drops the video track into a second refusal.
*/
private val vorbisSource = InputProbe(
videoCodec = null,
audioCodec = "vorbis",
hasVideo = false,
kind = InputKind.AUDIO_ONLY,
container = Container.OGG,
)
/** The same shape, for the other codec MP4 refuses. One case is a coincidence; two is the rule. */
private val pcmSource = InputProbe(
videoCodec = null,
audioCodec = "pcm_s16le",
hasVideo = false,
kind = InputKind.AUDIO_ONLY,
container = Container.WAV,
)
// --- copy and encode are different questions ----------------------------
/**
@@ -97,7 +125,17 @@ class ContainerCapabilitiesTest {
}
}
/** A suggestion that is itself invalid is worse than no suggestion. */
/**
* A suggestion that is itself invalid is worse than no suggestion.
*
* Only a branch that assembles its own suggestion list can break that promise: [suggestions]
* ends by filtering on `validate(...).isValid`, so everything routed through it is valid by
* construction. Those branches are what this table has to cover — the image output, and copy
* the video from a file that has none, which built its list by hand and came back refused for
* a Vorbis or PCM source into MP4 and an MP3 into WebM. The Advanced picker showed a one-tap
* fix that led straight to a second error, through two widenings of this table that never
* reached the branch.
*/
@Test
fun `every suggestion is itself valid`() {
val cases = listOf(
@@ -109,15 +147,31 @@ class ContainerCapabilitiesTest {
OutputSpec(Container.MP4, VideoCodec.COPY, AudioCodec.NONE) to mp3Source,
OutputSpec(Container.MP4, VideoCodec.NONE, AudioCodec.NONE) to mp3Source,
OutputSpec(Container.MP4, VideoCodec.COPY, AudioCodec.AAC) to mp3Source,
// Copy-the-video-from-a-file-with-no-video, the last branch that built its offer by
// hand. It escaped the five rows above because `spec.copy(videoCodec = NONE)` is valid
// exactly when the audio axis happens to be fine — true for the AAC and MP3 sources
// used there, false for any audio the target container cannot carry.
OutputSpec(Container.MP4, VideoCodec.COPY, AudioCodec.COPY) to vorbisSource,
OutputSpec(Container.MP4, VideoCodec.COPY, AudioCodec.COPY) to pcmSource,
OutputSpec(Container.WEBM, VideoCodec.COPY, AudioCodec.COPY) to mp3Source,
// The same branch with audio the container *can* hold, which is the half that already
// worked and must keep working: the repair here is a copy, so the offer is the very
// spec the caller handed to `suggestions`. It survives only because the exclusion is
// against what the user asked for rather than against the repair.
OutputSpec(Container.MP4, VideoCodec.COPY, AudioCodec.COPY) to mp3Source,
// The one branch that still builds its list by hand, so that it is asserted rather
// than merely reasoned about: an image container takes `None + None` and nothing else,
// which makes its single offer valid by construction.
OutputSpec(Container.GIF, VideoCodec.H264, AudioCodec.AAC) to h264Source,
)
cases.forEach { (spec, probe) ->
val invalid = ContainerCapabilities.validate(spec, probe) as? Validation.Invalid
?: throw AssertionError("expected $spec to be rejected")
assertTrue("no alternatives offered for $spec", invalid.suggestions.isNotEmpty())
assertTrue("no alternatives offered for $spec on $probe", invalid.suggestions.isNotEmpty())
invalid.suggestions.forEach { suggestion ->
assertTrue(
"suggested $suggestion for $spec is itself invalid",
"suggested $suggestion for $spec on $probe is itself invalid",
ContainerCapabilities.validate(suggestion, probe).isValid,
)
}